{
  "id": 331388,
  "title": "Caution of using \"publicly shared codes\"",
  "url": "/competitions/amex-default-prediction/discussion/331388",
  "author_name": "",
  "post_date": "2022-06-17T04:02:42.551209700Z",
  "votes": 26,
  "comment_count": 4,
  "views": 0,
  "content": "<h1><strong>Caution!</strong></h1>\n<p>From my past experience joining Kaggle Competition &amp; the shared code files I observe in the ongoing Kaggle Competitions:</p>\n<p>I notice that many shared codes are neither reproducible, nor applicable in the real \"competition testing period\" after the \"submission period\" passes.</p>\n<p>For example, the current notebook of the best score of the shared code (&gt;=0.796) in this AMEX competition has a deleted dataset that used to be available (<strong><em>update: the notebook I referred to can’t be seen public now since I made comment for that notebook. Probably author just deleted it or made it private. There used to be 3 publicly shared notebooks of score \"&gt;=0796\", but currently there're only two, as of July 19th, 2022. Nevertheless, I will use it as an example to address the issue</em></strong>). And due to the deletion, I think there will be error once we run the file now or in the future (<strong><em>unless author is willing to reshare the deleted dataset in the future</em></strong>), so a score \"0.796\" may not be reproducible.</p>\n<p>Even when the deleted dataset was previously shared, we don't know the methodology of getting output/submission (It probably came from another private methodology which the author may not be willing to share).</p>\n<p>I think it should be better to just show the score purely generated from the shared methodology in the code file here (in this case, the CatBoost model). The score reflected publicly is not a real reflection of the methodology in the code file. A false sense of \"high score\" may be misleading and detrimental for all of the Kagglers who are willing to learn from it. It is also not helpful for those who just want a higher score since the current submission file may include a submission of prediction that is \"fixed\" (I will term it \"fixed submission) and will only work for the current hidden test set. </p>\n<p>However, when the submission is run against the actual \"testing dataset\" (for example, in this competition, the other 49% of the test dataset), I think the notebook will still fail since the “fixed submission” is not a methodology that is shared in the code file, and the final submission largely depends on the “fixed submission” since it usually has a large weight assigned to a \"fixed submission\".</p>\n<h5><strong><em>Current Examples: publicly shared notebooks in the \"AMEX Competition\"</em></strong></h5>\n<p>You can check that the notebook of \"the highest score\" in AMEX competition currently (use filter to sort notebooks by \"best score\"), the \"fixed submission\" is given a weight of 0.955 &amp; the one generated from methodology is only given a weight of 0.045 (see the last chunk of the code file &amp; a \"fixed submission\" uploaded under \"Input\" \"Data Source“ under \"Data\" Section). leading me to question if the methodology described in the code file will have a decent score by itself alone. The link to the code file is: <a href=\"https://www.kaggle.com/code/jazivxt/expressions-of-gluttony\" target=\"_blank\">https://www.kaggle.com/code/jazivxt/expressions-of-gluttony</a></p>\n<p>In addition, it's not hard to see that the notebook of \"the second highest score\" in AMEX competition has a \"submission.csv\" that is generated purely based on the ensemble of 5 \"fixed submissions\". I doubt whether Kagglers can learn much from the notebook except a simple idea of ensembling multiple predictions. The link to this code file is:<br>\n<a href=\"https://www.kaggle.com/code/beezus666/ensemble-weighted-average/data\" target=\"_blank\">https://www.kaggle.com/code/beezus666/ensemble-weighted-average/data</a></p>\n<p>Both files I mention above have a score of &gt;= 0.796 for current hidden test set.</p>\n<p>A macro-view of the two notebooks with the \"highest scores\" can be accessed via the link:  <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/code?competitionId=35332&amp;sortBy=scoreDescending\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/code?competitionId=35332&amp;sortBy=scoreDescending</a></p>\n<h3><em>I've seen these kinds of tricks a lot in the previous Kaggle Competitions as well.</em></h3>\n<h5><strong><em>Past Examples: publicly shared notebooks in the \"Ubiquant Market Competition\"</em></strong></h5>\n<p>I once referenced a notebook of \"ensemble of NN models\". There used to be one NN model they use (uploaded in the input file) that was available to the public. However, several days later, they deleted the model in the input files. So, when I ran it again, it failed since the \"originally available model\" became unavailable. </p>\n<p>Many of the \"public shared notebooks\" in Ubiquant Market Competition that have achieved Pearson correlation coefficient &gt;=0.1550 have either \"deleted dataset\"/\"dataset no longer available\" which is the \"model weights\" made private by the author or \"private dataset\".</p>\n<p>See the following links for the reference (the six notebooks are the ones with the \"highest scores\"):</p>\n<p>You can easily verify by clicking the links below, go to ”Data\" Section, under \"Input\", under \"Data Sources\", You can see it.<br>\nYou can further verify whether they use the \"deleted or private datasets\" by running their notebooks &amp; examining their code lines to see where the problem occurs due to the \"deleted or private datasets\".</p>\n<p><a href=\"https://www.kaggle.com/code/zlhaaaph/ubiquant-zlh-0-159-solution\" target=\"_blank\">https://www.kaggle.com/code/zlhaaaph/ubiquant-zlh-0-159-solution</a><br>\nA score of 0.159 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/renatoreggiani/dnn-model-ensemble-weights-adjusted/data\" target=\"_blank\">https://www.kaggle.com/code/renatoreggiani/dnn-model-ensemble-weights-adjusted/data</a><br>\nA score of 0.1561 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/shigeeeru/add-model-ubiquant-model-ensemble\" target=\"_blank\">https://www.kaggle.com/code/shigeeeru/add-model-ubiquant-model-ensemble</a><br>\nA score of 0.1559 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/diedioskuren/ubiquant-more-models-ensemble\" target=\"_blank\">https://www.kaggle.com/code/diedioskuren/ubiquant-more-models-ensemble</a><br>\nA score of 0.1555 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/tanushka816/dnn-model-ensemble-with-diff-params\" target=\"_blank\">https://www.kaggle.com/code/tanushka816/dnn-model-ensemble-with-diff-params</a><br>\nA score of 0.1554 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/jillanisofttech/ubiquant-market-prediction-model-ensemble\" target=\"_blank\">https://www.kaggle.com/code/jillanisofttech/ubiquant-market-prediction-model-ensemble</a><br>\nA score of 0.1552 (the score here is \"the highest score\" among all the public versions of the code file)</p>\n<p>A macro-view of these \"highest score\" notebooks above can be accessed via the link: <a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/code?competitionId=32053&amp;sortBy=scoreDescending\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/code?competitionId=32053&amp;sortBy=scoreDescending</a></p>\n<p>Therefore, please do read the code carefully before incorporating the methodology in your own notebook! Simply \"copy and edit\" may work temporarily, but it is susceptible to many outside factors (for example, author deletes a dataset or model that used to be available).</p>\n<p>And even if the shared codes work for now (for the current hidden test set), it's always better to check if the final \"submission.csv\" includes a \"fixed submission\" that is not applicable to the real \"testing period\" (ex. another 49% of the test set in AMEX default prediction). </p>\n<p>Finally, I really appreciate those who shared the methodology and insights that will shed light on how to improve the predictive power of the model/features and etc. !<br>\nIn the end, it is really the methodology and knowledge that help us improve, not a pure \"high score\".</p>",
  "messages": [
    {
      "id": "1823107",
      "postDate": "06/17/2022 04:02:42",
      "content": "<h1><strong>Caution!</strong></h1>\n<p>From my past experience joining Kaggle Competition &amp; the shared code files I observe in the ongoing Kaggle Competitions:</p>\n<p>I notice that many shared codes are neither reproducible, nor applicable in the real \"competition testing period\" after the \"submission period\" passes.</p>\n<p>For example, the current notebook of the best score of the shared code (&gt;=0.796) in this AMEX competition has a deleted dataset that used to be available (<strong><em>update: the notebook I referred to can’t be seen public now since I made comment for that notebook. Probably author just deleted it or made it private. There used to be 3 publicly shared notebooks of score \"&gt;=0796\", but currently there're only two, as of July 19th, 2022. Nevertheless, I will use it as an example to address the issue</em></strong>). And due to the deletion, I think there will be error once we run the file now or in the future (<strong><em>unless author is willing to reshare the deleted dataset in the future</em></strong>), so a score \"0.796\" may not be reproducible.</p>\n<p>Even when the deleted dataset was previously shared, we don't know the methodology of getting output/submission (It probably came from another private methodology which the author may not be willing to share).</p>\n<p>I think it should be better to just show the score purely generated from the shared methodology in the code file here (in this case, the CatBoost model). The score reflected publicly is not a real reflection of the methodology in the code file. A false sense of \"high score\" may be misleading and detrimental for all of the Kagglers who are willing to learn from it. It is also not helpful for those who just want a higher score since the current submission file may include a submission of prediction that is \"fixed\" (I will term it \"fixed submission) and will only work for the current hidden test set. </p>\n<p>However, when the submission is run against the actual \"testing dataset\" (for example, in this competition, the other 49% of the test dataset), I think the notebook will still fail since the “fixed submission” is not a methodology that is shared in the code file, and the final submission largely depends on the “fixed submission” since it usually has a large weight assigned to a \"fixed submission\".</p>\n<h5><strong><em>Current Examples: publicly shared notebooks in the \"AMEX Competition\"</em></strong></h5>\n<p>You can check that the notebook of \"the highest score\" in AMEX competition currently (use filter to sort notebooks by \"best score\"), the \"fixed submission\" is given a weight of 0.955 &amp; the one generated from methodology is only given a weight of 0.045 (see the last chunk of the code file &amp; a \"fixed submission\" uploaded under \"Input\" \"Data Source“ under \"Data\" Section). leading me to question if the methodology described in the code file will have a decent score by itself alone. The link to the code file is: <a href=\"https://www.kaggle.com/code/jazivxt/expressions-of-gluttony\" target=\"_blank\">https://www.kaggle.com/code/jazivxt/expressions-of-gluttony</a></p>\n<p>In addition, it's not hard to see that the notebook of \"the second highest score\" in AMEX competition has a \"submission.csv\" that is generated purely based on the ensemble of 5 \"fixed submissions\". I doubt whether Kagglers can learn much from the notebook except a simple idea of ensembling multiple predictions. The link to this code file is:<br>\n<a href=\"https://www.kaggle.com/code/beezus666/ensemble-weighted-average/data\" target=\"_blank\">https://www.kaggle.com/code/beezus666/ensemble-weighted-average/data</a></p>\n<p>Both files I mention above have a score of &gt;= 0.796 for current hidden test set.</p>\n<p>A macro-view of the two notebooks with the \"highest scores\" can be accessed via the link:  <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/code?competitionId=35332&amp;sortBy=scoreDescending\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/code?competitionId=35332&amp;sortBy=scoreDescending</a></p>\n<h3><em>I've seen these kinds of tricks a lot in the previous Kaggle Competitions as well.</em></h3>\n<h5><strong><em>Past Examples: publicly shared notebooks in the \"Ubiquant Market Competition\"</em></strong></h5>\n<p>I once referenced a notebook of \"ensemble of NN models\". There used to be one NN model they use (uploaded in the input file) that was available to the public. However, several days later, they deleted the model in the input files. So, when I ran it again, it failed since the \"originally available model\" became unavailable. </p>\n<p>Many of the \"public shared notebooks\" in Ubiquant Market Competition that have achieved Pearson correlation coefficient &gt;=0.1550 have either \"deleted dataset\"/\"dataset no longer available\" which is the \"model weights\" made private by the author or \"private dataset\".</p>\n<p>See the following links for the reference (the six notebooks are the ones with the \"highest scores\"):</p>\n<p>You can easily verify by clicking the links below, go to ”Data\" Section, under \"Input\", under \"Data Sources\", You can see it.<br>\nYou can further verify whether they use the \"deleted or private datasets\" by running their notebooks &amp; examining their code lines to see where the problem occurs due to the \"deleted or private datasets\".</p>\n<p><a href=\"https://www.kaggle.com/code/zlhaaaph/ubiquant-zlh-0-159-solution\" target=\"_blank\">https://www.kaggle.com/code/zlhaaaph/ubiquant-zlh-0-159-solution</a><br>\nA score of 0.159 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/renatoreggiani/dnn-model-ensemble-weights-adjusted/data\" target=\"_blank\">https://www.kaggle.com/code/renatoreggiani/dnn-model-ensemble-weights-adjusted/data</a><br>\nA score of 0.1561 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/shigeeeru/add-model-ubiquant-model-ensemble\" target=\"_blank\">https://www.kaggle.com/code/shigeeeru/add-model-ubiquant-model-ensemble</a><br>\nA score of 0.1559 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/diedioskuren/ubiquant-more-models-ensemble\" target=\"_blank\">https://www.kaggle.com/code/diedioskuren/ubiquant-more-models-ensemble</a><br>\nA score of 0.1555 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/tanushka816/dnn-model-ensemble-with-diff-params\" target=\"_blank\">https://www.kaggle.com/code/tanushka816/dnn-model-ensemble-with-diff-params</a><br>\nA score of 0.1554 (the score here is \"the highest score\" among all the public versions of the code file)<br>\n<a href=\"https://www.kaggle.com/code/jillanisofttech/ubiquant-market-prediction-model-ensemble\" target=\"_blank\">https://www.kaggle.com/code/jillanisofttech/ubiquant-market-prediction-model-ensemble</a><br>\nA score of 0.1552 (the score here is \"the highest score\" among all the public versions of the code file)</p>\n<p>A macro-view of these \"highest score\" notebooks above can be accessed via the link: <a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/code?competitionId=32053&amp;sortBy=scoreDescending\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/code?competitionId=32053&amp;sortBy=scoreDescending</a></p>\n<p>Therefore, please do read the code carefully before incorporating the methodology in your own notebook! Simply \"copy and edit\" may work temporarily, but it is susceptible to many outside factors (for example, author deletes a dataset or model that used to be available).</p>\n<p>And even if the shared codes work for now (for the current hidden test set), it's always better to check if the final \"submission.csv\" includes a \"fixed submission\" that is not applicable to the real \"testing period\" (ex. another 49% of the test set in AMEX default prediction). </p>\n<p>Finally, I really appreciate those who shared the methodology and insights that will shed light on how to improve the predictive power of the model/features and etc. !<br>\nIn the end, it is really the methodology and knowledge that help us improve, not a pure \"high score\".</p>",
      "rawMarkdown": "#**Caution!**\n\nFrom my past experience joining Kaggle Competition & the shared code files I observe in the ongoing Kaggle Competitions:\n\nI notice that many shared codes are neither reproducible, nor applicable in the real \"competition testing period\" after the \"submission period\" passes.\n\nFor example, the current notebook of the best score of the shared code (>=0.796) in this AMEX competition has a deleted dataset that used to be available (***update: the notebook I referred to can’t be seen public now since I made comment for that notebook. Probably author just deleted it or made it private. There used to be 3 publicly shared notebooks of score \">=0796\", but currently there're only two, as of July 19th, 2022. Nevertheless, I will use it as an example to address the issue***). And due to the deletion, I think there will be error once we run the file now or in the future (***unless author is willing to reshare the deleted dataset in the future***), so a score \"0.796\" may not be reproducible.\n\nEven when the deleted dataset was previously shared, we don't know the methodology of getting output/submission (It probably came from another private methodology which the author may not be willing to share).\n\nI think it should be better to just show the score purely generated from the shared methodology in the code file here (in this case, the CatBoost model). The score reflected publicly is not a real reflection of the methodology in the code file. A false sense of \"high score\" may be misleading and detrimental for all of the Kagglers who are willing to learn from it. It is also not helpful for those who just want a higher score since the current submission file may include a submission of prediction that is \"fixed\" (I will term it \"fixed submission) and will only work for the current hidden test set. \n\nHowever, when the submission is run against the actual \"testing dataset\" (for example, in this competition, the other 49% of the test dataset), I think the notebook will still fail since the “fixed submission” is not a methodology that is shared in the code file, and the final submission largely depends on the “fixed submission” since it usually has a large weight assigned to a \"fixed submission\".\n\n##### ***Current Examples: publicly shared notebooks in the \"AMEX Competition\"***\n\nYou can check that the notebook of \"the highest score\" in AMEX competition currently (use filter to sort notebooks by \"best score\"), the \"fixed submission\" is given a weight of 0.955 & the one generated from methodology is only given a weight of 0.045 (see the last chunk of the code file & a \"fixed submission\" uploaded under \"Input\" \"Data Source“ under \"Data\" Section). leading me to question if the methodology described in the code file will have a decent score by itself alone. The link to the code file is: https://www.kaggle.com/code/jazivxt/expressions-of-gluttony\n\nIn addition, it's not hard to see that the notebook of \"the second highest score\" in AMEX competition has a \"submission.csv\" that is generated purely based on the ensemble of 5 \"fixed submissions\". I doubt whether Kagglers can learn much from the notebook except a simple idea of ensembling multiple predictions. The link to this code file is:\nhttps://www.kaggle.com/code/beezus666/ensemble-weighted-average/data\n\nBoth files I mention above have a score of >= 0.796 for current hidden test set.\n\nA macro-view of the two notebooks with the \"highest scores\" can be accessed via the link:  https://www.kaggle.com/competitions/amex-default-prediction/code?competitionId=35332&sortBy=scoreDescending\n\n### *I've seen these kinds of tricks a lot in the previous Kaggle Competitions as well.*\n\n##### ***Past Examples: publicly shared notebooks in the \"Ubiquant Market Competition\"***\n\nI once referenced a notebook of \"ensemble of NN models\". There used to be one NN model they use (uploaded in the input file) that was available to the public. However, several days later, they deleted the model in the input files. So, when I ran it again, it failed since the \"originally available model\" became unavailable. \n\nMany of the \"public shared notebooks\" in Ubiquant Market Competition that have achieved Pearson correlation coefficient >=0.1550 have either \"deleted dataset\"/\"dataset no longer available\" which is the \"model weights\" made private by the author or \"private dataset\".\n\nSee the following links for the reference (the six notebooks are the ones with the \"highest scores\"):\n\nYou can easily verify by clicking the links below, go to ”Data\" Section, under \"Input\", under \"Data Sources\", You can see it.\nYou can further verify whether they use the \"deleted or private datasets\" by running their notebooks & examining their code lines to see where the problem occurs due to the \"deleted or private datasets\".\n\nhttps://www.kaggle.com/code/zlhaaaph/ubiquant-zlh-0-159-solution\nA score of 0.159 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/renatoreggiani/dnn-model-ensemble-weights-adjusted/data\nA score of 0.1561 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/shigeeeru/add-model-ubiquant-model-ensemble\nA score of 0.1559 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/diedioskuren/ubiquant-more-models-ensemble\nA score of 0.1555 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/tanushka816/dnn-model-ensemble-with-diff-params\nA score of 0.1554 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/jillanisofttech/ubiquant-market-prediction-model-ensemble\nA score of 0.1552 (the score here is \"the highest score\" among all the public versions of the code file)\n\nA macro-view of these \"highest score\" notebooks above can be accessed via the link: https://www.kaggle.com/competitions/ubiquant-market-prediction/code?competitionId=32053&sortBy=scoreDescending\n\nTherefore, please do read the code carefully before incorporating the methodology in your own notebook! Simply \"copy and edit\" may work temporarily, but it is susceptible to many outside factors (for example, author deletes a dataset or model that used to be available).\n\nAnd even if the shared codes work for now (for the current hidden test set), it's always better to check if the final \"submission.csv\" includes a \"fixed submission\" that is not applicable to the real \"testing period\" (ex. another 49% of the test set in AMEX default prediction). \n\nFinally, I really appreciate those who shared the methodology and insights that will shed light on how to improve the predictive power of the model/features and etc. !\nIn the end, it is really the methodology and knowledge that help us improve, not a pure \"high score\".",
      "votes": null
    },
    {
      "id": "1823580",
      "postDate": "06/17/2022 13:57:19",
      "content": "<p>This is a very good and detailed post.<br>\nI would like to add that, in general, it is better to use \"publicly available code\" as a learning resource, and not as a starting point for your own submission.<br>\nThe reason is that, as you correctly point out, the original author may have used a different data set, or may have had access to a private data set. In addition, the original author may have used a different approach that is not reproducible.<br>\nSo, my advice is to use \"publicly available code\" to learn from, and to create your own code from</p>",
      "rawMarkdown": "This is a very good and detailed post.\nI would like to add that, in general, it is better to use \"publicly available code\" as a learning resource, and not as a starting point for your own submission.\nThe reason is that, as you correctly point out, the original author may have used a different data set, or may have had access to a private data set. In addition, the original author may have used a different approach that is not reproducible.\nSo, my advice is to use \"publicly available code\" to learn from, and to create your own code from",
      "votes": null
    },
    {
      "id": "1823725",
      "postDate": "06/17/2022 16:08:27",
      "content": "<p>Thanks for the information sir</p>",
      "rawMarkdown": "Thanks for the information sir",
      "votes": null
    },
    {
      "id": "1824011",
      "postDate": "06/17/2022 21:55:49",
      "content": "<p>This is a very informative post that really helps me save a lot of time in selecting the \"instructive posts\" that have no \"fixed submission\" or \"private, deleted datasets or models\". </p>\n<p>But one thing I want to comment is that while there are a significant number of \"publicly shared codes\" having the problem addressed in this post, we can still learn a bit from the methodology they propose, though the result generated purely from the methodology may be inferior to that of \"fixed submission\" or \"private, deleted datasets or models\" for most of the times.</p>\n<p>I think what Kagglers care the most is really the \"learning process\" from the \"publicly shared codes\", not a pure “high score\" that gives us a false sense of usefulness and usually makes us waste a lot of time reading them.</p>\n<p>I don't want to necessarily say that the creators of uninstructive \"high scores\" intentionally prevent us from learning a superior technique while deceive us giving them upvotes, but those codes are indeed not that helpful, especially for those who are new to the area (whether it be the topic, for example, credit card default, or data science and machine learning in general) and are willing to learn from it.</p>\n<p>Kagglers, please spend your time on picking \"publicly shared codes\" wisely!</p>\n<p>Thanks again for sharing this useful piece of info!</p>",
      "rawMarkdown": "This is a very informative post that really helps me save a lot of time in selecting the \"instructive posts\" that have no \"fixed submission\" or \"private, deleted datasets or models\". \n\nBut one thing I want to comment is that while there are a significant number of \"publicly shared codes\" having the problem addressed in this post, we can still learn a bit from the methodology they propose, though the result generated purely from the methodology may be inferior to that of \"fixed submission\" or \"private, deleted datasets or models\" for most of the times.\n\nI think what Kagglers care the most is really the \"learning process\" from the \"publicly shared codes\", not a pure “high score\" that gives us a false sense of usefulness and usually makes us waste a lot of time reading them.\n\nI don't want to necessarily say that the creators of uninstructive \"high scores\" intentionally prevent us from learning a superior technique while deceive us giving them upvotes, but those codes are indeed not that helpful, especially for those who are new to the area (whether it be the topic, for example, credit card default, or data science and machine learning in general) and are willing to learn from it.\n\nKagglers, please spend your time on picking \"publicly shared codes\" wisely!\n\nThanks again for sharing this useful piece of info!",
      "votes": null
    },
    {
      "id": "1824895",
      "postDate": "06/18/2022 17:52:36",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/franciscofeng\" target=\"_blank\">@franciscofeng</a> , important tips</p>",
      "rawMarkdown": "Thank you @franciscofeng , important tips",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1823580,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "06/17/2022 13:57:19",
      "content": "<p>This is a very good and detailed post.<br>\nI would like to add that, in general, it is better to use \"publicly available code\" as a learning resource, and not as a starting point for your own submission.<br>\nThe reason is that, as you correctly point out, the original author may have used a different data set, or may have had access to a private data set. In addition, the original author may have used a different approach that is not reproducible.<br>\nSo, my advice is to use \"publicly available code\" to learn from, and to create your own code from</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1823725,
      "author_name": "ryoferzz",
      "author_url": "",
      "post_date": "06/17/2022 16:08:27",
      "content": "<p>Thanks for the information sir</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1824011,
      "author_name": "wentinglu",
      "author_url": "",
      "post_date": "06/17/2022 21:55:49",
      "content": "<p>This is a very informative post that really helps me save a lot of time in selecting the \"instructive posts\" that have no \"fixed submission\" or \"private, deleted datasets or models\". </p>\n<p>But one thing I want to comment is that while there are a significant number of \"publicly shared codes\" having the problem addressed in this post, we can still learn a bit from the methodology they propose, though the result generated purely from the methodology may be inferior to that of \"fixed submission\" or \"private, deleted datasets or models\" for most of the times.</p>\n<p>I think what Kagglers care the most is really the \"learning process\" from the \"publicly shared codes\", not a pure “high score\" that gives us a false sense of usefulness and usually makes us waste a lot of time reading them.</p>\n<p>I don't want to necessarily say that the creators of uninstructive \"high scores\" intentionally prevent us from learning a superior technique while deceive us giving them upvotes, but those codes are indeed not that helpful, especially for those who are new to the area (whether it be the topic, for example, credit card default, or data science and machine learning in general) and are willing to learn from it.</p>\n<p>Kagglers, please spend your time on picking \"publicly shared codes\" wisely!</p>\n<p>Thanks again for sharing this useful piece of info!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1824895,
      "author_name": "saberghaderi",
      "author_url": "",
      "post_date": "06/18/2022 17:52:36",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/franciscofeng\" target=\"_blank\">@franciscofeng</a> , important tips</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1823107": "#**Caution!**\n\nFrom my past experience joining Kaggle Competition & the shared code files I observe in the ongoing Kaggle Competitions:\n\nI notice that many shared codes are neither reproducible, nor applicable in the real \"competition testing period\" after the \"submission period\" passes.\n\nFor example, the current notebook of the best score of the shared code (>=0.796) in this AMEX competition has a deleted dataset that used to be available (***update: the notebook I referred to can’t be seen public now since I made comment for that notebook. Probably author just deleted it or made it private. There used to be 3 publicly shared notebooks of score \">=0796\", but currently there're only two, as of July 19th, 2022. Nevertheless, I will use it as an example to address the issue***). And due to the deletion, I think there will be error once we run the file now or in the future (***unless author is willing to reshare the deleted dataset in the future***), so a score \"0.796\" may not be reproducible.\n\nEven when the deleted dataset was previously shared, we don't know the methodology of getting output/submission (It probably came from another private methodology which the author may not be willing to share).\n\nI think it should be better to just show the score purely generated from the shared methodology in the code file here (in this case, the CatBoost model). The score reflected publicly is not a real reflection of the methodology in the code file. A false sense of \"high score\" may be misleading and detrimental for all of the Kagglers who are willing to learn from it. It is also not helpful for those who just want a higher score since the current submission file may include a submission of prediction that is \"fixed\" (I will term it \"fixed submission) and will only work for the current hidden test set. \n\nHowever, when the submission is run against the actual \"testing dataset\" (for example, in this competition, the other 49% of the test dataset), I think the notebook will still fail since the “fixed submission” is not a methodology that is shared in the code file, and the final submission largely depends on the “fixed submission” since it usually has a large weight assigned to a \"fixed submission\".\n\n##### ***Current Examples: publicly shared notebooks in the \"AMEX Competition\"***\n\nYou can check that the notebook of \"the highest score\" in AMEX competition currently (use filter to sort notebooks by \"best score\"), the \"fixed submission\" is given a weight of 0.955 & the one generated from methodology is only given a weight of 0.045 (see the last chunk of the code file & a \"fixed submission\" uploaded under \"Input\" \"Data Source“ under \"Data\" Section). leading me to question if the methodology described in the code file will have a decent score by itself alone. The link to the code file is: https://www.kaggle.com/code/jazivxt/expressions-of-gluttony\n\nIn addition, it's not hard to see that the notebook of \"the second highest score\" in AMEX competition has a \"submission.csv\" that is generated purely based on the ensemble of 5 \"fixed submissions\". I doubt whether Kagglers can learn much from the notebook except a simple idea of ensembling multiple predictions. The link to this code file is:\nhttps://www.kaggle.com/code/beezus666/ensemble-weighted-average/data\n\nBoth files I mention above have a score of >= 0.796 for current hidden test set.\n\nA macro-view of the two notebooks with the \"highest scores\" can be accessed via the link:  https://www.kaggle.com/competitions/amex-default-prediction/code?competitionId=35332&sortBy=scoreDescending\n\n### *I've seen these kinds of tricks a lot in the previous Kaggle Competitions as well.*\n\n##### ***Past Examples: publicly shared notebooks in the \"Ubiquant Market Competition\"***\n\nI once referenced a notebook of \"ensemble of NN models\". There used to be one NN model they use (uploaded in the input file) that was available to the public. However, several days later, they deleted the model in the input files. So, when I ran it again, it failed since the \"originally available model\" became unavailable. \n\nMany of the \"public shared notebooks\" in Ubiquant Market Competition that have achieved Pearson correlation coefficient >=0.1550 have either \"deleted dataset\"/\"dataset no longer available\" which is the \"model weights\" made private by the author or \"private dataset\".\n\nSee the following links for the reference (the six notebooks are the ones with the \"highest scores\"):\n\nYou can easily verify by clicking the links below, go to ”Data\" Section, under \"Input\", under \"Data Sources\", You can see it.\nYou can further verify whether they use the \"deleted or private datasets\" by running their notebooks & examining their code lines to see where the problem occurs due to the \"deleted or private datasets\".\n\nhttps://www.kaggle.com/code/zlhaaaph/ubiquant-zlh-0-159-solution\nA score of 0.159 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/renatoreggiani/dnn-model-ensemble-weights-adjusted/data\nA score of 0.1561 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/shigeeeru/add-model-ubiquant-model-ensemble\nA score of 0.1559 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/diedioskuren/ubiquant-more-models-ensemble\nA score of 0.1555 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/tanushka816/dnn-model-ensemble-with-diff-params\nA score of 0.1554 (the score here is \"the highest score\" among all the public versions of the code file)\nhttps://www.kaggle.com/code/jillanisofttech/ubiquant-market-prediction-model-ensemble\nA score of 0.1552 (the score here is \"the highest score\" among all the public versions of the code file)\n\nA macro-view of these \"highest score\" notebooks above can be accessed via the link: https://www.kaggle.com/competitions/ubiquant-market-prediction/code?competitionId=32053&sortBy=scoreDescending\n\nTherefore, please do read the code carefully before incorporating the methodology in your own notebook! Simply \"copy and edit\" may work temporarily, but it is susceptible to many outside factors (for example, author deletes a dataset or model that used to be available).\n\nAnd even if the shared codes work for now (for the current hidden test set), it's always better to check if the final \"submission.csv\" includes a \"fixed submission\" that is not applicable to the real \"testing period\" (ex. another 49% of the test set in AMEX default prediction). \n\nFinally, I really appreciate those who shared the methodology and insights that will shed light on how to improve the predictive power of the model/features and etc. !\nIn the end, it is really the methodology and knowledge that help us improve, not a pure \"high score\".",
    "1823580": "This is a very good and detailed post.\nI would like to add that, in general, it is better to use \"publicly available code\" as a learning resource, and not as a starting point for your own submission.\nThe reason is that, as you correctly point out, the original author may have used a different data set, or may have had access to a private data set. In addition, the original author may have used a different approach that is not reproducible.\nSo, my advice is to use \"publicly available code\" to learn from, and to create your own code from",
    "1823725": "Thanks for the information sir",
    "1824011": "This is a very informative post that really helps me save a lot of time in selecting the \"instructive posts\" that have no \"fixed submission\" or \"private, deleted datasets or models\". \n\nBut one thing I want to comment is that while there are a significant number of \"publicly shared codes\" having the problem addressed in this post, we can still learn a bit from the methodology they propose, though the result generated purely from the methodology may be inferior to that of \"fixed submission\" or \"private, deleted datasets or models\" for most of the times.\n\nI think what Kagglers care the most is really the \"learning process\" from the \"publicly shared codes\", not a pure “high score\" that gives us a false sense of usefulness and usually makes us waste a lot of time reading them.\n\nI don't want to necessarily say that the creators of uninstructive \"high scores\" intentionally prevent us from learning a superior technique while deceive us giving them upvotes, but those codes are indeed not that helpful, especially for those who are new to the area (whether it be the topic, for example, credit card default, or data science and machine learning in general) and are willing to learn from it.\n\nKagglers, please spend your time on picking \"publicly shared codes\" wisely!\n\nThanks again for sharing this useful piece of info!",
    "1824895": "Thank you @franciscofeng , important tips"
  },
  "source": "meta"
}