{
  "id": 174920,
  "title": "Where is the warning banner for notebooks?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174920",
  "author_name": "Gilles Vandewiele",
  "post_date": "2020-08-16T08:23:18.388000",
  "votes": 44,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Seems like with the new front-end changes, the banner that warns people to <strong>NOT</strong> share high-scoring notebooks in the final week of the competition has vanished.</p>\n<p>As a result, <a href=\"https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing\" target=\"_blank\">someone thought it was a fantastic idea to post a blending kernel</a> that scores 0.9648 and scores you an easy bronze medal. Two days before the end of the competition, this punishes everyone that worked hard and cannot log in by the end of the competition.</p>",
  "messages": [
    {
      "id": 972090,
      "postDate": "2020-08-16T08:23:18.387Z",
      "content": "<p>Seems like with the new front-end changes, the banner that warns people to <strong>NOT</strong> share high-scoring notebooks in the final week of the competition has vanished.</p>\n<p>As a result, <a href=\"https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing\" target=\"_blank\">someone thought it was a fantastic idea to post a blending kernel</a> that scores 0.9648 and scores you an easy bronze medal. Two days before the end of the competition, this punishes everyone that worked hard and cannot log in by the end of the competition.</p>",
      "rawMarkdown": "Seems like with the new front-end changes, the banner that warns people to **NOT** share high-scoring notebooks in the final week of the competition has vanished.\n\nAs a result, [someone thought it was a fantastic idea to post a blending kernel](https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing) that scores 0.9648 and scores you an easy bronze medal. Two days before the end of the competition, this punishes everyone that worked hard and cannot log in by the end of the competition.",
      "votes": 44
    },
    {
      "id": 972099,
      "postDate": "2020-08-16T08:30:23.463Z",
      "content": "<p>No surprise he disabled comments. We finally need downvote for kernels.</p>",
      "rawMarkdown": "No surprise he disabled comments. We finally need downvote for kernels.",
      "votes": 19,
      "replies": [
        {
          "id": 972108,
          "postDate": "2020-08-16T08:40:09.727Z",
          "content": "<p>It is time that Kaggle organizes an analytics competition (where participants just make notebooks and analyses) around flagging of toxic things for which it can be difficult to obtain objective evidence. This includes: upvoting schemes for notebooks/topics, forum spamming, private sharing, and so on.</p>\n<p>As in some games, we could use a concept of trusted \"gamemasters\" that can follow up suspicious activity by e.g. asking questions or asking participants to (somewhat) reproduce approaches. These could then punish people (and their activities can be publicly known).</p>\n<p>Kaggle, can we at least have an elaborate discussion on this? Because it is clear that the current progression system, as is, encourages bad behaviour and stimulates gaming the system.</p>",
          "rawMarkdown": "It is time that Kaggle organizes an analytics competition (where participants just make notebooks and analyses) around flagging of toxic things for which it can be difficult to obtain objective evidence. This includes: upvoting schemes for notebooks/topics, forum spamming, private sharing, and so on.\n\nAs in some games, we could use a concept of trusted \"gamemasters\" that can follow up suspicious activity by e.g. asking questions or asking participants to (somewhat) reproduce approaches. These could then punish people (and their activities can be publicly known).\n\nKaggle, can we at least have an elaborate discussion on this? Because it is clear that the current progression system, as is, encourages bad behaviour and stimulates gaming the system.",
          "votes": 2
        }
      ]
    },
    {
      "id": 972739,
      "postDate": "2020-08-16T19:11:49.700Z",
      "content": "<p>I feel sad for the people who upvote such notebooks.</p>\n<p>As <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> mentioned, I've never understood why Kaggle has downvotes for discussions and not for notebooks and datasets. Sure, there are some downvoting abusers and debateable-cases but most of the times the net votes on discussion posts give a fair view of their value.</p>\n<p>Kaggle is unfairly partial to notebooks (and datasets).</p>",
      "rawMarkdown": "I feel sad for the people who upvote such notebooks.\n\nAs @philippsinger mentioned, I've never understood why Kaggle has downvotes for discussions and not for notebooks and datasets. Sure, there are some downvoting abusers and debateable-cases but most of the times the net votes on discussion posts give a fair view of their value.\n\nKaggle is unfairly partial to notebooks (and datasets).",
      "votes": 12
    },
    {
      "id": 972359,
      "postDate": "2020-08-16T13:42:02.823Z",
      "content": "<p>Permanent ban of the account is necessary. </p>\n<p>He doesn't even hide his will to screw up the LB. </p>",
      "rawMarkdown": "Permanent ban of the account is necessary. \n\nHe doesn't even hide his will to screw up the LB. ",
      "votes": 7,
      "replies": [
        {
          "id": 972804,
          "postDate": "2020-08-16T20:44:54.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> agreed consequences, or at least a sanction (banned from comp for 1 year or something).  Also, anyone that uses that kernel should be disqualified.</p>",
          "rawMarkdown": "@serigne agreed consequences, or at least a sanction (banned from comp for 1 year or something).  Also, anyone that uses that kernel should be disqualified.",
          "votes": -2
        }
      ]
    },
    {
      "id": 973065,
      "postDate": "2020-08-17T04:58:43.910Z",
      "content": "<p>My reaction in the last week of every competition where this happens<br>\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/973065/16606/IMG_1333.JPG\" alt=\"\"></p>",
      "rawMarkdown": "My reaction in the last week of every competition where this happens\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/973065/16606/IMG_1333.JPG)",
      "votes": 6,
      "replies": [
        {
          "id": 973832,
          "postDate": "2020-08-17T14:43:54.403Z",
          "content": "<p>Best reaction so far :')</p>",
          "rawMarkdown": "Best reaction so far :')",
          "votes": 2
        },
        {
          "id": 974578,
          "postDate": "2020-08-18T01:41:50.683Z",
          "content": "<p>Catching up on the forums today and this one made me lol. </p>",
          "rawMarkdown": "Catching up on the forums today and this one made me lol. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 972664,
      "postDate": "2020-08-16T18:10:58.387Z",
      "content": "<p>I have a 0.9593 submission that scores 0.95 on CV. Pearson correlation between my predictions and that submission is 0.48 so hopefully it's overfitted to public LB. 😬😬</p>",
      "rawMarkdown": "I have a 0.9593 submission that scores 0.95 on CV. Pearson correlation between my predictions and that submission is 0.48 so hopefully it's overfitted to public LB. 😬😬",
      "votes": 6,
      "replies": [
        {
          "id": 972771,
          "postDate": "2020-08-16T20:12:40.590Z",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congratulations on reaching very good CV/LB scores! Did you rank your predictions and the submission in question before computing the correlation coefficient? Just curious… (I would compute it myself but the kernel has been removed already).</p>",
          "rawMarkdown": "@vaillant Congratulations on reaching very good CV/LB scores! Did you rank your predictions and the submission in question before computing the correlation coefficient? Just curious... (I would compute it myself but the kernel has been removed already)."
        },
        {
          "id": 972787,
          "postDate": "2020-08-16T20:26:41.533Z",
          "content": "<p>I did rank both of them before computing the correlation.</p>",
          "rawMarkdown": "I did rank both of them before computing the correlation.",
          "votes": 1
        },
        {
          "id": 972825,
          "postDate": "2020-08-16T21:26:10.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Great! Thank you for clarifying this.</p>",
          "rawMarkdown": "@vaillant Great! Thank you for clarifying this."
        },
        {
          "id": 973040,
          "postDate": "2020-08-17T04:08:10.733Z",
          "content": "<p>For AUC metric, does rank co-relation matter if the target classes are well separated?</p>",
          "rawMarkdown": "For AUC metric, does rank co-relation matter if the target classes are well separated?",
          "votes": 1
        },
        {
          "id": 973134,
          "postDate": "2020-08-17T06:04:11.897Z",
          "content": "<p><a href=\"https://www.kaggle.com/manjeshg03\" target=\"_blank\">@manjeshg03</a> AUC is not affected by the ranking operation. But when you are computing a correlation coefficient between two different sets of predictions generated in two very different ways, it is generally a good idea to rank them first. This way you are calibrating them to the same scale, so to speak.</p>",
          "rawMarkdown": "@manjeshg03 AUC is not affected by the ranking operation. But when you are computing a correlation coefficient between two different sets of predictions generated in two very different ways, it is generally a good idea to rank them first. This way you are calibrating them to the same scale, so to speak.",
          "votes": 1
        },
        {
          "id": 973144,
          "postDate": "2020-08-17T06:15:52.630Z",
          "content": "<p>Spearman correlation is better suited for AUC</p>",
          "rawMarkdown": "Spearman correlation is better suited for AUC",
          "votes": 1
        }
      ]
    },
    {
      "id": 972630,
      "postDate": "2020-08-16T17:34:14.973Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2861255%2Fd1fe2c5cbed84cf8320c03f170a258f1%2Fclass.jpg?generation=1597599248620054&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2861255%2Fd1fe2c5cbed84cf8320c03f170a258f1%2Fclass.jpg?generation=1597599248620054&alt=media)",
      "votes": 6
    },
    {
      "id": 972917,
      "postDate": "2020-08-17T01:17:16.463Z",
      "content": "<p>I guess I will be quitting kaggle again :( I really love the competitive nature. Gets me motivated. But when there are notebooks like this and also the 0.9619 that pads out like 400 places starting from low bronze.. I just feel cheated. Especially when my best personal submissions is ever so slightly worse than it.</p>\n<p>Fortunately, that new kernel is down already. Unfortunately I didn't get to run it :(</p>",
      "rawMarkdown": "I guess I will be quitting kaggle again :( I really love the competitive nature. Gets me motivated. But when there are notebooks like this and also the 0.9619 that pads out like 400 places starting from low bronze.. I just feel cheated. Especially when my best personal submissions is ever so slightly worse than it.\n\nFortunately, that new kernel is down already. Unfortunately I didn't get to run it :(",
      "votes": 3,
      "replies": [
        {
          "id": 972931,
          "postDate": "2020-08-17T01:30:41.683Z",
          "content": "<p>well better for you not to have run it………….thats my opinion.  I stick to my own work.</p>",
          "rawMarkdown": "well better for you not to have run it.............thats my opinion.  I stick to my own work.",
          "votes": 2
        },
        {
          "id": 972938,
          "postDate": "2020-08-17T01:35:19.083Z",
          "content": "<p><a href=\"https://www.kaggle.com/cherring\" target=\"_blank\">@cherring</a> Don't feel discouraged. Wait until after private Lb is revealed before making a conclusion. Many teams are overfitting public LB. (And I think the LB 0.965 notebook that you didn't see was very overfitted). Just keep working on your own work and I think you will finish well on private LB. Good luck.</p>",
          "rawMarkdown": "@cherring Don't feel discouraged. Wait until after private Lb is revealed before making a conclusion. Many teams are overfitting public LB. (And I think the LB 0.965 notebook that you didn't see was very overfitted). Just keep working on your own work and I think you will finish well on private LB. Good luck.",
          "votes": 4
        }
      ]
    },
    {
      "id": 972313,
      "postDate": "2020-08-16T13:07:59.870Z",
      "content": "<p>Kaggle have made it a <strong><em>thing</em></strong> to make changes on their platform during the last week of competitions 🙈</p>",
      "rawMarkdown": "Kaggle have made it a ***thing*** to make changes on their platform during the last week of competitions 🙈",
      "votes": 3
    },
    {
      "id": 972206,
      "postDate": "2020-08-16T10:32:07.853Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F755831%2Fe4c8220bba61a095021982d1d804109a%2FScreenshot%202020-08-16%20at%2012.31.00.png?generation=1597573901401979&amp;alt=media\" alt=\"\">I still got the banner… just had to click on the \"your work\" tab</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F755831%2Fe4c8220bba61a095021982d1d804109a%2FScreenshot%202020-08-16%20at%2012.31.00.png?generation=1597573901401979&alt=media)I still got the banner... just had to click on the \"your work\" tab",
      "votes": 3
    },
    {
      "id": 972150,
      "postDate": "2020-08-16T09:14:53.180Z",
      "content": "<p>Finally, I get the reasons why my LB decreasing faster than the <strong>Bugatti</strong> car.  <br>\n<a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> , you help me to find the reasons. The reasons are :</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing\" target=\"_blank\">The High Scoring  Kernel</a></li>\n<li>The kaggle accounts <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635\" target=\"_blank\">you described</a></li>\n</ol>",
      "rawMarkdown": "Finally, I get the reasons why my LB decreasing faster than the **Bugatti** car.  \n@group16 , you help me to find the reasons. The reasons are :\n1.  [The High Scoring  Kernel](https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing)\n2. The kaggle accounts [you described](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635)",
      "votes": 3
    },
    {
      "id": 972153,
      "postDate": "2020-08-16T09:16:54.347Z",
      "content": "<p>The only reason for doing this in my view is Notebook medal. I think that there should be this mechanism by which the notebooks having a score more than a particular threshold (different for every competition accordingly) after a time limit (say Rules acceptance deadline) should not get qualified for a medal and the user gets punished if this happens. This will help prevent people from posting high scoring notebooks (either legitimate or traps). </p>\n<p>Now, one could workaround a bit here, by not creating a submission file, while doing everything and mentioning in the notebook about this, this way they can again get medals, so for this, I will second <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>'s suggestion, we do need notebook downvotes.</p>",
      "rawMarkdown": "The only reason for doing this in my view is Notebook medal. I think that there should be this mechanism by which the notebooks having a score more than a particular threshold (different for every competition accordingly) after a time limit (say Rules acceptance deadline) should not get qualified for a medal and the user gets punished if this happens. This will help prevent people from posting high scoring notebooks (either legitimate or traps). \n\nNow, one could workaround a bit here, by not creating a submission file, while doing everything and mentioning in the notebook about this, this way they can again get medals, so for this, I will second @philippsinger's suggestion, we do need notebook downvotes.",
      "votes": 4
    },
    {
      "id": 973358,
      "postDate": "2020-08-17T09:09:21.837Z",
      "content": "<p>**Update: **<a href=\"https://www.kaggle.com/deepakd14/efficient-ensembling-highest-public-lb-0-9646?scriptVersionId=40893586\" target=\"_blank\">https://www.kaggle.com/deepakd14/efficient-ensembling-highest-public-lb-0-9646?scriptVersionId=40893586</a><br>\nOne more guy who got inspired by that 0.9648 kernel owner, published  .9646 public score kernel on the every last day of the competition.<br>\nFrom cheaters to high score kernel, These malpractices are really ruining the competition environment.</p>",
      "rawMarkdown": "\n**Update: **https://www.kaggle.com/deepakd14/efficient-ensembling-highest-public-lb-0-9646?scriptVersionId=40893586\nOne more guy who got inspired by that 0.9648 kernel owner, published  .9646 public score kernel on the every last day of the competition.\nFrom cheaters to high score kernel, These malpractices are really ruining the competition environment.\n",
      "votes": 2,
      "replies": [
        {
          "id": 973372,
          "postDate": "2020-08-17T09:15:43.377Z",
          "content": "<p>Seems like some people have set a bad example here. Everybody thinks publishing a high LB notebook is a milestone and for betterment of the community.</p>\n<hr>\n<p>Update: Notebook taken down!</p>",
          "rawMarkdown": "Seems like some people have set a bad example here. Everybody thinks publishing a high LB notebook is a milestone and for betterment of the community.\n*****\n\nUpdate: Notebook taken down!",
          "votes": 1
        }
      ]
    },
    {
      "id": 973091,
      "postDate": "2020-08-17T05:19:41.300Z",
      "content": "<p>And the person deleted the kernel. This is more ridiculous. Now, some have an unfair advantage.</p>",
      "rawMarkdown": "And the person deleted the kernel. This is more ridiculous. Now, some have an unfair advantage.",
      "votes": 2,
      "replies": [
        {
          "id": 973292,
          "postDate": "2020-08-17T08:12:45.897Z",
          "content": "<p>And in case of a shake up .. an unfair disadvantage .. just kidding :D</p>",
          "rawMarkdown": "And in case of a shake up .. an unfair disadvantage .. just kidding :D",
          "votes": 5
        }
      ]
    },
    {
      "id": 974833,
      "postDate": "2020-08-18T04:18:44.570Z",
      "content": "<p>That kernel is heavily overfit, it only scored 0.8992 on private LB and surpassed by many single / kfold model.</p>",
      "rawMarkdown": "That kernel is heavily overfit, it only scored 0.8992 on private LB and surpassed by many single / kfold model."
    },
    {
      "id": 973027,
      "postDate": "2020-08-17T03:51:27.207Z",
      "content": "<p>I think Kaggle should automatically take down(or not <em>allow</em> to publish in the first place) the kernels which break into top 10% during the last week of competitions. It is heartbreaking to see your rank fall hundreds of places due to a blending kernel :(</p>",
      "rawMarkdown": "I think Kaggle should automatically take down(or not *allow* to publish in the first place) the kernels which break into top 10% during the last week of competitions. It is heartbreaking to see your rank fall hundreds of places due to a blending kernel :("
    },
    {
      "id": 972379,
      "postDate": "2020-08-16T13:58:48.607Z",
      "content": "<p>But, would it perform on the private lb? In previous competitions also there have been many public kernels with high score on public lb. I nevertheless focused on my own solutions which languished towards the bottom of the public lb and me wondering how unjust this world is, but at the end of the competition I see my solution moving hundreds of places higher on the private lb. This happened with me in Bengali.ai and M5 Forecasting Accuracy.</p>",
      "rawMarkdown": "But, would it perform on the private lb? In previous competitions also there have been many public kernels with high score on public lb. I nevertheless focused on my own solutions which languished towards the bottom of the public lb and me wondering how unjust this world is, but at the end of the competition I see my solution moving hundreds of places higher on the private lb. This happened with me in Bengali.ai and M5 Forecasting Accuracy.",
      "replies": [
        {
          "id": 972382,
          "postDate": "2020-08-16T14:03:41.027Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 972386,
          "postDate": "2020-08-16T14:06:57.430Z",
          "content": "<p>If it works on private LB, that will destroy the work of hundreds people.<br>\nElse, it is still misleading for Kaggle beginners and make the LB useless.</p>",
          "rawMarkdown": "If it works on private LB, that will destroy the work of hundreds people.\nElse, it is still misleading for Kaggle beginners and make the LB useless.",
          "votes": 1
        },
        {
          "id": 972396,
          "postDate": "2020-08-16T14:22:53.527Z",
          "content": "<p>Myself being a beginner, I feel the presence of such public kernels that score highly on public lb has taught me a very critical lesson in data science - don't overfit, no matter how tempting it might seem. It teaches a lesson in avoiding honey traps in data science and maybe in life in general :P .  A public LB is anyways useless in comparing solutions from different competitors. Its only an indicator that our submission pipeline doesn't have any major errors.   </p>",
          "rawMarkdown": "Myself being a beginner, I feel the presence of such public kernels that score highly on public lb has taught me a very critical lesson in data science - don't overfit, no matter how tempting it might seem. It teaches a lesson in avoiding honey traps in data science and maybe in life in general :P .  A public LB is anyways useless in comparing solutions from different competitors. Its only an indicator that our submission pipeline doesn't have any major errors.   ",
          "votes": 1
        },
        {
          "id": 972422,
          "postDate": "2020-08-16T14:43:11.913Z",
          "content": "<p>How can you know if it is overfitting as you cannot calculate the oof CV ?</p>",
          "rawMarkdown": "How can you know if it is overfitting as you cannot calculate the oof CV ?"
        },
        {
          "id": 972577,
          "postDate": "2020-08-16T16:46:04.057Z",
          "content": "<p>Yes, we don't know if its overfitting and neither do we know if its not overfitting. Overfitting is a very critical aspect in designing any solution. So, in general, when these high scoring public kernels  turn out to be damp squib on the private lb, it very strongly drives home this very important aspect of having a solid validation strategy. So overall, I feel the onus is on the competition hosts to have the train test split such that there are substantial differences between the train-test and between public-private parts of the test set, which would help ensure the insignificance of high score on public lb. Based on my recent experiences I can say that the most competition hosts are pretty good at this.</p>",
          "rawMarkdown": "Yes, we don't know if its overfitting and neither do we know if its not overfitting. Overfitting is a very critical aspect in designing any solution. So, in general, when these high scoring public kernels  turn out to be damp squib on the private lb, it very strongly drives home this very important aspect of having a solid validation strategy. So overall, I feel the onus is on the competition hosts to have the train test split such that there are substantial differences between the train-test and between public-private parts of the test set, which would help ensure the insignificance of high score on public lb. Based on my recent experiences I can say that the most competition hosts are pretty good at this."
        },
        {
          "id": 972685,
          "postDate": "2020-08-16T18:28:29.637Z",
          "content": "<p>Hey Pratik, i have seen people getting bronze or even silver medals on a badly blended set of models. So it is not a guarantee that the solution will overfit. and my team has been denied some medals in the past. because of last minute release of high scoring kernel.</p>",
          "rawMarkdown": "Hey Pratik, i have seen people getting bronze or even silver medals on a badly blended set of models. So it is not a guarantee that the solution will overfit. and my team has been denied some medals in the past. because of last minute release of high scoring kernel.",
          "votes": 3
        },
        {
          "id": 972746,
          "postDate": "2020-08-16T19:27:16.060Z",
          "content": "<p>I agree about learning is very important. <br>\nHowever, I think it is unfair to give the prizes &amp; top places deserved by honest competitors to cheators.</p>",
          "rawMarkdown": "I agree about learning is very important. \nHowever, I think it is unfair to give the prizes & top places deserved by honest competitors to cheators.",
          "votes": 1
        },
        {
          "id": 972999,
          "postDate": "2020-08-17T03:05:39.883Z",
          "content": "<p>Hi Ashish, thanks for sharing your experience. <br>\nIndeed it could be quite heart wrenching to see one's honest hard work not getting the medal/top_spot that it deserves due to high scoring public kernels. </p>",
          "rawMarkdown": "Hi Ashish, thanks for sharing your experience. \nIndeed it could be quite heart wrenching to see one's honest hard work not getting the medal/top_spot that it deserves due to high scoring public kernels. "
        }
      ]
    },
    {
      "id": 973916,
      "postDate": "2020-08-17T15:45:08.013Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 972343,
      "postDate": "2020-08-16T13:31:26.357Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 972333,
      "postDate": "2020-08-16T13:23:39.380Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 972269,
      "postDate": "2020-08-16T12:12:54.967Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true,
      "replies": [
        {
          "id": 972688,
          "postDate": "2020-08-16T18:29:35.950Z",
          "content": "<p>i have seen people getting bronze or even silver medals on a badly blended set of models. So it is not a guarantee that the solution will overfit. and my team has been denied some medals in the past. because of last minute release of high scoring kernel.</p>",
          "rawMarkdown": "i have seen people getting bronze or even silver medals on a badly blended set of models. So it is not a guarantee that the solution will overfit. and my team has been denied some medals in the past. because of last minute release of high scoring kernel.",
          "votes": 1
        },
        {
          "id": 972695,
          "postDate": "2020-08-16T18:35:51.263Z",
          "content": "<p>I have no answer to that, the discussion explain my point of view. I wish you luck in this competition. </p>",
          "rawMarkdown": "I have no answer to that, the discussion explain my point of view. I wish you luck in this competition. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 972099,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-16T08:30:23.463000",
      "content": "<p>No surprise he disabled comments. We finally need downvote for kernels.</p>",
      "votes": 19,
      "replies": [
        {
          "id": 972108,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-16T08:40:09.727000",
          "content": "<p>It is time that Kaggle organizes an analytics competition (where participants just make notebooks and analyses) around flagging of toxic things for which it can be difficult to obtain objective evidence. This includes: upvoting schemes for notebooks/topics, forum spamming, private sharing, and so on.</p>\n<p>As in some games, we could use a concept of trusted \"gamemasters\" that can follow up suspicious activity by e.g. asking questions or asking participants to (somewhat) reproduce approaches. These could then punish people (and their activities can be publicly known).</p>\n<p>Kaggle, can we at least have an elaborate discussion on this? Because it is clear that the current progression system, as is, encourages bad behaviour and stimulates gaming the system.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 972739,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-08-16T19:11:49.700000",
      "content": "<p>I feel sad for the people who upvote such notebooks.</p>\n<p>As <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> mentioned, I've never understood why Kaggle has downvotes for discussions and not for notebooks and datasets. Sure, there are some downvoting abusers and debateable-cases but most of the times the net votes on discussion posts give a fair view of their value.</p>\n<p>Kaggle is unfairly partial to notebooks (and datasets).</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 972359,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-08-16T13:42:02.823000",
      "content": "<p>Permanent ban of the account is necessary. </p>\n<p>He doesn't even hide his will to screw up the LB. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 972804,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-16T20:44:54.320000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> agreed consequences, or at least a sanction (banned from comp for 1 year or something).  Also, anyone that uses that kernel should be disqualified.</p>",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 973065,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2020-08-17T04:58:43.910000",
      "content": "<p>My reaction in the last week of every competition where this happens<br>\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/973065/16606/IMG_1333.JPG\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 973832,
          "author_name": "Dimas Munoz",
          "author_url": "",
          "post_date": "2020-08-17T14:43:54.403000",
          "content": "<p>Best reaction so far :')</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 974578,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-18T01:41:50.683000",
          "content": "<p>Catching up on the forums today and this one made me lol. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 972664,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-08-16T18:10:58.387000",
      "content": "<p>I have a 0.9593 submission that scores 0.95 on CV. Pearson correlation between my predictions and that submission is 0.48 so hopefully it's overfitted to public LB. 😬😬</p>",
      "votes": 6,
      "replies": [
        {
          "id": 972771,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-16T20:12:40.590000",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congratulations on reaching very good CV/LB scores! Did you rank your predictions and the submission in question before computing the correlation coefficient? Just curious… (I would compute it myself but the kernel has been removed already).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972787,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-08-16T20:26:41.533000",
          "content": "<p>I did rank both of them before computing the correlation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972825,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-16T21:26:10.043000",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Great! Thank you for clarifying this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 973040,
          "author_name": "Manjesh Gupta",
          "author_url": "",
          "post_date": "2020-08-17T04:08:10.733000",
          "content": "<p>For AUC metric, does rank co-relation matter if the target classes are well separated?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 973134,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-17T06:04:11.897000",
          "content": "<p><a href=\"https://www.kaggle.com/manjeshg03\" target=\"_blank\">@manjeshg03</a> AUC is not affected by the ranking operation. But when you are computing a correlation coefficient between two different sets of predictions generated in two very different ways, it is generally a good idea to rank them first. This way you are calibrating them to the same scale, so to speak.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 973144,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2020-08-17T06:15:52.630000",
          "content": "<p>Spearman correlation is better suited for AUC</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 972630,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2020-08-16T17:34:14.973000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2861255%2Fd1fe2c5cbed84cf8320c03f170a258f1%2Fclass.jpg?generation=1597599248620054&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 972917,
      "author_name": "cherring",
      "author_url": "",
      "post_date": "2020-08-17T01:17:16.463000",
      "content": "<p>I guess I will be quitting kaggle again :( I really love the competitive nature. Gets me motivated. But when there are notebooks like this and also the 0.9619 that pads out like 400 places starting from low bronze.. I just feel cheated. Especially when my best personal submissions is ever so slightly worse than it.</p>\n<p>Fortunately, that new kernel is down already. Unfortunately I didn't get to run it :(</p>",
      "votes": 3,
      "replies": [
        {
          "id": 972931,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-17T01:30:41.683000",
          "content": "<p>well better for you not to have run it………….thats my opinion.  I stick to my own work.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 972938,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-17T01:35:19.083000",
          "content": "<p><a href=\"https://www.kaggle.com/cherring\" target=\"_blank\">@cherring</a> Don't feel discouraged. Wait until after private Lb is revealed before making a conclusion. Many teams are overfitting public LB. (And I think the LB 0.965 notebook that you didn't see was very overfitted). Just keep working on your own work and I think you will finish well on private LB. Good luck.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 972313,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-08-16T13:07:59.870000",
      "content": "<p>Kaggle have made it a <strong><em>thing</em></strong> to make changes on their platform during the last week of competitions 🙈</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 972206,
      "author_name": "Bram Steenwinckel",
      "author_url": "",
      "post_date": "2020-08-16T10:32:07.853000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F755831%2Fe4c8220bba61a095021982d1d804109a%2FScreenshot%202020-08-16%20at%2012.31.00.png?generation=1597573901401979&amp;alt=media\" alt=\"\">I still got the banner… just had to click on the \"your work\" tab</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 972150,
      "author_name": "Md Fahim",
      "author_url": "",
      "post_date": "2020-08-16T09:14:53.180000",
      "content": "<p>Finally, I get the reasons why my LB decreasing faster than the <strong>Bugatti</strong> car.  <br>\n<a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> , you help me to find the reasons. The reasons are :</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing\" target=\"_blank\">The High Scoring  Kernel</a></li>\n<li>The kaggle accounts <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635\" target=\"_blank\">you described</a></li>\n</ol>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 972153,
      "author_name": "Gajendra Saraswat",
      "author_url": "",
      "post_date": "2020-08-16T09:16:54.347000",
      "content": "<p>The only reason for doing this in my view is Notebook medal. I think that there should be this mechanism by which the notebooks having a score more than a particular threshold (different for every competition accordingly) after a time limit (say Rules acceptance deadline) should not get qualified for a medal and the user gets punished if this happens. This will help prevent people from posting high scoring notebooks (either legitimate or traps). </p>\n<p>Now, one could workaround a bit here, by not creating a submission file, while doing everything and mentioning in the notebook about this, this way they can again get medals, so for this, I will second <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>'s suggestion, we do need notebook downvotes.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 973358,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2020-08-17T09:09:21.837000",
      "content": "<p>**Update: **<a href=\"https://www.kaggle.com/deepakd14/efficient-ensembling-highest-public-lb-0-9646?scriptVersionId=40893586\" target=\"_blank\">https://www.kaggle.com/deepakd14/efficient-ensembling-highest-public-lb-0-9646?scriptVersionId=40893586</a><br>\nOne more guy who got inspired by that 0.9648 kernel owner, published  .9646 public score kernel on the every last day of the competition.<br>\nFrom cheaters to high score kernel, These malpractices are really ruining the competition environment.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 973372,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-08-17T09:15:43.377000",
          "content": "<p>Seems like some people have set a bad example here. Everybody thinks publishing a high LB notebook is a milestone and for betterment of the community.</p>\n<hr>\n<p>Update: Notebook taken down!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 973091,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2020-08-17T05:19:41.300000",
      "content": "<p>And the person deleted the kernel. This is more ridiculous. Now, some have an unfair advantage.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 973292,
          "author_name": "Pratik",
          "author_url": "",
          "post_date": "2020-08-17T08:12:45.897000",
          "content": "<p>And in case of a shake up .. an unfair disadvantage .. just kidding :D</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 974833,
      "author_name": "Sandy Khosasi",
      "author_url": "",
      "post_date": "2020-08-18T04:18:44.570000",
      "content": "<p>That kernel is heavily overfit, it only scored 0.8992 on private LB and surpassed by many single / kfold model.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 973027,
      "author_name": "spoon spoon",
      "author_url": "",
      "post_date": "2020-08-17T03:51:27.207000",
      "content": "<p>I think Kaggle should automatically take down(or not <em>allow</em> to publish in the first place) the kernels which break into top 10% during the last week of competitions. It is heartbreaking to see your rank fall hundreds of places due to a blending kernel :(</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972379,
      "author_name": "Pratik",
      "author_url": "",
      "post_date": "2020-08-16T13:58:48.607000",
      "content": "<p>But, would it perform on the private lb? In previous competitions also there have been many public kernels with high score on public lb. I nevertheless focused on my own solutions which languished towards the bottom of the public lb and me wondering how unjust this world is, but at the end of the competition I see my solution moving hundreds of places higher on the private lb. This happened with me in Bengali.ai and M5 Forecasting Accuracy.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 972382,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-16T14:03:41.027000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 972386,
          "author_name": "Changyi",
          "author_url": "",
          "post_date": "2020-08-16T14:06:57.430000",
          "content": "<p>If it works on private LB, that will destroy the work of hundreds people.<br>\nElse, it is still misleading for Kaggle beginners and make the LB useless.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972396,
          "author_name": "Pratik",
          "author_url": "",
          "post_date": "2020-08-16T14:22:53.527000",
          "content": "<p>Myself being a beginner, I feel the presence of such public kernels that score highly on public lb has taught me a very critical lesson in data science - don't overfit, no matter how tempting it might seem. It teaches a lesson in avoiding honey traps in data science and maybe in life in general :P .  A public LB is anyways useless in comparing solutions from different competitors. Its only an indicator that our submission pipeline doesn't have any major errors.   </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972422,
          "author_name": "Changyi",
          "author_url": "",
          "post_date": "2020-08-16T14:43:11.913000",
          "content": "<p>How can you know if it is overfitting as you cannot calculate the oof CV ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972577,
          "author_name": "Pratik",
          "author_url": "",
          "post_date": "2020-08-16T16:46:04.057000",
          "content": "<p>Yes, we don't know if its overfitting and neither do we know if its not overfitting. Overfitting is a very critical aspect in designing any solution. So, in general, when these high scoring public kernels  turn out to be damp squib on the private lb, it very strongly drives home this very important aspect of having a solid validation strategy. So overall, I feel the onus is on the competition hosts to have the train test split such that there are substantial differences between the train-test and between public-private parts of the test set, which would help ensure the insignificance of high score on public lb. Based on my recent experiences I can say that the most competition hosts are pretty good at this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972685,
          "author_name": "Ashish Gupta",
          "author_url": "",
          "post_date": "2020-08-16T18:28:29.637000",
          "content": "<p>Hey Pratik, i have seen people getting bronze or even silver medals on a badly blended set of models. So it is not a guarantee that the solution will overfit. and my team has been denied some medals in the past. because of last minute release of high scoring kernel.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 972746,
          "author_name": "Changyi",
          "author_url": "",
          "post_date": "2020-08-16T19:27:16.060000",
          "content": "<p>I agree about learning is very important. <br>\nHowever, I think it is unfair to give the prizes &amp; top places deserved by honest competitors to cheators.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972999,
          "author_name": "Pratik",
          "author_url": "",
          "post_date": "2020-08-17T03:05:39.883000",
          "content": "<p>Hi Ashish, thanks for sharing your experience. <br>\nIndeed it could be quite heart wrenching to see one's honest hard work not getting the medal/top_spot that it deserves due to high scoring public kernels. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 973916,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-17T15:45:08.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972343,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-16T13:31:26.357000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 972333,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-16T13:23:39.380000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 972269,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-16T12:12:54.967000",
      "content": "",
      "votes": -2,
      "replies": [
        {
          "id": 972688,
          "author_name": "Ashish Gupta",
          "author_url": "",
          "post_date": "2020-08-16T18:29:35.950000",
          "content": "<p>i have seen people getting bronze or even silver medals on a badly blended set of models. So it is not a guarantee that the solution will overfit. and my team has been denied some medals in the past. because of last minute release of high scoring kernel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972695,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T18:35:51.263000",
          "content": "<p>I have no answer to that, the discussion explain my point of view. I wish you luck in this competition. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "972090": "Seems like with the new front-end changes, the banner that warns people to **NOT** share high-scoring notebooks in the final week of the competition has vanished.\n\nAs a result, [someone thought it was a fantastic idea to post a blending kernel](https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing) that scores 0.9648 and scores you an easy bronze medal. Two days before the end of the competition, this punishes everyone that worked hard and cannot log in by the end of the competition.",
    "972099": "No surprise he disabled comments. We finally need downvote for kernels.",
    "972739": "I feel sad for the people who upvote such notebooks.\n\nAs @philippsinger mentioned, I've never understood why Kaggle has downvotes for discussions and not for notebooks and datasets. Sure, there are some downvoting abusers and debateable-cases but most of the times the net votes on discussion posts give a fair view of their value.\n\nKaggle is unfairly partial to notebooks (and datasets).",
    "972359": "Permanent ban of the account is necessary. \n\nHe doesn't even hide his will to screw up the LB. ",
    "973065": "My reaction in the last week of every competition where this happens\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/973065/16606/IMG_1333.JPG)",
    "972664": "I have a 0.9593 submission that scores 0.95 on CV. Pearson correlation between my predictions and that submission is 0.48 so hopefully it's overfitted to public LB. 😬😬",
    "972630": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2861255%2Fd1fe2c5cbed84cf8320c03f170a258f1%2Fclass.jpg?generation=1597599248620054&alt=media)",
    "972917": "I guess I will be quitting kaggle again :( I really love the competitive nature. Gets me motivated. But when there are notebooks like this and also the 0.9619 that pads out like 400 places starting from low bronze.. I just feel cheated. Especially when my best personal submissions is ever so slightly worse than it.\n\nFortunately, that new kernel is down already. Unfortunately I didn't get to run it :(",
    "972313": "Kaggle have made it a ***thing*** to make changes on their platform during the last week of competitions 🙈",
    "972206": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F755831%2Fe4c8220bba61a095021982d1d804109a%2FScreenshot%202020-08-16%20at%2012.31.00.png?generation=1597573901401979&alt=media)I still got the banner... just had to click on the \"your work\" tab",
    "972150": "Finally, I get the reasons why my LB decreasing faster than the **Bugatti** car.  \n@group16 , you help me to find the reasons. The reasons are :\n1.  [The High Scoring  Kernel](https://www.kaggle.com/mekhdigakhramanian/top-7-lb-0-9648-post-processing)\n2. The kaggle accounts [you described](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174635)",
    "972153": "The only reason for doing this in my view is Notebook medal. I think that there should be this mechanism by which the notebooks having a score more than a particular threshold (different for every competition accordingly) after a time limit (say Rules acceptance deadline) should not get qualified for a medal and the user gets punished if this happens. This will help prevent people from posting high scoring notebooks (either legitimate or traps). \n\nNow, one could workaround a bit here, by not creating a submission file, while doing everything and mentioning in the notebook about this, this way they can again get medals, so for this, I will second @philippsinger's suggestion, we do need notebook downvotes.",
    "973358": "\n**Update: **https://www.kaggle.com/deepakd14/efficient-ensembling-highest-public-lb-0-9646?scriptVersionId=40893586\nOne more guy who got inspired by that 0.9648 kernel owner, published  .9646 public score kernel on the every last day of the competition.\nFrom cheaters to high score kernel, These malpractices are really ruining the competition environment.\n",
    "973091": "And the person deleted the kernel. This is more ridiculous. Now, some have an unfair advantage.",
    "974833": "That kernel is heavily overfit, it only scored 0.8992 on private LB and surpassed by many single / kfold model.",
    "973027": "I think Kaggle should automatically take down(or not *allow* to publish in the first place) the kernels which break into top 10% during the last week of competitions. It is heartbreaking to see your rank fall hundreds of places due to a blending kernel :(",
    "972379": "But, would it perform on the private lb? In previous competitions also there have been many public kernels with high score on public lb. I nevertheless focused on my own solutions which languished towards the bottom of the public lb and me wondering how unjust this world is, but at the end of the competition I see my solution moving hundreds of places higher on the private lb. This happened with me in Bengali.ai and M5 Forecasting Accuracy.",
    "973916": "",
    "972343": "",
    "972333": "",
    "972269": ""
  }
}