{
  "id": 174956,
  "title": "How High-Score Kernels Doesn't Affect Us",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174956",
  "author_name": "Hiram Coria 🧬",
  "post_date": "2020-08-16T13:19:50.738000",
  "votes": 1,
  "comment_count": 29,
  "views": 0,
  "content": "<p><strong>Before you down-vote me please, read carefully</strong></p>\n<p>Let me first offer to you evidence that the <strong>Private Leaderboard</strong> might be very different from the <strong>Public</strong>:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard\" target=\"_blank\">Prostate cANcer graDe Assessment (PANDA) Challenge</a></li>\n<li><a href=\"https://www.kaggle.com/c/m5-forecasting-accuracy/leaderboard\" target=\"_blank\">M5 Forecasting - Accuracy</a></li>\n<li><a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/leaderboard\" target=\"_blank\">Jigsaw Unintended Bias in Toxicity Classification</a></li>\n</ul>\n<p>======================================================<br>\nUPDATE: <strong>Now this competition</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/leaderboard\" target=\"_blank\">SIIM-ISIC Melanoma Classification</a></li>\n</ul>\n<p>======================================================</p>\n<p>After that but before my argument i must tell something if you are <strong>GrandMaster</strong>, <strong>Kaggle Veteran</strong> or whoever that achieve a gold medal you are strictly immune to the <strong>high-score kernels</strong>. </p>\n<p>Kaggle it's a platform/community for people whom are or wants to be a <strong>Data Scientist</strong>. Therefore, there are a lot of content about this topic. Comments, notebooks and datasets are plenty of healthy content for novices, sometimes notebooks with good solutions for a competition. Talking about competitions, there is a kind of way to know if your solution it's better than a previous one or another's solution, <strong>score</strong>, but the score it's not 100% reliable. Sometimes the solution overfits the public test set. Recently people ask to Kaggle to highly regulate <strong>high-score kernels</strong>, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of <strong>chilling effects</strong>.  Long-term this might poisoned this community. </p>\n<p>As novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a <strong>Data Scienctist</strong>.  At the end i'm not saying that people who share <strong>high-score kernels</strong> are fine but if we forbid that, we are gonna decrease the quality of the content shared (by a kind of<strong>chilling effects</strong>).</p>\n<p>======================================================<br>\nUPDATE: <strong>CONCLUSION</strong></p>\n<p>Shared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.</p>\n<p>======================================================</p>\n<p>======================================================<br>\nUPDATE: <strong>Paper</strong></p>\n<p>The paper <a href=\"https://arxiv.org/pdf/1612.01345.pdf\" target=\"_blank\">Human-In-The-Loop Person Re-Identification</a> it's the reason of why i belive that insane ensembles are not a good solution and therefore harmless in this competition.</p>\n<p>======================================================</p>\n<p><strong>My apollogies for my level of English</strong>, please think about what you have read.</p>",
  "messages": [
    {
      "id": 972554,
      "postDate": "2020-08-16T16:23:22.837Z",
      "content": "<p>I wonder how you can learn anything from blending kernels that dont even cite the sources they have the models from?</p>",
      "rawMarkdown": "I wonder how you can learn anything from blending kernels that dont even cite the sources they have the models from?",
      "votes": 12,
      "replies": [
        {
          "id": 972563,
          "postDate": "2020-08-16T16:33:31.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> As i said in my post i'm defending the good solution that also has a high-score. And again, as i said there're notebooks that overfit the public test set. With all the respect that you deserve, you are a GrandMaster, you are not affected. Novices are the more affected to just fork the kind of notebook that you mentioned because we don't learn anything.</p>",
          "rawMarkdown": "@philippsinger As i said in my post i'm defending the good solution that also has a high-score. And again, as i said there're notebooks that overfit the public test set. With all the respect that you deserve, you are a GrandMaster, you are not affected. Novices are the more affected to just fork the kind of notebook that you mentioned because we don't learn anything."
        },
        {
          "id": 972579,
          "postDate": "2020-08-16T16:49:15.320Z",
          "content": "<p>So shouldn't it then be also us complaining about them?</p>",
          "rawMarkdown": "So shouldn't it then be also us complaining about them?",
          "votes": 3
        },
        {
          "id": 972604,
          "postDate": "2020-08-16T17:13:47.080Z",
          "content": "<p>DISCLAIMER: I don't down-vote you <strong>i respect your point Mr</strong>.</p>\n<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Yes you can, but in my post i explain:</p>\n<blockquote>\n  <p>As novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a Data Scienctist.</p>\n</blockquote>\n<p><strong>I NEVER SAID</strong>:</p>\n<blockquote>\n  <p>[…] I learned a lot seeing <strong>blending kernels that dont even cite the sources they have the models from</strong>.</p>\n</blockquote>",
          "rawMarkdown": "DISCLAIMER: I don't down-vote you **i respect your point Mr**.\n\n@philippsinger Yes you can, but in my post i explain:\n\n> As novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a Data Scienctist.\n\n**I NEVER SAID**:\n\n> [...] I learned a lot seeing **blending kernels that dont even cite the sources they have the models from**.",
          "votes": -4
        }
      ]
    },
    {
      "id": 972393,
      "postDate": "2020-08-16T14:15:07.563Z",
      "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> it's not about sharing high-score kernels. It's about sharing high-score kernels during the last week of a competition, ignoring the standard Kaggle request not to do it. In this competition it happened 2 days before the competition end. It's disrespect to all Kaggle community for the sake of one notebook medal.</p>",
      "rawMarkdown": "@hiramcho it's not about sharing high-score kernels. It's about sharing high-score kernels during the last week of a competition, ignoring the standard Kaggle request not to do it. In this competition it happened 2 days before the competition end. It's disrespect to all Kaggle community for the sake of one notebook medal.",
      "votes": 8,
      "replies": [
        {
          "id": 972394,
          "postDate": "2020-08-16T14:18:04.950Z",
          "content": "<p>Both of them, i saw people more than one month ago complaining about the high-scor kernels.</p>",
          "rawMarkdown": "Both of them, i saw people more than one month ago complaining about the high-scor kernels.",
          "votes": -1
        }
      ]
    },
    {
      "id": 972401,
      "postDate": "2020-08-16T14:27:25.030Z",
      "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> ,<br>\n<a href=\"https://www.kaggle.com/raininbox/blend-different-models-with-different-n-tile\" target=\"_blank\">This Public Kernel</a> in <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard\" target=\"_blank\">PANDA</a> competition won the silver medal. Some people won the silver medal just forking the kernel. <br>\nDoes not it affect the others ? </p>",
      "rawMarkdown": "@hiramcho ,\n[This Public Kernel](https://www.kaggle.com/raininbox/blend-different-models-with-different-n-tile) in [PANDA](https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard) competition won the silver medal. Some people won the silver medal just forking the kernel. \nDoes not it affect the others ? ",
      "votes": 4,
      "replies": [
        {
          "id": 972406,
          "postDate": "2020-08-16T14:30:56.123Z",
          "content": "<p><a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> he shared one mor before the timeline end. We have the time to learn and improve our models with his work, personally i don't use it because i used ResNet and he EfficientNet but again people could learn and improve their works with his.</p>",
          "rawMarkdown": "@mdfahimreshm he shared one mor before the timeline end. We have the time to learn and improve our models with his work, personally i don't use it because i used ResNet and he EfficientNet but again people could learn and improve their works with his.",
          "votes": -1
        },
        {
          "id": 972417,
          "postDate": "2020-08-16T14:37:53.130Z",
          "content": "<p>UPDATE:<br>\n<a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> he doesn't achive silver, he achive bronze (cause he has the 75 position).</p>",
          "rawMarkdown": "UPDATE:\n@mdfahimreshm he doesn't achive silver, he achive bronze (cause he has the 75 position)."
        },
        {
          "id": 972431,
          "postDate": "2020-08-16T14:51:37.267Z",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> , <br>\nYes, he doesn't achieve the silver medal but his public kernel does. You may see <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169115\" target=\"_blank\">my post</a> for more clearification.  </p>",
          "rawMarkdown": "@hiramcho , \nYes, he doesn't achieve the silver medal but his public kernel does. You may see [my post](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169115) for more clearification.  ",
          "votes": 1
        },
        {
          "id": 972442,
          "postDate": "2020-08-16T14:57:39.080Z",
          "content": "<p>I saw it <a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a>. You said:</p>\n<blockquote>\n  <p>Amazing this public kernel got the silver medal</p>\n</blockquote>\n<p>And nothing more, so i have two questions, would you kindly explain why his wonderful kernel harms anyone if he shared one month before the close of the PANDA competition? and why matters that he achive a silver medal (by votes)?</p>",
          "rawMarkdown": "I saw it @mdfahimreshm. You said:\n\n> Amazing this public kernel got the silver medal\n\n And nothing more, so i have two questions, would you kindly explain why his wonderful kernel harms anyone if he shared one month before the close of the PANDA competition? and why matters that he achive a silver medal (by votes)?",
          "votes": 1
        },
        {
          "id": 972472,
          "postDate": "2020-08-16T15:13:47.147Z",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> , <br>\nYou understand me wrong. I does not claim the beautiful kernel and the other who works on this. But you know, some of people just forked it and did not work so much. Fortunately, those people got the medal.    I am telling about them. </p>",
          "rawMarkdown": "@hiramcho , \nYou understand me wrong. I does not claim the beautiful kernel and the other who works on this. But you know, some of people just forked it and did not work so much. Fortunately, those people got the medal.    I am telling about them. ",
          "votes": 2
        },
        {
          "id": 972503,
          "postDate": "2020-08-16T15:40:59.800Z",
          "content": "<p><a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> thanks for clarify your idea, now i understand. About the people who achives bronze medal there're too few of them (just looking the private leaderboard i say that), and we must allow ourselves to lose to learn, now we know that an ensemble it's better generalizing than one single model because here comes my point, during PANDA competition his work hasn't the highest score (that's the reason of why his notebook only achives silver), many use the highest score kernel and not the better. During that competition whe didn't know that we could achive a bronze medal with that notebook. In resume, you are sayin that people achive a bronze medal just forking his work but i'm saying there were few who made them, because he hasn't the highest score, the people who just fork that notebook \"took a leap of faith\".</p>",
          "rawMarkdown": "@mdfahimreshm thanks for clarify your idea, now i understand. About the people who achives bronze medal there're too few of them (just looking the private leaderboard i say that), and we must allow ourselves to lose to learn, now we know that an ensemble it's better generalizing than one single model because here comes my point, during PANDA competition his work hasn't the highest score (that's the reason of why his notebook only achives silver), many use the highest score kernel and not the better. During that competition whe didn't know that we could achive a bronze medal with that notebook. In resume, you are sayin that people achive a bronze medal just forking his work but i'm saying there were few who made them, because he hasn't the highest score, the people who just fork that notebook \"took a leap of faith\".",
          "votes": -1
        },
        {
          "id": 973022,
          "postDate": "2020-08-17T03:44:45.287Z",
          "content": "<p>In such competitions like PANDA with very unstable LB, if 100 people publish their kernels, several of them could achieve quite high score at private LB just by coincidence. I'll just give an example. In <a href=\"https://www.kaggle.com/c/bengaliai-cv19/overview\" target=\"_blank\">Bengali competition</a> in a first few days after the start I published a series of kernels with a full pipeline: data cleaning and preprocessing, training, and submission to give new participant a good starting point in that challenge. The public score of <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference\" target=\"_blank\">my kernel</a> is quite low, just 0.964, which corresponds to <strong>~1200 at public LB</strong> at the end of the competition, and there were a number of kernels with 0.97+ score posted during the competition. However, at the end of the competition this kernel appeared to work quite well at private LB, giving <strong>~100th rank</strong>, and <strong>~30 people got bronze with nearly a single submission</strong>. Also a slight modification of the kernel could easily bring to silver, as several people contacted me to say \"thank you\". But it is clearly just a coincidence, and if the private LB was more predictable (or if I payed more attention to that organizers want), our team would get gold, not dropped to 200+ place because of a wrong selection( It happens, and cannot be fully avoided. </p>\n<p>There are just several things that people may follow to minimize appearance of strong public kernels at private LB. (1) Publish quite basic pipeline. For example, the amazing TF TPU pipeline published in this competition with huge models, external data, proper CV split, smooth labels, etc. is far away from being the basic one. (2) Do not upload models trained externally, share the code that could run on kaggle.  Otherwise people who don't have their own hardware and working only at kaggle would get screwed. And also such uploading facilitates creation of crazy blending kernels. (3) upvote only the content that you want to see at kaggle. If you think that blending kernels are bad, just don't upvote them. (4) I think the best time to publish a full pipeline, if it is strong, is the first month of a competition, to give people time to work on it, explore new ideas, and learn something. So you can promote some progress in the competition, not just make people angry. Some cool ideas certainly could be published later, but it is really bad if one is publishing something game changing in several weeks before the end of the competition, or like we saw today appearance of a high scoring kernel just in one day before the competition end. It may really ruin the competition. All above it it just a recommendation for being a good member of our diverse kaggle community. <br>\nThe thing I just wanted to say is that if people really try to help others and share their ideas and code, help to the progress in the competition, it should not be discouraged. Did u think how this competition would look like without the contribution made by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and other people shared their ideas?</p>",
          "rawMarkdown": "In such competitions like PANDA with very unstable LB, if 100 people publish their kernels, several of them could achieve quite high score at private LB just by coincidence. I'll just give an example. In [Bengali competition](https://www.kaggle.com/c/bengaliai-cv19/overview) in a first few days after the start I published a series of kernels with a full pipeline: data cleaning and preprocessing, training, and submission to give new participant a good starting point in that challenge. The public score of [my kernel](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference) is quite low, just 0.964, which corresponds to **~1200 at public LB** at the end of the competition, and there were a number of kernels with 0.97+ score posted during the competition. However, at the end of the competition this kernel appeared to work quite well at private LB, giving **~100th rank**, and **~30 people got bronze with nearly a single submission**. Also a slight modification of the kernel could easily bring to silver, as several people contacted me to say \"thank you\". But it is clearly just a coincidence, and if the private LB was more predictable (or if I payed more attention to that organizers want), our team would get gold, not dropped to 200+ place because of a wrong selection( It happens, and cannot be fully avoided. \n\nThere are just several things that people may follow to minimize appearance of strong public kernels at private LB. (1) Publish quite basic pipeline. For example, the amazing TF TPU pipeline published in this competition with huge models, external data, proper CV split, smooth labels, etc. is far away from being the basic one. (2) Do not upload models trained externally, share the code that could run on kaggle.  Otherwise people who don't have their own hardware and working only at kaggle would get screwed. And also such uploading facilitates creation of crazy blending kernels. (3) upvote only the content that you want to see at kaggle. If you think that blending kernels are bad, just don't upvote them. (4) I think the best time to publish a full pipeline, if it is strong, is the first month of a competition, to give people time to work on it, explore new ideas, and learn something. So you can promote some progress in the competition, not just make people angry. Some cool ideas certainly could be published later, but it is really bad if one is publishing something game changing in several weeks before the end of the competition, or like we saw today appearance of a high scoring kernel just in one day before the competition end. It may really ruin the competition. All above it it just a recommendation for being a good member of our diverse kaggle community. \nThe thing I just wanted to say is that if people really try to help others and share their ideas and code, help to the progress in the competition, it should not be discouraged. Did u think how this competition would look like without the contribution made by @cdeotte and other people shared their ideas?",
          "votes": 6
        },
        {
          "id": 973641,
          "postDate": "2020-08-17T12:22:42.380Z",
          "content": "<p>Well Said <a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a> , Without help of others kaggle wouldnt be this amazing and fun :D</p>",
          "rawMarkdown": "Well Said @lafoss , Without help of others kaggle wouldnt be this amazing and fun :D"
        }
      ]
    },
    {
      "id": 972697,
      "postDate": "2020-08-16T18:36:55.373Z",
      "content": "<p>You are certainly not right with the Jigsaw, I saw people getting bronze and silvers with a bad blend solution.</p>",
      "rawMarkdown": "You are certainly not right with the Jigsaw, I saw people getting bronze and silvers with a bad blend solution.",
      "votes": 3,
      "replies": [
        {
          "id": 972705,
          "postDate": "2020-08-16T18:41:47.030Z",
          "content": "<p>Would you share one work like you described (i'm not being challenging i'm just curious)?</p>",
          "rawMarkdown": "Would you share one work like you described (i'm not being challenging i'm just curious)?"
        },
        {
          "id": 972750,
          "postDate": "2020-08-16T19:36:27.583Z",
          "content": "<p>Take Tarek Hamdi's work for example. His imbalance weight ensembling fetched him a bronze medal on Private lb. </p>",
          "rawMarkdown": "Take Tarek Hamdi's work for example. His imbalance weight ensembling fetched him a bronze medal on Private lb. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 972512,
      "postDate": "2020-08-16T15:51:52.890Z",
      "content": "<p>I totally agree your learning part of the argument, but I don't see the point in making a last minute dash to a good leaderboard position.I mean yes we can make adjustments and improve our methods but then again, Sharing kernel with medal level accuracies should be prohibited at least a week before the deadline.</p>",
      "rawMarkdown": "I totally agree your learning part of the argument, but I don't see the point in making a last minute dash to a good leaderboard position.I mean yes we can make adjustments and improve our methods but then again, Sharing kernel with medal level accuracies should be prohibited at least a week before the deadline.",
      "votes": 4,
      "replies": [
        {
          "id": 972518,
          "postDate": "2020-08-16T15:58:58.280Z",
          "content": "<p>I understand your point. But as i said this could trigger <strong>chilling effects</strong> in this community. The risk is high enough for me to think that it's a bad idea.</p>",
          "rawMarkdown": "I understand your point. But as i said this could trigger **chilling effects** in this community. The risk is high enough for me to think that it's a bad idea.",
          "votes": -3
        },
        {
          "id": 972710,
          "postDate": "2020-08-16T18:47:39.337Z",
          "content": "<p>After the competition is over maybe the high scoring kernel can be made public, then people can analyse and improve models to better their scores.So learning part would be covered that way, And having said that anyone who wants to learn will learn, No matter what!</p>",
          "rawMarkdown": "After the competition is over maybe the high scoring kernel can be made public, then people can analyse and improve models to better their scores.So learning part would be covered that way, And having said that anyone who wants to learn will learn, No matter what!",
          "votes": 1
        },
        {
          "id": 972723,
          "postDate": "2020-08-16T18:58:34.557Z",
          "content": "<p>I'm going to stop answering because no matters how much evidence and arguments i made i just get downvotes \"because reasons\". Just think about what i write in the post, just that, don't change your opinion just think what could happend if Kaggle don't to publish high-score kernels (at any time).</p>",
          "rawMarkdown": "I'm going to stop answering because no matters how much evidence and arguments i made i just get downvotes \"because reasons\". Just think about what i write in the post, just that, don't change your opinion just think what could happend if Kaggle don't to publish high-score kernels (at any time)."
        }
      ]
    },
    {
      "id": 972326,
      "postDate": "2020-08-16T13:19:50.740Z",
      "content": "<p><strong>Before you down-vote me please, read carefully</strong></p>\n<p>Let me first offer to you evidence that the <strong>Private Leaderboard</strong> might be very different from the <strong>Public</strong>:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard\" target=\"_blank\">Prostate cANcer graDe Assessment (PANDA) Challenge</a></li>\n<li><a href=\"https://www.kaggle.com/c/m5-forecasting-accuracy/leaderboard\" target=\"_blank\">M5 Forecasting - Accuracy</a></li>\n<li><a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/leaderboard\" target=\"_blank\">Jigsaw Unintended Bias in Toxicity Classification</a></li>\n</ul>\n<p>======================================================<br>\nUPDATE: <strong>Now this competition</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/leaderboard\" target=\"_blank\">SIIM-ISIC Melanoma Classification</a></li>\n</ul>\n<p>======================================================</p>\n<p>After that but before my argument i must tell something if you are <strong>GrandMaster</strong>, <strong>Kaggle Veteran</strong> or whoever that achieve a gold medal you are strictly immune to the <strong>high-score kernels</strong>. </p>\n<p>Kaggle it's a platform/community for people whom are or wants to be a <strong>Data Scientist</strong>. Therefore, there are a lot of content about this topic. Comments, notebooks and datasets are plenty of healthy content for novices, sometimes notebooks with good solutions for a competition. Talking about competitions, there is a kind of way to know if your solution it's better than a previous one or another's solution, <strong>score</strong>, but the score it's not 100% reliable. Sometimes the solution overfits the public test set. Recently people ask to Kaggle to highly regulate <strong>high-score kernels</strong>, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of <strong>chilling effects</strong>.  Long-term this might poisoned this community. </p>\n<p>As novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a <strong>Data Scienctist</strong>.  At the end i'm not saying that people who share <strong>high-score kernels</strong> are fine but if we forbid that, we are gonna decrease the quality of the content shared (by a kind of<strong>chilling effects</strong>).</p>\n<p>======================================================<br>\nUPDATE: <strong>CONCLUSION</strong></p>\n<p>Shared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.</p>\n<p>======================================================</p>\n<p>======================================================<br>\nUPDATE: <strong>Paper</strong></p>\n<p>The paper <a href=\"https://arxiv.org/pdf/1612.01345.pdf\" target=\"_blank\">Human-In-The-Loop Person Re-Identification</a> it's the reason of why i belive that insane ensembles are not a good solution and therefore harmless in this competition.</p>\n<p>======================================================</p>\n<p><strong>My apollogies for my level of English</strong>, please think about what you have read.</p>",
      "rawMarkdown": "**Before you down-vote me please, read carefully**\n\nLet me first offer to you evidence that the **Private Leaderboard** might be very different from the **Public**:\n- [Prostate cANcer graDe Assessment (PANDA) Challenge](https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard)\n- [M5 Forecasting - Accuracy](https://www.kaggle.com/c/m5-forecasting-accuracy/leaderboard)\n- [Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/leaderboard)\n\n======================================================\nUPDATE: **Now this competition**\n\n- [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification/leaderboard)\n\n======================================================\n\nAfter that but before my argument i must tell something if you are **GrandMaster**, **Kaggle Veteran** or whoever that achieve a gold medal you are strictly immune to the **high-score kernels**. \n\nKaggle it's a platform/community for people whom are or wants to be a **Data Scientist**. Therefore, there are a lot of content about this topic. Comments, notebooks and datasets are plenty of healthy content for novices, sometimes notebooks with good solutions for a competition. Talking about competitions, there is a kind of way to know if your solution it's better than a previous one or another's solution, **score**, but the score it's not 100% reliable. Sometimes the solution overfits the public test set. Recently people ask to Kaggle to highly regulate **high-score kernels**, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of **chilling effects**.  Long-term this might poisoned this community. \n\nAs novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a **Data Scienctist**.  At the end i'm not saying that people who share **high-score kernels** are fine but if we forbid that, we are gonna decrease the quality of the content shared (by a kind of**chilling effects**).\n\n======================================================\nUPDATE: **CONCLUSION**\n\nShared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.\n\n======================================================\n\n======================================================\nUPDATE: **Paper**\n\nThe paper [Human-In-The-Loop Person Re-Identification](https://arxiv.org/pdf/1612.01345.pdf) it's the reason of why i belive that insane ensembles are not a good solution and therefore harmless in this competition.\n\n======================================================\n\n**My apollogies for my level of English**, please think about what you have read."
    },
    {
      "id": 973759,
      "postDate": "2020-08-17T13:53:58.410Z",
      "content": "<p>yeah i am one of those novices too. This is actually my first competition where i worked hard and undoubtedly with the help of other fellow kagglers i managed to learn many new things in this competition.<br>\nBut encountering high public kernels in the final days of the competition is discouraging when you work hard to get to the same place. Its better to share ideas rather csv files.<br>\nVast learning community of kaggle is the reason to be here in the first place so its better if its challenging than getting toxic. </p>",
      "rawMarkdown": "yeah i am one of those novices too. This is actually my first competition where i worked hard and undoubtedly with the help of other fellow kagglers i managed to learn many new things in this competition.\nBut encountering high public kernels in the final days of the competition is discouraging when you work hard to get to the same place. Its better to share ideas rather csv files.\nVast learning community of kaggle is the reason to be here in the first place so its better if its challenging than getting toxic. ",
      "votes": 1
    },
    {
      "id": 972702,
      "postDate": "2020-08-16T18:40:08.913Z",
      "content": "<p>To your statement,  \"Recently people ask to Kaggle to highly regulate high-score kernels, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of chilling effects. Long-term this might poisoned this community.\"</p>\n<p>We have asked for regulating high scoring kernels in the end days of the competition. Specially when it is a blend kernel.</p>",
      "rawMarkdown": "To your statement,  \"Recently people ask to Kaggle to highly regulate high-score kernels, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of chilling effects. Long-term this might poisoned this community.\"\n\nWe have asked for regulating high scoring kernels in the end days of the competition. Specially when it is a blend kernel.",
      "votes": 1,
      "replies": [
        {
          "id": 972707,
          "postDate": "2020-08-16T18:43:46.373Z",
          "content": "<p><a href=\"https://www.kaggle.com/roydatascience\" target=\"_blank\">@roydatascience</a> as i said in <strong>CONCLUSION</strong>:</p>\n<blockquote>\n  <p>Shared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.</p>\n</blockquote>\n<p>That's my point of view.</p>",
          "rawMarkdown": "@roydatascience as i said in **CONCLUSION**:\n\n> Shared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.\n\nThat's my point of view.",
          "votes": -2
        }
      ]
    },
    {
      "id": 972891,
      "postDate": "2020-08-17T00:20:09.367Z",
      "content": "<p>Sure sharing high-score kernels is fine with me as long as they don't pop up <strong>in the last few weeks before competition ends</strong>.</p>",
      "rawMarkdown": "Sure sharing high-score kernels is fine with me as long as they don't pop up **in the last few weeks before competition ends**.",
      "votes": 2
    },
    {
      "id": 972847,
      "postDate": "2020-08-16T22:12:05.760Z",
      "content": "<p>You have to take the time to try and fail to actually \"become a data scientist\". Nothing is delivered on a silver plate in the real world.</p>\n<p>Reading other codes is a good growth opportunity yes. I still do it on kaggle or on github for paper implementation. It doesn't matter if you're a kaggle veteran, an industry senior or a total beginner. That said, forking somebody's work and playing with parameters (or worse, fork + submit with no change) is hardly learning anything.</p>",
      "rawMarkdown": "You have to take the time to try and fail to actually \"become a data scientist\". Nothing is delivered on a silver plate in the real world.\n\nReading other codes is a good growth opportunity yes. I still do it on kaggle or on github for paper implementation. It doesn't matter if you're a kaggle veteran, an industry senior or a total beginner. That said, forking somebody's work and playing with parameters (or worse, fork + submit with no change) is hardly learning anything.",
      "votes": 2
    },
    {
      "id": 973327,
      "postDate": "2020-08-17T08:47:42.700Z",
      "content": "<p>There is a difference between working to learn vs working to win a competition.</p>\n<p>Also, the paper you cite is a case of \"getting more ground truth is the best way to improve performance\". Here in competition, we are in a situation where there won't be any new data; so that's one unrealistic non-real-world situation that can't be easily fixed in a \"public competition\" setting. </p>",
      "rawMarkdown": "There is a difference between working to learn vs working to win a competition.\n\nAlso, the paper you cite is a case of \"getting more ground truth is the best way to improve performance\". Here in competition, we are in a situation where there won't be any new data; so that's one unrealistic non-real-world situation that can't be easily fixed in a \"public competition\" setting. "
    },
    {
      "id": 973271,
      "postDate": "2020-08-17T07:53:58.363Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 972554,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-16T16:23:22.837000",
      "content": "<p>I wonder how you can learn anything from blending kernels that dont even cite the sources they have the models from?</p>",
      "votes": 12,
      "replies": [
        {
          "id": 972563,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T16:33:31.063000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> As i said in my post i'm defending the good solution that also has a high-score. And again, as i said there're notebooks that overfit the public test set. With all the respect that you deserve, you are a GrandMaster, you are not affected. Novices are the more affected to just fork the kind of notebook that you mentioned because we don't learn anything.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972579,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-16T16:49:15.320000",
          "content": "<p>So shouldn't it then be also us complaining about them?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 972604,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T17:13:47.080000",
          "content": "<p>DISCLAIMER: I don't down-vote you <strong>i respect your point Mr</strong>.</p>\n<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Yes you can, but in my post i explain:</p>\n<blockquote>\n  <p>As novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a Data Scienctist.</p>\n</blockquote>\n<p><strong>I NEVER SAID</strong>:</p>\n<blockquote>\n  <p>[…] I learned a lot seeing <strong>blending kernels that dont even cite the sources they have the models from</strong>.</p>\n</blockquote>",
          "votes": -4,
          "replies": []
        }
      ]
    },
    {
      "id": 972393,
      "author_name": "ELEVEN",
      "author_url": "",
      "post_date": "2020-08-16T14:15:07.563000",
      "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> it's not about sharing high-score kernels. It's about sharing high-score kernels during the last week of a competition, ignoring the standard Kaggle request not to do it. In this competition it happened 2 days before the competition end. It's disrespect to all Kaggle community for the sake of one notebook medal.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 972394,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T14:18:04.950000",
          "content": "<p>Both of them, i saw people more than one month ago complaining about the high-scor kernels.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 972401,
      "author_name": "Md Fahim",
      "author_url": "",
      "post_date": "2020-08-16T14:27:25.030000",
      "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> ,<br>\n<a href=\"https://www.kaggle.com/raininbox/blend-different-models-with-different-n-tile\" target=\"_blank\">This Public Kernel</a> in <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard\" target=\"_blank\">PANDA</a> competition won the silver medal. Some people won the silver medal just forking the kernel. <br>\nDoes not it affect the others ? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 972406,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T14:30:56.123000",
          "content": "<p><a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> he shared one mor before the timeline end. We have the time to learn and improve our models with his work, personally i don't use it because i used ResNet and he EfficientNet but again people could learn and improve their works with his.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 972417,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T14:37:53.130000",
          "content": "<p>UPDATE:<br>\n<a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> he doesn't achive silver, he achive bronze (cause he has the 75 position).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972431,
          "author_name": "Md Fahim",
          "author_url": "",
          "post_date": "2020-08-16T14:51:37.267000",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> , <br>\nYes, he doesn't achieve the silver medal but his public kernel does. You may see <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169115\" target=\"_blank\">my post</a> for more clearification.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972442,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T14:57:39.080000",
          "content": "<p>I saw it <a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a>. You said:</p>\n<blockquote>\n  <p>Amazing this public kernel got the silver medal</p>\n</blockquote>\n<p>And nothing more, so i have two questions, would you kindly explain why his wonderful kernel harms anyone if he shared one month before the close of the PANDA competition? and why matters that he achive a silver medal (by votes)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972472,
          "author_name": "Md Fahim",
          "author_url": "",
          "post_date": "2020-08-16T15:13:47.147000",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> , <br>\nYou understand me wrong. I does not claim the beautiful kernel and the other who works on this. But you know, some of people just forked it and did not work so much. Fortunately, those people got the medal.    I am telling about them. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 972503,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T15:40:59.800000",
          "content": "<p><a href=\"https://www.kaggle.com/mdfahimreshm\" target=\"_blank\">@mdfahimreshm</a> thanks for clarify your idea, now i understand. About the people who achives bronze medal there're too few of them (just looking the private leaderboard i say that), and we must allow ourselves to lose to learn, now we know that an ensemble it's better generalizing than one single model because here comes my point, during PANDA competition his work hasn't the highest score (that's the reason of why his notebook only achives silver), many use the highest score kernel and not the better. During that competition whe didn't know that we could achive a bronze medal with that notebook. In resume, you are sayin that people achive a bronze medal just forking his work but i'm saying there were few who made them, because he hasn't the highest score, the people who just fork that notebook \"took a leap of faith\".</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 973022,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-08-17T03:44:45.287000",
          "content": "<p>In such competitions like PANDA with very unstable LB, if 100 people publish their kernels, several of them could achieve quite high score at private LB just by coincidence. I'll just give an example. In <a href=\"https://www.kaggle.com/c/bengaliai-cv19/overview\" target=\"_blank\">Bengali competition</a> in a first few days after the start I published a series of kernels with a full pipeline: data cleaning and preprocessing, training, and submission to give new participant a good starting point in that challenge. The public score of <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference\" target=\"_blank\">my kernel</a> is quite low, just 0.964, which corresponds to <strong>~1200 at public LB</strong> at the end of the competition, and there were a number of kernels with 0.97+ score posted during the competition. However, at the end of the competition this kernel appeared to work quite well at private LB, giving <strong>~100th rank</strong>, and <strong>~30 people got bronze with nearly a single submission</strong>. Also a slight modification of the kernel could easily bring to silver, as several people contacted me to say \"thank you\". But it is clearly just a coincidence, and if the private LB was more predictable (or if I payed more attention to that organizers want), our team would get gold, not dropped to 200+ place because of a wrong selection( It happens, and cannot be fully avoided. </p>\n<p>There are just several things that people may follow to minimize appearance of strong public kernels at private LB. (1) Publish quite basic pipeline. For example, the amazing TF TPU pipeline published in this competition with huge models, external data, proper CV split, smooth labels, etc. is far away from being the basic one. (2) Do not upload models trained externally, share the code that could run on kaggle.  Otherwise people who don't have their own hardware and working only at kaggle would get screwed. And also such uploading facilitates creation of crazy blending kernels. (3) upvote only the content that you want to see at kaggle. If you think that blending kernels are bad, just don't upvote them. (4) I think the best time to publish a full pipeline, if it is strong, is the first month of a competition, to give people time to work on it, explore new ideas, and learn something. So you can promote some progress in the competition, not just make people angry. Some cool ideas certainly could be published later, but it is really bad if one is publishing something game changing in several weeks before the end of the competition, or like we saw today appearance of a high scoring kernel just in one day before the competition end. It may really ruin the competition. All above it it just a recommendation for being a good member of our diverse kaggle community. <br>\nThe thing I just wanted to say is that if people really try to help others and share their ideas and code, help to the progress in the competition, it should not be discouraged. Did u think how this competition would look like without the contribution made by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and other people shared their ideas?</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 973641,
          "author_name": "Haider Ali Shuvo",
          "author_url": "",
          "post_date": "2020-08-17T12:22:42.380000",
          "content": "<p>Well Said <a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a> , Without help of others kaggle wouldnt be this amazing and fun :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 972697,
      "author_name": "Ashish Gupta",
      "author_url": "",
      "post_date": "2020-08-16T18:36:55.373000",
      "content": "<p>You are certainly not right with the Jigsaw, I saw people getting bronze and silvers with a bad blend solution.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 972705,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T18:41:47.030000",
          "content": "<p>Would you share one work like you described (i'm not being challenging i'm just curious)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972750,
          "author_name": "Ashish Gupta",
          "author_url": "",
          "post_date": "2020-08-16T19:36:27.583000",
          "content": "<p>Take Tarek Hamdi's work for example. His imbalance weight ensembling fetched him a bronze medal on Private lb. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 972512,
      "author_name": "Pranav Singh",
      "author_url": "",
      "post_date": "2020-08-16T15:51:52.890000",
      "content": "<p>I totally agree your learning part of the argument, but I don't see the point in making a last minute dash to a good leaderboard position.I mean yes we can make adjustments and improve our methods but then again, Sharing kernel with medal level accuracies should be prohibited at least a week before the deadline.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 972518,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T15:58:58.280000",
          "content": "<p>I understand your point. But as i said this could trigger <strong>chilling effects</strong> in this community. The risk is high enough for me to think that it's a bad idea.</p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 972710,
          "author_name": "Pranav Singh",
          "author_url": "",
          "post_date": "2020-08-16T18:47:39.337000",
          "content": "<p>After the competition is over maybe the high scoring kernel can be made public, then people can analyse and improve models to better their scores.So learning part would be covered that way, And having said that anyone who wants to learn will learn, No matter what!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 972723,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T18:58:34.557000",
          "content": "<p>I'm going to stop answering because no matters how much evidence and arguments i made i just get downvotes \"because reasons\". Just think about what i write in the post, just that, don't change your opinion just think what could happend if Kaggle don't to publish high-score kernels (at any time).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 973759,
      "author_name": "Pawan KS",
      "author_url": "",
      "post_date": "2020-08-17T13:53:58.410000",
      "content": "<p>yeah i am one of those novices too. This is actually my first competition where i worked hard and undoubtedly with the help of other fellow kagglers i managed to learn many new things in this competition.<br>\nBut encountering high public kernels in the final days of the competition is discouraging when you work hard to get to the same place. Its better to share ideas rather csv files.<br>\nVast learning community of kaggle is the reason to be here in the first place so its better if its challenging than getting toxic. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 972702,
      "author_name": "Ashish Gupta",
      "author_url": "",
      "post_date": "2020-08-16T18:40:08.913000",
      "content": "<p>To your statement,  \"Recently people ask to Kaggle to highly regulate high-score kernels, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of chilling effects. Long-term this might poisoned this community.\"</p>\n<p>We have asked for regulating high scoring kernels in the end days of the competition. Specially when it is a blend kernel.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 972707,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-08-16T18:43:46.373000",
          "content": "<p><a href=\"https://www.kaggle.com/roydatascience\" target=\"_blank\">@roydatascience</a> as i said in <strong>CONCLUSION</strong>:</p>\n<blockquote>\n  <p>Shared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.</p>\n</blockquote>\n<p>That's my point of view.</p>",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 972891,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-08-17T00:20:09.367000",
      "content": "<p>Sure sharing high-score kernels is fine with me as long as they don't pop up <strong>in the last few weeks before competition ends</strong>.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 972847,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-08-16T22:12:05.760000",
      "content": "<p>You have to take the time to try and fail to actually \"become a data scientist\". Nothing is delivered on a silver plate in the real world.</p>\n<p>Reading other codes is a good growth opportunity yes. I still do it on kaggle or on github for paper implementation. It doesn't matter if you're a kaggle veteran, an industry senior or a total beginner. That said, forking somebody's work and playing with parameters (or worse, fork + submit with no change) is hardly learning anything.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 973327,
      "author_name": "George Rey",
      "author_url": "",
      "post_date": "2020-08-17T08:47:42.700000",
      "content": "<p>There is a difference between working to learn vs working to win a competition.</p>\n<p>Also, the paper you cite is a case of \"getting more ground truth is the best way to improve performance\". Here in competition, we are in a situation where there won't be any new data; so that's one unrealistic non-real-world situation that can't be easily fixed in a \"public competition\" setting. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 973271,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-17T07:53:58.363000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "972554": "I wonder how you can learn anything from blending kernels that dont even cite the sources they have the models from?",
    "972393": "@hiramcho it's not about sharing high-score kernels. It's about sharing high-score kernels during the last week of a competition, ignoring the standard Kaggle request not to do it. In this competition it happened 2 days before the competition end. It's disrespect to all Kaggle community for the sake of one notebook medal.",
    "972401": "@hiramcho ,\n[This Public Kernel](https://www.kaggle.com/raininbox/blend-different-models-with-different-n-tile) in [PANDA](https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard) competition won the silver medal. Some people won the silver medal just forking the kernel. \nDoes not it affect the others ? ",
    "972697": "You are certainly not right with the Jigsaw, I saw people getting bronze and silvers with a bad blend solution.",
    "972512": "I totally agree your learning part of the argument, but I don't see the point in making a last minute dash to a good leaderboard position.I mean yes we can make adjustments and improve our methods but then again, Sharing kernel with medal level accuracies should be prohibited at least a week before the deadline.",
    "972326": "**Before you down-vote me please, read carefully**\n\nLet me first offer to you evidence that the **Private Leaderboard** might be very different from the **Public**:\n- [Prostate cANcer graDe Assessment (PANDA) Challenge](https://www.kaggle.com/c/prostate-cancer-grade-assessment/leaderboard)\n- [M5 Forecasting - Accuracy](https://www.kaggle.com/c/m5-forecasting-accuracy/leaderboard)\n- [Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/leaderboard)\n\n======================================================\nUPDATE: **Now this competition**\n\n- [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification/leaderboard)\n\n======================================================\n\nAfter that but before my argument i must tell something if you are **GrandMaster**, **Kaggle Veteran** or whoever that achieve a gold medal you are strictly immune to the **high-score kernels**. \n\nKaggle it's a platform/community for people whom are or wants to be a **Data Scientist**. Therefore, there are a lot of content about this topic. Comments, notebooks and datasets are plenty of healthy content for novices, sometimes notebooks with good solutions for a competition. Talking about competitions, there is a kind of way to know if your solution it's better than a previous one or another's solution, **score**, but the score it's not 100% reliable. Sometimes the solution overfits the public test set. Recently people ask to Kaggle to highly regulate **high-score kernels**, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of **chilling effects**.  Long-term this might poisoned this community. \n\nAs novice i learned a lot seeing other peoples code and even winners' code. We, the novices are highly dependent of previous works to learn from those. If we only see poor solutions there's no opportunity to us to become a **Data Scienctist**.  At the end i'm not saying that people who share **high-score kernels** are fine but if we forbid that, we are gonna decrease the quality of the content shared (by a kind of**chilling effects**).\n\n======================================================\nUPDATE: **CONCLUSION**\n\nShared high-score kernels, at any time of a competition it's a necessary evil, because we can learn during or after the competition new techniques or algorithms. Forbid that kind of notebooks or harass the people who made them it's going be worse than just let them be.\n\n======================================================\n\n======================================================\nUPDATE: **Paper**\n\nThe paper [Human-In-The-Loop Person Re-Identification](https://arxiv.org/pdf/1612.01345.pdf) it's the reason of why i belive that insane ensembles are not a good solution and therefore harmless in this competition.\n\n======================================================\n\n**My apollogies for my level of English**, please think about what you have read.",
    "973759": "yeah i am one of those novices too. This is actually my first competition where i worked hard and undoubtedly with the help of other fellow kagglers i managed to learn many new things in this competition.\nBut encountering high public kernels in the final days of the competition is discouraging when you work hard to get to the same place. Its better to share ideas rather csv files.\nVast learning community of kaggle is the reason to be here in the first place so its better if its challenging than getting toxic. ",
    "972702": "To your statement,  \"Recently people ask to Kaggle to highly regulate high-score kernels, kernels that could be overfitting. Then, kagglers could be uncomfortable to share great solutions for fear of being rejected by the community, creating a kind of chilling effects. Long-term this might poisoned this community.\"\n\nWe have asked for regulating high scoring kernels in the end days of the competition. Specially when it is a blend kernel.",
    "972891": "Sure sharing high-score kernels is fine with me as long as they don't pop up **in the last few weeks before competition ends**.",
    "972847": "You have to take the time to try and fail to actually \"become a data scientist\". Nothing is delivered on a silver plate in the real world.\n\nReading other codes is a good growth opportunity yes. I still do it on kaggle or on github for paper implementation. It doesn't matter if you're a kaggle veteran, an industry senior or a total beginner. That said, forking somebody's work and playing with parameters (or worse, fork + submit with no change) is hardly learning anything.",
    "973327": "There is a difference between working to learn vs working to win a competition.\n\nAlso, the paper you cite is a case of \"getting more ground truth is the best way to improve performance\". Here in competition, we are in a situation where there won't be any new data; so that's one unrealistic non-real-world situation that can't be easily fixed in a \"public competition\" setting. ",
    "973271": ""
  }
}