{
  "id": 175318,
  "title": "I said it: high score public kernels are dangerous",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175318",
  "author_name": "CPMP",
  "post_date": "2020-08-18T00:13:53.825000",
  "votes": 33,
  "comment_count": 32,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fce0b4af993372f7c52477a423182e536%2FScreenshot_2020-08-18%20MinMax%20highest%20public%20LB%209619.png?generation=1597709631496097&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 974445,
      "postDate": "2020-08-18T00:26:28.747Z",
      "content": "<p>maybe this kind of challenge is good in kaggle. focus on \"good generalization\", rather then training 100 models for ensemble or engineering tricks.</p>\n<p>\"good generalization\" is real issue in industrial application.</p>\n<p>like to see more discussion on improving generalization in discussion.</p>\n<p>most kaggler focus on training and building models and less time on analyse results and thinking of robustness and shakeup.</p>\n<p>the real magic of data scientist is not just to get good accuracy, but also able to know and predict accuracy the reliability and generalistion power of their model and solution. this is the art and magic.</p>",
      "rawMarkdown": "maybe this kind of challenge is good in kaggle. focus on \"good generalization\", rather then training 100 models for ensemble or engineering tricks.\n\n \"good generalization\" is real issue in industrial application.\n\nlike to see more discussion on improving generalization in discussion.\n\nmost kaggler focus on training and building models and less time on analyse results and thinking of robustness and shakeup.\n\nthe real magic of data scientist is not just to get good accuracy, but also able to know and predict accuracy the reliability and generalistion power of their model and solution. this is the art and magic.",
      "votes": 29
    },
    {
      "id": 974410,
      "postDate": "2020-08-18T00:13:53.827Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fce0b4af993372f7c52477a423182e536%2FScreenshot_2020-08-18%20MinMax%20highest%20public%20LB%209619.png?generation=1597709631496097&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fce0b4af993372f7c52477a423182e536%2FScreenshot_2020-08-18%20MinMax%20highest%20public%20LB%209619.png?generation=1597709631496097&alt=media)",
      "votes": 33
    },
    {
      "id": 977143,
      "postDate": "2020-08-19T10:14:39.950Z",
      "content": "<p>This is funny.  During the competition, I wrote that selecting a high score public kernel is like gambling if you don't reproduce its results with your own cross validation.  And some other Kagglers said the same.</p>\n<p>Then the top scoring kernel performs badly in private LB.    Frankly, what does it take to show people that blindly reusing kernels is dangerous if this example is not enough?  I'm lost here.</p>\n<p>I see a number of comments arguing that some high scoring kernels did well in private.  Well, that's what gambling is about: some win, some lose, at random.  But overall people lose.</p>\n<p>Gamble if you wish.  But don't say you were not warned.  And no need to argue with the reality I show in the post.  </p>",
      "rawMarkdown": "This is funny.  During the competition, I wrote that selecting a high score public kernel is like gambling if you don't reproduce its results with your own cross validation.  And some other Kagglers said the same.\n\nThen the top scoring kernel performs badly in private LB.    Frankly, what does it take to show people that blindly reusing kernels is dangerous if this example is not enough?  I'm lost here.\n\nI see a number of comments arguing that some high scoring kernels did well in private.  Well, that's what gambling is about: some win, some lose, at random.  But overall people lose.\n\nGamble if you wish.  But don't say you were not warned.  And no need to argue with the reality I show in the post.  \n",
      "votes": 3,
      "replies": [
        {
          "id": 977207,
          "postDate": "2020-08-19T10:53:42.333Z",
          "content": "<p>First of all Thanks.<br>\nIn one discussion, I asked you about why not believe in these high scoring kernels and you said to trust CV. I followed your advice( and others too who said to belive in CV) And here is the outcome. I'm in the bronze medal zone. <br>\nThanks again.</p>",
          "rawMarkdown": "First of all Thanks.\nIn one discussion, I asked you about why not believe in these high scoring kernels and you said to trust CV. I followed your advice( and others too who said to belive in CV) And here is the outcome. I'm in the bronze medal zone. \nThanks again.",
          "votes": 2
        },
        {
          "id": 977215,
          "postDate": "2020-08-19T10:59:00.343Z",
          "content": "<p>You made my day, very happy to have been useful to you.</p>",
          "rawMarkdown": "You made my day, very happy to have been useful to you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 978103,
      "postDate": "2020-08-20T00:20:22.347Z",
      "content": "<p>Especially when they are basically all tuning their parameters based on the LB response. I made the argument that this is similar to hyperparameter searching on the (public) test set where there is enough models (in the ensemble) and parameters to tune to allow overfitting.</p>\n<p>The mass of people all forking and trying different combination then create the slow convergence to high scores (overfitted).</p>",
      "rawMarkdown": "Especially when they are basically all tuning their parameters based on the LB response. I made the argument that this is similar to hyperparameter searching on the (public) test set where there is enough models (in the ensemble) and parameters to tune to allow overfitting.\n\nThe mass of people all forking and trying different combination then create the slow convergence to high scores (overfitted).",
      "votes": 2,
      "replies": [
        {
          "id": 978275,
          "postDate": "2020-08-20T04:34:08.580Z",
          "content": "<p><a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> I have a question, as you are saying that one should not tune his/her model on the basis of public leaderboard result. So should we tune our model on the basis of local CV score right ? </p>",
          "rawMarkdown": "@arroqc I have a question, as you are saying that one should not tune his/her model on the basis of public leaderboard result. So should we tune our model on the basis of local CV score right ? "
        },
        {
          "id": 979478,
          "postDate": "2020-08-20T21:49:40.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/prateek0x\" target=\"_blank\">@prateek0x</a> What I'm saying is that the LB should be used carefully. If you are using it as your guide then you are basically making that your CV and you should use your parameters in a way that has a chance to generalize.</p>\n<p>What I am warning on is that this competition had both a lot of users and a rather \"easy\" LB where the goal was to correctly rank only 78 cases. This makes it very easy to overfit by just random selections of models/paramaters and so you are taking a gamble when using such kernels as submission.</p>\n<p>The best is often to use models that both increase your CV and the LB. If one improve and not the other, one should try to guess why and act accordingly (noisy LB =&gt; focus on CV, very different distribution between train and test =&gt; focus on LB).</p>",
          "rawMarkdown": "@prateek0x What I'm saying is that the LB should be used carefully. If you are using it as your guide then you are basically making that your CV and you should use your parameters in a way that has a chance to generalize.\n\nWhat I am warning on is that this competition had both a lot of users and a rather \"easy\" LB where the goal was to correctly rank only 78 cases. This makes it very easy to overfit by just random selections of models/paramaters and so you are taking a gamble when using such kernels as submission.\n\nThe best is often to use models that both increase your CV and the LB. If one improve and not the other, one should try to guess why and act accordingly (noisy LB => focus on CV, very different distribution between train and test => focus on LB).",
          "votes": 1
        }
      ]
    },
    {
      "id": 975454,
      "postDate": "2020-08-18T09:55:55.297Z",
      "content": "<p>To win a big lottery you may visit our game just pay and in it . <a href=\"https://www.kbclucky.com/\" target=\"_blank\">Click</a></p>",
      "rawMarkdown": "To win a big lottery you may visit our game just pay and in it . [Click](https://www.kbclucky.com/)",
      "votes": -13
    },
    {
      "id": 979498,
      "postDate": "2020-08-20T22:19:14.510Z",
      "content": "<p>Fair point put by you</p>",
      "rawMarkdown": "Fair point put by you"
    },
    {
      "id": 979496,
      "postDate": "2020-08-20T22:18:54.707Z",
      "content": "<p>Fair point</p>",
      "rawMarkdown": "Fair point"
    },
    {
      "id": 975478,
      "postDate": "2020-08-18T10:06:14.807Z",
      "content": "<p>I wouldn't say the same. It all comes down to how the kernels are used. I saw Chris's kernels on using effnets and started there by reimplementing the strategy and working out experiments from there. It all depends on intentions. If you're interested to learn from the best and study their methods they outline in these kernels it is actually very beneficial. </p>",
      "rawMarkdown": "I wouldn't say the same. It all comes down to how the kernels are used. I saw Chris's kernels on using effnets and started there by reimplementing the strategy and working out experiments from there. It all depends on intentions. If you're interested to learn from the best and study their methods they outline in these kernels it is actually very beneficial. ",
      "replies": [
        {
          "id": 976516,
          "postDate": "2020-08-18T23:20:58.310Z",
          "content": "<p>I think he was talking about kernels which use only blending with random coefficient. </p>",
          "rawMarkdown": "I think he was talking about kernels which use only blending with random coefficient. ",
          "votes": 2
        },
        {
          "id": 977271,
          "postDate": "2020-08-19T11:38:59.337Z",
          "content": "<p>Fair point. </p>",
          "rawMarkdown": "Fair point. "
        },
        {
          "id": 978728,
          "postDate": "2020-08-20T10:56:59.277Z",
          "content": "<blockquote>\n  <p>I think he was talking about kernels which use only blending with random coefficient. </p>\n</blockquote>\n<p>It's not even random…It is, may be, more \"dangerous\" as these coeffs are mainly based on public LB feedback.  Which increase the overfitting : If initial kernel overfit, then blend with high coeff will overfit more , same for blend of blends etc. </p>",
          "rawMarkdown": "> I think he was talking about kernels which use only blending with random coefficient. \n\nIt's not even random...It is, may be, more \"dangerous\" as these coeffs are mainly based on public LB feedback.  Which increase the overfitting : If initial kernel overfit, then blend with high coeff will overfit more , same for blend of blends etc. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 974424,
      "postDate": "2020-08-18T00:19:42.957Z",
      "content": "<p>Just a quick look through most public kernels, none of them did well in private LB.</p>",
      "rawMarkdown": "Just a quick look through most public kernels, none of them did well in private LB.",
      "replies": [
        {
          "id": 974430,
          "postDate": "2020-08-18T00:20:49.993Z",
          "content": "<p>I would have been surprised if any of them did well given they only relied on public LB feedback for tuning.</p>",
          "rawMarkdown": "I would have been surprised if any of them did well given they only relied on public LB feedback for tuning."
        },
        {
          "id": 974442,
          "postDate": "2020-08-18T00:25:16.430Z",
          "content": "<p>There's a few random ones that did well - ex. <a href=\"https://www.kaggle.com/ajaykumar7778/fork-of-ensemble-melanoma-ac9964?scriptVersionId=40334665\" target=\"_blank\">https://www.kaggle.com/ajaykumar7778/fork-of-ensemble-melanoma-ac9964?scriptVersionId=40334665</a></p>",
          "rawMarkdown": "There's a few random ones that did well - ex. https://www.kaggle.com/ajaykumar7778/fork-of-ensemble-melanoma-ac9964?scriptVersionId=40334665"
        },
        {
          "id": 977148,
          "postDate": "2020-08-19T10:18:47.820Z",
          "content": "<p>Yes but you have no clue to know if this one can perform well in the private LB before the end of the competitioon. There is no crossval available in these kernels.</p>",
          "rawMarkdown": "Yes but you have no clue to know if this one can perform well in the private LB before the end of the competitioon. There is no crossval available in these kernels.",
          "votes": 1
        },
        {
          "id": 978093,
          "postDate": "2020-08-19T23:52:39.400Z",
          "content": "<p>Agreed, its just a lottery if you're submitting random public kernels without any scrutiny / combination of your own models.</p>",
          "rawMarkdown": "Agreed, its just a lottery if you're submitting random public kernels without any scrutiny / combination of your own models."
        }
      ]
    },
    {
      "id": 978973,
      "postDate": "2020-08-20T14:46:34.713Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 979015,
          "postDate": "2020-08-20T15:11:07.097Z",
          "content": "<p>your training data is way larger than public test.  Why not use it for validation?</p>",
          "rawMarkdown": "your training data is way larger than public test.  Why not use it for validation?"
        },
        {
          "id": 979121,
          "postDate": "2020-08-20T16:21:58.603Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 979153,
          "postDate": "2020-08-20T16:56:35.267Z",
          "content": "<p>I'm discussing model performance evaluation as well.  </p>",
          "rawMarkdown": "I'm discussing model performance evaluation as well.  "
        },
        {
          "id": 979162,
          "postDate": "2020-08-20T17:01:38.020Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 979472,
          "postDate": "2020-08-20T21:34:38.700Z",
          "content": "<p>I know ;)</p>\n<p>As soon as you use test data to evaluate your model performance you fit to it, which most likely mean you overfit to it.</p>\n<p>This is made worse because public test data is small.</p>\n<p>What you should do is to use kfold cross validation.  It is more reliable as it is based on k train/validation split, and validation data is larger.  You can use public test score as well, but it should at most be used as one of your folds.</p>",
          "rawMarkdown": "I know ;)\n\nAs soon as you use test data to evaluate your model performance you fit to it, which most likely mean you overfit to it.\n\nThis is made worse because public test data is small.\n\nWhat you should do is to use kfold cross validation.  It is more reliable as it is based on k train/validation split, and validation data is larger.  You can use public test score as well, but it should at most be used as one of your folds.\n\n\n",
          "votes": 1
        },
        {
          "id": 979893,
          "postDate": "2020-08-21T07:26:38.003Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 980098,
          "postDate": "2020-08-21T10:27:28.603Z",
          "content": "<p>kfold cross validation is not a machine learning algorithm.  It is a way to evaluate any machine learning algorithm.</p>",
          "rawMarkdown": "kfold cross validation is not a machine learning algorithm.  It is a way to evaluate any machine learning algorithm."
        },
        {
          "id": 980127,
          "postDate": "2020-08-21T10:46:20.233Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 980198,
          "postDate": "2020-08-21T11:46:26.513Z",
          "content": "<blockquote>\n  <p>kflod method will be generally more beneficial than any other one, strange!</p>\n</blockquote>\n<p>Why strange?</p>",
          "rawMarkdown": "> kflod method will be generally more beneficial than any other one, strange!\n\nWhy strange?"
        },
        {
          "id": 980224,
          "postDate": "2020-08-21T12:22:56.463Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 974436,
      "postDate": "2020-08-18T00:23:20.217Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 974433,
      "postDate": "2020-08-18T00:22:22.787Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 974445,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-08-18T00:26:28.747000",
      "content": "<p>maybe this kind of challenge is good in kaggle. focus on \"good generalization\", rather then training 100 models for ensemble or engineering tricks.</p>\n<p>\"good generalization\" is real issue in industrial application.</p>\n<p>like to see more discussion on improving generalization in discussion.</p>\n<p>most kaggler focus on training and building models and less time on analyse results and thinking of robustness and shakeup.</p>\n<p>the real magic of data scientist is not just to get good accuracy, but also able to know and predict accuracy the reliability and generalistion power of their model and solution. this is the art and magic.</p>",
      "votes": 29,
      "replies": []
    },
    {
      "id": 977143,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-19T10:14:39.950000",
      "content": "<p>This is funny.  During the competition, I wrote that selecting a high score public kernel is like gambling if you don't reproduce its results with your own cross validation.  And some other Kagglers said the same.</p>\n<p>Then the top scoring kernel performs badly in private LB.    Frankly, what does it take to show people that blindly reusing kernels is dangerous if this example is not enough?  I'm lost here.</p>\n<p>I see a number of comments arguing that some high scoring kernels did well in private.  Well, that's what gambling is about: some win, some lose, at random.  But overall people lose.</p>\n<p>Gamble if you wish.  But don't say you were not warned.  And no need to argue with the reality I show in the post.  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 977207,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-08-19T10:53:42.333000",
          "content": "<p>First of all Thanks.<br>\nIn one discussion, I asked you about why not believe in these high scoring kernels and you said to trust CV. I followed your advice( and others too who said to belive in CV) And here is the outcome. I'm in the bronze medal zone. <br>\nThanks again.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 977215,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-19T10:59:00.343000",
          "content": "<p>You made my day, very happy to have been useful to you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 978103,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-08-20T00:20:22.347000",
      "content": "<p>Especially when they are basically all tuning their parameters based on the LB response. I made the argument that this is similar to hyperparameter searching on the (public) test set where there is enough models (in the ensemble) and parameters to tune to allow overfitting.</p>\n<p>The mass of people all forking and trying different combination then create the slow convergence to high scores (overfitted).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 978275,
          "author_name": "Prateek Mishra",
          "author_url": "",
          "post_date": "2020-08-20T04:34:08.580000",
          "content": "<p><a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> I have a question, as you are saying that one should not tune his/her model on the basis of public leaderboard result. So should we tune our model on the basis of local CV score right ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979478,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-20T21:49:40.993000",
          "content": "<p><a href=\"https://www.kaggle.com/prateek0x\" target=\"_blank\">@prateek0x</a> What I'm saying is that the LB should be used carefully. If you are using it as your guide then you are basically making that your CV and you should use your parameters in a way that has a chance to generalize.</p>\n<p>What I am warning on is that this competition had both a lot of users and a rather \"easy\" LB where the goal was to correctly rank only 78 cases. This makes it very easy to overfit by just random selections of models/paramaters and so you are taking a gamble when using such kernels as submission.</p>\n<p>The best is often to use models that both increase your CV and the LB. If one improve and not the other, one should try to guess why and act accordingly (noisy LB =&gt; focus on CV, very different distribution between train and test =&gt; focus on LB).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 975454,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T09:55:55.297000",
      "content": "<p>To win a big lottery you may visit our game just pay and in it . <a href=\"https://www.kbclucky.com/\" target=\"_blank\">Click</a></p>",
      "votes": -13,
      "replies": []
    },
    {
      "id": 979498,
      "author_name": "aadish agrawal",
      "author_url": "",
      "post_date": "2020-08-20T22:19:14.510000",
      "content": "<p>Fair point put by you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 979496,
      "author_name": "aadish agrawal",
      "author_url": "",
      "post_date": "2020-08-20T22:18:54.707000",
      "content": "<p>Fair point</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975478,
      "author_name": "DragonPG",
      "author_url": "",
      "post_date": "2020-08-18T10:06:14.807000",
      "content": "<p>I wouldn't say the same. It all comes down to how the kernels are used. I saw Chris's kernels on using effnets and started there by reimplementing the strategy and working out experiments from there. It all depends on intentions. If you're interested to learn from the best and study their methods they outline in these kernels it is actually very beneficial. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 976516,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-08-18T23:20:58.310000",
          "content": "<p>I think he was talking about kernels which use only blending with random coefficient. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 977271,
          "author_name": "DragonPG",
          "author_url": "",
          "post_date": "2020-08-19T11:38:59.337000",
          "content": "<p>Fair point. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978728,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-20T10:56:59.277000",
          "content": "<blockquote>\n  <p>I think he was talking about kernels which use only blending with random coefficient. </p>\n</blockquote>\n<p>It's not even random…It is, may be, more \"dangerous\" as these coeffs are mainly based on public LB feedback.  Which increase the overfitting : If initial kernel overfit, then blend with high coeff will overfit more , same for blend of blends etc. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 974424,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-08-18T00:19:42.957000",
      "content": "<p>Just a quick look through most public kernels, none of them did well in private LB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 974430,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-18T00:20:49.993000",
          "content": "<p>I would have been surprised if any of them did well given they only relied on public LB feedback for tuning.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 974442,
          "author_name": "Ash Jha",
          "author_url": "",
          "post_date": "2020-08-18T00:25:16.430000",
          "content": "<p>There's a few random ones that did well - ex. <a href=\"https://www.kaggle.com/ajaykumar7778/fork-of-ensemble-melanoma-ac9964?scriptVersionId=40334665\" target=\"_blank\">https://www.kaggle.com/ajaykumar7778/fork-of-ensemble-melanoma-ac9964?scriptVersionId=40334665</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 977148,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-08-19T10:18:47.820000",
          "content": "<p>Yes but you have no clue to know if this one can perform well in the private LB before the end of the competitioon. There is no crossval available in these kernels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 978093,
          "author_name": "Ash Jha",
          "author_url": "",
          "post_date": "2020-08-19T23:52:39.400000",
          "content": "<p>Agreed, its just a lottery if you're submitting random public kernels without any scrutiny / combination of your own models.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 978973,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-20T14:46:34.713000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 979015,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-20T15:11:07.097000",
          "content": "<p>your training data is way larger than public test.  Why not use it for validation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979121,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-20T16:21:58.603000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979153,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-20T16:56:35.267000",
          "content": "<p>I'm discussing model performance evaluation as well.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979162,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-20T17:01:38.020000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979472,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-20T21:34:38.700000",
          "content": "<p>I know ;)</p>\n<p>As soon as you use test data to evaluate your model performance you fit to it, which most likely mean you overfit to it.</p>\n<p>This is made worse because public test data is small.</p>\n<p>What you should do is to use kfold cross validation.  It is more reliable as it is based on k train/validation split, and validation data is larger.  You can use public test score as well, but it should at most be used as one of your folds.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979893,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-21T07:26:38.003000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980098,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-21T10:27:28.603000",
          "content": "<p>kfold cross validation is not a machine learning algorithm.  It is a way to evaluate any machine learning algorithm.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980127,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-21T10:46:20.233000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980198,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-21T11:46:26.513000",
          "content": "<blockquote>\n  <p>kflod method will be generally more beneficial than any other one, strange!</p>\n</blockquote>\n<p>Why strange?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980224,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-21T12:22:56.463000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974436,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:23:20.217000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974433,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:22:22.787000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974445": "maybe this kind of challenge is good in kaggle. focus on \"good generalization\", rather then training 100 models for ensemble or engineering tricks.\n\n \"good generalization\" is real issue in industrial application.\n\nlike to see more discussion on improving generalization in discussion.\n\nmost kaggler focus on training and building models and less time on analyse results and thinking of robustness and shakeup.\n\nthe real magic of data scientist is not just to get good accuracy, but also able to know and predict accuracy the reliability and generalistion power of their model and solution. this is the art and magic.",
    "974410": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fce0b4af993372f7c52477a423182e536%2FScreenshot_2020-08-18%20MinMax%20highest%20public%20LB%209619.png?generation=1597709631496097&alt=media)",
    "977143": "This is funny.  During the competition, I wrote that selecting a high score public kernel is like gambling if you don't reproduce its results with your own cross validation.  And some other Kagglers said the same.\n\nThen the top scoring kernel performs badly in private LB.    Frankly, what does it take to show people that blindly reusing kernels is dangerous if this example is not enough?  I'm lost here.\n\nI see a number of comments arguing that some high scoring kernels did well in private.  Well, that's what gambling is about: some win, some lose, at random.  But overall people lose.\n\nGamble if you wish.  But don't say you were not warned.  And no need to argue with the reality I show in the post.  \n",
    "978103": "Especially when they are basically all tuning their parameters based on the LB response. I made the argument that this is similar to hyperparameter searching on the (public) test set where there is enough models (in the ensemble) and parameters to tune to allow overfitting.\n\nThe mass of people all forking and trying different combination then create the slow convergence to high scores (overfitted).",
    "975454": "To win a big lottery you may visit our game just pay and in it . [Click](https://www.kbclucky.com/)",
    "979498": "Fair point put by you",
    "979496": "Fair point",
    "975478": "I wouldn't say the same. It all comes down to how the kernels are used. I saw Chris's kernels on using effnets and started there by reimplementing the strategy and working out experiments from there. It all depends on intentions. If you're interested to learn from the best and study their methods they outline in these kernels it is actually very beneficial. ",
    "974424": "Just a quick look through most public kernels, none of them did well in private LB.",
    "978973": "",
    "974436": "",
    "974433": ""
  }
}