{
  "id": 175365,
  "title": "Kagglers acting like overfitting hyperparameter search",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175365",
  "author_name": "",
  "post_date": "2020-08-18T02:22:48.407858Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So all these debates about people using all the same public kernels that are blends of blends with some diverse ensembling technique made me realize something funny:</p>\n<blockquote>\n  <p>Kaggle users are, in this situation, acting like someone doing hyper parameter search directly on the (public) test set.</p>\n</blockquote>\n<p>Basically we have a rather small public leader board dataset where the goal is to correctly order \"only\" 78 malignant cases. Then you have a first wave of users trying the public kernels (changing bits of it) and some publishing their fork when the result is a bit better than the first notebook. But rather than using training data it uses test data (hidden) to make that call. </p>\n<p>Considering the number of users doing experiments on forks you are bound to have someone that finds a higher score solution ! It is very similar to doing Hyperparameter search directly on the test set: Choose params, submit, get score, repeat. Here the search is random.</p>\n<p>Then you have a second wave, that will use multiple kernels as described above and they will also have a set of hyperparameters like in min-max solution or weighted blends. We can easily imagine that with a completed enough pyramid of blends you get a final function that is flexible enough to really manipulate the order of predicted probabilities. When do they publish ? Once again, when the public LB increased. And once again they are numerous. Neural network get strong by putting together a lot of dumb elements. Here we have a lot of intelligent people acting dumb together.</p>\n<p>I would really warn beginners against this. It gives you a false impression of what is supposed to matter for your understanding of the field. Picking ideas from public kernels is good. Using the public LB as an optimization objective is risky. It also has not much to do with real life.</p>\n<p>Conclusion: We're humans. Let's share individually good ideas and not collectively act like HP random search :)</p>\n<p><em>Note: At least just doing a mean blend makes more sense. It is not using the LB as feedback.</em></p>",
  "messages": [
    {
      "id": "974664",
      "postDate": "08/18/2020 02:22:48",
      "content": "<p>So all these debates about people using all the same public kernels that are blends of blends with some diverse ensembling technique made me realize something funny:</p>\n<blockquote>\n  <p>Kaggle users are, in this situation, acting like someone doing hyper parameter search directly on the (public) test set.</p>\n</blockquote>\n<p>Basically we have a rather small public leader board dataset where the goal is to correctly order \"only\" 78 malignant cases. Then you have a first wave of users trying the public kernels (changing bits of it) and some publishing their fork when the result is a bit better than the first notebook. But rather than using training data it uses test data (hidden) to make that call. </p>\n<p>Considering the number of users doing experiments on forks you are bound to have someone that finds a higher score solution ! It is very similar to doing Hyperparameter search directly on the test set: Choose params, submit, get score, repeat. Here the search is random.</p>\n<p>Then you have a second wave, that will use multiple kernels as described above and they will also have a set of hyperparameters like in min-max solution or weighted blends. We can easily imagine that with a completed enough pyramid of blends you get a final function that is flexible enough to really manipulate the order of predicted probabilities. When do they publish ? Once again, when the public LB increased. And once again they are numerous. Neural network get strong by putting together a lot of dumb elements. Here we have a lot of intelligent people acting dumb together.</p>\n<p>I would really warn beginners against this. It gives you a false impression of what is supposed to matter for your understanding of the field. Picking ideas from public kernels is good. Using the public LB as an optimization objective is risky. It also has not much to do with real life.</p>\n<p>Conclusion: We're humans. Let's share individually good ideas and not collectively act like HP random search :)</p>\n<p><em>Note: At least just doing a mean blend makes more sense. It is not using the LB as feedback.</em></p>",
      "rawMarkdown": "So all these debates about people using all the same public kernels that are blends of blends with some diverse ensembling technique made me realize something funny:\n\n> Kaggle users are, in this situation, acting like someone doing hyper parameter search directly on the (public) test set.\n\nBasically we have a rather small public leader board dataset where the goal is to correctly order \"only\" 78 malignant cases. Then you have a first wave of users trying the public kernels (changing bits of it) and some publishing their fork when the result is a bit better than the first notebook. But rather than using training data it uses test data (hidden) to make that call. \n\nConsidering the number of users doing experiments on forks you are bound to have someone that finds a higher score solution ! It is very similar to doing Hyperparameter search directly on the test set: Choose params, submit, get score, repeat. Here the search is random.\n\nThen you have a second wave, that will use multiple kernels as described above and they will also have a set of hyperparameters like in min-max solution or weighted blends. We can easily imagine that with a completed enough pyramid of blends you get a final function that is flexible enough to really manipulate the order of predicted probabilities. When do they publish ? Once again, when the public LB increased. And once again they are numerous. Neural network get strong by putting together a lot of dumb elements. Here we have a lot of intelligent people acting dumb together.\n\nI would really warn beginners against this. It gives you a false impression of what is supposed to matter for your understanding of the field. Picking ideas from public kernels is good. Using the public LB as an optimization objective is risky. It also has not much to do with real life.\n\nConclusion: We're humans. Let's share individually good ideas and not collectively act like HP random search :)\n\n*Note: At least just doing a mean blend makes more sense. It is not using the LB as feedback.*",
      "votes": null
    },
    {
      "id": "978678",
      "postDate": "08/20/2020 10:13:45",
      "content": "<p>I missed this.  Yes, LB climbing is a national sport.  And result is almost always disappointment in private LB.</p>\n<p>Why are people doing it?  I was puzzled for a while till I got this: </p>\n<p>Overfitters become thought leaders in every competition.  </p>\n<p>Indeed, people trust what is written in the forum based on the current public LB.  If you overfit to public LB like a mad person and reach the top of LB then you get lots of votes, followers, and attention.  Whatever you do or publish is seen as gold.  You then climb dataset, notebooks, or discussion ranks as a result.  </p>\n<p>Cherry on the cake:  it does not matter if you drop 1000 ranks in private LB, people will remember how they \"learned from you\" during the competition, for years.  I know of a competition where I survived shakeup and even moved to gold, while every other team in top 50 of public LB dropped a lot.  Yet people still remember how they learned from those who dropped.  </p>\n<p>Another reason was given to me in Tweet Sentiment competition forum: if you overfit to get a high score, then it is easier to get people willing to team with you.  I had not thought of that but it is yet another compelling reason to shine on public LB.</p>\n<p>I know some will read this as frustration from my result in this competition.  Well, that's not the case, I'm pretty happy by my result knowing I spent 13 days on it and my relative lack of experience in computer vision.  My only frustration is to have entered way too late.</p>\n<p>I rant above because I' unhappy to see all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation. </p>",
      "rawMarkdown": "I missed this.  Yes, LB climbing is a national sport.  And result is almost always disappointment in private LB.\n\nWhy are people doing it?  I was puzzled for a while till I got this: \n\nOverfitters become thought leaders in every competition.  \n\nIndeed, people trust what is written in the forum based on the current public LB.  If you overfit to public LB like a mad person and reach the top of LB then you get lots of votes, followers, and attention.  Whatever you do or publish is seen as gold.  You then climb dataset, notebooks, or discussion ranks as a result.  \n\nCherry on the cake:  it does not matter if you drop 1000 ranks in private LB, people will remember how they \"learned from you\" during the competition, for years.  I know of a competition where I survived shakeup and even moved to gold, while every other team in top 50 of public LB dropped a lot.  Yet people still remember how they learned from those who dropped.  \n\nAnother reason was given to me in Tweet Sentiment competition forum: if you overfit to get a high score, then it is easier to get people willing to team with you.  I had not thought of that but it is yet another compelling reason to shine on public LB.\n\nI know some will read this as frustration from my result in this competition.  Well, that's not the case, I'm pretty happy by my result knowing I spent 13 days on it and my relative lack of experience in computer vision.  My only frustration is to have entered way too late.\n\nI rant above because I' unhappy to see all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation.",
      "votes": null
    },
    {
      "id": "978708",
      "postDate": "08/20/2020 10:38:15",
      "content": "<p>mean, cruel but damn true ;)</p>",
      "rawMarkdown": "mean, cruel but damn true ;)",
      "votes": null
    },
    {
      "id": "978871",
      "postDate": "08/20/2020 13:13:41",
      "content": "<p>Unfortunately its true.</p>",
      "rawMarkdown": "Unfortunately its true.",
      "votes": null
    },
    {
      "id": "979417",
      "postDate": "08/20/2020 20:28:27",
      "content": "<p>Yes, an astute and correct observation.</p>\n<p>I wonder who the grumpy one is??</p>\n<p>…from a recent post of mine</p>\n<blockquote>\n  <p>I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;) </p>\n</blockquote>",
      "rawMarkdown": "Yes, an astute and correct observation.\n\nI wonder who the grumpy one is??\n\n...from a recent post of mine\n> I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;)",
      "votes": null
    },
    {
      "id": "979467",
      "postDate": "08/20/2020 21:26:14",
      "content": "<blockquote>\n  <p>all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation.</p>\n</blockquote>\n<p>I think everyone who reviews this competition will learn the importance of cross validation. All the top solutions prevailed because of cross validation, specifically optimizing their CV instead of LB. (For those that missed the discussions about cross validation, i posted a discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a> and starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a>).</p>",
      "rawMarkdown": "> all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation.\n\nI think everyone who reviews this competition will learn the importance of cross validation. All the top solutions prevailed because of cross validation, specifically optimizing their CV instead of LB. (For those that missed the discussions about cross validation, i posted a discussion [here][1] and starter notebook [here][2]).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\n[2]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
      "votes": null
    },
    {
      "id": "979468",
      "postDate": "08/20/2020 21:29:41",
      "content": "<p>Chris, I agree. Those who read writeups will learn useful things indeed. I'm not worried about them.  </p>",
      "rawMarkdown": "Chris, I agree. Those who read writeups will learn useful things indeed. I'm not worried about them.",
      "votes": null
    },
    {
      "id": "979499",
      "postDate": "08/20/2020 22:22:04",
      "content": "<blockquote>\n  <p>I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;) </p>\n</blockquote>\n<p>I don't know who the grumpy ones are, but I bet that if you thank them explicitly then they'll be less grumpy ;)</p>",
      "rawMarkdown": "> I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;) \n\nI don't know who the grumpy ones are, but I bet that if you thank them explicitly then they'll be less grumpy ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 978678,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/20/2020 10:13:45",
      "content": "<p>I missed this.  Yes, LB climbing is a national sport.  And result is almost always disappointment in private LB.</p>\n<p>Why are people doing it?  I was puzzled for a while till I got this: </p>\n<p>Overfitters become thought leaders in every competition.  </p>\n<p>Indeed, people trust what is written in the forum based on the current public LB.  If you overfit to public LB like a mad person and reach the top of LB then you get lots of votes, followers, and attention.  Whatever you do or publish is seen as gold.  You then climb dataset, notebooks, or discussion ranks as a result.  </p>\n<p>Cherry on the cake:  it does not matter if you drop 1000 ranks in private LB, people will remember how they \"learned from you\" during the competition, for years.  I know of a competition where I survived shakeup and even moved to gold, while every other team in top 50 of public LB dropped a lot.  Yet people still remember how they learned from those who dropped.  </p>\n<p>Another reason was given to me in Tweet Sentiment competition forum: if you overfit to get a high score, then it is easier to get people willing to team with you.  I had not thought of that but it is yet another compelling reason to shine on public LB.</p>\n<p>I know some will read this as frustration from my result in this competition.  Well, that's not the case, I'm pretty happy by my result knowing I spent 13 days on it and my relative lack of experience in computer vision.  My only frustration is to have entered way too late.</p>\n<p>I rant above because I' unhappy to see all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 978708,
          "author_name": "fiyeroleung",
          "author_url": "",
          "post_date": "08/20/2020 10:38:15",
          "content": "<p>mean, cruel but damn true ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 978871,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "08/20/2020 13:13:41",
          "content": "<p>Unfortunately its true.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979417,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "08/20/2020 20:28:27",
          "content": "<p>Yes, an astute and correct observation.</p>\n<p>I wonder who the grumpy one is??</p>\n<p>…from a recent post of mine</p>\n<blockquote>\n  <p>I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;) </p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979467,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/20/2020 21:26:14",
          "content": "<blockquote>\n  <p>all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation.</p>\n</blockquote>\n<p>I think everyone who reviews this competition will learn the importance of cross validation. All the top solutions prevailed because of cross validation, specifically optimizing their CV instead of LB. (For those that missed the discussions about cross validation, i posted a discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">here</a> and starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a>).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979468,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/20/2020 21:29:41",
          "content": "<p>Chris, I agree. Those who read writeups will learn useful things indeed. I'm not worried about them.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979499,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/20/2020 22:22:04",
          "content": "<blockquote>\n  <p>I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;) </p>\n</blockquote>\n<p>I don't know who the grumpy ones are, but I bet that if you thank them explicitly then they'll be less grumpy ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "974664": "So all these debates about people using all the same public kernels that are blends of blends with some diverse ensembling technique made me realize something funny:\n\n> Kaggle users are, in this situation, acting like someone doing hyper parameter search directly on the (public) test set.\n\nBasically we have a rather small public leader board dataset where the goal is to correctly order \"only\" 78 malignant cases. Then you have a first wave of users trying the public kernels (changing bits of it) and some publishing their fork when the result is a bit better than the first notebook. But rather than using training data it uses test data (hidden) to make that call. \n\nConsidering the number of users doing experiments on forks you are bound to have someone that finds a higher score solution ! It is very similar to doing Hyperparameter search directly on the test set: Choose params, submit, get score, repeat. Here the search is random.\n\nThen you have a second wave, that will use multiple kernels as described above and they will also have a set of hyperparameters like in min-max solution or weighted blends. We can easily imagine that with a completed enough pyramid of blends you get a final function that is flexible enough to really manipulate the order of predicted probabilities. When do they publish ? Once again, when the public LB increased. And once again they are numerous. Neural network get strong by putting together a lot of dumb elements. Here we have a lot of intelligent people acting dumb together.\n\nI would really warn beginners against this. It gives you a false impression of what is supposed to matter for your understanding of the field. Picking ideas from public kernels is good. Using the public LB as an optimization objective is risky. It also has not much to do with real life.\n\nConclusion: We're humans. Let's share individually good ideas and not collectively act like HP random search :)\n\n*Note: At least just doing a mean blend makes more sense. It is not using the LB as feedback.*",
    "978678": "I missed this.  Yes, LB climbing is a national sport.  And result is almost always disappointment in private LB.\n\nWhy are people doing it?  I was puzzled for a while till I got this: \n\nOverfitters become thought leaders in every competition.  \n\nIndeed, people trust what is written in the forum based on the current public LB.  If you overfit to public LB like a mad person and reach the top of LB then you get lots of votes, followers, and attention.  Whatever you do or publish is seen as gold.  You then climb dataset, notebooks, or discussion ranks as a result.  \n\nCherry on the cake:  it does not matter if you drop 1000 ranks in private LB, people will remember how they \"learned from you\" during the competition, for years.  I know of a competition where I survived shakeup and even moved to gold, while every other team in top 50 of public LB dropped a lot.  Yet people still remember how they learned from those who dropped.  \n\nAnother reason was given to me in Tweet Sentiment competition forum: if you overfit to get a high score, then it is easier to get people willing to team with you.  I had not thought of that but it is yet another compelling reason to shine on public LB.\n\nI know some will read this as frustration from my result in this competition.  Well, that's not the case, I'm pretty happy by my result knowing I spent 13 days on it and my relative lack of experience in computer vision.  My only frustration is to have entered way too late.\n\nI rant above because I' unhappy to see all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation.",
    "978708": "mean, cruel but damn true ;)",
    "978871": "Unfortunately its true.",
    "979417": "Yes, an astute and correct observation.\n\nI wonder who the grumpy one is??\n\n...from a recent post of mine\n> I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;)",
    "979467": "> all these people missing out a great opportunity to learn how to asses model performance using training data, for instance with cross validation.\n\nI think everyone who reviews this competition will learn the importance of cross validation. All the top solutions prevailed because of cross validation, specifically optimizing their CV instead of LB. (For those that missed the discussions about cross validation, i posted a discussion [here][1] and starter notebook [here][2]).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\n[2]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
    "979468": "Chris, I agree. Those who read writeups will learn useful things indeed. I'm not worried about them.",
    "979499": "> I'd still like to thank Chris, AgentAuers and the many others who have given freely and openly to help beginners and provide insight, even the \"grumpy\" ones ;) \n\nI don't know who the grumpy ones are, but I bet that if you thank them explicitly then they'll be less grumpy ;)"
  },
  "source": "meta"
}