{
  "id": 549611,
  "title": "Why constrain the overfitting gives worse results?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/549611",
  "author_name": "",
  "post_date": "2024-12-03T03:28:14.615973200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In the popular notebook <a href=\"https://www.kaggle.com/code/cchangyyy/0-494-notebook\" target=\"_blank\">https://www.kaggle.com/code/cchangyyy/0-494-notebook</a><br>\nWhich ensembles 3 ensembled model results, the models are clearly overfit, especially the last model, where the train QWK is 0.9+. So I tried</p>\n<ol>\n<li><p>decrease the depth and add regulations, although the result in the train val split improves, it decrease on the public LB.</p></li>\n<li><p>Add more features, for the first model with extra features it actually improves on both offline and public LB(with single sub, from 0.344 -&gt; 0.368), but when ensembled together, the results worsened.</p></li>\n</ol>\n<p>Overall, I got really confused on the offline evaluation vs public LB. Any one have insights?</p>",
  "messages": [
    {
      "id": "3061866",
      "postDate": "12/03/2024 03:28:14",
      "content": "<p>In the popular notebook <a href=\"https://www.kaggle.com/code/cchangyyy/0-494-notebook\" target=\"_blank\">https://www.kaggle.com/code/cchangyyy/0-494-notebook</a><br>\nWhich ensembles 3 ensembled model results, the models are clearly overfit, especially the last model, where the train QWK is 0.9+. So I tried</p>\n<ol>\n<li><p>decrease the depth and add regulations, although the result in the train val split improves, it decrease on the public LB.</p></li>\n<li><p>Add more features, for the first model with extra features it actually improves on both offline and public LB(with single sub, from 0.344 -&gt; 0.368), but when ensembled together, the results worsened.</p></li>\n</ol>\n<p>Overall, I got really confused on the offline evaluation vs public LB. Any one have insights?</p>",
      "rawMarkdown": "In the popular notebook https://www.kaggle.com/code/cchangyyy/0-494-notebook\nWhich ensembles 3 ensembled model results, the models are clearly overfit, especially the last model, where the train QWK is 0.9+. So I tried\n\n1. decrease the depth and add regulations, although the result in the train val split improves, it decrease on the public LB.\n\n2. Add more features, for the first model with extra features it actually improves on both offline and public LB(with single sub, from 0.344 -> 0.368), but when ensembled together, the results worsened.\n\nOverall, I got really confused on the offline evaluation vs public LB. Any one have insights?",
      "votes": null
    },
    {
      "id": "3062926",
      "postDate": "12/04/2024 03:27:22",
      "content": "<p>This is an issue that runs throughout the entire competition, and I had already brought it up for discussion quite some time ago. <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535458\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535458</a> <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536441，I\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536441，I</a> believe the fundamental reason is the poor data quality. There are too many missing values and unreliable data points. If you don't perform enough or sufficiently good data cleaning, cross-validation results become unreliable, and it's even harder for us to grasp the data distribution on the leaderboard. As a result, this competition has essentially turned into a lottery game for me. However, I have learned a lot of data processing techniques along the way!</p>",
      "rawMarkdown": "This is an issue that runs throughout the entire competition, and I had already brought it up for discussion quite some time ago. https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535458 https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536441，I believe the fundamental reason is the poor data quality. There are too many missing values and unreliable data points. If you don't perform enough or sufficiently good data cleaning, cross-validation results become unreliable, and it's even harder for us to grasp the data distribution on the leaderboard. As a result, this competition has essentially turned into a lottery game for me. However, I have learned a lot of data processing techniques along the way!",
      "votes": null
    },
    {
      "id": "3063888",
      "postDate": "12/05/2024 00:48:44",
      "content": "<p>Well, here is my intuition in the ensembles method we are applying here:<br>\nWhen we have multiple models trying to learn something with different feature spaces, the different ensembled versions learn (hopefully) different perspectives, improving the diversity. However, this appears to be highly sensitive to the perturbations, as even changing only the seed will make different groups of these ensembles possibly look from similar perspectives, skewing the prediction towards one group rather than combining their collective knowledge.</p>\n<p>So based on luck, it sometimes gives diversity and improves, sometimes it makes different ensemble groups to be aligned, causing the prediction power to be as good as a single group, actually degrading the performance. This is my explanation at the very least and this is the reason why I stopped trying to tune these ensembles models after some point.</p>",
      "rawMarkdown": "Well, here is my intuition in the ensembles method we are applying here:\nWhen we have multiple models trying to learn something with different feature spaces, the different ensembled versions learn (hopefully) different perspectives, improving the diversity. However, this appears to be highly sensitive to the perturbations, as even changing only the seed will make different groups of these ensembles possibly look from similar perspectives, skewing the prediction towards one group rather than combining their collective knowledge.\n\nSo based on luck, it sometimes gives diversity and improves, sometimes it makes different ensemble groups to be aligned, causing the prediction power to be as good as a single group, actually degrading the performance. This is my explanation at the very least and this is the reason why I stopped trying to tune these ensembles models after some point.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3062926,
      "author_name": "wayne127",
      "author_url": "",
      "post_date": "12/04/2024 03:27:22",
      "content": "<p>This is an issue that runs throughout the entire competition, and I had already brought it up for discussion quite some time ago. <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535458\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535458</a> <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536441，I\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536441，I</a> believe the fundamental reason is the poor data quality. There are too many missing values and unreliable data points. If you don't perform enough or sufficiently good data cleaning, cross-validation results become unreliable, and it's even harder for us to grasp the data distribution on the leaderboard. As a result, this competition has essentially turned into a lottery game for me. However, I have learned a lot of data processing techniques along the way!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3063888,
      "author_name": "alperenduru",
      "author_url": "",
      "post_date": "12/05/2024 00:48:44",
      "content": "<p>Well, here is my intuition in the ensembles method we are applying here:<br>\nWhen we have multiple models trying to learn something with different feature spaces, the different ensembled versions learn (hopefully) different perspectives, improving the diversity. However, this appears to be highly sensitive to the perturbations, as even changing only the seed will make different groups of these ensembles possibly look from similar perspectives, skewing the prediction towards one group rather than combining their collective knowledge.</p>\n<p>So based on luck, it sometimes gives diversity and improves, sometimes it makes different ensemble groups to be aligned, causing the prediction power to be as good as a single group, actually degrading the performance. This is my explanation at the very least and this is the reason why I stopped trying to tune these ensembles models after some point.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3061866": "In the popular notebook https://www.kaggle.com/code/cchangyyy/0-494-notebook\nWhich ensembles 3 ensembled model results, the models are clearly overfit, especially the last model, where the train QWK is 0.9+. So I tried\n\n1. decrease the depth and add regulations, although the result in the train val split improves, it decrease on the public LB.\n\n2. Add more features, for the first model with extra features it actually improves on both offline and public LB(with single sub, from 0.344 -> 0.368), but when ensembled together, the results worsened.\n\nOverall, I got really confused on the offline evaluation vs public LB. Any one have insights?",
    "3062926": "This is an issue that runs throughout the entire competition, and I had already brought it up for discussion quite some time ago. https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535458 https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536441，I believe the fundamental reason is the poor data quality. There are too many missing values and unreliable data points. If you don't perform enough or sufficiently good data cleaning, cross-validation results become unreliable, and it's even harder for us to grasp the data distribution on the leaderboard. As a result, this competition has essentially turned into a lottery game for me. However, I have learned a lot of data processing techniques along the way!",
    "3063888": "Well, here is my intuition in the ensembles method we are applying here:\nWhen we have multiple models trying to learn something with different feature spaces, the different ensembled versions learn (hopefully) different perspectives, improving the diversity. However, this appears to be highly sensitive to the perturbations, as even changing only the seed will make different groups of these ensembles possibly look from similar perspectives, skewing the prediction towards one group rather than combining their collective knowledge.\n\nSo based on luck, it sometimes gives diversity and improves, sometimes it makes different ensemble groups to be aligned, causing the prediction power to be as good as a single group, actually degrading the performance. This is my explanation at the very least and this is the reason why I stopped trying to tune these ensembles models after some point."
  },
  "source": "meta"
}