{
  "id": 539798,
  "title": "Overfitting model",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/539798",
  "author_name": "",
  "post_date": "2024-10-10T19:56:38.578232100Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>I have built a ensemble model that is acheiving cv scores of 0.48 in training, however when I use the model on the leaderboard - I am finding that is overfitting 0.45. I have added various features into my model to help reduce the overfitting including cv, regularization, feature selection and removed outliers.  Does anyone know any other techniques I can use to reduce the overfitting?</p>",
  "messages": [
    {
      "id": "3014042",
      "postDate": "10/10/2024 19:56:38",
      "content": "<p>Hi all, </p>\n<p>I have built a ensemble model that is acheiving cv scores of 0.48 in training, however when I use the model on the leaderboard - I am finding that is overfitting 0.45. I have added various features into my model to help reduce the overfitting including cv, regularization, feature selection and removed outliers.  Does anyone know any other techniques I can use to reduce the overfitting?</p>",
      "rawMarkdown": "Hi all, \n\nI have built a ensemble model that is acheiving cv scores of 0.48 in training, however when I use the model on the leaderboard - I am finding that is overfitting 0.45. I have added various features into my model to help reduce the overfitting including cv, regularization, feature selection and removed outliers.  Does anyone know any other techniques I can use to reduce the overfitting?",
      "votes": null
    },
    {
      "id": "3014063",
      "postDate": "10/10/2024 20:41:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/peterhopkinson\" target=\"_blank\">@peterhopkinson</a>,</p>\n<p>Your techniques to reduce overfitting are the good ones (cross-validation, regularization, feature selection …)</p>\n<p>What makes you think you are overfitting ? </p>\n<p>You will start to overfit when you will make decisions based on your score on the public leaderboard ; as long as you don't do it, and as long as your CV score is .48 or higher even when changing your random seed, you will not overfit.  </p>",
      "rawMarkdown": "Hi @peterhopkinson,\n\nYour techniques to reduce overfitting are the good ones (cross-validation, regularization, feature selection ...)\n\nWhat makes you think you are overfitting ? \n\nYou will start to overfit when you will make decisions based on your score on the public leaderboard ; as long as you don't do it, and as long as your CV score is .48 or higher even when changing your random seed, you will not overfit.",
      "votes": null
    },
    {
      "id": "3014082",
      "postDate": "10/10/2024 21:53:57",
      "content": "<p>Just the difference between the cross validation score and the leaderboard score suggests the model is overfitting - I am evaluating the hyperparameter ranges to adjust them to reduce the chance of overfitting. </p>",
      "rawMarkdown": "Just the difference between the cross validation score and the leaderboard score suggests the model is overfitting - I am evaluating the hyperparameter ranges to adjust them to reduce the chance of overfitting.",
      "votes": null
    },
    {
      "id": "3014242",
      "postDate": "10/11/2024 05:02:58",
      "content": "<p>So you are not overfitting the public LB.</p>\n<p>Another thing to verify, is the difference between your score during traing and your validation score : as long as difference is less than .1, you are not overfitting training samples too much (my opinion, in this comp). Best public notebooks could have a .3 difference : their training score generalize poprly.</p>\n<p>We can compare those two scores because it should not be any difference between train and validation samples (it depends on your CV strategy and regularization)</p>\n<p>But we must not compare CV score and score on public LB, because we don't known anything about samples on  public LB.</p>\n<p>There is an interesting discussion about scores differences to bookmark <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "So you are not overfitting the public LB.\n\nAnother thing to verify, is the difference between your score during traing and your validation score : as long as difference is less than .1, you are not overfitting training samples too much (my opinion, in this comp). Best public notebooks could have a .3 difference : their training score generalize poprly.\n\nWe can compare those two scores because it should not be any difference between train and validation samples (it depends on your CV strategy and regularization)\n\nBut we must not compare CV score and score on public LB, because we don't known anything about samples on  public LB.\n\nThere is an interesting discussion about scores differences to bookmark [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523)",
      "votes": null
    },
    {
      "id": "3014401",
      "postDate": "10/11/2024 07:57:33",
      "content": "<p>hey <a href=\"https://www.kaggle.com/peterhopkinson\" target=\"_blank\">@peterhopkinson</a> , it's not overfitting . It's completely natural to have a small score difference between cv and Public test data.</p>",
      "rawMarkdown": "hey @peterhopkinson , it's not overfitting . It's completely natural to have a small score difference between cv and Public test data.",
      "votes": null
    },
    {
      "id": "3016013",
      "postDate": "10/13/2024 08:57:07",
      "content": "<p>You are not overfitting. The data available to us is not enough to overfit unless you do something too far.</p>",
      "rawMarkdown": "You are not overfitting. The data available to us is not enough to overfit unless you do something too far.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3014063,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "10/10/2024 20:41:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/peterhopkinson\" target=\"_blank\">@peterhopkinson</a>,</p>\n<p>Your techniques to reduce overfitting are the good ones (cross-validation, regularization, feature selection …)</p>\n<p>What makes you think you are overfitting ? </p>\n<p>You will start to overfit when you will make decisions based on your score on the public leaderboard ; as long as you don't do it, and as long as your CV score is .48 or higher even when changing your random seed, you will not overfit.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 3014082,
          "author_name": "peterhopkinson",
          "author_url": "",
          "post_date": "10/10/2024 21:53:57",
          "content": "<p>Just the difference between the cross validation score and the leaderboard score suggests the model is overfitting - I am evaluating the hyperparameter ranges to adjust them to reduce the chance of overfitting. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3014242,
              "author_name": "adaubas",
              "author_url": "",
              "post_date": "10/11/2024 05:02:58",
              "content": "<p>So you are not overfitting the public LB.</p>\n<p>Another thing to verify, is the difference between your score during traing and your validation score : as long as difference is less than .1, you are not overfitting training samples too much (my opinion, in this comp). Best public notebooks could have a .3 difference : their training score generalize poprly.</p>\n<p>We can compare those two scores because it should not be any difference between train and validation samples (it depends on your CV strategy and regularization)</p>\n<p>But we must not compare CV score and score on public LB, because we don't known anything about samples on  public LB.</p>\n<p>There is an interesting discussion about scores differences to bookmark <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523\" target=\"_blank\">here</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3014401,
      "author_name": "mohammedahmedxx12",
      "author_url": "",
      "post_date": "10/11/2024 07:57:33",
      "content": "<p>hey <a href=\"https://www.kaggle.com/peterhopkinson\" target=\"_blank\">@peterhopkinson</a> , it's not overfitting . It's completely natural to have a small score difference between cv and Public test data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3016013,
      "author_name": "chanpreetsingh07",
      "author_url": "",
      "post_date": "10/13/2024 08:57:07",
      "content": "<p>You are not overfitting. The data available to us is not enough to overfit unless you do something too far.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3014042": "Hi all, \n\nI have built a ensemble model that is acheiving cv scores of 0.48 in training, however when I use the model on the leaderboard - I am finding that is overfitting 0.45. I have added various features into my model to help reduce the overfitting including cv, regularization, feature selection and removed outliers.  Does anyone know any other techniques I can use to reduce the overfitting?",
    "3014063": "Hi @peterhopkinson,\n\nYour techniques to reduce overfitting are the good ones (cross-validation, regularization, feature selection ...)\n\nWhat makes you think you are overfitting ? \n\nYou will start to overfit when you will make decisions based on your score on the public leaderboard ; as long as you don't do it, and as long as your CV score is .48 or higher even when changing your random seed, you will not overfit.",
    "3014082": "Just the difference between the cross validation score and the leaderboard score suggests the model is overfitting - I am evaluating the hyperparameter ranges to adjust them to reduce the chance of overfitting.",
    "3014242": "So you are not overfitting the public LB.\n\nAnother thing to verify, is the difference between your score during traing and your validation score : as long as difference is less than .1, you are not overfitting training samples too much (my opinion, in this comp). Best public notebooks could have a .3 difference : their training score generalize poprly.\n\nWe can compare those two scores because it should not be any difference between train and validation samples (it depends on your CV strategy and regularization)\n\nBut we must not compare CV score and score on public LB, because we don't known anything about samples on  public LB.\n\nThere is an interesting discussion about scores differences to bookmark [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523)",
    "3014401": "hey @peterhopkinson , it's not overfitting . It's completely natural to have a small score difference between cv and Public test data.",
    "3016013": "You are not overfitting. The data available to us is not enough to overfit unless you do something too far."
  },
  "source": "meta"
}