{
  "id": 205565,
  "title": "Local Cv Or Public Lb whom to trust ? Am I doing something wrong here ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205565",
  "author_name": "",
  "post_date": "2020-12-20T16:41:22.742768800Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was using CV Strategy described in <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this </a> notebook , With 2.9 Crores data points in Train Data and 7.5 lakhs in validation set . Earlier I used some simple feature engineering described in <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">this </a> notebook and some more tags and bundle based features .  The Cv and Lb correlation was good for this as when Cv was 0.76X LB was 0.75X or something .</p>\n<p>I added some time series based features , I used Loop Feature Engineering Approach and Local Cv improved to 0.8X , I got over joyed , but was struggling to solve submission scoring error , after lot of struggle I managed to made an inference kernel which can score . <strong>Note I didn't Fitted my Model on Train + Val Set ,  I only used Model fitted on Train Set</strong>. After My results on Public LB was out it came out to be 0.63X  , that is score on Public LB dropped drastically. </p>\n<p>I wanted to ask am I doing something wrong here or my model is overfitting on validation set , what am I doing wrong here . Can anyone please help me out ?<br>\n<a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>  <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  , If you guys can provide some suggestion over this it would be great.</p>\n<p>Thanks and regards,<br>\nAthar.</p>",
  "messages": [
    {
      "id": "1120194",
      "postDate": "12/20/2020 16:41:22",
      "content": "<p>I was using CV Strategy described in <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this </a> notebook , With 2.9 Crores data points in Train Data and 7.5 lakhs in validation set . Earlier I used some simple feature engineering described in <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">this </a> notebook and some more tags and bundle based features .  The Cv and Lb correlation was good for this as when Cv was 0.76X LB was 0.75X or something .</p>\n<p>I added some time series based features , I used Loop Feature Engineering Approach and Local Cv improved to 0.8X , I got over joyed , but was struggling to solve submission scoring error , after lot of struggle I managed to made an inference kernel which can score . <strong>Note I didn't Fitted my Model on Train + Val Set ,  I only used Model fitted on Train Set</strong>. After My results on Public LB was out it came out to be 0.63X  , that is score on Public LB dropped drastically. </p>\n<p>I wanted to ask am I doing something wrong here or my model is overfitting on validation set , what am I doing wrong here . Can anyone please help me out ?<br>\n<a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>  <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  , If you guys can provide some suggestion over this it would be great.</p>\n<p>Thanks and regards,<br>\nAthar.</p>",
      "rawMarkdown": "I was using CV Strategy described in [this ](https://www.kaggle.com/its7171/cv-strategy) notebook , With 2.9 Crores data points in Train Data and 7.5 lakhs in validation set . Earlier I used some simple feature engineering described in [this ](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering) notebook and some more tags and bundle based features .  The Cv and Lb correlation was good for this as when Cv was 0.76X LB was 0.75X or something .\n\nI added some time series based features , I used Loop Feature Engineering Approach and Local Cv improved to 0.8X , I got over joyed , but was struggling to solve submission scoring error , after lot of struggle I managed to made an inference kernel which can score . **Note I didn't Fitted my Model on Train + Val Set ,  I only used Model fitted on Train Set**. After My results on Public LB was out it came out to be 0.63X  , that is score on Public LB dropped drastically. \n\nI wanted to ask am I doing something wrong here or my model is overfitting on validation set , what am I doing wrong here . Can anyone please help me out ?\n@mamasinkgs  @its7171 @nyanpn  @cdeotte  , If you guys can provide some suggestion over this it would be great.\n\nThanks and regards,\nAthar.",
      "votes": null
    },
    {
      "id": "1120198",
      "postDate": "12/20/2020 16:47:31",
      "content": "<p>Check your code, my best gues is that you have leak in some of your features</p>",
      "rawMarkdown": "Check your code, my best gues is that you have leak in some of your features",
      "votes": null
    },
    {
      "id": "1120213",
      "postDate": "12/20/2020 17:09:24",
      "content": "<p>Ok thanks Ragnar will check that out . </p>",
      "rawMarkdown": "Ok thanks Ragnar will check that out .",
      "votes": null
    },
    {
      "id": "1120844",
      "postDate": "12/21/2020 06:42:38",
      "content": "<p>Great Post…..</p>",
      "rawMarkdown": "Great Post…..",
      "votes": null
    },
    {
      "id": "1120998",
      "postDate": "12/21/2020 09:26:59",
      "content": "<p>You need to check your features. It's leak. Our team use the same CV strategy, the difference is less than 0.002. </p>",
      "rawMarkdown": "You need to check your features. It's leak. Our team use the same CV strategy, the difference is less than 0.002.",
      "votes": null
    },
    {
      "id": "1121818",
      "postDate": "12/21/2020 23:52:00",
      "content": "<p>trust  my cv</p>",
      "rawMarkdown": "trust  my cv",
      "votes": null
    },
    {
      "id": "1121876",
      "postDate": "12/22/2020 01:43:27",
      "content": "<p>If LB and CV are close to each other, trust your cv. But such a huge difference indicates there is something wrong in your code.(e.g. feature leakage as mentioned above) Try to plot the feature importance of your model to find is there some features have extremely high importances( for LGB, use <code>lgb.plot_importance(model, importance_type='gain')</code>).</p>",
      "rawMarkdown": "If LB and CV are close to each other, trust your cv. But such a huge difference indicates there is something wrong in your code.(e.g. feature leakage as mentioned above) Try to plot the feature importance of your model to find is there some features have extremely high importances( for LGB, use `lgb.plot_importance(model, importance_type='gain')`).",
      "votes": null
    },
    {
      "id": "1121879",
      "postDate": "12/22/2020 01:48:00",
      "content": "<p>I guess there probably should be a bug in your code for loop feature engineering.</p>",
      "rawMarkdown": "I guess there probably should be a bug in your code for loop feature engineering.",
      "votes": null
    },
    {
      "id": "1121947",
      "postDate": "12/22/2020 03:39:44",
      "content": "<p>Thanks will check that out 🙂 .</p>",
      "rawMarkdown": "Thanks will check that out 🙂 .",
      "votes": null
    },
    {
      "id": "1121949",
      "postDate": "12/22/2020 03:41:08",
      "content": "<p>Ok thanks and The features having extremely high importance will most probably have leak  ?</p>",
      "rawMarkdown": "Ok thanks and The features having extremely high importance will most probably have leak  ?",
      "votes": null
    },
    {
      "id": "1122446",
      "postDate": "12/22/2020 12:52:10",
      "content": "<p>[Update] : It turns out there was a small bug in my Feature engineering Code ☹️☹️ ! The order of some statements were incorrect and because of that it was causing leakage (Most Probably) , My Cv score is looking more realistic now after finding out the correct order of statements and fixing it Thanks <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/mamas\" target=\"_blank\">@mamas</a> !</p>",
      "rawMarkdown": "[Update] : It turns out there was a small bug in my Feature engineering Code ☹️☹️ ! The order of some statements were incorrect and because of that it was causing leakage (Most Probably) , My Cv score is looking more realistic now after finding out the correct order of statements and fixing it Thanks @ragnar123 @mamas !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1120198,
      "author_name": "ragnar123",
      "author_url": "",
      "post_date": "12/20/2020 16:47:31",
      "content": "<p>Check your code, my best gues is that you have leak in some of your features</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120213,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "12/20/2020 17:09:24",
          "content": "<p>Ok thanks Ragnar will check that out . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120844,
      "author_name": "datawarriors",
      "author_url": "",
      "post_date": "12/21/2020 06:42:38",
      "content": "<p>Great Post…..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1120998,
      "author_name": "m10515009",
      "author_url": "",
      "post_date": "12/21/2020 09:26:59",
      "content": "<p>You need to check your features. It's leak. Our team use the same CV strategy, the difference is less than 0.002. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1121818,
      "author_name": "mynewlife",
      "author_url": "",
      "post_date": "12/21/2020 23:52:00",
      "content": "<p>trust  my cv</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1121876,
      "author_name": "louieshao",
      "author_url": "",
      "post_date": "12/22/2020 01:43:27",
      "content": "<p>If LB and CV are close to each other, trust your cv. But such a huge difference indicates there is something wrong in your code.(e.g. feature leakage as mentioned above) Try to plot the feature importance of your model to find is there some features have extremely high importances( for LGB, use <code>lgb.plot_importance(model, importance_type='gain')</code>).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121949,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "12/22/2020 03:41:08",
          "content": "<p>Ok thanks and The features having extremely high importance will most probably have leak  ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121879,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/22/2020 01:48:00",
      "content": "<p>I guess there probably should be a bug in your code for loop feature engineering.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121947,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "12/22/2020 03:39:44",
          "content": "<p>Thanks will check that out 🙂 .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122446,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "12/22/2020 12:52:10",
      "content": "<p>[Update] : It turns out there was a small bug in my Feature engineering Code ☹️☹️ ! The order of some statements were incorrect and because of that it was causing leakage (Most Probably) , My Cv score is looking more realistic now after finding out the correct order of statements and fixing it Thanks <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/mamas\" target=\"_blank\">@mamas</a> !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1120194": "I was using CV Strategy described in [this ](https://www.kaggle.com/its7171/cv-strategy) notebook , With 2.9 Crores data points in Train Data and 7.5 lakhs in validation set . Earlier I used some simple feature engineering described in [this ](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering) notebook and some more tags and bundle based features .  The Cv and Lb correlation was good for this as when Cv was 0.76X LB was 0.75X or something .\n\nI added some time series based features , I used Loop Feature Engineering Approach and Local Cv improved to 0.8X , I got over joyed , but was struggling to solve submission scoring error , after lot of struggle I managed to made an inference kernel which can score . **Note I didn't Fitted my Model on Train + Val Set ,  I only used Model fitted on Train Set**. After My results on Public LB was out it came out to be 0.63X  , that is score on Public LB dropped drastically. \n\nI wanted to ask am I doing something wrong here or my model is overfitting on validation set , what am I doing wrong here . Can anyone please help me out ?\n@mamasinkgs  @its7171 @nyanpn  @cdeotte  , If you guys can provide some suggestion over this it would be great.\n\nThanks and regards,\nAthar.",
    "1120198": "Check your code, my best gues is that you have leak in some of your features",
    "1120213": "Ok thanks Ragnar will check that out .",
    "1120844": "Great Post…..",
    "1120998": "You need to check your features. It's leak. Our team use the same CV strategy, the difference is less than 0.002.",
    "1121818": "trust  my cv",
    "1121876": "If LB and CV are close to each other, trust your cv. But such a huge difference indicates there is something wrong in your code.(e.g. feature leakage as mentioned above) Try to plot the feature importance of your model to find is there some features have extremely high importances( for LGB, use `lgb.plot_importance(model, importance_type='gain')`).",
    "1121879": "I guess there probably should be a bug in your code for loop feature engineering.",
    "1121947": "Thanks will check that out 🙂 .",
    "1121949": "Ok thanks and The features having extremely high importance will most probably have leak  ?",
    "1122446": "[Update] : It turns out there was a small bug in my Feature engineering Code ☹️☹️ ! The order of some statements were incorrect and because of that it was causing leakage (Most Probably) , My Cv score is looking more realistic now after finding out the correct order of statements and fixing it Thanks @ragnar123 @mamas !"
  },
  "source": "meta"
}