{
  "id": 475244,
  "title": "CV vs LB - large difference?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475244",
  "author_name": "",
  "post_date": "2024-02-07T16:01:47.312667100Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>my last submission: CV=0.596, LB=0.507.</p>\n<p>CV is calculated by training on weeks 0-46 and validating on weeks 46-91 (1 fold, so not actually CV).</p>\n<p>Submission trained on all data - weeks 0-91.</p>\n<p>The difference is around 0.09 - seems very large to me. Does anybody else see this pattern? <br>\nThis could mean that test data has a different distribution from train data. Possibly caused by new values of categorical variables in test that are not present in train. That would present a very challenging problem.</p>",
  "messages": [
    {
      "id": "2641681",
      "postDate": "02/07/2024 16:01:47",
      "content": "<p>my last submission: CV=0.596, LB=0.507.</p>\n<p>CV is calculated by training on weeks 0-46 and validating on weeks 46-91 (1 fold, so not actually CV).</p>\n<p>Submission trained on all data - weeks 0-91.</p>\n<p>The difference is around 0.09 - seems very large to me. Does anybody else see this pattern? <br>\nThis could mean that test data has a different distribution from train data. Possibly caused by new values of categorical variables in test that are not present in train. That would present a very challenging problem.</p>",
      "rawMarkdown": "my last submission: CV=0.596, LB=0.507.\n\nCV is calculated by training on weeks 0-46 and validating on weeks 46-91 (1 fold, so not actually CV).\n\nSubmission trained on all data - weeks 0-91.\n\nThe difference is around 0.09 - seems very large to me. Does anybody else see this pattern? \nThis could mean that test data has a different distribution from train data. Possibly caused by new values of categorical variables in test that are not present in train. That would present a very challenging problem.",
      "votes": null
    },
    {
      "id": "2641705",
      "postDate": "02/07/2024 16:16:04",
      "content": "<p>It could be attributed to the test set covers relatively large period.<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764</a></p>",
      "rawMarkdown": "It could be attributed to the test set covers relatively large period.\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764",
      "votes": null
    },
    {
      "id": "2641833",
      "postDate": "02/07/2024 17:42:19",
      "content": "<p>make sure ure actually using the correct metric, ure training on pre covid and validating with covid, so a score like that for that split would be pretty extraordinary i'd imagine</p>",
      "rawMarkdown": "make sure ure actually using the correct metric, ure training on pre covid and validating with covid, so a score like that for that split would be pretty extraordinary i'd imagine",
      "votes": null
    },
    {
      "id": "2641913",
      "postDate": "02/07/2024 18:43:00",
      "content": "<table>\n<thead>\n<tr>\n<th>Description</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>depth 0</td>\n<td>0.578</td>\n<td>0.507</td>\n</tr>\n<tr>\n<td>depth 0 &amp; 1</td>\n<td>0.668</td>\n<td>0.516</td>\n</tr>\n<tr>\n<td>depth 0 &amp; 1 &amp; 2</td>\n<td>0.675</td>\n<td>0.506</td>\n</tr>\n<tr>\n<td>depth 0 &amp; 1 &amp; 2 without categorical features</td>\n<td>0.643</td>\n<td>0.5</td>\n</tr>\n</tbody>\n</table>\n<p>StratifiedGroupKFold(5) is used here. So yeah, naively adding more features with <code>num_group1==0</code> and <code>num_group2==0</code> improves CV, but no significant impact on LB. </p>",
      "rawMarkdown": "| Description | CV | LB |\n| --- | --- | --- |\n| depth 0 | 0.578 | 0.507 |\n| depth 0 & 1 | 0.668 | 0.516 |\n| depth 0 & 1 & 2 | 0.675 | 0.506 |\n| depth 0 & 1 & 2 without categorical features |  0.643 | 0.5 |\n\n\nStratifiedGroupKFold(5) is used here. So yeah, naively adding more features with `num_group1==0` and `num_group2==0` improves CV, but no significant impact on LB.",
      "votes": null
    },
    {
      "id": "2641924",
      "postDate": "02/07/2024 18:53:44",
      "content": "<p>your cv has leakage ure using information from the future which could be causing this</p>",
      "rawMarkdown": "your cv has leakage ure using information from the future which could be causing this",
      "votes": null
    },
    {
      "id": "2641940",
      "postDate": "02/07/2024 19:07:34",
      "content": "<p>Well, I used StratifiedGroupKFold and not just StratifiedKFold, to eliminate the leakage by not mixing up data from different weeks, but I agree this does not solve the problem completely. Train / valid split by time and training on whole dataset also sounds better, but I have not tried it yet</p>",
      "rawMarkdown": "Well, I used StratifiedGroupKFold and not just StratifiedKFold, to eliminate the leakage by not mixing up data from different weeks, but I agree this does not solve the problem completely. Train / valid split by time and training on whole dataset also sounds better, but I have not tried it yet",
      "votes": null
    },
    {
      "id": "2642619",
      "postDate": "02/08/2024 09:40:49",
      "content": "<blockquote>\n  <p>This could mean that test data has a different distribution from train data.</p>\n</blockquote>\n<p>You are on the right track. This is obviously caused by covid. Remember the important part of the competition is also stability. </p>",
      "rawMarkdown": ">This could mean that test data has a different distribution from train data.\n\nYou are on the right track. This is obviously caused by covid. Remember the important part of the competition is also stability.",
      "votes": null
    },
    {
      "id": "2642751",
      "postDate": "02/08/2024 11:38:25",
      "content": "<p>train on 0-46, val on 47-73, test on 74-91, anyone same as me?</p>",
      "rawMarkdown": "train on 0-46, val on 47-73, test on 74-91, anyone same as me?",
      "votes": null
    },
    {
      "id": "2642755",
      "postDate": "02/08/2024 11:43:02",
      "content": "<table>\n<thead>\n<tr>\n<th>data</th>\n<th>cv(test, competition metric)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>depth=0 only filltered</td>\n<td>0.6749</td>\n</tr>\n</tbody>\n</table>\n<p>early stop on val</p>",
      "rawMarkdown": "| data | cv(test, competition metric) |\n| --- | --- |\n| depth=0 only filltered | 0.6749 |\n\nearly stop on val",
      "votes": null
    },
    {
      "id": "2644146",
      "postDate": "02/09/2024 09:54:56",
      "content": "<p>Although the LB (Leaderboard) score is low, the correlation between CV (Cross Validation) and LB is strong. I do not believe that COVID will significantly impact the model's AUC (Area Under the Curve) performance. Based on the evidence, <strong>training the model exclusively on data from pre-COVID times does not seem to deteriorate its performance on test sets from during the COVID era.</strong></p>",
      "rawMarkdown": "Although the LB (Leaderboard) score is low, the correlation between CV (Cross Validation) and LB is strong. I do not believe that COVID will significantly impact the model's AUC (Area Under the Curve) performance. Based on the evidence, **training the model exclusively on data from pre-COVID times does not seem to deteriorate its performance on test sets from during the COVID era.**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2641705,
      "author_name": "dongyk",
      "author_url": "",
      "post_date": "02/07/2024 16:16:04",
      "content": "<p>It could be attributed to the test set covers relatively large period.<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2641833,
      "author_name": "at7459",
      "author_url": "",
      "post_date": "02/07/2024 17:42:19",
      "content": "<p>make sure ure actually using the correct metric, ure training on pre covid and validating with covid, so a score like that for that split would be pretty extraordinary i'd imagine</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2641913,
      "author_name": "darynarr",
      "author_url": "",
      "post_date": "02/07/2024 18:43:00",
      "content": "<table>\n<thead>\n<tr>\n<th>Description</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>depth 0</td>\n<td>0.578</td>\n<td>0.507</td>\n</tr>\n<tr>\n<td>depth 0 &amp; 1</td>\n<td>0.668</td>\n<td>0.516</td>\n</tr>\n<tr>\n<td>depth 0 &amp; 1 &amp; 2</td>\n<td>0.675</td>\n<td>0.506</td>\n</tr>\n<tr>\n<td>depth 0 &amp; 1 &amp; 2 without categorical features</td>\n<td>0.643</td>\n<td>0.5</td>\n</tr>\n</tbody>\n</table>\n<p>StratifiedGroupKFold(5) is used here. So yeah, naively adding more features with <code>num_group1==0</code> and <code>num_group2==0</code> improves CV, but no significant impact on LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2641924,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "02/07/2024 18:53:44",
          "content": "<p>your cv has leakage ure using information from the future which could be causing this</p>",
          "votes": null,
          "replies": [
            {
              "id": 2641940,
              "author_name": "darynarr",
              "author_url": "",
              "post_date": "02/07/2024 19:07:34",
              "content": "<p>Well, I used StratifiedGroupKFold and not just StratifiedKFold, to eliminate the leakage by not mixing up data from different weeks, but I agree this does not solve the problem completely. Train / valid split by time and training on whole dataset also sounds better, but I have not tried it yet</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2642619,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/08/2024 09:40:49",
      "content": "<blockquote>\n  <p>This could mean that test data has a different distribution from train data.</p>\n</blockquote>\n<p>You are on the right track. This is obviously caused by covid. Remember the important part of the competition is also stability. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2642751,
      "author_name": "justin1357",
      "author_url": "",
      "post_date": "02/08/2024 11:38:25",
      "content": "<p>train on 0-46, val on 47-73, test on 74-91, anyone same as me?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2642755,
          "author_name": "justin1357",
          "author_url": "",
          "post_date": "02/08/2024 11:43:02",
          "content": "<table>\n<thead>\n<tr>\n<th>data</th>\n<th>cv(test, competition metric)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>depth=0 only filltered</td>\n<td>0.6749</td>\n</tr>\n</tbody>\n</table>\n<p>early stop on val</p>",
          "votes": null,
          "replies": [
            {
              "id": 2644146,
              "author_name": "justin1357",
              "author_url": "",
              "post_date": "02/09/2024 09:54:56",
              "content": "<p>Although the LB (Leaderboard) score is low, the correlation between CV (Cross Validation) and LB is strong. I do not believe that COVID will significantly impact the model's AUC (Area Under the Curve) performance. Based on the evidence, <strong>training the model exclusively on data from pre-COVID times does not seem to deteriorate its performance on test sets from during the COVID era.</strong></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2641681": "my last submission: CV=0.596, LB=0.507.\n\nCV is calculated by training on weeks 0-46 and validating on weeks 46-91 (1 fold, so not actually CV).\n\nSubmission trained on all data - weeks 0-91.\n\nThe difference is around 0.09 - seems very large to me. Does anybody else see this pattern? \nThis could mean that test data has a different distribution from train data. Possibly caused by new values of categorical variables in test that are not present in train. That would present a very challenging problem.",
    "2641705": "It could be attributed to the test set covers relatively large period.\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764",
    "2641833": "make sure ure actually using the correct metric, ure training on pre covid and validating with covid, so a score like that for that split would be pretty extraordinary i'd imagine",
    "2641913": "| Description | CV | LB |\n| --- | --- | --- |\n| depth 0 | 0.578 | 0.507 |\n| depth 0 & 1 | 0.668 | 0.516 |\n| depth 0 & 1 & 2 | 0.675 | 0.506 |\n| depth 0 & 1 & 2 without categorical features |  0.643 | 0.5 |\n\n\nStratifiedGroupKFold(5) is used here. So yeah, naively adding more features with `num_group1==0` and `num_group2==0` improves CV, but no significant impact on LB.",
    "2641924": "your cv has leakage ure using information from the future which could be causing this",
    "2641940": "Well, I used StratifiedGroupKFold and not just StratifiedKFold, to eliminate the leakage by not mixing up data from different weeks, but I agree this does not solve the problem completely. Train / valid split by time and training on whole dataset also sounds better, but I have not tried it yet",
    "2642619": ">This could mean that test data has a different distribution from train data.\n\nYou are on the right track. This is obviously caused by covid. Remember the important part of the competition is also stability.",
    "2642751": "train on 0-46, val on 47-73, test on 74-91, anyone same as me?",
    "2642755": "| data | cv(test, competition metric) |\n| --- | --- |\n| depth=0 only filltered | 0.6749 |\n\nearly stop on val",
    "2644146": "Although the LB (Leaderboard) score is low, the correlation between CV (Cross Validation) and LB is strong. I do not believe that COVID will significantly impact the model's AUC (Area Under the Curve) performance. Based on the evidence, **training the model exclusively on data from pre-COVID times does not seem to deteriorate its performance on test sets from during the COVID era.**"
  },
  "source": "meta"
}