{
  "id": 352567,
  "title": "Trust CV or LB？",
  "url": "/competitions/open-problems-multimodal/discussion/352567",
  "author_name": "",
  "post_date": "2022-09-15T01:25:17.754240300Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I use 3KFold for cite and get 0.905 CV score,  but its LB score is no better than the model with CV score of 0.886. <br>\nWhich one should I trust?</p>",
  "messages": [
    {
      "id": "1939733",
      "postDate": "09/15/2022 01:25:17",
      "content": "<p>I use 3KFold for cite and get 0.905 CV score,  but its LB score is no better than the model with CV score of 0.886. <br>\nWhich one should I trust?</p>",
      "rawMarkdown": "I use 3KFold for cite and get 0.905 CV score,  but its LB score is no better than the model with CV score of 0.886. \nWhich one should I trust?",
      "votes": null
    },
    {
      "id": "1939829",
      "postDate": "09/15/2022 03:48:22",
      "content": "<p>Since the train dataset is much larger than the test, trust cv is a good strategy, and you may could submit 1 max cv and 1 max lb in the final. </p>",
      "rawMarkdown": "Since the train dataset is much larger than the test, trust cv is a good strategy, and you may could submit 1 max cv and 1 max lb in the final.",
      "votes": null
    },
    {
      "id": "1939866",
      "postDate": "09/15/2022 05:03:59",
      "content": "<p>I use multiple random seeds and get the same score, so I think the CV score is reliable. But what puzzles me is that in my other submissions, CV scores and LB scores are all positively correlated</p>",
      "rawMarkdown": "I use multiple random seeds and get the same score, so I think the CV score is reliable. But what puzzles me is that in my other submissions, CV scores and LB scores are all positively correlated",
      "votes": null
    },
    {
      "id": "1941239",
      "postDate": "09/15/2022 21:36:41",
      "content": "<p>I think you could trust cv more in this case, lb has less data, making it easier to be influenced by the random effects</p>",
      "rawMarkdown": "I think you could trust cv more in this case, lb has less data, making it easier to be influenced by the random effects",
      "votes": null
    },
    {
      "id": "1942030",
      "postDate": "09/16/2022 11:57:14",
      "content": "<p>I found that it seems that the decision tree models (LGBM, Catboost) can easily get a high CV score (higher than 0.9), but its performance on LB is not so good. What's the matter?</p>",
      "rawMarkdown": "I found that it seems that the decision tree models (LGBM, Catboost) can easily get a high CV score (higher than 0.9), but its performance on LB is not so good. What's the matter?",
      "votes": null
    },
    {
      "id": "1943717",
      "postDate": "09/17/2022 18:00:23",
      "content": "<p>Same here.  I'm getting LB scores between 0.5-0.6 on decision tree models with CV greater than 0.9.  I thought for sure I have an error in my process.  Now I'm not so sure.  How bad is your LB score on these?</p>",
      "rawMarkdown": "Same here.  I'm getting LB scores between 0.5-0.6 on decision tree models with CV greater than 0.9.  I thought for sure I have an error in my process.  Now I'm not so sure.  How bad is your LB score on these?",
      "votes": null
    },
    {
      "id": "1944012",
      "postDate": "09/18/2022 01:43:06",
      "content": "<p>I never tried to submit only the citeseq partial model, and after merging my best multiome model, the LB scores of lgbm and catboost were 0.810 and 0.808, which were lower than the NN model of 0.811. </p>",
      "rawMarkdown": "I never tried to submit only the citeseq partial model, and after merging my best multiome model, the LB scores of lgbm and catboost were 0.810 and 0.808, which were lower than the NN model of 0.811.",
      "votes": null
    },
    {
      "id": "1944251",
      "postDate": "09/18/2022 07:07:45",
      "content": "<p>Ah ok. Your differences are subtle.  I submitted multi NN + cite catboost to get the .5-.6.  I'm reusing the NN multi that contributed to a .804.  So my cite catboost portion is probably messed up with leakage, mixing up rows, or a bad join.  Or maybe I broke my combining script. Your response gives me hope that I can figure out what's wrong with it. Thanks</p>",
      "rawMarkdown": "Ah ok. Your differences are subtle.  I submitted multi NN + cite catboost to get the .5-.6.  I'm reusing the NN multi that contributed to a .804.  So my cite catboost portion is probably messed up with leakage, mixing up rows, or a bad join.  Or maybe I broke my combining script. Your response gives me hope that I can figure out what's wrong with it. Thanks",
      "votes": null
    },
    {
      "id": "1944361",
      "postDate": "09/18/2022 08:34:01",
      "content": "<p>Public LB is different from private LB. There will be a shake up.<br>\nSo you should trust your CV, but only if you do it correctly. <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202</a></p>",
      "rawMarkdown": "Public LB is different from private LB. There will be a shake up.\nSo you should trust your CV, but only if you do it correctly. \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202",
      "votes": null
    },
    {
      "id": "1944960",
      "postDate": "09/18/2022 17:33:01",
      "content": "<p>I realized I just didn't have my submission sorted by row_id.  Sorting fixed it.</p>",
      "rawMarkdown": "I realized I just didn't have my submission sorted by row_id.  Sorting fixed it.",
      "votes": null
    },
    {
      "id": "1945760",
      "postDate": "09/19/2022 11:27:48",
      "content": "<p>Thank you for your advice. <br>\n:)</p>",
      "rawMarkdown": "Thank you for your advice. \n:)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1939829,
      "author_name": "learnmore1",
      "author_url": "",
      "post_date": "09/15/2022 03:48:22",
      "content": "<p>Since the train dataset is much larger than the test, trust cv is a good strategy, and you may could submit 1 max cv and 1 max lb in the final. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1939866,
          "author_name": "qqzzxxdd",
          "author_url": "",
          "post_date": "09/15/2022 05:03:59",
          "content": "<p>I use multiple random seeds and get the same score, so I think the CV score is reliable. But what puzzles me is that in my other submissions, CV scores and LB scores are all positively correlated</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1941239,
          "author_name": "learnmore1",
          "author_url": "",
          "post_date": "09/15/2022 21:36:41",
          "content": "<p>I think you could trust cv more in this case, lb has less data, making it easier to be influenced by the random effects</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1942030,
      "author_name": "qqzzxxdd",
      "author_url": "",
      "post_date": "09/16/2022 11:57:14",
      "content": "<p>I found that it seems that the decision tree models (LGBM, Catboost) can easily get a high CV score (higher than 0.9), but its performance on LB is not so good. What's the matter?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1943717,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "09/17/2022 18:00:23",
          "content": "<p>Same here.  I'm getting LB scores between 0.5-0.6 on decision tree models with CV greater than 0.9.  I thought for sure I have an error in my process.  Now I'm not so sure.  How bad is your LB score on these?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1944012,
          "author_name": "qqzzxxdd",
          "author_url": "",
          "post_date": "09/18/2022 01:43:06",
          "content": "<p>I never tried to submit only the citeseq partial model, and after merging my best multiome model, the LB scores of lgbm and catboost were 0.810 and 0.808, which were lower than the NN model of 0.811. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1944251,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "09/18/2022 07:07:45",
          "content": "<p>Ah ok. Your differences are subtle.  I submitted multi NN + cite catboost to get the .5-.6.  I'm reusing the NN multi that contributed to a .804.  So my cite catboost portion is probably messed up with leakage, mixing up rows, or a bad join.  Or maybe I broke my combining script. Your response gives me hope that I can figure out what's wrong with it. Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1944960,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "09/18/2022 17:33:01",
          "content": "<p>I realized I just didn't have my submission sorted by row_id.  Sorting fixed it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1944361,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/18/2022 08:34:01",
      "content": "<p>Public LB is different from private LB. There will be a shake up.<br>\nSo you should trust your CV, but only if you do it correctly. <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1945760,
          "author_name": "qqzzxxdd",
          "author_url": "",
          "post_date": "09/19/2022 11:27:48",
          "content": "<p>Thank you for your advice. <br>\n:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1939733": "I use 3KFold for cite and get 0.905 CV score,  but its LB score is no better than the model with CV score of 0.886. \nWhich one should I trust?",
    "1939829": "Since the train dataset is much larger than the test, trust cv is a good strategy, and you may could submit 1 max cv and 1 max lb in the final.",
    "1939866": "I use multiple random seeds and get the same score, so I think the CV score is reliable. But what puzzles me is that in my other submissions, CV scores and LB scores are all positively correlated",
    "1941239": "I think you could trust cv more in this case, lb has less data, making it easier to be influenced by the random effects",
    "1942030": "I found that it seems that the decision tree models (LGBM, Catboost) can easily get a high CV score (higher than 0.9), but its performance on LB is not so good. What's the matter?",
    "1943717": "Same here.  I'm getting LB scores between 0.5-0.6 on decision tree models with CV greater than 0.9.  I thought for sure I have an error in my process.  Now I'm not so sure.  How bad is your LB score on these?",
    "1944012": "I never tried to submit only the citeseq partial model, and after merging my best multiome model, the LB scores of lgbm and catboost were 0.810 and 0.808, which were lower than the NN model of 0.811.",
    "1944251": "Ah ok. Your differences are subtle.  I submitted multi NN + cite catboost to get the .5-.6.  I'm reusing the NN multi that contributed to a .804.  So my cite catboost portion is probably messed up with leakage, mixing up rows, or a bad join.  Or maybe I broke my combining script. Your response gives me hope that I can figure out what's wrong with it. Thanks",
    "1944361": "Public LB is different from private LB. There will be a shake up.\nSo you should trust your CV, but only if you do it correctly. \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/347202",
    "1944960": "I realized I just didn't have my submission sorted by row_id.  Sorting fixed it.",
    "1945760": "Thank you for your advice. \n:)"
  },
  "source": "meta"
}