{
  "id": 234886,
  "title": "How to prevent overfitting the public leaderboard test data?",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/234886",
  "author_name": "",
  "post_date": "2021-04-26T17:41:45.295929700Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We can see from previous similar competitions that there's a huge <strong>shakeup</strong> between <strong>Public</strong> and <strong>Private</strong> leaderboard. Some teams did a good job on the Public leaderboard, but fail in the Private leaderboard. That's because of the overfitting to the part of the test data. So is there an effective way to prevent that kind of overfitting? </p>\n<p>BTW, I know that we should always trust our CV score, not LB score. However, in this competition, the CV and LB are so unstable and uncorrelated as is also posted in this thread:<br>\nStrange, best model, cv and lb? <a href=\"url\" target=\"_blank\"> https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/230099</a></p>\n<p>So how should we do with our CV score? Any method to avoid overfitting? I'm looking for some advice.</p>\n<p>Private leaderboards of previous competitions:<br>\nCassava Leaf Disease Classification 2021: <a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/leaderboard\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/leaderboard</a>  <br>\nPlant Pathology 2020 <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2020-fgvc7/leaderboard</a></p>",
  "messages": [
    {
      "id": "1285221",
      "postDate": "04/26/2021 17:41:45",
      "content": "<p>We can see from previous similar competitions that there's a huge <strong>shakeup</strong> between <strong>Public</strong> and <strong>Private</strong> leaderboard. Some teams did a good job on the Public leaderboard, but fail in the Private leaderboard. That's because of the overfitting to the part of the test data. So is there an effective way to prevent that kind of overfitting? </p>\n<p>BTW, I know that we should always trust our CV score, not LB score. However, in this competition, the CV and LB are so unstable and uncorrelated as is also posted in this thread:<br>\nStrange, best model, cv and lb? <a href=\"url\" target=\"_blank\"> https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/230099</a></p>\n<p>So how should we do with our CV score? Any method to avoid overfitting? I'm looking for some advice.</p>\n<p>Private leaderboards of previous competitions:<br>\nCassava Leaf Disease Classification 2021: <a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/leaderboard\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/leaderboard</a>  <br>\nPlant Pathology 2020 <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2020-fgvc7/leaderboard</a></p>",
      "rawMarkdown": "We can see from previous similar competitions that there's a huge **shakeup** between **Public** and **Private** leaderboard. Some teams did a good job on the Public leaderboard, but fail in the Private leaderboard. That's because of the overfitting to the part of the test data. So is there an effective way to prevent that kind of overfitting? \n\nBTW, I know that we should always trust our CV score, not LB score. However, in this competition, the CV and LB are so unstable and uncorrelated as is also posted in this thread:\nStrange, best model, cv and lb? [ https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/230099](url)\n\nSo how should we do with our CV score? Any method to avoid overfitting? I'm looking for some advice.\n\nPrivate leaderboards of previous competitions:\nCassava Leaf Disease Classification 2021: [https://www.kaggle.com/c/cassava-leaf-disease-classification/leaderboard  ](url)\nPlant Pathology 2020 [https://www.kaggle.com/c/plant-pathology-2020-fgvc7/leaderboard](url)",
      "votes": null
    },
    {
      "id": "1312169",
      "postDate": "05/17/2021 21:49:16",
      "content": "<p>i do have a question.<br>\nThere exists other combination of categories than in train.csv.<br>\nWe need to predict those as well or just unique labels from train.csv?</p>",
      "rawMarkdown": "i do have a question.\nThere exists other combination of categories than in train.csv.\nWe need to predict those as well or just unique labels from train.csv?",
      "votes": null
    },
    {
      "id": "1312348",
      "postDate": "05/18/2021 02:13:12",
      "content": "<p>It's multi-label problem, so there might be more than one possible labels in our prediction. </p>",
      "rawMarkdown": "It's multi-label problem, so there might be more than one possible labels in our prediction.",
      "votes": null
    },
    {
      "id": "1313493",
      "postDate": "05/18/2021 15:26:52",
      "content": "<p>To be honest, I don't see a way except checking your CV and using multi-folds.</p>",
      "rawMarkdown": "To be honest, I don't see a way except checking your CV and using multi-folds.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1312169,
      "author_name": "micheomaano",
      "author_url": "",
      "post_date": "05/17/2021 21:49:16",
      "content": "<p>i do have a question.<br>\nThere exists other combination of categories than in train.csv.<br>\nWe need to predict those as well or just unique labels from train.csv?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1312348,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/18/2021 02:13:12",
          "content": "<p>It's multi-label problem, so there might be more than one possible labels in our prediction. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1313493,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "05/18/2021 15:26:52",
      "content": "<p>To be honest, I don't see a way except checking your CV and using multi-folds.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1285221": "We can see from previous similar competitions that there's a huge **shakeup** between **Public** and **Private** leaderboard. Some teams did a good job on the Public leaderboard, but fail in the Private leaderboard. That's because of the overfitting to the part of the test data. So is there an effective way to prevent that kind of overfitting? \n\nBTW, I know that we should always trust our CV score, not LB score. However, in this competition, the CV and LB are so unstable and uncorrelated as is also posted in this thread:\nStrange, best model, cv and lb? [ https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/230099](url)\n\nSo how should we do with our CV score? Any method to avoid overfitting? I'm looking for some advice.\n\nPrivate leaderboards of previous competitions:\nCassava Leaf Disease Classification 2021: [https://www.kaggle.com/c/cassava-leaf-disease-classification/leaderboard  ](url)\nPlant Pathology 2020 [https://www.kaggle.com/c/plant-pathology-2020-fgvc7/leaderboard](url)",
    "1312169": "i do have a question.\nThere exists other combination of categories than in train.csv.\nWe need to predict those as well or just unique labels from train.csv?",
    "1312348": "It's multi-label problem, so there might be more than one possible labels in our prediction.",
    "1313493": "To be honest, I don't see a way except checking your CV and using multi-folds."
  },
  "source": "meta"
}