{
  "id": 324304,
  "title": "Should I trust the CV score or the public score?",
  "url": "/competitions/birdclef-2022/discussion/324304",
  "author_name": "",
  "post_date": "2022-05-11T00:39:59.736090Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In this competition, some of my models had better local CV F1 scores, but poor public scores. This makes me wonder whether to trust the CV score or the public score, after all, the key to the competition is the 84% of the test data on the private ranking.</p>\n<p>I currently want to give higher weights to the CV score:</p>\n<ol>\n<li>The proportion of test data in the private ranking is very large, so it is necessary to examine the generalization ability and the degree of overfitting of the model.</li>\n<li>As <a href=\"https://www.kaggle.com/kotanoda\" target=\"_blank\">@kotanoda</a> said in the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/319582\" target=\"_blank\">discussion</a>, There is a big dissociation between the validation score and the public score. Possibly, domain shift between train data and test data cause this dissociation.</li>\n</ol>\n<p>I'm not sure if my simple idea is correct and would appreciate advice from fellow Kagglers in the Kaggle community. 😂</p>\n<p>Thanks for your help! 😁😁</p>",
  "messages": [
    {
      "id": "1784135",
      "postDate": "05/11/2022 00:39:59",
      "content": "<p>In this competition, some of my models had better local CV F1 scores, but poor public scores. This makes me wonder whether to trust the CV score or the public score, after all, the key to the competition is the 84% of the test data on the private ranking.</p>\n<p>I currently want to give higher weights to the CV score:</p>\n<ol>\n<li>The proportion of test data in the private ranking is very large, so it is necessary to examine the generalization ability and the degree of overfitting of the model.</li>\n<li>As <a href=\"https://www.kaggle.com/kotanoda\" target=\"_blank\">@kotanoda</a> said in the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/319582\" target=\"_blank\">discussion</a>, There is a big dissociation between the validation score and the public score. Possibly, domain shift between train data and test data cause this dissociation.</li>\n</ol>\n<p>I'm not sure if my simple idea is correct and would appreciate advice from fellow Kagglers in the Kaggle community. 😂</p>\n<p>Thanks for your help! 😁😁</p>",
      "rawMarkdown": "In this competition, some of my models had better local CV F1 scores, but poor public scores. This makes me wonder whether to trust the CV score or the public score, after all, the key to the competition is the 84% of the test data on the private ranking.\n\nI currently want to give higher weights to the CV score:\n1. The proportion of test data in the private ranking is very large, so it is necessary to examine the generalization ability and the degree of overfitting of the model.\n2. As @kotanoda said in the [discussion](https://www.kaggle.com/competitions/birdclef-2022/discussion/319582), There is a big dissociation between the validation score and the public score. Possibly, domain shift between train data and test data cause this dissociation.\n\nI'm not sure if my simple idea is correct and would appreciate advice from fellow Kagglers in the Kaggle community. 😂\n\nThanks for your help! 😁😁",
      "votes": null
    },
    {
      "id": "1784431",
      "postDate": "05/11/2022 07:14:55",
      "content": "<p>In fact, I found out in Cornell Birdcall Identification competition, <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183208\" target=\"_blank\">1st Place Solution</a> believes more in LB, so I don't know how to decide now.</p>",
      "rawMarkdown": "In fact, I found out in Cornell Birdcall Identification competition, @taggatle [1st Place Solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183208) believes more in LB, so I don't know how to decide now.",
      "votes": null
    },
    {
      "id": "1784539",
      "postDate": "05/11/2022 08:32:44",
      "content": "<p>It's definitely worth checking your models' performance on the public score to see if there is a significant difference. If there is a big discrepancy, it could be due to overfitting or a domain shift between the training and test data. In either case, it's important to try to improve your model's generalization ability.</p>",
      "rawMarkdown": "It's definitely worth checking your models' performance on the public score to see if there is a significant difference. If there is a big discrepancy, it could be due to overfitting or a domain shift between the training and test data. In either case, it's important to try to improve your model's generalization ability.",
      "votes": null
    },
    {
      "id": "1784565",
      "postDate": "05/11/2022 09:07:25",
      "content": "<p>You are right. However, for now, this makes me unsure whether I should choose to trust the model with a good LB or the model with a good CV when submitting a private list.</p>",
      "rawMarkdown": "You are right. However, for now, this makes me unsure whether I should choose to trust the model with a good LB or the model with a good CV when submitting a private list.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1784431,
      "author_name": "jiedengsc",
      "author_url": "",
      "post_date": "05/11/2022 07:14:55",
      "content": "<p>In fact, I found out in Cornell Birdcall Identification competition, <a href=\"https://www.kaggle.com/taggatle\" target=\"_blank\">@taggatle</a> <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183208\" target=\"_blank\">1st Place Solution</a> believes more in LB, so I don't know how to decide now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784539,
      "author_name": "",
      "author_url": "",
      "post_date": "05/11/2022 08:32:44",
      "content": "<p>It's definitely worth checking your models' performance on the public score to see if there is a significant difference. If there is a big discrepancy, it could be due to overfitting or a domain shift between the training and test data. In either case, it's important to try to improve your model's generalization ability.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784565,
          "author_name": "jiedengsc",
          "author_url": "",
          "post_date": "05/11/2022 09:07:25",
          "content": "<p>You are right. However, for now, this makes me unsure whether I should choose to trust the model with a good LB or the model with a good CV when submitting a private list.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1784135": "In this competition, some of my models had better local CV F1 scores, but poor public scores. This makes me wonder whether to trust the CV score or the public score, after all, the key to the competition is the 84% of the test data on the private ranking.\n\nI currently want to give higher weights to the CV score:\n1. The proportion of test data in the private ranking is very large, so it is necessary to examine the generalization ability and the degree of overfitting of the model.\n2. As @kotanoda said in the [discussion](https://www.kaggle.com/competitions/birdclef-2022/discussion/319582), There is a big dissociation between the validation score and the public score. Possibly, domain shift between train data and test data cause this dissociation.\n\nI'm not sure if my simple idea is correct and would appreciate advice from fellow Kagglers in the Kaggle community. 😂\n\nThanks for your help! 😁😁",
    "1784431": "In fact, I found out in Cornell Birdcall Identification competition, @taggatle [1st Place Solution](https://www.kaggle.com/c/birdsong-recognition/discussion/183208) believes more in LB, so I don't know how to decide now.",
    "1784539": "It's definitely worth checking your models' performance on the public score to see if there is a significant difference. If there is a big discrepancy, it could be due to overfitting or a domain shift between the training and test data. In either case, it's important to try to improve your model's generalization ability.",
    "1784565": "You are right. However, for now, this makes me unsure whether I should choose to trust the model with a good LB or the model with a good CV when submitting a private list."
  },
  "source": "meta"
}