{
  "id": 299804,
  "title": "Why relation between CV and LB is not linear ? ",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/299804",
  "author_name": "",
  "post_date": "2022-01-10T03:01:00.846589300Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I'm using yolov5 in this competition. But, I'm facing two problem.</p>\n<p>First, difference between CV - 0.705 and LB - 0.505 is very huge. I'm referenced CV method <a href=\"https://www.kaggle.com/hyunmingu/evaluate-f2-score-for-yolov5-model\" target=\"_blank\">in this notebook.</a> What do you use CV calculation method in your case ? Is your method reflect linear relation between CV and LB ? </p>\n<p>Second, even if CV is improves, LB is not improve always. I think that because I use incorrect CV calculation method or incorrect data split method.</p>\n<p>I'm trying to change data split method from 80:20 with only labeled data to 5fold StratifiedKFold with subsequence <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">in this notebook.</a></p>\n<p>LB improved, but there are still remain two problem above. For improve CV, I adjust hyper parameter. Even if CV improved, LB is not improved. I'm really confused. </p>\n<p>How do I address these problem ? </p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1644066",
      "postDate": "01/10/2022 03:01:00",
      "content": "<p>Hi all,</p>\n<p>I'm using yolov5 in this competition. But, I'm facing two problem.</p>\n<p>First, difference between CV - 0.705 and LB - 0.505 is very huge. I'm referenced CV method <a href=\"https://www.kaggle.com/hyunmingu/evaluate-f2-score-for-yolov5-model\" target=\"_blank\">in this notebook.</a> What do you use CV calculation method in your case ? Is your method reflect linear relation between CV and LB ? </p>\n<p>Second, even if CV is improves, LB is not improve always. I think that because I use incorrect CV calculation method or incorrect data split method.</p>\n<p>I'm trying to change data split method from 80:20 with only labeled data to 5fold StratifiedKFold with subsequence <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">in this notebook.</a></p>\n<p>LB improved, but there are still remain two problem above. For improve CV, I adjust hyper parameter. Even if CV improved, LB is not improved. I'm really confused. </p>\n<p>How do I address these problem ? </p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi all,\n\nI'm using yolov5 in this competition. But, I'm facing two problem.\n\nFirst, difference between CV - 0.705 and LB - 0.505 is very huge. I'm referenced CV method [in this notebook.](https://www.kaggle.com/hyunmingu/evaluate-f2-score-for-yolov5-model) What do you use CV calculation method in your case ? Is your method reflect linear relation between CV and LB ? \n\nSecond, even if CV is improves, LB is not improve always. I think that because I use incorrect CV calculation method or incorrect data split method.\n\nI'm trying to change data split method from 80:20 with only labeled data to 5fold StratifiedKFold with subsequence [in this notebook.](https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences)\n\nLB improved, but there are still remain two problem above. For improve CV, I adjust hyper parameter. Even if CV improved, LB is not improved. I'm really confused. \n\nHow do I address these problem ? \n\nThanks.",
      "votes": null
    },
    {
      "id": "1644700",
      "postDate": "01/10/2022 12:23:16",
      "content": "<p>LB does not always have to improve with CV. Because the data may be from a slightly (or not slightly 🙁) different distribution. </p>",
      "rawMarkdown": "LB does not always have to improve with CV. Because the data may be from a slightly (or not slightly 🙁) different distribution.",
      "votes": null
    },
    {
      "id": "1644830",
      "postDate": "01/10/2022 14:11:28",
      "content": "<p>Hmm.. Yes you are right. But, in my experience, the good model improve together.</p>",
      "rawMarkdown": "Hmm.. Yes you are right. But, in my experience, the good model improve together.",
      "votes": null
    },
    {
      "id": "1645151",
      "postDate": "01/10/2022 18:14:16",
      "content": "<p>I've found that even using the notebook for the splits still lead to overfitting as sequences from the same video also tend to be quite similar. I myself have prototyped models using 2 videos as train and 1 video as validation and then retrained the model on all data after I found a good training pipeline. Doing this helped my generalization onto the LB even with less accurate resnet50 based models.</p>",
      "rawMarkdown": "I've found that even using the notebook for the splits still lead to overfitting as sequences from the same video also tend to be quite similar. I myself have prototyped models using 2 videos as train and 1 video as validation and then retrained the model on all data after I found a good training pipeline. Doing this helped my generalization onto the LB even with less accurate resnet50 based models.",
      "votes": null
    },
    {
      "id": "1651314",
      "postDate": "01/15/2022 17:12:15",
      "content": "<p>There are CV data and LB data differences. Also Each video is so different in train dataset</p>",
      "rawMarkdown": "There are CV data and LB data differences. Also Each video is so different in train dataset",
      "votes": null
    },
    {
      "id": "1653032",
      "postDate": "01/17/2022 07:19:33",
      "content": "<p>If so, how do we determine a good model ? A correct CV value is one of the judgment materials. To prevent shake down in the end of competition, what can I do ? It is very dangerous to follow only the public LB score in my experience.</p>",
      "rawMarkdown": "If so, how do we determine a good model ? A correct CV value is one of the judgment materials. To prevent shake down in the end of competition, what can I do ? It is very dangerous to follow only the public LB score in my experience.",
      "votes": null
    },
    {
      "id": "1661695",
      "postDate": "01/23/2022 16:12:52",
      "content": "<p>\"CV - 0.705 and LB - 0.505 is very huge.\", I think you should narrow the gap between cv and lb. also you need to check your metrics is correct. </p>",
      "rawMarkdown": "\"CV - 0.705 and LB - 0.505 is very huge.\", I think you should narrow the gap between cv and lb. also you need to check your metrics is correct.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1644700,
      "author_name": "vadbeg",
      "author_url": "",
      "post_date": "01/10/2022 12:23:16",
      "content": "<p>LB does not always have to improve with CV. Because the data may be from a slightly (or not slightly 🙁) different distribution. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1644830,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "01/10/2022 14:11:28",
          "content": "<p>Hmm.. Yes you are right. But, in my experience, the good model improve together.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1645151,
      "author_name": "maxvandijck",
      "author_url": "",
      "post_date": "01/10/2022 18:14:16",
      "content": "<p>I've found that even using the notebook for the splits still lead to overfitting as sequences from the same video also tend to be quite similar. I myself have prototyped models using 2 videos as train and 1 video as validation and then retrained the model on all data after I found a good training pipeline. Doing this helped my generalization onto the LB even with less accurate resnet50 based models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1651314,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "01/15/2022 17:12:15",
      "content": "<p>There are CV data and LB data differences. Also Each video is so different in train dataset</p>",
      "votes": null,
      "replies": [
        {
          "id": 1653032,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "01/17/2022 07:19:33",
          "content": "<p>If so, how do we determine a good model ? A correct CV value is one of the judgment materials. To prevent shake down in the end of competition, what can I do ? It is very dangerous to follow only the public LB score in my experience.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1661695,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "01/23/2022 16:12:52",
          "content": "<p>\"CV - 0.705 and LB - 0.505 is very huge.\", I think you should narrow the gap between cv and lb. also you need to check your metrics is correct. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1644066": "Hi all,\n\nI'm using yolov5 in this competition. But, I'm facing two problem.\n\nFirst, difference between CV - 0.705 and LB - 0.505 is very huge. I'm referenced CV method [in this notebook.](https://www.kaggle.com/hyunmingu/evaluate-f2-score-for-yolov5-model) What do you use CV calculation method in your case ? Is your method reflect linear relation between CV and LB ? \n\nSecond, even if CV is improves, LB is not improve always. I think that because I use incorrect CV calculation method or incorrect data split method.\n\nI'm trying to change data split method from 80:20 with only labeled data to 5fold StratifiedKFold with subsequence [in this notebook.](https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences)\n\nLB improved, but there are still remain two problem above. For improve CV, I adjust hyper parameter. Even if CV improved, LB is not improved. I'm really confused. \n\nHow do I address these problem ? \n\nThanks.",
    "1644700": "LB does not always have to improve with CV. Because the data may be from a slightly (or not slightly 🙁) different distribution.",
    "1644830": "Hmm.. Yes you are right. But, in my experience, the good model improve together.",
    "1645151": "I've found that even using the notebook for the splits still lead to overfitting as sequences from the same video also tend to be quite similar. I myself have prototyped models using 2 videos as train and 1 video as validation and then retrained the model on all data after I found a good training pipeline. Doing this helped my generalization onto the LB even with less accurate resnet50 based models.",
    "1651314": "There are CV data and LB data differences. Also Each video is so different in train dataset",
    "1653032": "If so, how do we determine a good model ? A correct CV value is one of the judgment materials. To prevent shake down in the end of competition, what can I do ? It is very dangerous to follow only the public LB score in my experience.",
    "1661695": "\"CV - 0.705 and LB - 0.505 is very huge.\", I think you should narrow the gap between cv and lb. also you need to check your metrics is correct."
  },
  "source": "meta"
}