{
  "id": 286693,
  "title": "Trust CV or LB?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/286693",
  "author_name": "cool_rabbit",
  "post_date": "2021-11-10T09:14:05.023000",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi. I'm relatively new here at this competition.</p>\n<p>After reading some discussion, I realize train and test data are different in terms of overlapping.<br>\nIn my experiment, I use stratified groupkfold (n_splits=5), and CV and LB correlation of single fold is not so strong (not weak though).<br>\nConsidering public/private split ratio, I guess it is acceptable to trust LB with paying attention to CV score.<br>\nWhat's your opinion and strategy?</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 1577647,
      "postDate": "2021-11-10T09:14:05.023Z",
      "content": "<p>Hi. I'm relatively new here at this competition.</p>\n<p>After reading some discussion, I realize train and test data are different in terms of overlapping.<br>\nIn my experiment, I use stratified groupkfold (n_splits=5), and CV and LB correlation of single fold is not so strong (not weak though).<br>\nConsidering public/private split ratio, I guess it is acceptable to trust LB with paying attention to CV score.<br>\nWhat's your opinion and strategy?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi. I'm relatively new here at this competition.\n\nAfter reading some discussion, I realize train and test data are different in terms of overlapping.\nIn my experiment, I use stratified groupkfold (n_splits=5), and CV and LB correlation of single fold is not so strong (not weak though).\nConsidering public/private split ratio, I guess it is acceptable to trust LB with paying attention to CV score.\nWhat's your opinion and strategy?\n\nThank you.",
      "votes": 10
    },
    {
      "id": 1581774,
      "postDate": "2021-11-14T08:33:22.257Z",
      "content": "<p>what you probably want to do is test out your CV: </p>\n<ol>\n<li>Train multiple [different] models and get various CV scores.</li>\n<li>Submit them all to the LB. </li>\n<li>Plot [scatter] the CV scores you got vs the LB scores you got.</li>\n<li>Optional: measure [spearman / maybe also person] correlation of the LB scores to the CV scores.</li>\n</ol>\n<p>If you find that there is a strong correlation between the LB and the CV, you can trust your CV.</p>\n<p><strong>BUT!!</strong></p>\n<p>Only up to a point.<br>\nUsually, you will find out that as you get to the highest ranks on the leaderboard, your CV-LB correlation starts to lower. from there, you probably should look closely into the data to try and find the samples that are \"not aligned\" or the samples that mess up your models. </p>\n<p>It all depends on the competition itself and the data. sometimes [often] if you did everything correct: you should trust your CV. whoever, there are some situations where coming up with a meaningful CV that won't break until the end of the competition is really really hard. when this is the case: you probably want to treat the LB as \"another fold\". </p>",
      "rawMarkdown": "what you probably want to do is test out your CV: \n1.  Train multiple [different] models and get various CV scores.\n2. Submit them all to the LB. \n3. Plot [scatter] the CV scores you got vs the LB scores you got.\n4. Optional: measure [spearman / maybe also person] correlation of the LB scores to the CV scores.\n\nIf you find that there is a strong correlation between the LB and the CV, you can trust your CV.\n\n\n**BUT!!**\n\nOnly up to a point.\nUsually, you will find out that as you get to the highest ranks on the leaderboard, your CV-LB correlation starts to lower. from there, you probably should look closely into the data to try and find the samples that are \"not aligned\" or the samples that mess up your models. \n\nIt all depends on the competition itself and the data. sometimes [often] if you did everything correct: you should trust your CV. whoever, there are some situations where coming up with a meaningful CV that won't break until the end of the competition is really really hard. when this is the case: you probably want to treat the LB as \"another fold\". ",
      "votes": 5
    },
    {
      "id": 1581839,
      "postDate": "2021-11-14T09:27:37.203Z",
      "content": "<p>In my case, ｗhen the CV (stratified kfold for cell_type) increased by 0.01, the LB also increased by 0.01.<br>\nCV 0.2717 -&gt; LB 0.295<br>\nCV 0.2817 -&gt; LB 0.304</p>",
      "rawMarkdown": "In my case, ｗhen the CV (stratified kfold for cell_type) increased by 0.01, the LB also increased by 0.01.\nCV 0.2717 -> LB 0.295\nCV 0.2817 -> LB 0.304",
      "votes": 1,
      "replies": [
        {
          "id": 1581853,
          "postDate": "2021-11-14T09:44:31.773Z",
          "content": "<p>Great correlation!!<br>\nI'll try kfolds (now single fold).</p>",
          "rawMarkdown": "Great correlation!!\nI'll try kfolds (now single fold)."
        },
        {
          "id": 1581889,
          "postDate": "2021-11-14T10:40:11.070Z",
          "content": "<p>The above score is the result of single fold out of 10 folds.</p>",
          "rawMarkdown": "The above score is the result of single fold out of 10 folds.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1578216,
      "postDate": "2021-11-10T18:52:08.353Z",
      "content": "<p>My strategy:<br>\n1- High-risk submission ( Not CV but High in LB)<br>\n2- Low-risk submission ( CV but good in LB)</p>",
      "rawMarkdown": "My strategy:\n1- High-risk submission ( Not CV but High in LB)\n2- Low-risk submission ( CV but good in LB)",
      "votes": 2,
      "replies": [
        {
          "id": 1578763,
          "postDate": "2021-11-11T10:16:54.853Z",
          "content": "<p>This is the strategy I often take.</p>",
          "rawMarkdown": "This is the strategy I often take."
        },
        {
          "id": 1578767,
          "postDate": "2021-11-11T10:20:57.700Z",
          "rawMarkdown": "",
          "votes": 7,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1578434,
      "postDate": "2021-11-11T03:26:37.183Z",
      "content": "<p>Always trust CV!</p>",
      "rawMarkdown": "Always trust CV!",
      "replies": [
        {
          "id": 1578639,
          "postDate": "2021-11-11T08:28:30.307Z",
          "content": "<p>Not always correct, depends on the data setup. Some competitions have different test data than train data and then you might need to trust LB. I have no idea what's the case in this competition.</p>",
          "rawMarkdown": "Not always correct, depends on the data setup. Some competitions have different test data than train data and then you might need to trust LB. I have no idea what's the case in this competition.",
          "votes": 16
        },
        {
          "id": 1578677,
          "postDate": "2021-11-11T09:07:13.177Z",
          "content": "<p>Thnak you <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <br>\nI have learned something new from your answer. </p>",
          "rawMarkdown": "Thnak you @philippsinger \nI have learned something new from your answer. "
        },
        {
          "id": 1578735,
          "postDate": "2021-11-11T09:52:44.710Z",
          "content": "<p>Always trust philippsinger!</p>",
          "rawMarkdown": "Always trust philippsinger!",
          "votes": 16
        },
        {
          "id": 1579012,
          "postDate": "2021-11-11T13:03:09.657Z",
          "content": "<p>I'm very well aware of the uncertainties in this data set.</p>\n<ol>\n<li>The ground truth values in the train set have overlapping instances whereas those in the test set do not. Even though the percentage of them is small (it was reported to be a couple of percent), this is just not OK. This means there will be randomness in the CV/LB relationship. The funny thing is we are not given how the ground truth values of the test set are assigned to mitigate the overlaps.</li>\n<li>The labels are not accurate enough for the given evaluation metric. This creates another uncertainty with the small instances' IoU threshold values.</li>\n</ol>\n<p>I don't see a straightforward way how to construct the cross validation. Once the problems above are addressed one should be able to trust the CV score.</p>\n<p>I have seen some suggestions for the second one. The host needs to use conditional IoU thresholds instead of the entire range of IoU values.<br>\nFor the first one, one can perhaps identify the overlapping regions just by burning some predictions to see the changes in the public LB score.  But then what's gonna happen for the private LB data?</p>",
          "rawMarkdown": "I'm very well aware of the uncertainties in this data set.\n\n1. The ground truth values in the train set have overlapping instances whereas those in the test set do not. Even though the percentage of them is small (it was reported to be a couple of percent), this is just not OK. This means there will be randomness in the CV/LB relationship. The funny thing is we are not given how the ground truth values of the test set are assigned to mitigate the overlaps.\n2. The labels are not accurate enough for the given evaluation metric. This creates another uncertainty with the small instances' IoU threshold values.\n\nI don't see a straightforward way how to construct the cross validation. Once the problems above are addressed one should be able to trust the CV score.\n\nI have seen some suggestions for the second one. The host needs to use conditional IoU thresholds instead of the entire range of IoU values.\nFor the first one, one can perhaps identify the overlapping regions just by burning some predictions to see the changes in the public LB score.  But then what's gonna happen for the private LB data?",
          "votes": 2
        },
        {
          "id": 1579023,
          "postDate": "2021-11-11T13:18:25.823Z",
          "content": "<p>I agree. Shake will happen to some extent, and more with non-random public/private split.<br>\nOne easy way is that we can use more less-overlapping images for validation dataset.</p>",
          "rawMarkdown": "I agree. Shake will happen to some extent, and more with non-random public/private split.\nOne easy way is that we can use more less-overlapping images for validation dataset.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1581774,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-14T08:33:22.257000",
      "content": "<p>what you probably want to do is test out your CV: </p>\n<ol>\n<li>Train multiple [different] models and get various CV scores.</li>\n<li>Submit them all to the LB. </li>\n<li>Plot [scatter] the CV scores you got vs the LB scores you got.</li>\n<li>Optional: measure [spearman / maybe also person] correlation of the LB scores to the CV scores.</li>\n</ol>\n<p>If you find that there is a strong correlation between the LB and the CV, you can trust your CV.</p>\n<p><strong>BUT!!</strong></p>\n<p>Only up to a point.<br>\nUsually, you will find out that as you get to the highest ranks on the leaderboard, your CV-LB correlation starts to lower. from there, you probably should look closely into the data to try and find the samples that are \"not aligned\" or the samples that mess up your models. </p>\n<p>It all depends on the competition itself and the data. sometimes [often] if you did everything correct: you should trust your CV. whoever, there are some situations where coming up with a meaningful CV that won't break until the end of the competition is really really hard. when this is the case: you probably want to treat the LB as \"another fold\". </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1581839,
      "author_name": "Yamame🐟",
      "author_url": "",
      "post_date": "2021-11-14T09:27:37.203000",
      "content": "<p>In my case, ｗhen the CV (stratified kfold for cell_type) increased by 0.01, the LB also increased by 0.01.<br>\nCV 0.2717 -&gt; LB 0.295<br>\nCV 0.2817 -&gt; LB 0.304</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1581853,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-11-14T09:44:31.773000",
          "content": "<p>Great correlation!!<br>\nI'll try kfolds (now single fold).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1581889,
          "author_name": "Yamame🐟",
          "author_url": "",
          "post_date": "2021-11-14T10:40:11.070000",
          "content": "<p>The above score is the result of single fold out of 10 folds.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1578216,
      "author_name": "Faisal Alsrheed",
      "author_url": "",
      "post_date": "2021-11-10T18:52:08.353000",
      "content": "<p>My strategy:<br>\n1- High-risk submission ( Not CV but High in LB)<br>\n2- Low-risk submission ( CV but good in LB)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1578763,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-11-11T10:16:54.853000",
          "content": "<p>This is the strategy I often take.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1578767,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-11-11T10:20:57.700000",
          "content": "",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1578434,
      "author_name": "Tolga",
      "author_url": "",
      "post_date": "2021-11-11T03:26:37.183000",
      "content": "<p>Always trust CV!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1578639,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-11-11T08:28:30.307000",
          "content": "<p>Not always correct, depends on the data setup. Some competitions have different test data than train data and then you might need to trust LB. I have no idea what's the case in this competition.</p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 1578677,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2021-11-11T09:07:13.177000",
          "content": "<p>Thnak you <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <br>\nI have learned something new from your answer. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1578735,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-11-11T09:52:44.710000",
          "content": "<p>Always trust philippsinger!</p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 1579012,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2021-11-11T13:03:09.657000",
          "content": "<p>I'm very well aware of the uncertainties in this data set.</p>\n<ol>\n<li>The ground truth values in the train set have overlapping instances whereas those in the test set do not. Even though the percentage of them is small (it was reported to be a couple of percent), this is just not OK. This means there will be randomness in the CV/LB relationship. The funny thing is we are not given how the ground truth values of the test set are assigned to mitigate the overlaps.</li>\n<li>The labels are not accurate enough for the given evaluation metric. This creates another uncertainty with the small instances' IoU threshold values.</li>\n</ol>\n<p>I don't see a straightforward way how to construct the cross validation. Once the problems above are addressed one should be able to trust the CV score.</p>\n<p>I have seen some suggestions for the second one. The host needs to use conditional IoU thresholds instead of the entire range of IoU values.<br>\nFor the first one, one can perhaps identify the overlapping regions just by burning some predictions to see the changes in the public LB score.  But then what's gonna happen for the private LB data?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1579023,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-11-11T13:18:25.823000",
          "content": "<p>I agree. Shake will happen to some extent, and more with non-random public/private split.<br>\nOne easy way is that we can use more less-overlapping images for validation dataset.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1577647": "Hi. I'm relatively new here at this competition.\n\nAfter reading some discussion, I realize train and test data are different in terms of overlapping.\nIn my experiment, I use stratified groupkfold (n_splits=5), and CV and LB correlation of single fold is not so strong (not weak though).\nConsidering public/private split ratio, I guess it is acceptable to trust LB with paying attention to CV score.\nWhat's your opinion and strategy?\n\nThank you.",
    "1581774": "what you probably want to do is test out your CV: \n1.  Train multiple [different] models and get various CV scores.\n2. Submit them all to the LB. \n3. Plot [scatter] the CV scores you got vs the LB scores you got.\n4. Optional: measure [spearman / maybe also person] correlation of the LB scores to the CV scores.\n\nIf you find that there is a strong correlation between the LB and the CV, you can trust your CV.\n\n\n**BUT!!**\n\nOnly up to a point.\nUsually, you will find out that as you get to the highest ranks on the leaderboard, your CV-LB correlation starts to lower. from there, you probably should look closely into the data to try and find the samples that are \"not aligned\" or the samples that mess up your models. \n\nIt all depends on the competition itself and the data. sometimes [often] if you did everything correct: you should trust your CV. whoever, there are some situations where coming up with a meaningful CV that won't break until the end of the competition is really really hard. when this is the case: you probably want to treat the LB as \"another fold\". ",
    "1581839": "In my case, ｗhen the CV (stratified kfold for cell_type) increased by 0.01, the LB also increased by 0.01.\nCV 0.2717 -> LB 0.295\nCV 0.2817 -> LB 0.304",
    "1578216": "My strategy:\n1- High-risk submission ( Not CV but High in LB)\n2- Low-risk submission ( CV but good in LB)",
    "1578434": "Always trust CV!"
  }
}