{
  "id": 232398,
  "title": "High CV but low LB? Something is wrong?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/232398",
  "author_name": "Yann Majewski",
  "post_date": "2021-04-13T15:14:08.373000",
  "votes": 10,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Im having this issue where I get 0.94 CV on 4 folds but only 0.908 LB. Im pretty sure everything is right and there is no bug in my code/validation strategy. </p>\n<p>Anyone getting these really big gaps?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 1272559,
      "postDate": "2021-04-13T15:14:08.373Z",
      "content": "<p>Im having this issue where I get 0.94 CV on 4 folds but only 0.908 LB. Im pretty sure everything is right and there is no bug in my code/validation strategy. </p>\n<p>Anyone getting these really big gaps?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Im having this issue where I get 0.94 CV on 4 folds but only 0.908 LB. Im pretty sure everything is right and there is no bug in my code/validation strategy. \n\nAnyone getting these really big gaps?\n\nThanks!",
      "votes": 10
    },
    {
      "id": 1276930,
      "postDate": "2021-04-18T07:14:10.080Z",
      "content": "<p>The difference between Public Score and Local score can also be due to differences in evaluation strategies. E.g Pubic score is calculated as avg of DICE of 5 test images. But if you calculate local cv as avg of DICE of individual crops used for cross-validation then many differences can occur. In theory, one must hold out few images entirely from training and consider taking avg of individual dice scores of hold-out images as local cv score.</p>\n<p><img src=\"https://i.ibb.co/68SzbMB/Picture1.png\" alt=\"Case 1\"></p>\n<p><img src=\"https://i.ibb.co/Dt9rH9m/Picture2.png\" alt=\"Case 2\"></p>",
      "rawMarkdown": "The difference between Public Score and Local score can also be due to differences in evaluation strategies. E.g Pubic score is calculated as avg of DICE of 5 test images. But if you calculate local cv as avg of DICE of individual crops used for cross-validation then many differences can occur. In theory, one must hold out few images entirely from training and consider taking avg of individual dice scores of hold-out images as local cv score.\n\n![Case 1](https://i.ibb.co/68SzbMB/Picture1.png)\n\n![Case 2](https://i.ibb.co/Dt9rH9m/Picture2.png)",
      "votes": 4,
      "replies": [
        {
          "id": 1277298,
          "postDate": "2021-04-18T15:45:00.767Z",
          "content": "<p>Thanks for this explaination, and yes my CV score is based on putting entire images (whole scan) in the validation so it cant be some type of leak.</p>",
          "rawMarkdown": "Thanks for this explaination, and yes my CV score is based on putting entire images (whole scan) in the validation so it cant be some type of leak.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1273038,
      "postDate": "2021-04-14T03:31:29.483Z",
      "content": "<p>We have such kind of score gap indeed.</p>",
      "rawMarkdown": "We have such kind of score gap indeed.",
      "votes": 1,
      "replies": [
        {
          "id": 1273600,
          "postDate": "2021-04-14T13:46:28.580Z",
          "content": "<p>Interesting.. Do you mean that your CV is that much higher that you can get the LB  that you got on the public test set? or just some submissions with high CV get low lb?</p>",
          "rawMarkdown": "Interesting.. Do you mean that your CV is that much higher that you can get the LB  that you got on the public test set? or just some submissions with high CV get low lb?"
        },
        {
          "id": 1275398,
          "postDate": "2021-04-16T09:10:25.977Z",
          "content": "<p>Yes, CV &gt; 0.94 and LB &lt; 0.93 without hand-labeling.</p>",
          "rawMarkdown": "Yes, CV > 0.94 and LB < 0.93 without hand-labeling."
        },
        {
          "id": 1275457,
          "postDate": "2021-04-16T10:10:09.540Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1272642,
      "postDate": "2021-04-13T16:34:03.370Z",
      "content": "<p>It seems to be a problem with everybody here, people are saying that its better to trust your CV than the LB.</p>\n<p>In my case, i'm getting .93+ on CV and currently at .915 at LB, are you ensembling these models for submission?</p>",
      "rawMarkdown": "It seems to be a problem with everybody here, people are saying that its better to trust your CV than the LB.\n\nIn my case, i'm getting .93+ on CV and currently at .915 at LB, are you ensembling these models for submission?",
      "votes": 1,
      "replies": [
        {
          "id": 1272653,
          "postDate": "2021-04-13T16:48:59.940Z",
          "content": "<p>Yeah reading other posts it seems there is some difference in the train dataset and public test set (maybe?). This score is just a single model (4 folds), i will eventually ensemble a couple of them to see how it perferms.</p>",
          "rawMarkdown": "Yeah reading other posts it seems there is some difference in the train dataset and public test set (maybe?). This score is just a single model (4 folds), i will eventually ensemble a couple of them to see how it perferms."
        }
      ]
    },
    {
      "id": 1283539,
      "postDate": "2021-04-25T03:11:54.263Z",
      "content": "<p>It seems to be a problem</p>",
      "rawMarkdown": "It seems to be a problem"
    },
    {
      "id": 1278342,
      "postDate": "2021-04-19T19:12:20.423Z",
      "content": "<p>Suffering from the same problem. Maybe the way, we use the rastario for patch extraction from test data causes the problem. I don't know exactly. I tried several models with various augmentation changes, the gap is always there. My best model achieved 0.944 on CV and even performed better than other models on Fold-3 which has some difficult images compared to the rest, but it got an LB of 0.915 (~3.0 % difference).</p>",
      "rawMarkdown": "Suffering from the same problem. Maybe the way, we use the rastario for patch extraction from test data causes the problem. I don't know exactly. I tried several models with various augmentation changes, the gap is always there. My best model achieved 0.944 on CV and even performed better than other models on Fold-3 which has some difficult images compared to the rest, but it got an LB of 0.915 (~3.0 % difference).\n\n ",
      "replies": [
        {
          "id": 1278427,
          "postDate": "2021-04-19T21:54:52.300Z",
          "content": "<p>Yep weird stuff for me as well, 0.94 CV and 0.911 LB</p>",
          "rawMarkdown": "Yep weird stuff for me as well, 0.94 CV and 0.911 LB"
        },
        {
          "id": 1278442,
          "postDate": "2021-04-19T22:33:23.363Z",
          "content": "<p>It seems there is almost a consistent 2-3% drop in CV. How did you improve your CV to the levels of 94%?</p>",
          "rawMarkdown": "It seems there is almost a consistent 2-3% drop in CV. How did you improve your CV to the levels of 94%?"
        }
      ]
    },
    {
      "id": 1276153,
      "postDate": "2021-04-17T07:15:36.740Z",
      "content": "<p>How do you perform your cv split? </p>\n<p>We do 5 folds with 12 training and 3 validation examples in each fold and the gap is much smaller.</p>",
      "rawMarkdown": "How do you perform your cv split? \n\nWe do 5 folds with 12 training and 3 validation examples in each fold and the gap is much smaller.",
      "replies": [
        {
          "id": 1276315,
          "postDate": "2021-04-17T11:42:29.947Z",
          "content": "<p>Was this just a KFold split or Stratified?</p>",
          "rawMarkdown": "Was this just a KFold split or Stratified?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1276930,
      "author_name": "makeitsimple",
      "author_url": "",
      "post_date": "2021-04-18T07:14:10.080000",
      "content": "<p>The difference between Public Score and Local score can also be due to differences in evaluation strategies. E.g Pubic score is calculated as avg of DICE of 5 test images. But if you calculate local cv as avg of DICE of individual crops used for cross-validation then many differences can occur. In theory, one must hold out few images entirely from training and consider taking avg of individual dice scores of hold-out images as local cv score.</p>\n<p><img src=\"https://i.ibb.co/68SzbMB/Picture1.png\" alt=\"Case 1\"></p>\n<p><img src=\"https://i.ibb.co/Dt9rH9m/Picture2.png\" alt=\"Case 2\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1277298,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-04-18T15:45:00.767000",
          "content": "<p>Thanks for this explaination, and yes my CV score is based on putting entire images (whole scan) in the validation so it cant be some type of leak.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1273038,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-04-14T03:31:29.483000",
      "content": "<p>We have such kind of score gap indeed.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1273600,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-04-14T13:46:28.580000",
          "content": "<p>Interesting.. Do you mean that your CV is that much higher that you can get the LB  that you got on the public test set? or just some submissions with high CV get low lb?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1275398,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-04-16T09:10:25.977000",
          "content": "<p>Yes, CV &gt; 0.94 and LB &lt; 0.93 without hand-labeling.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1275457,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-16T10:10:09.540000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1272642,
      "author_name": "victorasso",
      "author_url": "",
      "post_date": "2021-04-13T16:34:03.370000",
      "content": "<p>It seems to be a problem with everybody here, people are saying that its better to trust your CV than the LB.</p>\n<p>In my case, i'm getting .93+ on CV and currently at .915 at LB, are you ensembling these models for submission?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1272653,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-04-13T16:48:59.940000",
          "content": "<p>Yeah reading other posts it seems there is some difference in the train dataset and public test set (maybe?). This score is just a single model (4 folds), i will eventually ensemble a couple of them to see how it perferms.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1283539,
      "author_name": "Mark Mario",
      "author_url": "",
      "post_date": "2021-04-25T03:11:54.263000",
      "content": "<p>It seems to be a problem</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1278342,
      "author_name": "Aramis Vesal",
      "author_url": "",
      "post_date": "2021-04-19T19:12:20.423000",
      "content": "<p>Suffering from the same problem. Maybe the way, we use the rastario for patch extraction from test data causes the problem. I don't know exactly. I tried several models with various augmentation changes, the gap is always there. My best model achieved 0.944 on CV and even performed better than other models on Fold-3 which has some difficult images compared to the rest, but it got an LB of 0.915 (~3.0 % difference).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1278427,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-04-19T21:54:52.300000",
          "content": "<p>Yep weird stuff for me as well, 0.94 CV and 0.911 LB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278442,
          "author_name": "Andrew Shao",
          "author_url": "",
          "post_date": "2021-04-19T22:33:23.363000",
          "content": "<p>It seems there is almost a consistent 2-3% drop in CV. How did you improve your CV to the levels of 94%?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1276153,
      "author_name": "Philipp Sodmann",
      "author_url": "",
      "post_date": "2021-04-17T07:15:36.740000",
      "content": "<p>How do you perform your cv split? </p>\n<p>We do 5 folds with 12 training and 3 validation examples in each fold and the gap is much smaller.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1276315,
          "author_name": "Andrew Shao",
          "author_url": "",
          "post_date": "2021-04-17T11:42:29.947000",
          "content": "<p>Was this just a KFold split or Stratified?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1272559": "Im having this issue where I get 0.94 CV on 4 folds but only 0.908 LB. Im pretty sure everything is right and there is no bug in my code/validation strategy. \n\nAnyone getting these really big gaps?\n\nThanks!",
    "1276930": "The difference between Public Score and Local score can also be due to differences in evaluation strategies. E.g Pubic score is calculated as avg of DICE of 5 test images. But if you calculate local cv as avg of DICE of individual crops used for cross-validation then many differences can occur. In theory, one must hold out few images entirely from training and consider taking avg of individual dice scores of hold-out images as local cv score.\n\n![Case 1](https://i.ibb.co/68SzbMB/Picture1.png)\n\n![Case 2](https://i.ibb.co/Dt9rH9m/Picture2.png)",
    "1273038": "We have such kind of score gap indeed.",
    "1272642": "It seems to be a problem with everybody here, people are saying that its better to trust your CV than the LB.\n\nIn my case, i'm getting .93+ on CV and currently at .915 at LB, are you ensembling these models for submission?",
    "1283539": "It seems to be a problem",
    "1278342": "Suffering from the same problem. Maybe the way, we use the rastario for patch extraction from test data causes the problem. I don't know exactly. I tried several models with various augmentation changes, the gap is always there. My best model achieved 0.944 on CV and even performed better than other models on Fold-3 which has some difficult images compared to the rest, but it got an LB of 0.915 (~3.0 % difference).\n\n ",
    "1276153": "How do you perform your cv split? \n\nWe do 5 folds with 12 training and 3 validation examples in each fold and the gap is much smaller."
  }
}