{
  "id": 225138,
  "title": "What's your CV and LB?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/225138",
  "author_name": "Shihao Shao",
  "post_date": "2021-03-11T01:09:16.614000",
  "votes": 19,
  "comment_count": 49,
  "views": 0,
  "content": "<p>I noticed that several competitions has this thread so I created it. Hope it would help us judge the relevance between CV and LB.</p>\n<p>In my case, (NEW DATASET)<br>\nCV 0.9187<br>\nLB 0.905</p>\n<p>Updated, (OLD DATASET, MODEL STRUCTURE IS SAME AS THE PREVIOUS ONE)<br>\nCV 0.892<br>\nLB 0.915 (It's interesting that with old dataset, same model structure, I reached a higher score)</p>\n<p>Updated, (NEW DATASET)<br>\nCV 0.9260<br>\nLB 0.920</p>\n<p>========<br>\nSo as you know, hand labeling in public test dataset is permitted, which may mislead the LB score. So I think the CV we post here may be important to judge our ranking.</p>",
  "messages": [
    {
      "id": 1234129,
      "postDate": "2021-03-11T01:09:16.613Z",
      "content": "<p>I noticed that several competitions has this thread so I created it. Hope it would help us judge the relevance between CV and LB.</p>\n<p>In my case, (NEW DATASET)<br>\nCV 0.9187<br>\nLB 0.905</p>\n<p>Updated, (OLD DATASET, MODEL STRUCTURE IS SAME AS THE PREVIOUS ONE)<br>\nCV 0.892<br>\nLB 0.915 (It's interesting that with old dataset, same model structure, I reached a higher score)</p>\n<p>Updated, (NEW DATASET)<br>\nCV 0.9260<br>\nLB 0.920</p>\n<p>========<br>\nSo as you know, hand labeling in public test dataset is permitted, which may mislead the LB score. So I think the CV we post here may be important to judge our ranking.</p>",
      "rawMarkdown": "I noticed that several competitions has this thread so I created it. Hope it would help us judge the relevance between CV and LB.\n\nIn my case, (NEW DATASET)\nCV 0.9187\nLB 0.905\n\nUpdated, (OLD DATASET, MODEL STRUCTURE IS SAME AS THE PREVIOUS ONE)\nCV 0.892\nLB 0.915 (It's interesting that with old dataset, same model structure, I reached a higher score)\n\nUpdated, (NEW DATASET)\nCV 0.9260\nLB 0.920\n\n\n========\nSo as you know, hand labeling in public test dataset is permitted, which may mislead the LB score. So I think the CV we post here may be important to judge our ranking.",
      "votes": 18
    },
    {
      "id": 1247689,
      "postDate": "2021-03-22T00:53:12.857Z",
      "content": "<p>[new data]<br>\nresnet34-unet<br>\ncv0.931, lb0.924</p>\n<p>update:<br>\nmodify postprocess: 0.925 -&gt; 0.929</p>",
      "rawMarkdown": "[new data]\nresnet34-unet\ncv0.931, lb0.924\n\nupdate:\nmodify postprocess: 0.925 -> 0.929",
      "votes": 7,
      "replies": [
        {
          "id": 1247698,
          "postDate": "2021-03-22T01:25:44.573Z",
          "content": "<p>High score with simple model! Awesome! Did you try much bigger backbones and test the performances? I used efficientnetb7 but only got 0.923 (even with time-consuming postprocess).</p>",
          "rawMarkdown": "High score with simple model! Awesome! Did you try much bigger backbones and test the performances? I used efficientnetb7 but only got 0.923 (even with time-consuming postprocess).",
          "votes": 1
        },
        {
          "id": 1247701,
          "postDate": "2021-03-22T01:29:31.477Z",
          "content": "<p>Bigger model didn't improve validation score, so I didn't submit.</p>",
          "rawMarkdown": "Bigger model didn't improve validation score, so I didn't submit.",
          "votes": 1
        },
        {
          "id": 1247967,
          "postDate": "2021-03-22T08:22:16.770Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a>, i also use resnet34-unet, and cv get 0.934+, but lb just get 0.916, so i want to ask what's your train image size, and submit input image size?</p>",
          "rawMarkdown": "Hi @phalanx, i also use resnet34-unet, and cv get 0.934+, but lb just get 0.916, so i want to ask what's your train image size, and submit input image size?"
        },
        {
          "id": 1248730,
          "postDate": "2021-03-22T19:41:39.140Z",
          "content": "<p>For how many epochs did you train your model?</p>",
          "rawMarkdown": "For how many epochs did you train your model?"
        },
        {
          "id": 1249076,
          "postDate": "2021-03-23T04:48:19.347Z",
          "content": "<p><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> what is your cv-strategy? Kfold? GroupKFold?</p>",
          "rawMarkdown": "@phalanx what is your cv-strategy? Kfold? GroupKFold?"
        },
        {
          "id": 1249206,
          "postDate": "2021-03-23T07:01:03.893Z",
          "content": "<p>bigger model costs too much time to deal. I am trying resnet34 recently with 512 and it has got only LB 0.915 so far. </p>",
          "rawMarkdown": "bigger model costs too much time to deal. I am trying resnet34 recently with 512 and it has got only LB 0.915 so far. "
        }
      ]
    },
    {
      "id": 1234280,
      "postDate": "2021-03-11T05:24:13.053Z",
      "content": "<p>CV: 0.873<br>\nLB: 0.921</p>\n<p>I haven't re-trained the model yet with the new dataset </p>",
      "rawMarkdown": "CV: 0.873\nLB: 0.921\n\nI haven't re-trained the model yet with the new dataset ",
      "votes": 3,
      "replies": [
        {
          "id": 1234575,
          "postDate": "2021-03-11T11:27:22.393Z",
          "rawMarkdown": "",
          "votes": -3,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1236170,
      "postDate": "2021-03-12T20:39:42.347Z",
      "content": "<p>Resubmitted the best single model trained on the old data without retraining: CV was on old data 0.94 and new LB 0.926 (old LB 0.873).</p>",
      "rawMarkdown": "Resubmitted the best single model trained on the old data without retraining: CV was on old data 0.94 and new LB 0.926 (old LB 0.873).",
      "votes": 4,
      "replies": [
        {
          "id": 1264240,
          "postDate": "2021-04-06T02:31:45.217Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1235102,
      "postDate": "2021-03-11T20:20:53.703Z",
      "content": "<p>Single fold (12:3) retrained with new data<br>\nLocal Dice: 0.942 (on crops)<br>\npublic LB: 0.922<br>\nIt might be an easy fold for validation</p>",
      "rawMarkdown": "Single fold (12:3) retrained with new data\nLocal Dice: 0.942 (on crops)\npublic LB: 0.922\nIt might be an easy fold for validation",
      "votes": 4,
      "replies": [
        {
          "id": 1242061,
          "postDate": "2021-03-17T11:28:16.527Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1256216,
      "postDate": "2021-03-29T16:53:24.433Z",
      "content": "<p>Ecaresnet50-unet++<br>\n1024x1024<br>\nCV 0.9301<br>\nLB 0.903</p>\n<p>Im not sure why i have such a big difference between CV and LB </p>",
      "rawMarkdown": "Ecaresnet50-unet++\n1024x1024\nCV 0.9301\nLB 0.903\n\nIm not sure why i have such a big difference between CV and LB ",
      "votes": 1
    },
    {
      "id": 1255474,
      "postDate": "2021-03-28T20:46:46.593Z",
      "content": "<p>[new data] - no pseudo or handlabeling….yet…<br>\nEffNet-Unet <br>\nCV: 0.905<br>\nLB: 0.917</p>",
      "rawMarkdown": "[new data] - no pseudo or handlabeling....yet...\nEffNet-Unet \nCV: 0.905\nLB: 0.917",
      "votes": 1
    },
    {
      "id": 1234495,
      "postDate": "2021-03-11T09:51:35.883Z",
      "content": "<p>do you see correlation between CV and LB ?</p>",
      "rawMarkdown": "do you see correlation between CV and LB ?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1234612,
          "postDate": "2021-03-11T12:10:48.003Z",
          "content": "<p>Actually not that much.</p>",
          "rawMarkdown": "Actually not that much.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1238596,
      "postDate": "2021-03-15T06:13:25.400Z",
      "content": "<p>resubmit old tensorflow model:<br>\nCV:  0.8921<br>\nLB: 0.910</p>\n<p>=======================<br>\nI used pytorch to train new data<br>\nCV: 0.9112<br>\nLB: 0.869<br>\nWhy did I get this huge gap between cv and lb when using pytorch…</p>",
      "rawMarkdown": "resubmit old tensorflow model:\nCV:  0.8921\nLB: 0.910\n\n=======================\nI used pytorch to train new data\nCV: 0.9112\nLB: 0.869\nWhy did I get this huge gap between cv and lb when using pytorch...\n",
      "votes": 2,
      "replies": [
        {
          "id": 1264241,
          "postDate": "2021-04-06T02:32:25.230Z",
          "content": "<p>I encountered a similar problem, the gap between cv and lb is very big, how did you solve it?</p>",
          "rawMarkdown": "I encountered a similar problem, the gap between cv and lb is very big, how did you solve it?"
        }
      ]
    },
    {
      "id": 1234282,
      "postDate": "2021-03-11T05:26:48.763Z",
      "content": "<p>I used previous model with:<br>\nCV：0.890, pre-LB: 0.860<br>\nLB：0.914</p>",
      "rawMarkdown": "I used previous model with:\nCV：0.890, pre-LB: 0.860\nLB：0.914",
      "votes": 2,
      "replies": [
        {
          "id": 1253778,
          "postDate": "2021-03-27T03:08:36.500Z",
          "content": "<p>update: cv: 0.935 lb:0.920</p>",
          "rawMarkdown": "update: cv: 0.935 lb:0.920"
        }
      ]
    },
    {
      "id": 1262319,
      "postDate": "2021-04-04T05:44:49.847Z",
      "content": "<p><code>So as you know, hand labeling in public test dataset is permitted</code><br>\nReally?</p>",
      "rawMarkdown": "```So as you know, hand labeling in public test dataset is permitted```\nReally?",
      "replies": [
        {
          "id": 1262372,
          "postDate": "2021-04-04T07:19:50.453Z",
          "content": "<p>Indeed - see <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1250442\" target=\"_blank\">this post by Addison Howard</a><br>\nand rule changes A.3 supercedes B.5<br>\n\"Submissions may use or incorporate information from hand labeling or human prediction of the validation dataset or public test data records.\"</p>",
          "rawMarkdown": "Indeed - see [this post by Addison Howard](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1250442)\nand rule changes A.3 supercedes B.5\n\"Submissions may use or incorporate information from hand labeling or human prediction of the validation dataset or public test data records.\"",
          "votes": 2
        },
        {
          "id": 1262411,
          "postDate": "2021-04-04T08:30:16.377Z",
          "content": "<p>Thanks for the information. That's disheartening for many, as the leaderboard is not representative now.</p>",
          "rawMarkdown": "Thanks for the information. That's disheartening for many, as the leaderboard is not representative now.",
          "votes": 1
        },
        {
          "id": 1262428,
          "postDate": "2021-04-04T08:54:59.833Z",
          "content": "<p>Based on other comments posted, there can be differences in samples that are fresh frozen or Formalin Fixed Paraffin Embedded (FFPE).  The hosts said that no sclerosed glomeruli are annotated in train, public test, or private test. So what is in d488c759a is uncertain that benefits from hand labels, but trying to get models to generalise well, maybe via external data like <a href=\"https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024\" target=\"_blank\">this</a>, could be more informative than the LB. </p>",
          "rawMarkdown": "Based on other comments posted, there can be differences in samples that are fresh frozen or Formalin Fixed Paraffin Embedded (FFPE).  The hosts said that no sclerosed glomeruli are annotated in train, public test, or private test. So what is in d488c759a is uncertain that benefits from hand labels, but trying to get models to generalise well, maybe via external data like [this](https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024), could be more informative than the LB. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257339,
      "postDate": "2021-03-30T18:26:52.350Z",
      "content": "<p>latest: CV: 0.921, LB:0.927</p>",
      "rawMarkdown": "latest: CV: 0.921, LB:0.927"
    },
    {
      "id": 1246550,
      "postDate": "2021-03-20T20:49:55.540Z",
      "content": "<p>I don't really understand this.</p>\n<p>Model 1:<br>\nCV: 0.930<br>\nLB: 0.907</p>\n<p>Model 2:<br>\nCV: 0.926<br>\nLB: 0.912</p>\n<p>Validation scheme: full resolution masks (instead of mean of tiles) with 5 folds (12-3)</p>",
      "rawMarkdown": "I don't really understand this.\n\nModel 1:\nCV: 0.930\nLB: 0.907\n\nModel 2:\nCV: 0.926\nLB: 0.912\n\nValidation scheme: full resolution masks (instead of mean of tiles) with 5 folds (12-3)",
      "replies": [
        {
          "id": 1247238,
          "postDate": "2021-03-21T15:02:20.217Z",
          "content": "<p>I also faced issue like yours. I think we must find some ways to make our results stable. </p>",
          "rawMarkdown": "I also faced issue like yours. I think we must find some ways to make our results stable. "
        }
      ]
    },
    {
      "id": 1244690,
      "postDate": "2021-03-19T06:41:00.320Z",
      "content": "<p>After new data, can anyone confirm that CV and LB are correlated? </p>",
      "rawMarkdown": "After new data, can anyone confirm that CV and LB are correlated? ",
      "replies": [
        {
          "id": 1245257,
          "postDate": "2021-03-19T16:13:06.963Z",
          "content": "<p>actually not so correlated so far here. Same CV scores (within 0.001) lead to variations up to 0.005 on the LB for the few tests I've made.</p>",
          "rawMarkdown": "actually not so correlated so far here. Same CV scores (within 0.001) lead to variations up to 0.005 on the LB for the few tests I've made."
        },
        {
          "id": 1245758,
          "postDate": "2021-03-20T05:47:19.187Z",
          "content": "<p>I see, but if the distribution of train, test and the private test is more or less same then this might not be a big problem I suppose. But still it is risky LOL </p>",
          "rawMarkdown": "I see, but if the distribution of train, test and the private test is more or less same then this might not be a big problem I suppose. But still it is risky LOL "
        },
        {
          "id": 1247220,
          "postDate": "2021-03-21T14:37:13.263Z",
          "content": "<p>In a version of submission, I got a boost of about 0.005 but got -0.007 in LB</p>",
          "rawMarkdown": "In a version of submission, I got a boost of about 0.005 but got -0.007 in LB",
          "votes": 1
        },
        {
          "id": 1249077,
          "postDate": "2021-03-23T04:56:26.787Z",
          "content": "<p>I think that must be due to some sort of difference in the distribution of training, and public test data. So mostly there will be some sort of shake up in this competition. </p>",
          "rawMarkdown": "I think that must be due to some sort of difference in the distribution of training, and public test data. So mostly there will be some sort of shake up in this competition. "
        }
      ]
    },
    {
      "id": 1237970,
      "postDate": "2021-03-14T14:36:49.720Z",
      "content": "<p>CV 0.944<br>\nLB 0.924 (old model 0.920)</p>\n<p>The new data provides a big boost.</p>",
      "rawMarkdown": "CV 0.944\nLB 0.924 (old model 0.920)\n\nThe new data provides a big boost.",
      "replies": [
        {
          "id": 1287665,
          "postDate": "2021-04-29T09:44:48.343Z",
          "content": "<p>Oh，great job！ How can you do it?</p>",
          "rawMarkdown": "Oh，great job！ How can you do it?"
        }
      ]
    },
    {
      "id": 1237579,
      "postDate": "2021-03-14T09:21:41.630Z",
      "content": "<p>My mean oof cv at whole slide dice is about 0.933 and my lb is 0.910</p>",
      "rawMarkdown": "My mean oof cv at whole slide dice is about 0.933 and my lb is 0.910"
    },
    {
      "id": 1236837,
      "postDate": "2021-03-13T13:51:59.533Z",
      "content": "<p>Resubmitted old model with single fold, I get 0.920 on LB</p>",
      "rawMarkdown": "Resubmitted old model with single fold, I get 0.920 on LB"
    },
    {
      "id": 1236316,
      "postDate": "2021-03-13T02:15:39.787Z",
      "content": "<p>Did anyone find whether the new dataset, training the same model, performs better than the old one?</p>",
      "rawMarkdown": "Did anyone find whether the new dataset, training the same model, performs better than the old one?",
      "replies": [
        {
          "id": 1236597,
          "postDate": "2021-03-13T09:37:01.947Z",
          "content": "<p>I found a strange thing that old model performs nearly the same as new models with Effnet-4.  Did you use 256 size samples or others? I found different sizes will cause the differences of score between two models various.</p>",
          "rawMarkdown": "I found a strange thing that old model performs nearly the same as new models with Effnet-4.  Did you use 256 size samples or others? I found different sizes will cause the differences of score between two models various."
        },
        {
          "id": 1236857,
          "postDate": "2021-03-13T14:10:08.950Z",
          "content": "<p>I only use 512 for training for now. I found that different model suit for different datasest. Also, Did you try 512 images? Did it perform stable in both dataset in your case? In my case, the scores can be various and unstable.</p>",
          "rawMarkdown": "I only use 512 for training for now. I found that different model suit for different datasest. Also, Did you try 512 images? Did it perform stable in both dataset in your case? In my case, the scores can be various and unstable."
        },
        {
          "id": 1236928,
          "postDate": "2021-03-13T15:01:09.090Z",
          "content": "<p>I just tried 128 and 256 with comparing new and old datas. 256’s differences between old and new is bigger than 128’s. However, the score is not stable and some variables are not the same, it can’t prove much regularity. </p>",
          "rawMarkdown": "I just tried 128 and 256 with comparing new and old datas. 256’s differences between old and new is bigger than 128’s. However, the score is not stable and some variables are not the same, it can’t prove much regularity. ",
          "votes": 1
        },
        {
          "id": 1237030,
          "postDate": "2021-03-13T17:03:17.083Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1237272,
          "postDate": "2021-03-14T01:13:31.073Z",
          "content": "<p>yes, same here! I get better LB score with the old model. Is it possible some of the test images may be \"familiar\" to the old model? I cannot find other explanation. This topic may worth a separate post</p>",
          "rawMarkdown": "yes, same here! I get better LB score with the old model. Is it possible some of the test images may be \"familiar\" to the old model? I cannot find other explanation. This topic may worth a separate post",
          "votes": 1
        }
      ]
    },
    {
      "id": 1234210,
      "postDate": "2021-03-11T04:00:38.570Z",
      "content": "<p>is this score on the updated dataset？</p>",
      "rawMarkdown": "is this score on the updated dataset？",
      "replies": [
        {
          "id": 1234254,
          "postDate": "2021-03-11T04:45:59.480Z",
          "content": "<p>Yes. Using all of the new dataset. Single fold.</p>",
          "rawMarkdown": "Yes. Using all of the new dataset. Single fold."
        }
      ]
    },
    {
      "id": 1235414,
      "postDate": "2021-03-12T05:57:42.710Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1235548,
          "postDate": "2021-03-12T08:29:29.413Z",
          "content": "<p>In my case, some TTAs like contrast adjustment will ruin the performance. So I just try rotate and flip for now. Feel free to publish a public notebook you submitted :D. Maybe people will have proper advices for your code.</p>",
          "rawMarkdown": "In my case, some TTAs like contrast adjustment will ruin the performance. So I just try rotate and flip for now. Feel free to publish a public notebook you submitted :D. Maybe people will have proper advices for your code."
        },
        {
          "id": 1236343,
          "postDate": "2021-03-13T03:09:52Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1250727,
          "postDate": "2021-03-24T08:29:29.660Z",
          "content": "<p>hi fury,same as you,have you find any way to slove this problem?</p>",
          "rawMarkdown": "hi fury,same as you,have you find any way to slove this problem?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1247689,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2021-03-22T00:53:12.857000",
      "content": "<p>[new data]<br>\nresnet34-unet<br>\ncv0.931, lb0.924</p>\n<p>update:<br>\nmodify postprocess: 0.925 -&gt; 0.929</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1247698,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-22T01:25:44.573000",
          "content": "<p>High score with simple model! Awesome! Did you try much bigger backbones and test the performances? I used efficientnetb7 but only got 0.923 (even with time-consuming postprocess).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1247701,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2021-03-22T01:29:31.477000",
          "content": "<p>Bigger model didn't improve validation score, so I didn't submit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1247967,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2021-03-22T08:22:16.770000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a>, i also use resnet34-unet, and cv get 0.934+, but lb just get 0.916, so i want to ask what's your train image size, and submit input image size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248730,
          "author_name": "Rohil Pal",
          "author_url": "",
          "post_date": "2021-03-22T19:41:39.140000",
          "content": "<p>For how many epochs did you train your model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1249076,
          "author_name": "FGPC",
          "author_url": "",
          "post_date": "2021-03-23T04:48:19.347000",
          "content": "<p><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> what is your cv-strategy? Kfold? GroupKFold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1249206,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-23T07:01:03.893000",
          "content": "<p>bigger model costs too much time to deal. I am trying resnet34 recently with 512 and it has got only LB 0.915 so far. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1234280,
      "author_name": "andras",
      "author_url": "",
      "post_date": "2021-03-11T05:24:13.053000",
      "content": "<p>CV: 0.873<br>\nLB: 0.921</p>\n<p>I haven't re-trained the model yet with the new dataset </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1234575,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-11T11:27:22.393000",
          "content": "",
          "votes": -3,
          "replies": []
        }
      ]
    },
    {
      "id": 1236170,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "2021-03-12T20:39:42.347000",
      "content": "<p>Resubmitted the best single model trained on the old data without retraining: CV was on old data 0.94 and new LB 0.926 (old LB 0.873).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1264240,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-06T02:31:45.217000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235102,
      "author_name": "Darkate",
      "author_url": "",
      "post_date": "2021-03-11T20:20:53.703000",
      "content": "<p>Single fold (12:3) retrained with new data<br>\nLocal Dice: 0.942 (on crops)<br>\npublic LB: 0.922<br>\nIt might be an easy fold for validation</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1242061,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-17T11:28:16.527000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1256216,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2021-03-29T16:53:24.433000",
      "content": "<p>Ecaresnet50-unet++<br>\n1024x1024<br>\nCV 0.9301<br>\nLB 0.903</p>\n<p>Im not sure why i have such a big difference between CV and LB </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1255474,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2021-03-28T20:46:46.593000",
      "content": "<p>[new data] - no pseudo or handlabeling….yet…<br>\nEffNet-Unet <br>\nCV: 0.905<br>\nLB: 0.917</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1234495,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2021-03-11T09:51:35.883000",
      "content": "<p>do you see correlation between CV and LB ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1234612,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-11T12:10:48.003000",
          "content": "<p>Actually not that much.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1238596,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2021-03-15T06:13:25.400000",
      "content": "<p>resubmit old tensorflow model:<br>\nCV:  0.8921<br>\nLB: 0.910</p>\n<p>=======================<br>\nI used pytorch to train new data<br>\nCV: 0.9112<br>\nLB: 0.869<br>\nWhy did I get this huge gap between cv and lb when using pytorch…</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1264241,
          "author_name": "想躺平",
          "author_url": "",
          "post_date": "2021-04-06T02:32:25.230000",
          "content": "<p>I encountered a similar problem, the gap between cv and lb is very big, how did you solve it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1234282,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-03-11T05:26:48.763000",
      "content": "<p>I used previous model with:<br>\nCV：0.890, pre-LB: 0.860<br>\nLB：0.914</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1253778,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-27T03:08:36.500000",
          "content": "<p>update: cv: 0.935 lb:0.920</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1262319,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-04-04T05:44:49.847000",
      "content": "<p><code>So as you know, hand labeling in public test dataset is permitted</code><br>\nReally?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1262372,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2021-04-04T07:19:50.453000",
          "content": "<p>Indeed - see <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1250442\" target=\"_blank\">this post by Addison Howard</a><br>\nand rule changes A.3 supercedes B.5<br>\n\"Submissions may use or incorporate information from hand labeling or human prediction of the validation dataset or public test data records.\"</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1262411,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2021-04-04T08:30:16.377000",
          "content": "<p>Thanks for the information. That's disheartening for many, as the leaderboard is not representative now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1262428,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2021-04-04T08:54:59.833000",
          "content": "<p>Based on other comments posted, there can be differences in samples that are fresh frozen or Formalin Fixed Paraffin Embedded (FFPE).  The hosts said that no sclerosed glomeruli are annotated in train, public test, or private test. So what is in d488c759a is uncertain that benefits from hand labels, but trying to get models to generalise well, maybe via external data like <a href=\"https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024\" target=\"_blank\">this</a>, could be more informative than the LB. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257339,
      "author_name": "andras",
      "author_url": "",
      "post_date": "2021-03-30T18:26:52.350000",
      "content": "<p>latest: CV: 0.921, LB:0.927</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1246550,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2021-03-20T20:49:55.540000",
      "content": "<p>I don't really understand this.</p>\n<p>Model 1:<br>\nCV: 0.930<br>\nLB: 0.907</p>\n<p>Model 2:<br>\nCV: 0.926<br>\nLB: 0.912</p>\n<p>Validation scheme: full resolution masks (instead of mean of tiles) with 5 folds (12-3)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1247238,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-21T15:02:20.217000",
          "content": "<p>I also faced issue like yours. I think we must find some ways to make our results stable. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1244690,
      "author_name": "Urvish",
      "author_url": "",
      "post_date": "2021-03-19T06:41:00.320000",
      "content": "<p>After new data, can anyone confirm that CV and LB are correlated? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1245257,
          "author_name": "FabienDaniel",
          "author_url": "",
          "post_date": "2021-03-19T16:13:06.963000",
          "content": "<p>actually not so correlated so far here. Same CV scores (within 0.001) lead to variations up to 0.005 on the LB for the few tests I've made.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1245758,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2021-03-20T05:47:19.187000",
          "content": "<p>I see, but if the distribution of train, test and the private test is more or less same then this might not be a big problem I suppose. But still it is risky LOL </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1247220,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-21T14:37:13.263000",
          "content": "<p>In a version of submission, I got a boost of about 0.005 but got -0.007 in LB</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1249077,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2021-03-23T04:56:26.787000",
          "content": "<p>I think that must be due to some sort of difference in the distribution of training, and public test data. So mostly there will be some sort of shake up in this competition. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1237970,
      "author_name": "gakki",
      "author_url": "",
      "post_date": "2021-03-14T14:36:49.720000",
      "content": "<p>CV 0.944<br>\nLB 0.924 (old model 0.920)</p>\n<p>The new data provides a big boost.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1287665,
          "author_name": "LuzhangCV",
          "author_url": "",
          "post_date": "2021-04-29T09:44:48.343000",
          "content": "<p>Oh，great job！ How can you do it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1237579,
      "author_name": "Yineng Xiong",
      "author_url": "",
      "post_date": "2021-03-14T09:21:41.630000",
      "content": "<p>My mean oof cv at whole slide dice is about 0.933 and my lb is 0.910</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1236837,
      "author_name": "He",
      "author_url": "",
      "post_date": "2021-03-13T13:51:59.533000",
      "content": "<p>Resubmitted old model with single fold, I get 0.920 on LB</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1236316,
      "author_name": "Shihao Shao",
      "author_url": "",
      "post_date": "2021-03-13T02:15:39.787000",
      "content": "<p>Did anyone find whether the new dataset, training the same model, performs better than the old one?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1236597,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-13T09:37:01.947000",
          "content": "<p>I found a strange thing that old model performs nearly the same as new models with Effnet-4.  Did you use 256 size samples or others? I found different sizes will cause the differences of score between two models various.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1236857,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-13T14:10:08.950000",
          "content": "<p>I only use 512 for training for now. I found that different model suit for different datasest. Also, Did you try 512 images? Did it perform stable in both dataset in your case? In my case, the scores can be various and unstable.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1236928,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-13T15:01:09.090000",
          "content": "<p>I just tried 128 and 256 with comparing new and old datas. 256’s differences between old and new is bigger than 128’s. However, the score is not stable and some variables are not the same, it can’t prove much regularity. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1237030,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-13T17:03:17.083000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1237272,
          "author_name": "andras",
          "author_url": "",
          "post_date": "2021-03-14T01:13:31.073000",
          "content": "<p>yes, same here! I get better LB score with the old model. Is it possible some of the test images may be \"familiar\" to the old model? I cannot find other explanation. This topic may worth a separate post</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1234210,
      "author_name": "Yineng Xiong",
      "author_url": "",
      "post_date": "2021-03-11T04:00:38.570000",
      "content": "<p>is this score on the updated dataset？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1234254,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-11T04:45:59.480000",
          "content": "<p>Yes. Using all of the new dataset. Single fold.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235414,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-12T05:57:42.710000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1235548,
          "author_name": "Shihao Shao",
          "author_url": "",
          "post_date": "2021-03-12T08:29:29.413000",
          "content": "<p>In my case, some TTAs like contrast adjustment will ruin the performance. So I just try rotate and flip for now. Feel free to publish a public notebook you submitted :D. Maybe people will have proper advices for your code.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1236343,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-13T03:09:52",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1250727,
          "author_name": "FatLiuyun",
          "author_url": "",
          "post_date": "2021-03-24T08:29:29.660000",
          "content": "<p>hi fury,same as you,have you find any way to slove this problem?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1234129": "I noticed that several competitions has this thread so I created it. Hope it would help us judge the relevance between CV and LB.\n\nIn my case, (NEW DATASET)\nCV 0.9187\nLB 0.905\n\nUpdated, (OLD DATASET, MODEL STRUCTURE IS SAME AS THE PREVIOUS ONE)\nCV 0.892\nLB 0.915 (It's interesting that with old dataset, same model structure, I reached a higher score)\n\nUpdated, (NEW DATASET)\nCV 0.9260\nLB 0.920\n\n\n========\nSo as you know, hand labeling in public test dataset is permitted, which may mislead the LB score. So I think the CV we post here may be important to judge our ranking.",
    "1247689": "[new data]\nresnet34-unet\ncv0.931, lb0.924\n\nupdate:\nmodify postprocess: 0.925 -> 0.929",
    "1234280": "CV: 0.873\nLB: 0.921\n\nI haven't re-trained the model yet with the new dataset ",
    "1236170": "Resubmitted the best single model trained on the old data without retraining: CV was on old data 0.94 and new LB 0.926 (old LB 0.873).",
    "1235102": "Single fold (12:3) retrained with new data\nLocal Dice: 0.942 (on crops)\npublic LB: 0.922\nIt might be an easy fold for validation",
    "1256216": "Ecaresnet50-unet++\n1024x1024\nCV 0.9301\nLB 0.903\n\nIm not sure why i have such a big difference between CV and LB ",
    "1255474": "[new data] - no pseudo or handlabeling....yet...\nEffNet-Unet \nCV: 0.905\nLB: 0.917",
    "1234495": "do you see correlation between CV and LB ?\n",
    "1238596": "resubmit old tensorflow model:\nCV:  0.8921\nLB: 0.910\n\n=======================\nI used pytorch to train new data\nCV: 0.9112\nLB: 0.869\nWhy did I get this huge gap between cv and lb when using pytorch...\n",
    "1234282": "I used previous model with:\nCV：0.890, pre-LB: 0.860\nLB：0.914",
    "1262319": "```So as you know, hand labeling in public test dataset is permitted```\nReally?",
    "1257339": "latest: CV: 0.921, LB:0.927",
    "1246550": "I don't really understand this.\n\nModel 1:\nCV: 0.930\nLB: 0.907\n\nModel 2:\nCV: 0.926\nLB: 0.912\n\nValidation scheme: full resolution masks (instead of mean of tiles) with 5 folds (12-3)",
    "1244690": "After new data, can anyone confirm that CV and LB are correlated? ",
    "1237970": "CV 0.944\nLB 0.924 (old model 0.920)\n\nThe new data provides a big boost.",
    "1237579": "My mean oof cv at whole slide dice is about 0.933 and my lb is 0.910",
    "1236837": "Resubmitted old model with single fold, I get 0.920 on LB",
    "1236316": "Did anyone find whether the new dataset, training the same model, performs better than the old one?",
    "1234210": "is this score on the updated dataset？",
    "1235414": ""
  }
}