{
  "id": 232124,
  "title": "Better CV but worse LB on larger model",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/232124",
  "author_name": "Tsai29",
  "post_date": "2021-04-12T09:51:45.561000",
  "votes": 24,
  "comment_count": 26,
  "views": 0,
  "content": "<p>My current score was came from 5 folds efficientnetb2 unet. (No TTA)<br>\nLocal cv : 0.927<br>\nLB : 0.926</p>\n<p>And when I switch to larger model(efficientnetb4, b3), they always give me better local cv(B3:0.928, B4:0.93) but worse LB(B3:0.922, B4:0.920). <br>\nJust wondering if anyone have similar situation as mine.</p>",
  "messages": [
    {
      "id": 1271096,
      "postDate": "2021-04-12T09:51:45.560Z",
      "content": "<p>My current score was came from 5 folds efficientnetb2 unet. (No TTA)<br>\nLocal cv : 0.927<br>\nLB : 0.926</p>\n<p>And when I switch to larger model(efficientnetb4, b3), they always give me better local cv(B3:0.928, B4:0.93) but worse LB(B3:0.922, B4:0.920). <br>\nJust wondering if anyone have similar situation as mine.</p>",
      "rawMarkdown": "My current score was came from 5 folds efficientnetb2 unet. (No TTA)\nLocal cv : 0.927\nLB : 0.926\n\nAnd when I switch to larger model(efficientnetb4, b3), they always give me better local cv(B3:0.928, B4:0.93) but worse LB(B3:0.922, B4:0.920). \nJust wondering if anyone have similar situation as mine.",
      "votes": 24
    },
    {
      "id": 1275721,
      "postDate": "2021-04-16T15:56:18.263Z",
      "content": "<p>An interesting paper has recently been published on this topic.</p>\n<p>ArXiv: <a href=\"https://arxiv.org/abs/2103.14749\" target=\"_blank\">https://arxiv.org/abs/2103.14749</a><br>\nGithub: <a href=\"https://github.com/cgnorthcutt/label-errors\" target=\"_blank\">https://github.com/cgnorthcutt/label-errors</a></p>\n<p>The argument of the paper can be summarized as follows:<br>\nWhen the labels of the test data were wrong, models with larger capacity (such as NasNet) showed better accuracy.<br>\nHowever, after manually correcting the labels of the test data that seemed to be wrong and comparing the accuracy of the models again, the models with smaller capacity (such as ResNet18) showed better accuracy.<br>\nThe authors cited several reasons for this, for example, low-capacity models having an effect similar to regularization.</p>\n<p>Perhaps this competition is in a similar situation. It seems risky to me to train a large model with a special label like <code>d488c759a</code> as psuedo.</p>\n<p>In any case, I think there is no doubt that the difference in annotation methods between train data and test data has a significant impact on model selection.</p>\n<p><img src=\"https://f.easyuploader.app/eu-prd/upload/20210417003517_6f50624e.jpg\" alt=\"\"></p>",
      "rawMarkdown": "An interesting paper has recently been published on this topic.\n\nArXiv: https://arxiv.org/abs/2103.14749\nGithub: https://github.com/cgnorthcutt/label-errors\n\nThe argument of the paper can be summarized as follows:\nWhen the labels of the test data were wrong, models with larger capacity (such as NasNet) showed better accuracy.\nHowever, after manually correcting the labels of the test data that seemed to be wrong and comparing the accuracy of the models again, the models with smaller capacity (such as ResNet18) showed better accuracy.\nThe authors cited several reasons for this, for example, low-capacity models having an effect similar to regularization.\n\nPerhaps this competition is in a similar situation. It seems risky to me to train a large model with a special label like `d488c759a` as psuedo.\n\nIn any case, I think there is no doubt that the difference in annotation methods between train data and test data has a significant impact on model selection.\n\n![](https://f.easyuploader.app/eu-prd/upload/20210417003517_6f50624e.jpg)",
      "votes": 9,
      "replies": [
        {
          "id": 1276006,
          "postDate": "2021-04-17T01:38:16.043Z",
          "content": "<p>Thanks a lot for sharing this!</p>",
          "rawMarkdown": "Thanks a lot for sharing this!",
          "votes": 1
        },
        {
          "id": 1276016,
          "postDate": "2021-04-17T01:59:20.427Z",
          "content": "<p>my guess is that private dataset is equally badly labeled. Wouldn't in that case warrant usage of larger model? Here's my take:</p>\n<ul>\n<li>one sumbission using d5 preudo and light model</li>\n<li>other submission use no pseudo, but larger model (i.e. efficientnetb5+)</li>\n</ul>\n<p>Thoughts?</p>",
          "rawMarkdown": "my guess is that private dataset is equally badly labeled. Wouldn't in that case warrant usage of larger model? Here's my take:\n- one sumbission using d5 preudo and light model\n- other submission use no pseudo, but larger model (i.e. efficientnetb5+)\n\nThoughts?",
          "votes": 2
        },
        {
          "id": 1276049,
          "postDate": "2021-04-17T03:25:56.457Z",
          "content": "<p><a href=\"https://www.kaggle.com/andrasferenczi\" target=\"_blank\">@andrasferenczi</a> </p>\n<blockquote>\n  <p>my guess is that private dataset is equally badly labeled.</p>\n</blockquote>\n<p>As long as there is no guaranteed reply from the host, that is a possibility, and vice versa. I also think that the strategy you are considering is not a bad one, if you refer to the above paper.<br>\nThe other thing to consider will be how much the fact that the number of images for private evaluation is approximately doubled will affect the score.  </p>\n<p>Finally, one more important thing I forgot to mention.  <br>\n<strong>After you have done all you can do, pray.</strong></p>\n<p><img src=\"https://f.easyuploader.app/eu-prd/upload/20210417122353_46537164.jpg\" alt=\"\"></p>",
          "rawMarkdown": "@andrasferenczi \n\n> my guess is that private dataset is equally badly labeled.\n\nAs long as there is no guaranteed reply from the host, that is a possibility, and vice versa. I also think that the strategy you are considering is not a bad one, if you refer to the above paper.\nThe other thing to consider will be how much the fact that the number of images for private evaluation is approximately doubled will affect the score.  \n\nFinally, one more important thing I forgot to mention.  \n**After you have done all you can do, pray.**\n\n![](https://f.easyuploader.app/eu-prd/upload/20210417122353_46537164.jpg)"
        },
        {
          "id": 1277568,
          "postDate": "2021-04-18T23:56:43.023Z",
          "content": "<p>Well in many medical competitions the private set had better labeling because in the end they want to know how it will really perform even with noisy train data. Sometimes the private set is a consensus between multiple specialists</p>",
          "rawMarkdown": "Well in many medical competitions the private set had better labeling because in the end they want to know how it will really perform even with noisy train data. Sometimes the private set is a consensus between multiple specialists"
        },
        {
          "id": 1287914,
          "postDate": "2021-04-29T14:19:14.893Z",
          "content": "<p><a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a> , so you think this is by design? I'm not sure they would have allowed hand-labeling, had that been the case :(</p>",
          "rawMarkdown": "@yannmajewski , so you think this is by design? I'm not sure they would have allowed hand-labeling, had that been the case :("
        }
      ]
    },
    {
      "id": 1272097,
      "postDate": "2021-04-13T08:09:05.437Z",
      "content": "<p>Just an observation on the public LB test d488c759a.tiff is the only image for a patient not previously seen in train and is the only image that is 100 percent_cortex in the dataset_information.csv (and glomeruli are mainly found in the cortex). Something like a resnet34 seemed to do better on this one image than a larger efficientnet in testing different models for this one image.  So you could try your b2 on this one image prediction and b4 or b3 on the rest just to see if that is true for your models too.  Of course, this would not be a submission to select for final subs. But maybe indicates something about how models generalise to all cortex images or where a patient was not previously seen in train. </p>",
      "rawMarkdown": "Just an observation on the public LB test d488c759a.tiff is the only image for a patient not previously seen in train and is the only image that is 100 percent_cortex in the dataset_information.csv (and glomeruli are mainly found in the cortex). Something like a resnet34 seemed to do better on this one image than a larger efficientnet in testing different models for this one image.  So you could try your b2 on this one image prediction and b4 or b3 on the rest just to see if that is true for your models too.  Of course, this would not be a submission to select for final subs. But maybe indicates something about how models generalise to all cortex images or where a patient was not previously seen in train. ",
      "votes": 3,
      "replies": [
        {
          "id": 1272437,
          "postDate": "2021-04-13T13:24:49.220Z",
          "content": "<p><a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> Thanks a lot for the information, this is really helpful. I will post my result here after I done doing the experiments.</p>",
          "rawMarkdown": "@something4kag Thanks a lot for the information, this is really helpful. I will post my result here after I done doing the experiments.",
          "votes": 1
        },
        {
          "id": 1287505,
          "postDate": "2021-04-29T06:23:34.083Z",
          "content": "<p>Have you done your experiment?  How were results?</p>",
          "rawMarkdown": "Have you done your experiment?  How were results?"
        }
      ]
    },
    {
      "id": 1272168,
      "postDate": "2021-04-13T09:29:02.210Z",
      "content": "<p>if we got only CV score<br>\nhow can we know which efficientnet base number is better？<br>\nthis question always confuse me </p>",
      "rawMarkdown": "if we got only CV score\nhow can we know which efficientnet base number is better？\nthis question always confuse me ",
      "votes": 1,
      "replies": [
        {
          "id": 1272190,
          "postDate": "2021-04-13T09:50:09.383Z",
          "content": "<p>oh its base on image resolution <br>\n<a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/</a></p>",
          "rawMarkdown": "oh its base on image resolution \nhttps://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/"
        },
        {
          "id": 1272434,
          "postDate": "2021-04-13T13:23:48.287Z",
          "content": "<p>The CV scores I listed are the dice coefficient of validation set, so you can divide the original dataset from competition to training/validation set. Then train all kinds of models on the same training set and validate the local CV on the same validation set. This is how I judge which efficientnet is better.</p>",
          "rawMarkdown": "The CV scores I listed are the dice coefficient of validation set, so you can divide the original dataset from competition to training/validation set. Then train all kinds of models on the same training set and validate the local CV on the same validation set. This is how I judge which efficientnet is better."
        }
      ]
    },
    {
      "id": 1271899,
      "postDate": "2021-04-13T03:20:57.140Z",
      "content": "<p>can't agree more. </p>",
      "rawMarkdown": "can't agree more. ",
      "votes": 1
    },
    {
      "id": 1271475,
      "postDate": "2021-04-12T16:31:52.490Z",
      "content": "<p>Same thing happened to me for larger models, i'm getting better scores locally than on LB, i'm yet to finish training my last version and check it again.</p>\n<p>Thank you for the post, its good to know that its not just my model.</p>",
      "rawMarkdown": "Same thing happened to me for larger models, i'm getting better scores locally than on LB, i'm yet to finish training my last version and check it again.\n\nThank you for the post, its good to know that its not just my model.",
      "votes": 1
    },
    {
      "id": 1273060,
      "postDate": "2021-04-14T04:23:58.470Z",
      "content": "<p>I will trust my CV. Since LB can't be determined much as <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993#1272538\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993#1272538</a>. I thought public dataset still has some problems not solved.</p>",
      "rawMarkdown": "I will trust my CV. Since LB can't be determined much as https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993#1272538. I thought public dataset still has some problems not solved.",
      "votes": 2
    },
    {
      "id": 1272195,
      "postDate": "2021-04-13T09:51:48.300Z",
      "content": "<p>How are you ensembling the models in ur CV? I mean you cant just average the predicted masks right?Also there is no way to tell the score locally right? You have to submit the kernel code to get the score?</p>",
      "rawMarkdown": "How are you ensembling the models in ur CV? I mean you cant just average the predicted masks right?Also there is no way to tell the score locally right? You have to submit the kernel code to get the score?",
      "votes": 2,
      "replies": [
        {
          "id": 1272333,
          "postDate": "2021-04-13T12:06:43.307Z",
          "content": "<p>Averaging the predicted masks is exactly what i was planning to do, you will get the score per pixel, so you just go ahead and average them and filter the result with a specific threshold, why you say we can't do that?</p>\n<p>By local score its just the validation set score, if you are training your models on your own infrastructure, you should be able to do so.</p>",
          "rawMarkdown": "Averaging the predicted masks is exactly what i was planning to do, you will get the score per pixel, so you just go ahead and average them and filter the result with a specific threshold, why you say we can't do that?\n\nBy local score its just the validation set score, if you are training your models on your own infrastructure, you should be able to do so.",
          "votes": 2
        },
        {
          "id": 1272428,
          "postDate": "2021-04-13T13:19:50.393Z",
          "content": "<p>Yeah, I was doing the ensemble exactly how <a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a> described, so does the local score part.</p>",
          "rawMarkdown": "Yeah, I was doing the ensemble exactly how @victorasso described, so does the local score part.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1271911,
      "postDate": "2021-04-13T03:46:25.897Z",
      "content": "<p>I would trust my CV more than the LB</p>",
      "rawMarkdown": "I would trust my CV more than the LB",
      "votes": 2,
      "replies": [
        {
          "id": 1272439,
          "postDate": "2021-04-13T13:26:07.167Z",
          "content": "<p>I think we can select one best cv and one best lb submissions for two final submission, this is how I did in every competition I joined.</p>",
          "rawMarkdown": "I think we can select one best cv and one best lb submissions for two final submission, this is how I did in every competition I joined."
        }
      ]
    },
    {
      "id": 1287625,
      "postDate": "2021-04-29T08:37:13.607Z",
      "content": "<p>I rethink this problem. Host just said all the principles are the same. I just find some samples which are really hard to classify by me in the training set. Therefore, I think these might not be mistakes but just classifying by maybe some perfessional methods (e.g. staining reagent) which are difficult to judge by appearance? Thus, our test set could not be recognized as having mistakes.  </p>",
      "rawMarkdown": "I rethink this problem. Host just said all the principles are the same. I just find some samples which are really hard to classify by me in the training set. Therefore, I think these might not be mistakes but just classifying by maybe some perfessional methods (e.g. staining reagent) which are difficult to judge by appearance? Thus, our test set could not be recognized as having mistakes.  "
    },
    {
      "id": 1284100,
      "postDate": "2021-04-25T15:01:31.163Z",
      "content": "<p>emmm，2333.</p>",
      "rawMarkdown": "emmm，2333."
    },
    {
      "id": 1280755,
      "postDate": "2021-04-22T10:00:00.863Z",
      "content": "<p>interesting </p>",
      "rawMarkdown": " interesting "
    },
    {
      "id": 1271878,
      "postDate": "2021-04-13T02:35:20.910Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1271888,
          "postDate": "2021-04-13T02:53:24.670Z",
          "content": "<p>Yeah, that's my guess too. But I might still keep larger model submission for one of the final submissions. Since I always keep the final submissions with one best LB and one best CV.</p>",
          "rawMarkdown": "Yeah, that's my guess too. But I might still keep larger model submission for one of the final submissions. Since I always keep the final submissions with one best LB and one best CV.",
          "votes": 2
        },
        {
          "id": 1272090,
          "postDate": "2021-04-13T08:01:59.613Z",
          "content": "<p>yeah i agree i used effnetb7 its cv was ridiculosuly high , and lb is 0.921 mean while i just got a 0.927 lb score with a way smaller model </p>",
          "rawMarkdown": "yeah i agree i used effnetb7 its cv was ridiculosuly high , and lb is 0.921 mean while i just got a 0.927 lb score with a way smaller model ",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1275721,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2021-04-16T15:56:18.263000",
      "content": "<p>An interesting paper has recently been published on this topic.</p>\n<p>ArXiv: <a href=\"https://arxiv.org/abs/2103.14749\" target=\"_blank\">https://arxiv.org/abs/2103.14749</a><br>\nGithub: <a href=\"https://github.com/cgnorthcutt/label-errors\" target=\"_blank\">https://github.com/cgnorthcutt/label-errors</a></p>\n<p>The argument of the paper can be summarized as follows:<br>\nWhen the labels of the test data were wrong, models with larger capacity (such as NasNet) showed better accuracy.<br>\nHowever, after manually correcting the labels of the test data that seemed to be wrong and comparing the accuracy of the models again, the models with smaller capacity (such as ResNet18) showed better accuracy.<br>\nThe authors cited several reasons for this, for example, low-capacity models having an effect similar to regularization.</p>\n<p>Perhaps this competition is in a similar situation. It seems risky to me to train a large model with a special label like <code>d488c759a</code> as psuedo.</p>\n<p>In any case, I think there is no doubt that the difference in annotation methods between train data and test data has a significant impact on model selection.</p>\n<p><img src=\"https://f.easyuploader.app/eu-prd/upload/20210417003517_6f50624e.jpg\" alt=\"\"></p>",
      "votes": 9,
      "replies": [
        {
          "id": 1276006,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-04-17T01:38:16.043000",
          "content": "<p>Thanks a lot for sharing this!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1276016,
          "author_name": "andras",
          "author_url": "",
          "post_date": "2021-04-17T01:59:20.427000",
          "content": "<p>my guess is that private dataset is equally badly labeled. Wouldn't in that case warrant usage of larger model? Here's my take:</p>\n<ul>\n<li>one sumbission using d5 preudo and light model</li>\n<li>other submission use no pseudo, but larger model (i.e. efficientnetb5+)</li>\n</ul>\n<p>Thoughts?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1276049,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2021-04-17T03:25:56.457000",
          "content": "<p><a href=\"https://www.kaggle.com/andrasferenczi\" target=\"_blank\">@andrasferenczi</a> </p>\n<blockquote>\n  <p>my guess is that private dataset is equally badly labeled.</p>\n</blockquote>\n<p>As long as there is no guaranteed reply from the host, that is a possibility, and vice versa. I also think that the strategy you are considering is not a bad one, if you refer to the above paper.<br>\nThe other thing to consider will be how much the fact that the number of images for private evaluation is approximately doubled will affect the score.  </p>\n<p>Finally, one more important thing I forgot to mention.  <br>\n<strong>After you have done all you can do, pray.</strong></p>\n<p><img src=\"https://f.easyuploader.app/eu-prd/upload/20210417122353_46537164.jpg\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277568,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-04-18T23:56:43.023000",
          "content": "<p>Well in many medical competitions the private set had better labeling because in the end they want to know how it will really perform even with noisy train data. Sometimes the private set is a consensus between multiple specialists</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287914,
          "author_name": "andras",
          "author_url": "",
          "post_date": "2021-04-29T14:19:14.893000",
          "content": "<p><a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a> , so you think this is by design? I'm not sure they would have allowed hand-labeling, had that been the case :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1272097,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-04-13T08:09:05.437000",
      "content": "<p>Just an observation on the public LB test d488c759a.tiff is the only image for a patient not previously seen in train and is the only image that is 100 percent_cortex in the dataset_information.csv (and glomeruli are mainly found in the cortex). Something like a resnet34 seemed to do better on this one image than a larger efficientnet in testing different models for this one image.  So you could try your b2 on this one image prediction and b4 or b3 on the rest just to see if that is true for your models too.  Of course, this would not be a submission to select for final subs. But maybe indicates something about how models generalise to all cortex images or where a patient was not previously seen in train. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1272437,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-04-13T13:24:49.220000",
          "content": "<p><a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> Thanks a lot for the information, this is really helpful. I will post my result here after I done doing the experiments.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1287505,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-04-29T06:23:34.083000",
          "content": "<p>Have you done your experiment?  How were results?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1272168,
      "author_name": "Drzhuzhe",
      "author_url": "",
      "post_date": "2021-04-13T09:29:02.210000",
      "content": "<p>if we got only CV score<br>\nhow can we know which efficientnet base number is better？<br>\nthis question always confuse me </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1272190,
          "author_name": "Drzhuzhe",
          "author_url": "",
          "post_date": "2021-04-13T09:50:09.383000",
          "content": "<p>oh its base on image resolution <br>\n<a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1272434,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-04-13T13:23:48.287000",
          "content": "<p>The CV scores I listed are the dice coefficient of validation set, so you can divide the original dataset from competition to training/validation set. Then train all kinds of models on the same training set and validate the local CV on the same validation set. This is how I judge which efficientnet is better.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1271899,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2021-04-13T03:20:57.140000",
      "content": "<p>can't agree more. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1271475,
      "author_name": "victorasso",
      "author_url": "",
      "post_date": "2021-04-12T16:31:52.490000",
      "content": "<p>Same thing happened to me for larger models, i'm getting better scores locally than on LB, i'm yet to finish training my last version and check it again.</p>\n<p>Thank you for the post, its good to know that its not just my model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1273060,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-04-14T04:23:58.470000",
      "content": "<p>I will trust my CV. Since LB can't be determined much as <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993#1272538\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993#1272538</a>. I thought public dataset still has some problems not solved.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1272195,
      "author_name": "Mostafa Ibrahim",
      "author_url": "",
      "post_date": "2021-04-13T09:51:48.300000",
      "content": "<p>How are you ensembling the models in ur CV? I mean you cant just average the predicted masks right?Also there is no way to tell the score locally right? You have to submit the kernel code to get the score?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1272333,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "2021-04-13T12:06:43.307000",
          "content": "<p>Averaging the predicted masks is exactly what i was planning to do, you will get the score per pixel, so you just go ahead and average them and filter the result with a specific threshold, why you say we can't do that?</p>\n<p>By local score its just the validation set score, if you are training your models on your own infrastructure, you should be able to do so.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1272428,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-04-13T13:19:50.393000",
          "content": "<p>Yeah, I was doing the ensemble exactly how <a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a> described, so does the local score part.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1271911,
      "author_name": "andras",
      "author_url": "",
      "post_date": "2021-04-13T03:46:25.897000",
      "content": "<p>I would trust my CV more than the LB</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1272439,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-04-13T13:26:07.167000",
          "content": "<p>I think we can select one best cv and one best lb submissions for two final submission, this is how I did in every competition I joined.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1287625,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-04-29T08:37:13.607000",
      "content": "<p>I rethink this problem. Host just said all the principles are the same. I just find some samples which are really hard to classify by me in the training set. Therefore, I think these might not be mistakes but just classifying by maybe some perfessional methods (e.g. staining reagent) which are difficult to judge by appearance? Thus, our test set could not be recognized as having mistakes.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1284100,
      "author_name": "Ali Barry",
      "author_url": "",
      "post_date": "2021-04-25T15:01:31.163000",
      "content": "<p>emmm，2333.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1280755,
      "author_name": "PowerJose",
      "author_url": "",
      "post_date": "2021-04-22T10:00:00.863000",
      "content": "<p>interesting </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1271878,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-13T02:35:20.910000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1271888,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-04-13T02:53:24.670000",
          "content": "<p>Yeah, that's my guess too. But I might still keep larger model submission for one of the final submissions. Since I always keep the final submissions with one best LB and one best CV.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1272090,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2021-04-13T08:01:59.613000",
          "content": "<p>yeah i agree i used effnetb7 its cv was ridiculosuly high , and lb is 0.921 mean while i just got a 0.927 lb score with a way smaller model </p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1271096": "My current score was came from 5 folds efficientnetb2 unet. (No TTA)\nLocal cv : 0.927\nLB : 0.926\n\nAnd when I switch to larger model(efficientnetb4, b3), they always give me better local cv(B3:0.928, B4:0.93) but worse LB(B3:0.922, B4:0.920). \nJust wondering if anyone have similar situation as mine.",
    "1275721": "An interesting paper has recently been published on this topic.\n\nArXiv: https://arxiv.org/abs/2103.14749\nGithub: https://github.com/cgnorthcutt/label-errors\n\nThe argument of the paper can be summarized as follows:\nWhen the labels of the test data were wrong, models with larger capacity (such as NasNet) showed better accuracy.\nHowever, after manually correcting the labels of the test data that seemed to be wrong and comparing the accuracy of the models again, the models with smaller capacity (such as ResNet18) showed better accuracy.\nThe authors cited several reasons for this, for example, low-capacity models having an effect similar to regularization.\n\nPerhaps this competition is in a similar situation. It seems risky to me to train a large model with a special label like `d488c759a` as psuedo.\n\nIn any case, I think there is no doubt that the difference in annotation methods between train data and test data has a significant impact on model selection.\n\n![](https://f.easyuploader.app/eu-prd/upload/20210417003517_6f50624e.jpg)",
    "1272097": "Just an observation on the public LB test d488c759a.tiff is the only image for a patient not previously seen in train and is the only image that is 100 percent_cortex in the dataset_information.csv (and glomeruli are mainly found in the cortex). Something like a resnet34 seemed to do better on this one image than a larger efficientnet in testing different models for this one image.  So you could try your b2 on this one image prediction and b4 or b3 on the rest just to see if that is true for your models too.  Of course, this would not be a submission to select for final subs. But maybe indicates something about how models generalise to all cortex images or where a patient was not previously seen in train. ",
    "1272168": "if we got only CV score\nhow can we know which efficientnet base number is better？\nthis question always confuse me ",
    "1271899": "can't agree more. ",
    "1271475": "Same thing happened to me for larger models, i'm getting better scores locally than on LB, i'm yet to finish training my last version and check it again.\n\nThank you for the post, its good to know that its not just my model.",
    "1273060": "I will trust my CV. Since LB can't be determined much as https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993#1272538. I thought public dataset still has some problems not solved.",
    "1272195": "How are you ensembling the models in ur CV? I mean you cant just average the predicted masks right?Also there is no way to tell the score locally right? You have to submit the kernel code to get the score?",
    "1271911": "I would trust my CV more than the LB",
    "1287625": "I rethink this problem. Host just said all the principles are the same. I just find some samples which are really hard to classify by me in the training set. Therefore, I think these might not be mistakes but just classifying by maybe some perfessional methods (e.g. staining reagent) which are difficult to judge by appearance? Thus, our test set could not be recognized as having mistakes.  ",
    "1284100": "emmm，2333.",
    "1280755": " interesting ",
    "1271878": ""
  }
}