{
  "id": 222313,
  "title": "Does pseudo labeling on public test leads to public LB Overfit?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222313",
  "author_name": "Dracarys",
  "post_date": "2021-02-26T09:41:30.862000",
  "votes": 7,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Pseudo labeling is very interesting topic and has helped in lots of competitions. But there is an argument that pseudo labeling on public test will lead to overfitting (leakage) on public Lb. </p>\n<p>What i think is since validation set remains untouched, pseudo labeling should lead to better generalization. </p>\n<p>A nice quote from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\" target=\"_blank\">notebook</a></p>\n<blockquote>\n  <p>Pseudo labeling helps all types of models because all models can be visualized as finding shapes of target=1 and target=0 in p-dimensional space. See here for examples. More points allow for better estimation of shapes.</p>\n</blockquote>\n<p>I am really sorry if this topic is redundant, but i really want to understand if pseudo labeling leads to overfit on Public LB or not. Feel free to share your opinions.</p>",
  "messages": [
    {
      "id": 1218928,
      "postDate": "2021-02-26T09:41:30.863Z",
      "content": "<p>Pseudo labeling is very interesting topic and has helped in lots of competitions. But there is an argument that pseudo labeling on public test will lead to overfitting (leakage) on public Lb. </p>\n<p>What i think is since validation set remains untouched, pseudo labeling should lead to better generalization. </p>\n<p>A nice quote from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\" target=\"_blank\">notebook</a></p>\n<blockquote>\n  <p>Pseudo labeling helps all types of models because all models can be visualized as finding shapes of target=1 and target=0 in p-dimensional space. See here for examples. More points allow for better estimation of shapes.</p>\n</blockquote>\n<p>I am really sorry if this topic is redundant, but i really want to understand if pseudo labeling leads to overfit on Public LB or not. Feel free to share your opinions.</p>",
      "rawMarkdown": "Pseudo labeling is very interesting topic and has helped in lots of competitions. But there is an argument that pseudo labeling on public test will lead to overfitting (leakage) on public Lb. \n\nWhat i think is since validation set remains untouched, pseudo labeling should lead to better generalization. \n\nA nice quote from @cdeotte [notebook](https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969)\n\n> Pseudo labeling helps all types of models because all models can be visualized as finding shapes of target=1 and target=0 in p-dimensional space. See here for examples. More points allow for better estimation of shapes.\n\nI am really sorry if this topic is redundant, but i really want to understand if pseudo labeling leads to overfit on Public LB or not. Feel free to share your opinions.\n",
      "votes": 7
    },
    {
      "id": 1219333,
      "postDate": "2021-02-26T17:49:45.707Z",
      "content": "<p>Im not quite sure how to use pseudolabeling here since its a multilabel problem. If for example we take the usual threshold 0.95 and have this prediction on a test image (I used an actual prediction): [0.121778,0.509806,0.690276,0.013107,0.049202,0.139460,0.938884,0.231646,0.412316,0.822261,0.995229]<br>\nthe label will than be: [0,0,0,0,0,0,0,0,0,1]<br>\nWill it make it so the model will fit on a false label since there could be missed labels?</p>",
      "rawMarkdown": "Im not quite sure how to use pseudolabeling here since its a multilabel problem. If for example we take the usual threshold 0.95 and have this prediction on a test image (I used an actual prediction): [0.121778,0.509806,0.690276,0.013107,0.049202,0.139460,0.938884,0.231646,0.412316,0.822261,0.995229]\nthe label will than be: [0,0,0,0,0,0,0,0,0,1]\nWill it make it so the model will fit on a false label since there could be missed labels?",
      "votes": 1,
      "replies": [
        {
          "id": 1219350,
          "postDate": "2021-02-26T18:23:53.827Z",
          "content": "<p>I'm not sure that this is really that different than a single binary label situation. In any case, there's also the option of soft-labeling (i.e. use something like 0.93884 as the target rather than 0 or 1).</p>",
          "rawMarkdown": "I'm not sure that this is really that different than a single binary label situation. In any case, there's also the option of soft-labeling (i.e. use something like 0.93884 as the target rather than 0 or 1).",
          "votes": 2
        },
        {
          "id": 1219387,
          "postDate": "2021-02-26T18:57:12.893Z",
          "content": "<p>Well in a multiclass problem you would take the softmax predictions so if a prediction is over 0.95 that means that the other labels are really low, so you are confident that its the good label. The same thing in a binary classification, you train a model to predict one or the other so if the confidence is high on one label than the other prediction confidence will be low. But in our case we can be confident for one label and be less confident for an other one  but set only one label as positive. however you are right we could use soft labels and that would sort of resolve the problem.</p>",
          "rawMarkdown": "Well in a multiclass problem you would take the softmax predictions so if a prediction is over 0.95 that means that the other labels are really low, so you are confident that its the good label. The same thing in a binary classification, you train a model to predict one or the other so if the confidence is high on one label than the other prediction confidence will be low. But in our case we can be confident for one label and be less confident for an other one  but set only one label as positive. however you are right we could use soft labels and that would sort of resolve the problem."
        },
        {
          "id": 1219981,
          "postDate": "2021-02-27T12:28:10.030Z",
          "content": "<p>I am not sure whether it will overfit. But it's strange that a lot of res200 fold cv is 0.955 but can ensemble reach pb 0.966+ or higher, while efficient net cv is 0.962 and alot of get 0.963 pb, I don't know such a model which has a lower cv and higher pb is not overfitting.</p>",
          "rawMarkdown": "I am not sure whether it will overfit. But it's strange that a lot of res200 fold cv is 0.955 but can ensemble reach pb 0.966+ or higher, while efficient net cv is 0.962 and alot of get 0.963 pb, I don't know such a model which has a lower cv and higher pb is not overfitting."
        },
        {
          "id": 1220132,
          "postDate": "2021-02-27T15:48:00.863Z",
          "content": "<p>Trust your CV, the public leaderboard has a lot less examples so its less robust to trust it</p>",
          "rawMarkdown": "Trust your CV, the public leaderboard has a lot less examples so its less robust to trust it"
        },
        {
          "id": 1220161,
          "postDate": "2021-02-27T16:50:55.860Z",
          "content": "<p>If you use soft pseudo labels, might I suggest adding gaussian noise to the predictions or performing mixup augmentation to reduce confirmation bias. You might see some nice improvements ;)</p>\n<p>Also be careful of data leakage when using pseudo labels. For example, if you have 5 folds and you are validating on fold 0, then make sure that your pseudo labels come from a model that has not seen data from fold 0. </p>",
          "rawMarkdown": "If you use soft pseudo labels, might I suggest adding gaussian noise to the predictions or performing mixup augmentation to reduce confirmation bias. You might see some nice improvements ;)\n\nAlso be careful of data leakage when using pseudo labels. For example, if you have 5 folds and you are validating on fold 0, then make sure that your pseudo labels come from a model that has not seen data from fold 0. ",
          "votes": 4
        },
        {
          "id": 1220176,
          "postDate": "2021-02-27T17:17:17.937Z",
          "content": "<p>Yes good point on confirmation bias! And i was planning tu use 5 folds to predict on public test images.</p>",
          "rawMarkdown": "Yes good point on confirmation bias! And i was planning tu use 5 folds to predict on public test images.",
          "votes": 1
        },
        {
          "id": 1220189,
          "postDate": "2021-02-27T17:46:51.007Z",
          "content": "<p>If you use the average predictions from 5 folds, you will have some leakage. Kaggle isn't allowing images to be attached to forum posts right now, but if you go to <a href=\"https://www.youtube.com/watch?v=SsnWM1xWDu4\" target=\"_blank\">this video</a> around 12:00 and watch for a couple minutes, you will see what I mean. </p>",
          "rawMarkdown": "If you use the average predictions from 5 folds, you will have some leakage. Kaggle isn't allowing images to be attached to forum posts right now, but if you go to [this video](https://www.youtube.com/watch?v=SsnWM1xWDu4) around 12:00 and watch for a couple minutes, you will see what I mean. ",
          "votes": 4
        },
        {
          "id": 1220230,
          "postDate": "2021-02-27T18:43:26.253Z",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> thanks for the video link, it really helped.</p>",
          "rawMarkdown": "@tuckerarrants thanks for the video link, it really helped.",
          "votes": 1
        },
        {
          "id": 1220272,
          "postDate": "2021-02-27T19:37:01.420Z",
          "content": "<p>Im not quite sure why predicting on the public test images is data leakage? The public test images are unseen to the models regardless of the fold</p>",
          "rawMarkdown": "Im not quite sure why predicting on the public test images is data leakage? The public test images are unseen to the models regardless of the fold"
        },
        {
          "id": 1220275,
          "postDate": "2021-02-27T19:42:01.313Z",
          "content": "<p>There is subtle leakage: the way your model predicts is based on what it has seen. So predictions from a model that has seen fold 1 will contain information about the training samples in fold 1, so if you then use those predictions as pseudo labels when you are validating against fold 1, there is leakage. </p>",
          "rawMarkdown": "There is subtle leakage: the way your model predicts is based on what it has seen. So predictions from a model that has seen fold 1 will contain information about the training samples in fold 1, so if you then use those predictions as pseudo labels when you are validating against fold 1, there is leakage. ",
          "votes": 2
        },
        {
          "id": 1227704,
          "postDate": "2021-03-05T19:27:19.073Z",
          "content": "<p>so how we make groupkfold for unseen data as in how to properly split folds notebook it use both patient id and distribution of labels to make fold <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> </p>",
          "rawMarkdown": "so how we make groupkfold for unseen data as in how to properly split folds notebook it use both patient id and distribution of labels to make fold @tuckerarrants ",
          "votes": 1
        },
        {
          "id": 1227722,
          "postDate": "2021-03-05T19:45:42.663Z",
          "content": "<p>I pretrain on all unseen data at once and validate against the labeled data stratified by patient ID and labels. So I am not creating folds for the pseudo labeled data, I am just reusing the initial folds I created for the labeled data as my validation set. </p>",
          "rawMarkdown": "I pretrain on all unseen data at once and validate against the labeled data stratified by patient ID and labels. So I am not creating folds for the pseudo labeled data, I am just reusing the initial folds I created for the labeled data as my validation set. "
        },
        {
          "id": 1228446,
          "postDate": "2021-03-06T11:47:17.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/tucker\" target=\"_blank\">@tucker</a> would you mind to form a team together </p>",
          "rawMarkdown": "@tucker would you mind to form a team together "
        }
      ]
    },
    {
      "id": 1219050,
      "postDate": "2021-02-26T12:04:07.950Z",
      "content": "<p>I would guess that it would improve your performance on both public and private LB, but that it would do so more on the public LB and less on the private LB (but I don't have a good intuition for how much it would do that). In that sense, you might call that overfitting the public LB and it perhaps means we should be a little bit more cautious about letting the LB guide our decisions. On the other hand, I would have thought this would be correctly reflected by the impact it has on your out-of-fold predictions that you produce (after doing pseudo-labelling with a model trained on one fold and then re-training with the data from the fold + the pseudo-labelled data). Is the answer perhaps simply to trust your CV over the public LB (as is very, very often the case)?</p>",
      "rawMarkdown": "I would guess that it would improve your performance on both public and private LB, but that it would do so more on the public LB and less on the private LB (but I don't have a good intuition for how much it would do that). In that sense, you might call that overfitting the public LB and it perhaps means we should be a little bit more cautious about letting the LB guide our decisions. On the other hand, I would have thought this would be correctly reflected by the impact it has on your out-of-fold predictions that you produce (after doing pseudo-labelling with a model trained on one fold and then re-training with the data from the fold + the pseudo-labelled data). Is the answer perhaps simply to trust your CV over the public LB (as is very, very often the case)?",
      "votes": 2,
      "replies": [
        {
          "id": 1219082,
          "postDate": "2021-02-26T12:34:15.843Z",
          "content": "<p>I am not trusting LB here, if CV is improved using pseudo labeling and so does public LB, boost should reflect both on public and private LB and should not mean that model is overfitting on public LB.</p>",
          "rawMarkdown": "I am not trusting LB here, if CV is improved using pseudo labeling and so does public LB, boost should reflect both on public and private LB and should not mean that model is overfitting on public LB."
        },
        {
          "id": 1219255,
          "postDate": "2021-02-26T15:52:24.137Z",
          "content": "<p>I think LB is very trustworthy in this competition. So far everything we tried on CV worked on LB. there was also private LB data leakage at the start of the competition. I was able to test a few subs while it was there, and public/private LB correlation was perfect as well. </p>\n<p>On top of that, cross-validation st.dev is only 0.0014</p>",
          "rawMarkdown": "I think LB is very trustworthy in this competition. So far everything we tried on CV worked on LB. there was also private LB data leakage at the start of the competition. I was able to test a few subs while it was there, and public/private LB correlation was perfect as well. \n\nOn top of that, cross-validation st.dev is only 0.0014",
          "votes": 11
        }
      ]
    },
    {
      "id": 1227535,
      "postDate": "2021-03-05T16:19:48.040Z",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "rawMarkdown": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
      "votes": 1
    },
    {
      "id": 1227496,
      "postDate": "2021-03-05T15:36:26.760Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1219333,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2021-02-26T17:49:45.707000",
      "content": "<p>Im not quite sure how to use pseudolabeling here since its a multilabel problem. If for example we take the usual threshold 0.95 and have this prediction on a test image (I used an actual prediction): [0.121778,0.509806,0.690276,0.013107,0.049202,0.139460,0.938884,0.231646,0.412316,0.822261,0.995229]<br>\nthe label will than be: [0,0,0,0,0,0,0,0,0,1]<br>\nWill it make it so the model will fit on a false label since there could be missed labels?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1219350,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-02-26T18:23:53.827000",
          "content": "<p>I'm not sure that this is really that different than a single binary label situation. In any case, there's also the option of soft-labeling (i.e. use something like 0.93884 as the target rather than 0 or 1).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1219387,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-02-26T18:57:12.893000",
          "content": "<p>Well in a multiclass problem you would take the softmax predictions so if a prediction is over 0.95 that means that the other labels are really low, so you are confident that its the good label. The same thing in a binary classification, you train a model to predict one or the other so if the confidence is high on one label than the other prediction confidence will be low. But in our case we can be confident for one label and be less confident for an other one  but set only one label as positive. however you are right we could use soft labels and that would sort of resolve the problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1219981,
          "author_name": "Dongzhu Zhao",
          "author_url": "",
          "post_date": "2021-02-27T12:28:10.030000",
          "content": "<p>I am not sure whether it will overfit. But it's strange that a lot of res200 fold cv is 0.955 but can ensemble reach pb 0.966+ or higher, while efficient net cv is 0.962 and alot of get 0.963 pb, I don't know such a model which has a lower cv and higher pb is not overfitting.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220132,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-02-27T15:48:00.863000",
          "content": "<p>Trust your CV, the public leaderboard has a lot less examples so its less robust to trust it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220161,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-02-27T16:50:55.860000",
          "content": "<p>If you use soft pseudo labels, might I suggest adding gaussian noise to the predictions or performing mixup augmentation to reduce confirmation bias. You might see some nice improvements ;)</p>\n<p>Also be careful of data leakage when using pseudo labels. For example, if you have 5 folds and you are validating on fold 0, then make sure that your pseudo labels come from a model that has not seen data from fold 0. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1220176,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-02-27T17:17:17.937000",
          "content": "<p>Yes good point on confirmation bias! And i was planning tu use 5 folds to predict on public test images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220189,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-02-27T17:46:51.007000",
          "content": "<p>If you use the average predictions from 5 folds, you will have some leakage. Kaggle isn't allowing images to be attached to forum posts right now, but if you go to <a href=\"https://www.youtube.com/watch?v=SsnWM1xWDu4\" target=\"_blank\">this video</a> around 12:00 and watch for a couple minutes, you will see what I mean. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1220230,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2021-02-27T18:43:26.253000",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> thanks for the video link, it really helped.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220272,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-02-27T19:37:01.420000",
          "content": "<p>Im not quite sure why predicting on the public test images is data leakage? The public test images are unseen to the models regardless of the fold</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220275,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-02-27T19:42:01.313000",
          "content": "<p>There is subtle leakage: the way your model predicts is based on what it has seen. So predictions from a model that has seen fold 1 will contain information about the training samples in fold 1, so if you then use those predictions as pseudo labels when you are validating against fold 1, there is leakage. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1227704,
          "author_name": "Mohamed abdelrazik",
          "author_url": "",
          "post_date": "2021-03-05T19:27:19.073000",
          "content": "<p>so how we make groupkfold for unseen data as in how to properly split folds notebook it use both patient id and distribution of labels to make fold <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1227722,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-03-05T19:45:42.663000",
          "content": "<p>I pretrain on all unseen data at once and validate against the labeled data stratified by patient ID and labels. So I am not creating folds for the pseudo labeled data, I am just reusing the initial folds I created for the labeled data as my validation set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1228446,
          "author_name": "Mohamed abdelrazik",
          "author_url": "",
          "post_date": "2021-03-06T11:47:17.137000",
          "content": "<p><a href=\"https://www.kaggle.com/tucker\" target=\"_blank\">@tucker</a> would you mind to form a team together </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1219050,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-26T12:04:07.950000",
      "content": "<p>I would guess that it would improve your performance on both public and private LB, but that it would do so more on the public LB and less on the private LB (but I don't have a good intuition for how much it would do that). In that sense, you might call that overfitting the public LB and it perhaps means we should be a little bit more cautious about letting the LB guide our decisions. On the other hand, I would have thought this would be correctly reflected by the impact it has on your out-of-fold predictions that you produce (after doing pseudo-labelling with a model trained on one fold and then re-training with the data from the fold + the pseudo-labelled data). Is the answer perhaps simply to trust your CV over the public LB (as is very, very often the case)?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1219082,
          "author_name": "Dracarys",
          "author_url": "",
          "post_date": "2021-02-26T12:34:15.843000",
          "content": "<p>I am not trusting LB here, if CV is improved using pseudo labeling and so does public LB, boost should reflect both on public and private LB and should not mean that model is overfitting on public LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1219255,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-26T15:52:24.137000",
          "content": "<p>I think LB is very trustworthy in this competition. So far everything we tried on CV worked on LB. there was also private LB data leakage at the start of the competition. I was able to test a few subs while it was there, and public/private LB correlation was perfect as well. </p>\n<p>On top of that, cross-validation st.dev is only 0.0014</p>",
          "votes": 11,
          "replies": []
        }
      ]
    },
    {
      "id": 1227535,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-03-05T16:19:48.040000",
      "content": "<p>To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.</p>\n<p>Link to Independent validation dataset - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788\" target=\"_blank\">link</a><br>\nI hope it helps.</p>\n<p><img src=\"https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&amp;X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&amp;X-Goog-Date=20210305T161613Z&amp;X-Goog-Expires=172799&amp;X-Goog-SignedHeaders=host&amp;X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1227496,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-05T15:36:26.760000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1218928": "Pseudo labeling is very interesting topic and has helped in lots of competitions. But there is an argument that pseudo labeling on public test will lead to overfitting (leakage) on public Lb. \n\nWhat i think is since validation set remains untouched, pseudo labeling should lead to better generalization. \n\nA nice quote from @cdeotte [notebook](https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969)\n\n> Pseudo labeling helps all types of models because all models can be visualized as finding shapes of target=1 and target=0 in p-dimensional space. See here for examples. More points allow for better estimation of shapes.\n\nI am really sorry if this topic is redundant, but i really want to understand if pseudo labeling leads to overfit on Public LB or not. Feel free to share your opinions.\n",
    "1219333": "Im not quite sure how to use pseudolabeling here since its a multilabel problem. If for example we take the usual threshold 0.95 and have this prediction on a test image (I used an actual prediction): [0.121778,0.509806,0.690276,0.013107,0.049202,0.139460,0.938884,0.231646,0.412316,0.822261,0.995229]\nthe label will than be: [0,0,0,0,0,0,0,0,0,1]\nWill it make it so the model will fit on a false label since there could be missed labels?",
    "1219050": "I would guess that it would improve your performance on both public and private LB, but that it would do so more on the public LB and less on the private LB (but I don't have a good intuition for how much it would do that). In that sense, you might call that overfitting the public LB and it perhaps means we should be a little bit more cautious about letting the LB guide our decisions. On the other hand, I would have thought this would be correctly reflected by the impact it has on your out-of-fold predictions that you produce (after doing pseudo-labelling with a model trained on one fold and then re-training with the data from the fold + the pseudo-labelled data). Is the answer perhaps simply to trust your CV over the public LB (as is very, very often the case)?",
    "1227535": "To check for overfitting you can also check the model on an independent validation dataset. I have created an Independent validation dataset, and as per guidelines of the competition am sharing the model to the public. Since labelling Test-set isnt allowed, so i came up with some smart ways to use other data.\n\nLink to Independent validation dataset - [link](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/223788)\nI hope it helps.\n\n![](https://storage.googleapis.com/kagglesdsdata/datasets/1194466/1996989/Ranzcr%20-%20Frame%206.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210305%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210305T161613Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=2180e9bff41c98ebe908e479b19d780d7b315ea183d11ffebd31415c82cbe1c45ab5b3a0d5632e7cd7a718f9bdaa969093acad050c2618058f7136f3d19337b2c0d15f66e420ace748c11e861cb2510ba62d7b3579ecef2c9985ec8c4ac20245d4a68d3647c17d3156bcf96fe21702acb8cf9a0c3c34cee956bdb82892696629e3068e4324c8b1e9c0af826240680adeea2c2842f3e732e420f8a69d35bd6818b68e5c62546388441f75b12c9e80487006f2ccead15737e4cecc28abd398b322362c4de0b50741cda4e7395f33c7fd9c14ebda88c7b177cdd8bce67f3c977a7086383bdbbcb655343d735df876d451454e0071bb1962148cbc9f5f643d8df30c)",
    "1227496": ""
  }
}