{
  "id": 226421,
  "title": "Large CV-LB Gap , the reason for our downfall",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226421",
  "author_name": "",
  "post_date": "2021-03-16T12:30:21.350006400Z",
  "votes": 15,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi all ,<br>\nNow that 12 hours is left in this competition , I wish everyone the best , mark your last submissions carefully and may the shakeup be in your favour </p>\n<p>The last week has been very frustrating for us , we have reached as high as 0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968 . And due to the strong correlation between our models we have not been able to break the 0.970 barrier . We are very curious to know how will our high cv blend perform on private when on public it even fails to touch our best</p>\n<p>Is it just us or anyone else facing the similar kind of gap? </p>",
  "messages": [
    {
      "id": "1240460",
      "postDate": "03/16/2021 12:30:21",
      "content": "<p>Hi all ,<br>\nNow that 12 hours is left in this competition , I wish everyone the best , mark your last submissions carefully and may the shakeup be in your favour </p>\n<p>The last week has been very frustrating for us , we have reached as high as 0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968 . And due to the strong correlation between our models we have not been able to break the 0.970 barrier . We are very curious to know how will our high cv blend perform on private when on public it even fails to touch our best</p>\n<p>Is it just us or anyone else facing the similar kind of gap? </p>",
      "rawMarkdown": "Hi all ,\nNow that 12 hours is left in this competition , I wish everyone the best , mark your last submissions carefully and may the shakeup be in your favour \n\nThe last week has been very frustrating for us , we have reached as high as 0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968 . And due to the strong correlation between our models we have not been able to break the 0.970 barrier . We are very curious to know how will our high cv blend perform on private when on public it even fails to touch our best\n\nIs it just us or anyone else facing the similar kind of gap?",
      "votes": null
    },
    {
      "id": "1240487",
      "postDate": "03/16/2021 12:47:22",
      "content": "<blockquote>\n  <p>0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968</p>\n</blockquote>\n<p>I agree there were lots of model which crossed 0.971 in CV but all end up in 0.97 for us</p>",
      "rawMarkdown": "> 0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968\n\nI agree there were lots of model which crossed 0.971 in CV but all end up in 0.97 for us",
      "votes": null
    },
    {
      "id": "1240578",
      "postDate": "03/16/2021 14:06:34",
      "content": "<p>If there are no leak of your CV, I suggest modify your title to 'ascending to the top'.</p>",
      "rawMarkdown": "If there are no leak of your CV, I suggest modify your title to 'ascending to the top'.",
      "votes": null
    },
    {
      "id": "1240588",
      "postDate": "03/16/2021 14:14:24",
      "content": "<p>I have a somewhat consistent difference of around 0.005 to 0.02 (mostly close to 0.01) by which public LB score are <strong>higher than my CV ROC AuCs</strong> (with a certain tendency for a larger gap for lower CV values like 0.92 and a smaller gap at value above 0.96). That may of course just be my random number seed or something - although I always assumed this was more the particular random (?) choice of a small number of images for the public LB (and for the private LB this difference may look very different) and looked the same for most. Admittedly, I never got any CV result above 0.97, so I don't know how that would look like with my CV scheme.</p>",
      "rawMarkdown": "I have a somewhat consistent difference of around 0.005 to 0.02 (mostly close to 0.01) by which public LB score are **higher than my CV ROC AuCs** (with a certain tendency for a larger gap for lower CV values like 0.92 and a smaller gap at value above 0.96). That may of course just be my random number seed or something - although I always assumed this was more the particular random (?) choice of a small number of images for the public LB (and for the private LB this difference may look very different) and looked the same for most. Admittedly, I never got any CV result above 0.97, so I don't know how that would look like with my CV scheme.",
      "votes": null
    },
    {
      "id": "1240590",
      "postDate": "03/16/2021 14:15:01",
      "content": "<p>I don't think there is any leak in our cv. We used the same splits as of <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> everywhere. Although, we are still unsure how well they are capable to perform on private dataset, as everyone predicting very low possibilities of shakeup. <br>\nEdit:  There is a leakage in cv.😬</p>",
      "rawMarkdown": "I don't think there is any leak in our cv. We used the same splits as of @underwearfitting everywhere. Although, we are still unsure how well they are capable to perform on private dataset, as everyone predicting very low possibilities of shakeup. \nEdit:  There is a leakage in cv.😬",
      "votes": null
    },
    {
      "id": "1240594",
      "postDate": "03/16/2021 14:16:37",
      "content": "<p>If you're using the NIH dataset, it's very hard to avoid this type of leakage</p>",
      "rawMarkdown": "If you're using the NIH dataset, it's very hard to avoid this type of leakage",
      "votes": null
    },
    {
      "id": "1240603",
      "postDate": "03/16/2021 14:25:20",
      "content": "<p>Correct, but we haven't used NIH pseudo labels anywhere in training except for using ammarwali pretrained weights(might have been trained on NIH?). Do you think it could still lead to leakage in cv?</p>",
      "rawMarkdown": "Correct, but we haven't used NIH pseudo labels anywhere in training except for using ammarwali pretrained weights(might have been trained on NIH?). Do you think it could still lead to leakage in cv?",
      "votes": null
    },
    {
      "id": "1240605",
      "postDate": "03/16/2021 14:27:53",
      "content": "<p>Do you use four stage pretrained weight? I guess that's the source of the leakage.</p>",
      "rawMarkdown": "Do you use four stage pretrained weight? I guess that's the source of the leakage.",
      "votes": null
    },
    {
      "id": "1240607",
      "postDate": "03/16/2021 14:28:54",
      "content": "<p>If you used some publicly available pretrained models, aka the chestx pretrained, there is high chance of leak😂</p>",
      "rawMarkdown": "If you used some publicly available pretrained models, aka the chestx pretrained, there is high chance of leak😂",
      "votes": null
    },
    {
      "id": "1240609",
      "postDate": "03/16/2021 14:30:13",
      "content": "<p>Yes, it was used. I think you're right. 👍</p>",
      "rawMarkdown": "Yes, it was used. I think you're right. 👍",
      "votes": null
    },
    {
      "id": "1240611",
      "postDate": "03/16/2021 14:31:31",
      "content": "<p>That might be partly true <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> but now we have come way too far , we also think there is something wrong with our cv</p>",
      "rawMarkdown": "That might be partly true @steamedsheep but now we have come way too far , we also think there is something wrong with our cv",
      "votes": null
    },
    {
      "id": "1240617",
      "postDate": "03/16/2021 14:34:24",
      "content": "<p>When I read about how the weight was generated, I realized it was super dangerous. “Took my best model weight and trained on Chestx” or sth along the line.</p>",
      "rawMarkdown": "When I read about how the weight was generated, I realized it was super dangerous. “Took my best model weight and trained on Chestx” or sth along the line.",
      "votes": null
    },
    {
      "id": "1240637",
      "postDate": "03/16/2021 14:45:49",
      "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> I guess its too late for us now haha , should have caught the mistake earlier</p>",
      "rawMarkdown": "underwearfitting I guess its too late for us now haha , should have caught the mistake earlier",
      "votes": null
    },
    {
      "id": "1240829",
      "postDate": "03/16/2021 16:43:19",
      "content": "<p>Same for me <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <br>\nI am facing a large gap in CV-LB as well… Maybe because I used NIH data set as well for pre-training. I am constantly hitting ~0.972 on my CV but only managed 0.967 on LB. I have been in this loop since last week 😜<br>\nI thought I was the only one and was afraid to ask. Glad that you shared…<br>\nToo late for me, but best of luck to you and your team 👍</p>",
      "rawMarkdown": "Same for me @tanulsingh077 \nI am facing a large gap in CV-LB as well... Maybe because I used NIH data set as well for pre-training. I am constantly hitting ~0.972 on my CV but only managed 0.967 on LB. I have been in this loop since last week 😜\nI thought I was the only one and was afraid to ask. Glad that you shared...\nToo late for me, but best of luck to you and your team 👍",
      "votes": null
    },
    {
      "id": "1240858",
      "postDate": "03/16/2021 17:01:47",
      "content": "<p>i cannot wait the read the top solutions. </p>",
      "rawMarkdown": "i cannot wait the read the top solutions.",
      "votes": null
    },
    {
      "id": "1240866",
      "postDate": "03/16/2021 17:13:02",
      "content": "<p>We tried so hard to find what makes people above 0.97. I am expecting something special </p>",
      "rawMarkdown": "We tried so hard to find what makes people above 0.97. I am expecting something special",
      "votes": null
    },
    {
      "id": "1240881",
      "postDate": "03/16/2021 17:31:49",
      "content": "<p>I cannot bump this enough. CNNs are really good at cheating.</p>",
      "rawMarkdown": "I cannot bump this enough. CNNs are really good at cheating.",
      "votes": null
    },
    {
      "id": "1241025",
      "postDate": "03/16/2021 20:24:15",
      "content": "<p>I didn't get any stable CV score. Anyway, I've entered too late. Juste here to read the top solutions once it is over. Best of luck for the final hours!</p>",
      "rawMarkdown": "I didn't get any stable CV score. Anyway, I've entered too late. Juste here to read the top solutions once it is over. Best of luck for the final hours!",
      "votes": null
    },
    {
      "id": "1241091",
      "postDate": "03/16/2021 21:48:37",
      "content": "<p>I haved suffered from this kind of gap, so I decided not to use NIH external datasets. That might affect my private score.</p>",
      "rawMarkdown": "I haved suffered from this kind of gap, so I decided not to use NIH external datasets. That might affect my private score.",
      "votes": null
    },
    {
      "id": "1241128",
      "postDate": "03/16/2021 22:58:25",
      "content": "<p>I suffered this gap on my models too. and it seems the top solutions have a magic / secret sauce to their recipe. Looking forward to learn from the solutions.</p>",
      "rawMarkdown": "I suffered this gap on my models too. and it seems the top solutions have a magic / secret sauce to their recipe. Looking forward to learn from the solutions.",
      "votes": null
    },
    {
      "id": "1241136",
      "postDate": "03/16/2021 23:10:55",
      "content": "<p>I don't know about other teams but our pipeline is intuitive. We made good use of annotations. </p>",
      "rawMarkdown": "I don't know about other teams but our pipeline is intuitive. We made good use of annotations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1240487,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "03/16/2021 12:47:22",
      "content": "<blockquote>\n  <p>0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968</p>\n</blockquote>\n<p>I agree there were lots of model which crossed 0.971 in CV but all end up in 0.97 for us</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1240578,
      "author_name": "steamedsheep",
      "author_url": "",
      "post_date": "03/16/2021 14:06:34",
      "content": "<p>If there are no leak of your CV, I suggest modify your title to 'ascending to the top'.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1240590,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/16/2021 14:15:01",
          "content": "<p>I don't think there is any leak in our cv. We used the same splits as of <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> everywhere. Although, we are still unsure how well they are capable to perform on private dataset, as everyone predicting very low possibilities of shakeup. <br>\nEdit:  There is a leakage in cv.😬</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240607,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "03/16/2021 14:28:54",
          "content": "<p>If you used some publicly available pretrained models, aka the chestx pretrained, there is high chance of leak😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1240588,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/16/2021 14:14:24",
      "content": "<p>I have a somewhat consistent difference of around 0.005 to 0.02 (mostly close to 0.01) by which public LB score are <strong>higher than my CV ROC AuCs</strong> (with a certain tendency for a larger gap for lower CV values like 0.92 and a smaller gap at value above 0.96). That may of course just be my random number seed or something - although I always assumed this was more the particular random (?) choice of a small number of images for the public LB (and for the private LB this difference may look very different) and looked the same for most. Admittedly, I never got any CV result above 0.97, so I don't know how that would look like with my CV scheme.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1240594,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "03/16/2021 14:16:37",
      "content": "<p>If you're using the NIH dataset, it's very hard to avoid this type of leakage</p>",
      "votes": null,
      "replies": [
        {
          "id": 1240603,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/16/2021 14:25:20",
          "content": "<p>Correct, but we haven't used NIH pseudo labels anywhere in training except for using ammarwali pretrained weights(might have been trained on NIH?). Do you think it could still lead to leakage in cv?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240605,
          "author_name": "steamedsheep",
          "author_url": "",
          "post_date": "03/16/2021 14:27:53",
          "content": "<p>Do you use four stage pretrained weight? I guess that's the source of the leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240609,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/16/2021 14:30:13",
          "content": "<p>Yes, it was used. I think you're right. 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240611,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "03/16/2021 14:31:31",
          "content": "<p>That might be partly true <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> but now we have come way too far , we also think there is something wrong with our cv</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240617,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "03/16/2021 14:34:24",
          "content": "<p>When I read about how the weight was generated, I realized it was super dangerous. “Took my best model weight and trained on Chestx” or sth along the line.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240637,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "03/16/2021 14:45:49",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> I guess its too late for us now haha , should have caught the mistake earlier</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240881,
          "author_name": "jy2tong",
          "author_url": "",
          "post_date": "03/16/2021 17:31:49",
          "content": "<p>I cannot bump this enough. CNNs are really good at cheating.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1240829,
      "author_name": "manabendrarout",
      "author_url": "",
      "post_date": "03/16/2021 16:43:19",
      "content": "<p>Same for me <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <br>\nI am facing a large gap in CV-LB as well… Maybe because I used NIH data set as well for pre-training. I am constantly hitting ~0.972 on my CV but only managed 0.967 on LB. I have been in this loop since last week 😜<br>\nI thought I was the only one and was afraid to ask. Glad that you shared…<br>\nToo late for me, but best of luck to you and your team 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1240858,
      "author_name": "yl1202",
      "author_url": "",
      "post_date": "03/16/2021 17:01:47",
      "content": "<p>i cannot wait the read the top solutions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1240866,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "03/16/2021 17:13:02",
          "content": "<p>We tried so hard to find what makes people above 0.97. I am expecting something special </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241025,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/16/2021 20:24:15",
      "content": "<p>I didn't get any stable CV score. Anyway, I've entered too late. Juste here to read the top solutions once it is over. Best of luck for the final hours!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241091,
      "author_name": "yoshitaka1105",
      "author_url": "",
      "post_date": "03/16/2021 21:48:37",
      "content": "<p>I haved suffered from this kind of gap, so I decided not to use NIH external datasets. That might affect my private score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241128,
      "author_name": "roydatascience",
      "author_url": "",
      "post_date": "03/16/2021 22:58:25",
      "content": "<p>I suffered this gap on my models too. and it seems the top solutions have a magic / secret sauce to their recipe. Looking forward to learn from the solutions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241136,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "03/16/2021 23:10:55",
          "content": "<p>I don't know about other teams but our pipeline is intuitive. We made good use of annotations. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1240460": "Hi all ,\nNow that 12 hours is left in this competition , I wish everyone the best , mark your last submissions carefully and may the shakeup be in your favour \n\nThe last week has been very frustrating for us , we have reached as high as 0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968 . And due to the strong correlation between our models we have not been able to break the 0.970 barrier . We are very curious to know how will our high cv blend perform on private when on public it even fails to touch our best\n\nIs it just us or anyone else facing the similar kind of gap?",
    "1240487": "> 0.973 on CV with single models but to our surprise , every last one of our models with good cv have failed us on LB with the max being 0.968\n\nI agree there were lots of model which crossed 0.971 in CV but all end up in 0.97 for us",
    "1240578": "If there are no leak of your CV, I suggest modify your title to 'ascending to the top'.",
    "1240588": "I have a somewhat consistent difference of around 0.005 to 0.02 (mostly close to 0.01) by which public LB score are **higher than my CV ROC AuCs** (with a certain tendency for a larger gap for lower CV values like 0.92 and a smaller gap at value above 0.96). That may of course just be my random number seed or something - although I always assumed this was more the particular random (?) choice of a small number of images for the public LB (and for the private LB this difference may look very different) and looked the same for most. Admittedly, I never got any CV result above 0.97, so I don't know how that would look like with my CV scheme.",
    "1240590": "I don't think there is any leak in our cv. We used the same splits as of @underwearfitting everywhere. Although, we are still unsure how well they are capable to perform on private dataset, as everyone predicting very low possibilities of shakeup. \nEdit:  There is a leakage in cv.😬",
    "1240594": "If you're using the NIH dataset, it's very hard to avoid this type of leakage",
    "1240603": "Correct, but we haven't used NIH pseudo labels anywhere in training except for using ammarwali pretrained weights(might have been trained on NIH?). Do you think it could still lead to leakage in cv?",
    "1240605": "Do you use four stage pretrained weight? I guess that's the source of the leakage.",
    "1240607": "If you used some publicly available pretrained models, aka the chestx pretrained, there is high chance of leak😂",
    "1240609": "Yes, it was used. I think you're right. 👍",
    "1240611": "That might be partly true @steamedsheep but now we have come way too far , we also think there is something wrong with our cv",
    "1240617": "When I read about how the weight was generated, I realized it was super dangerous. “Took my best model weight and trained on Chestx” or sth along the line.",
    "1240637": "underwearfitting I guess its too late for us now haha , should have caught the mistake earlier",
    "1240829": "Same for me @tanulsingh077 \nI am facing a large gap in CV-LB as well... Maybe because I used NIH data set as well for pre-training. I am constantly hitting ~0.972 on my CV but only managed 0.967 on LB. I have been in this loop since last week 😜\nI thought I was the only one and was afraid to ask. Glad that you shared...\nToo late for me, but best of luck to you and your team 👍",
    "1240858": "i cannot wait the read the top solutions.",
    "1240866": "We tried so hard to find what makes people above 0.97. I am expecting something special",
    "1240881": "I cannot bump this enough. CNNs are really good at cheating.",
    "1241025": "I didn't get any stable CV score. Anyway, I've entered too late. Juste here to read the top solutions once it is over. Best of luck for the final hours!",
    "1241091": "I haved suffered from this kind of gap, so I decided not to use NIH external datasets. That might affect my private score.",
    "1241128": "I suffered this gap on my models too. and it seems the top solutions have a magic / secret sauce to their recipe. Looking forward to learn from the solutions.",
    "1241136": "I don't know about other teams but our pipeline is intuitive. We made good use of annotations."
  },
  "source": "meta"
}