{
  "id": 223168,
  "title": "CV and LB mismatch",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/223168",
  "author_name": "",
  "post_date": "2021-03-02T18:19:38.776668900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>that seems to be a newbie's question but i just don't understand what's going on. I'm training a model with a holdout set (with train-test split) and get a score as high as 0.90 AUC. The i do a submission and a LB score is just 0.54.</p>\n<p>Ok, then i do cross-validation (both KFold and GroupKFold, K = 5), again get CV AUC = 0.91 and again LB score is only 0.53. </p>\n<p>I tried to find a leak or some kind of a stupid mistake in my scripts but it all seems alright.</p>\n<p>So my question is where can i read about proper CV strategies or find good examples of the strategies so i can compare with mine and finally find out what i am doing wrong? Or probably one can point me to potential pitfalls i could miss? </p>\n<p>Ant help is appreciated, thank you in advance!</p>",
  "messages": [
    {
      "id": "1224442",
      "postDate": "03/02/2021 18:19:38",
      "content": "<p>Hi all,</p>\n<p>that seems to be a newbie's question but i just don't understand what's going on. I'm training a model with a holdout set (with train-test split) and get a score as high as 0.90 AUC. The i do a submission and a LB score is just 0.54.</p>\n<p>Ok, then i do cross-validation (both KFold and GroupKFold, K = 5), again get CV AUC = 0.91 and again LB score is only 0.53. </p>\n<p>I tried to find a leak or some kind of a stupid mistake in my scripts but it all seems alright.</p>\n<p>So my question is where can i read about proper CV strategies or find good examples of the strategies so i can compare with mine and finally find out what i am doing wrong? Or probably one can point me to potential pitfalls i could miss? </p>\n<p>Ant help is appreciated, thank you in advance!</p>",
      "rawMarkdown": "Hi all,\n\nthat seems to be a newbie's question but i just don't understand what's going on. I'm training a model with a holdout set (with train-test split) and get a score as high as 0.90 AUC. The i do a submission and a LB score is just 0.54.\n\nOk, then i do cross-validation (both KFold and GroupKFold, K = 5), again get CV AUC = 0.91 and again LB score is only 0.53. \n\nI tried to find a leak or some kind of a stupid mistake in my scripts but it all seems alright.\n\nSo my question is where can i read about proper CV strategies or find good examples of the strategies so i can compare with mine and finally find out what i am doing wrong? Or probably one can point me to potential pitfalls i could miss? \n\nAnt help is appreciated, thank you in advance!",
      "votes": null
    },
    {
      "id": "1224453",
      "postDate": "03/02/2021 18:31:16",
      "content": "<p>Seeing a lb score that low I have to assume that something is wrong with your inference script. Maybe something is different between your training procedure and your inference code. That would be the first thing I would look at. There is very little possibility you are overfitting that badly. </p>",
      "rawMarkdown": "Seeing a lb score that low I have to assume that something is wrong with your inference script. Maybe something is different between your training procedure and your inference code. That would be the first thing I would look at. There is very little possibility you are overfitting that badly.",
      "votes": null
    },
    {
      "id": "1224455",
      "postDate": "03/02/2021 18:35:03",
      "content": "<p>I highly recommend you read <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a>'s notebook on splitting folds that you can find <a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">here</a>. I believe most competitors are using this strategy. (I am using a modified version of it). </p>\n<p>That being said, I think you have some sort of bug if you are getting 0.53 leaderboard as the <a href=\"https://www.kaggle.com/titericz/baseline-mean-average\" target=\"_blank\">mean average baseline</a> of AUC is 0.5</p>",
      "rawMarkdown": "I highly recommend you read [@underwearfitting](https://www.kaggle.com/underwearfitting)'s notebook on splitting folds that you can find [here](https://www.kaggle.com/underwearfitting/how-to-properly-split-folds). I believe most competitors are using this strategy. (I am using a modified version of it). \n\nThat being said, I think you have some sort of bug if you are getting 0.53 leaderboard as the [mean average baseline](https://www.kaggle.com/titericz/baseline-mean-average) of AUC is 0.5",
      "votes": null
    },
    {
      "id": "1224558",
      "postDate": "03/02/2021 20:55:33",
      "content": "<blockquote>\n  <p>Or probably one can point me to potential pitfalls i could miss?</p>\n</blockquote>\n<p>Predict the same image on your local validation scheme and your online notebook. Make sure that both images (right before they go into your model) and predictions are equal.</p>\n<p>You almost for sure have some differences between local and online inference. It could be as simple as forgetting to change the name of the model that you predict on or forgetting to normalize image and predicting on 0-255 images instead of 0.0-1.0 ones.</p>",
      "rawMarkdown": "> Or probably one can point me to potential pitfalls i could miss?\n\nPredict the same image on your local validation scheme and your online notebook. Make sure that both images (right before they go into your model) and predictions are equal.\n\nYou almost for sure have some differences between local and online inference. It could be as simple as forgetting to change the name of the model that you predict on or forgetting to normalize image and predicting on 0-255 images instead of 0.0-1.0 ones.",
      "votes": null
    },
    {
      "id": "1224613",
      "postDate": "03/02/2021 22:46:33",
      "content": "<p>I'm wondering if a GroupKFold strategy would be superior when averaging the predictions from the K classifiers, compared to using StratifiedGroupKFold. I might need to test this at the end of the competition.</p>",
      "rawMarkdown": "I'm wondering if a GroupKFold strategy would be superior when averaging the predictions from the K classifiers, compared to using StratifiedGroupKFold. I might need to test this at the end of the competition.",
      "votes": null
    },
    {
      "id": "1226342",
      "postDate": "03/04/2021 13:11:21",
      "content": "<p>I had the same issue. The problem was that I was getting images' paths as os.listdir(imagepath) instead of df[\"paths\"].values. In this case the paths are automatically sorted and you predict labels for the wrong images. </p>",
      "rawMarkdown": "I had the same issue. The problem was that I was getting images' paths as os.listdir(imagepath) instead of df[\"paths\"].values. In this case the paths are automatically sorted and you predict labels for the wrong images.",
      "votes": null
    },
    {
      "id": "1227184",
      "postDate": "03/05/2021 09:17:20",
      "content": "<p>That was particularly the case! Thank you!</p>",
      "rawMarkdown": "That was particularly the case! Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1224453,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/02/2021 18:31:16",
      "content": "<p>Seeing a lb score that low I have to assume that something is wrong with your inference script. Maybe something is different between your training procedure and your inference code. That would be the first thing I would look at. There is very little possibility you are overfitting that badly. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1224455,
      "author_name": "tuckerarrants",
      "author_url": "",
      "post_date": "03/02/2021 18:35:03",
      "content": "<p>I highly recommend you read <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a>'s notebook on splitting folds that you can find <a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">here</a>. I believe most competitors are using this strategy. (I am using a modified version of it). </p>\n<p>That being said, I think you have some sort of bug if you are getting 0.53 leaderboard as the <a href=\"https://www.kaggle.com/titericz/baseline-mean-average\" target=\"_blank\">mean average baseline</a> of AUC is 0.5</p>",
      "votes": null,
      "replies": [
        {
          "id": 1224613,
          "author_name": "gianlucarossi",
          "author_url": "",
          "post_date": "03/02/2021 22:46:33",
          "content": "<p>I'm wondering if a GroupKFold strategy would be superior when averaging the predictions from the K classifiers, compared to using StratifiedGroupKFold. I might need to test this at the end of the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1224558,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "03/02/2021 20:55:33",
      "content": "<blockquote>\n  <p>Or probably one can point me to potential pitfalls i could miss?</p>\n</blockquote>\n<p>Predict the same image on your local validation scheme and your online notebook. Make sure that both images (right before they go into your model) and predictions are equal.</p>\n<p>You almost for sure have some differences between local and online inference. It could be as simple as forgetting to change the name of the model that you predict on or forgetting to normalize image and predicting on 0-255 images instead of 0.0-1.0 ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226342,
      "author_name": "vadimtimakin",
      "author_url": "",
      "post_date": "03/04/2021 13:11:21",
      "content": "<p>I had the same issue. The problem was that I was getting images' paths as os.listdir(imagepath) instead of df[\"paths\"].values. In this case the paths are automatically sorted and you predict labels for the wrong images. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1227184,
          "author_name": "denisstenyushkin",
          "author_url": "",
          "post_date": "03/05/2021 09:17:20",
          "content": "<p>That was particularly the case! Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1224442": "Hi all,\n\nthat seems to be a newbie's question but i just don't understand what's going on. I'm training a model with a holdout set (with train-test split) and get a score as high as 0.90 AUC. The i do a submission and a LB score is just 0.54.\n\nOk, then i do cross-validation (both KFold and GroupKFold, K = 5), again get CV AUC = 0.91 and again LB score is only 0.53. \n\nI tried to find a leak or some kind of a stupid mistake in my scripts but it all seems alright.\n\nSo my question is where can i read about proper CV strategies or find good examples of the strategies so i can compare with mine and finally find out what i am doing wrong? Or probably one can point me to potential pitfalls i could miss? \n\nAnt help is appreciated, thank you in advance!",
    "1224453": "Seeing a lb score that low I have to assume that something is wrong with your inference script. Maybe something is different between your training procedure and your inference code. That would be the first thing I would look at. There is very little possibility you are overfitting that badly.",
    "1224455": "I highly recommend you read [@underwearfitting](https://www.kaggle.com/underwearfitting)'s notebook on splitting folds that you can find [here](https://www.kaggle.com/underwearfitting/how-to-properly-split-folds). I believe most competitors are using this strategy. (I am using a modified version of it). \n\nThat being said, I think you have some sort of bug if you are getting 0.53 leaderboard as the [mean average baseline](https://www.kaggle.com/titericz/baseline-mean-average) of AUC is 0.5",
    "1224558": "> Or probably one can point me to potential pitfalls i could miss?\n\nPredict the same image on your local validation scheme and your online notebook. Make sure that both images (right before they go into your model) and predictions are equal.\n\nYou almost for sure have some differences between local and online inference. It could be as simple as forgetting to change the name of the model that you predict on or forgetting to normalize image and predicting on 0-255 images instead of 0.0-1.0 ones.",
    "1224613": "I'm wondering if a GroupKFold strategy would be superior when averaging the predictions from the K classifiers, compared to using StratifiedGroupKFold. I might need to test this at the end of the competition.",
    "1226342": "I had the same issue. The problem was that I was getting images' paths as os.listdir(imagepath) instead of df[\"paths\"].values. In this case the paths are automatically sorted and you predict labels for the wrong images.",
    "1227184": "That was particularly the case! Thank you!"
  },
  "source": "meta"
}