{
  "id": 160312,
  "title": "Why are validation roc_auc_score and submission score so different?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160312",
  "author_name": "",
  "post_date": "2020-06-20T18:38:12.215664800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I use roc_auc_score at the end of each epoch as I do transfer learning on an Efficient Net B1 model.  When I do it at the end of the epoch, my score is around 0.85.  However, when I submit my model, I get a score of 0.7.  I'm not sure why this is the case and was wondering if somebody could explain it.  Thanks!</p>",
  "messages": [
    {
      "id": "894772",
      "postDate": "06/20/2020 18:38:12",
      "content": "<p>I use roc_auc_score at the end of each epoch as I do transfer learning on an Efficient Net B1 model.  When I do it at the end of the epoch, my score is around 0.85.  However, when I submit my model, I get a score of 0.7.  I'm not sure why this is the case and was wondering if somebody could explain it.  Thanks!</p>",
      "rawMarkdown": "I use roc_auc_score at the end of each epoch as I do transfer learning on an Efficient Net B1 model.  When I do it at the end of the epoch, my score is around 0.85.  However, when I submit my model, I get a score of 0.7.  I'm not sure why this is the case and was wondering if somebody could explain it.  Thanks!",
      "votes": null
    },
    {
      "id": "895505",
      "postDate": "06/21/2020 12:14:06",
      "content": "<p>That is quite a big difference, maybe something is wrong with your train/validation split or cross validation strategy or how you calculate the auc score during training. </p>\n\n<p>It is important to apply GroupKFold on ['patient_id'] in this competition, because the dataset contains many duplicate images per patient. If only StratifiedKFold on ['target'] is applied, you will have partially the same images in training and validation set (leak). </p>\n\n<p>```\nfrom sklearn.model_selection import GroupKFold</p>\n\n<p>N_SPLITS = 5 <br>\ncv = GroupKFold(n_splits=N_SPLITS) <br>\nsplits = list(cv.split(train_df, train_df['target'], groups=train_df['patient_id']))</p>\n\n<p>FOLD = 0    # [0-4] for N_SPLITS = 5 <br>\ntrain_data = train_df.iloc[splits[FOLD][0]]\nval_data = train_df.iloc[splits[FOLD][1]]\n```</p>\n\n<p>It would be also possible to stratify by a combination of ['age', 'anatom', 'sex', 'target'], but in my experience stratification just by ['target'] and using GroupKFold works well.</p>\n\n<p>Next is, what many others have already noticed, when external data is used, the difference between CV and LB increases. To get more reliable results for your CV and decrease the gap between CV and LB, only add external data for instance ISIC 2019 to your training split, but do not use it in your validation split. </p>",
      "rawMarkdown": "That is quite a big difference, maybe something is wrong with your train/validation split or cross validation strategy or how you calculate the auc score during training. \n\nIt is important to apply GroupKFold on ['patient_id'] in this competition, because the dataset contains many duplicate images per patient. If only StratifiedKFold on ['target'] is applied, you will have partially the same images in training and validation set (leak). \n\n```\nfrom sklearn.model_selection import GroupKFold\n\nN_SPLITS = 5       \ncv = GroupKFold(n_splits=N_SPLITS)    \nsplits = list(cv.split(train_df, train_df['target'], groups=train_df['patient_id']))\n            \nFOLD = 0    # [0-4] for N_SPLITS = 5    \ntrain_data = train_df.iloc[splits[FOLD][0]]\nval_data = train_df.iloc[splits[FOLD][1]]\n```\n\n\nIt would be also possible to stratify by a combination of ['age', 'anatom', 'sex', 'target'], but in my experience stratification just by ['target'] and using GroupKFold works well.\n\nNext is, what many others have already noticed, when external data is used, the difference between CV and LB increases. To get more reliable results for your CV and decrease the gap between CV and LB, only add external data for instance ISIC 2019 to your training split, but do not use it in your validation split.",
      "votes": null
    },
    {
      "id": "896200",
      "postDate": "06/22/2020 00:46:51",
      "content": "<p>Thank you for bringing this to my attention! I am not using any external data, but I will definitely use GroupKFold to prevent any leak.</p>",
      "rawMarkdown": "Thank you for bringing this to my attention! I am not using any external data, but I will definitely use GroupKFold to prevent any leak.",
      "votes": null
    },
    {
      "id": "918631",
      "postDate": "07/07/2020 11:22:51",
      "content": "<p><a href=\"/khabel\">@khabel</a> I am using <code>GroupKFold</code>and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?</p>",
      "rawMarkdown": "khabel I am using `GroupKFold `and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 895505,
      "author_name": "khabel",
      "author_url": "",
      "post_date": "06/21/2020 12:14:06",
      "content": "<p>That is quite a big difference, maybe something is wrong with your train/validation split or cross validation strategy or how you calculate the auc score during training. </p>\n\n<p>It is important to apply GroupKFold on ['patient_id'] in this competition, because the dataset contains many duplicate images per patient. If only StratifiedKFold on ['target'] is applied, you will have partially the same images in training and validation set (leak). </p>\n\n<p>```\nfrom sklearn.model_selection import GroupKFold</p>\n\n<p>N_SPLITS = 5 <br>\ncv = GroupKFold(n_splits=N_SPLITS) <br>\nsplits = list(cv.split(train_df, train_df['target'], groups=train_df['patient_id']))</p>\n\n<p>FOLD = 0    # [0-4] for N_SPLITS = 5 <br>\ntrain_data = train_df.iloc[splits[FOLD][0]]\nval_data = train_df.iloc[splits[FOLD][1]]\n```</p>\n\n<p>It would be also possible to stratify by a combination of ['age', 'anatom', 'sex', 'target'], but in my experience stratification just by ['target'] and using GroupKFold works well.</p>\n\n<p>Next is, what many others have already noticed, when external data is used, the difference between CV and LB increases. To get more reliable results for your CV and decrease the gap between CV and LB, only add external data for instance ISIC 2019 to your training split, but do not use it in your validation split. </p>",
      "votes": null,
      "replies": [
        {
          "id": 896200,
          "author_name": "alicia183",
          "author_url": "",
          "post_date": "06/22/2020 00:46:51",
          "content": "<p>Thank you for bringing this to my attention! I am not using any external data, but I will definitely use GroupKFold to prevent any leak.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918631,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "07/07/2020 11:22:51",
          "content": "<p><a href=\"/khabel\">@khabel</a> I am using <code>GroupKFold</code>and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "894772": "I use roc_auc_score at the end of each epoch as I do transfer learning on an Efficient Net B1 model.  When I do it at the end of the epoch, my score is around 0.85.  However, when I submit my model, I get a score of 0.7.  I'm not sure why this is the case and was wondering if somebody could explain it.  Thanks!",
    "895505": "That is quite a big difference, maybe something is wrong with your train/validation split or cross validation strategy or how you calculate the auc score during training. \n\nIt is important to apply GroupKFold on ['patient_id'] in this competition, because the dataset contains many duplicate images per patient. If only StratifiedKFold on ['target'] is applied, you will have partially the same images in training and validation set (leak). \n\n```\nfrom sklearn.model_selection import GroupKFold\n\nN_SPLITS = 5       \ncv = GroupKFold(n_splits=N_SPLITS)    \nsplits = list(cv.split(train_df, train_df['target'], groups=train_df['patient_id']))\n            \nFOLD = 0    # [0-4] for N_SPLITS = 5    \ntrain_data = train_df.iloc[splits[FOLD][0]]\nval_data = train_df.iloc[splits[FOLD][1]]\n```\n\n\nIt would be also possible to stratify by a combination of ['age', 'anatom', 'sex', 'target'], but in my experience stratification just by ['target'] and using GroupKFold works well.\n\nNext is, what many others have already noticed, when external data is used, the difference between CV and LB increases. To get more reliable results for your CV and decrease the gap between CV and LB, only add external data for instance ISIC 2019 to your training split, but do not use it in your validation split.",
    "896200": "Thank you for bringing this to my attention! I am not using any external data, but I will definitely use GroupKFold to prevent any leak.",
    "918631": "khabel I am using `GroupKFold `and observed that it also considering the percentage of target in each fold along with specified group as well. Am I getting it by chance or GroupKFold ensures stratification as well?"
  },
  "source": "meta"
}