{
  "id": 186683,
  "title": "Is your LB or your CV the best indication for selecting final submissions?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/186683",
  "author_name": "from coffee import *",
  "post_date": "2020-09-25T13:54:38.321000",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Dear fellow Kagglers,</p>\n<p>is your LB score or your CV score the best indication for selecting final submissions?</p>\n<p>My strong feeling is that due to the fact that <strong>only 15% of the hidden dataset</strong> is used for public LB calculation, many teams are currently strongly overfitting to those 15%.</p>\n<p>So we should better trust our CV!<br>\nTo get a proper CV, we need to mitigate any chance of data leakage: e.g. having the same patient in train and in test data. You will most probably have this kind of leakage if you are using (stratified)-K-Fold to split your train data in train and validation set.</p>\n<p>The best way for this competition is a sklearns <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold</a>, mitigating that a patient used for training is also part of the validation set and therefore has impact on your CV.</p>\n<p>Are you using GroupKFold, is your CV correlated with the LB?<br>\nFor me, it's mostly the case. <br>\nWhat are your results? Are you using GroupKFold? Or even GroupKFold <em>and</em> stratification?</p>",
  "messages": [
    {
      "id": 1026697,
      "postDate": "2020-09-25T13:54:38.320Z",
      "content": "<p>Dear fellow Kagglers,</p>\n<p>is your LB score or your CV score the best indication for selecting final submissions?</p>\n<p>My strong feeling is that due to the fact that <strong>only 15% of the hidden dataset</strong> is used for public LB calculation, many teams are currently strongly overfitting to those 15%.</p>\n<p>So we should better trust our CV!<br>\nTo get a proper CV, we need to mitigate any chance of data leakage: e.g. having the same patient in train and in test data. You will most probably have this kind of leakage if you are using (stratified)-K-Fold to split your train data in train and validation set.</p>\n<p>The best way for this competition is a sklearns <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold</a>, mitigating that a patient used for training is also part of the validation set and therefore has impact on your CV.</p>\n<p>Are you using GroupKFold, is your CV correlated with the LB?<br>\nFor me, it's mostly the case. <br>\nWhat are your results? Are you using GroupKFold? Or even GroupKFold <em>and</em> stratification?</p>",
      "rawMarkdown": "Dear fellow Kagglers,\n\nis your LB score or your CV score the best indication for selecting final submissions?\n\nMy strong feeling is that due to the fact that **only 15% of the hidden dataset** is used for public LB calculation, many teams are currently strongly overfitting to those 15%.\n\nSo we should better trust our CV!\nTo get a proper CV, we need to mitigate any chance of data leakage: e.g. having the same patient in train and in test data. You will most probably have this kind of leakage if you are using (stratified)-K-Fold to split your train data in train and validation set.\n\nThe best way for this competition is a sklearns [GroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html), mitigating that a patient used for training is also part of the validation set and therefore has impact on your CV.\n\nAre you using GroupKFold, is your CV correlated with the LB?\nFor me, it's mostly the case. \nWhat are your results? Are you using GroupKFold? Or even GroupKFold *and* stratification?\n",
      "votes": 7
    },
    {
      "id": 1026740,
      "postDate": "2020-09-25T14:40:55.850Z",
      "content": "<p>I agree the CV should be more representative of the model generalization than LB. There are 188 patients in the private set, and 15% are used to calculate LB (~ 28 patients). If you train using 5-groupkfold you will validate your model on 5 folds of 35 patients each. </p>\n<p>Using stratified groupkfold have worked well for me. </p>",
      "rawMarkdown": "I agree the CV should be more representative of the model generalization than LB. There are 188 patients in the private set, and 15% are used to calculate LB (~ 28 patients). If you train using 5-groupkfold you will validate your model on 5 folds of 35 patients each. \n\nUsing stratified groupkfold have worked well for me. ",
      "votes": 3
    },
    {
      "id": 1028130,
      "postDate": "2020-09-26T16:05:09.007Z",
      "content": "<p>Initially I did not find any correlation between CV and LB scores. Later, I started calculating the CV based on last 3 weeks only and now the CV and LB scores show a very good correlation. </p>",
      "rawMarkdown": "Initially I did not find any correlation between CV and LB scores. Later, I started calculating the CV based on last 3 weeks only and now the CV and LB scores show a very good correlation. ",
      "votes": 1
    },
    {
      "id": 1026880,
      "postDate": "2020-09-25T16:10:35.970Z",
      "content": "<p>Yes, I think <code>GroupKFold</code> is suitable. The LB is saturated is with high-scoring public kernels with hyperparameter tuning, which is <strong>NOT</strong> suitable for the private LB as a huge shake-up is imminent.</p>\n<p>Always <strong>trust your CV!</strong></p>",
      "rawMarkdown": "Yes, I think `GroupKFold` is suitable. The LB is saturated is with high-scoring public kernels with hyperparameter tuning, which is **NOT** suitable for the private LB as a huge shake-up is imminent.\n\nAlways **trust your CV!**",
      "votes": 2
    },
    {
      "id": 1027062,
      "postDate": "2020-09-25T19:36:57.017Z",
      "content": "<p>I agree with you!</p>\n<p>But group Fold CV score at hand is below -7 and LB will be the worst score, -24.7981. <br>\nIs there a similar person? Or do My notebook has some bugs?</p>\n<p>Can I trust CV? I don't know what to do.</p>",
      "rawMarkdown": "I agree with you!\n\nBut group Fold CV score at hand is below -7 and LB will be the worst score, -24.7981. \nIs there a similar person? Or do My notebook has some bugs?\n\nCan I trust CV? I don't know what to do.",
      "replies": [
        {
          "id": 1027273,
          "postDate": "2020-09-26T02:36:46.550Z",
          "content": "<p>I guess, -24.7981 is sample_submission_score.<br>\nThere must be some bugs or mistakes.</p>",
          "rawMarkdown": "I guess, -24.7981 is sample_submission_score.\nThere must be some bugs or mistakes.",
          "votes": 1
        },
        {
          "id": 1027299,
          "postDate": "2020-09-26T03:36:27.853Z",
          "content": "<p>Wonderful!! <br>\nI was wondering if the CT images couldn't be used for learning because I used CT images for NN when I got this score.</p>\n<p>I've been looking for it for a long time, but I haven't been able to find a bug. <br>\nAnyway, thank you for your kind advice!!</p>",
          "rawMarkdown": "Wonderful!! \nI was wondering if the CT images couldn't be used for learning because I used CT images for NN when I got this score.\n\nI've been looking for it for a long time, but I haven't been able to find a bug. \nAnyway, thank you for your kind advice!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1027270,
      "postDate": "2020-09-26T02:33:52.877Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1026740,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "2020-09-25T14:40:55.850000",
      "content": "<p>I agree the CV should be more representative of the model generalization than LB. There are 188 patients in the private set, and 15% are used to calculate LB (~ 28 patients). If you train using 5-groupkfold you will validate your model on 5 folds of 35 patients each. </p>\n<p>Using stratified groupkfold have worked well for me. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1028130,
      "author_name": "Abhishek Bhat",
      "author_url": "",
      "post_date": "2020-09-26T16:05:09.007000",
      "content": "<p>Initially I did not find any correlation between CV and LB scores. Later, I started calculating the CV based on last 3 weeks only and now the CV and LB scores show a very good correlation. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1026880,
      "author_name": "Aadhav Vignesh",
      "author_url": "",
      "post_date": "2020-09-25T16:10:35.970000",
      "content": "<p>Yes, I think <code>GroupKFold</code> is suitable. The LB is saturated is with high-scoring public kernels with hyperparameter tuning, which is <strong>NOT</strong> suitable for the private LB as a huge shake-up is imminent.</p>\n<p>Always <strong>trust your CV!</strong></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1027062,
      "author_name": "KFurudate",
      "author_url": "",
      "post_date": "2020-09-25T19:36:57.017000",
      "content": "<p>I agree with you!</p>\n<p>But group Fold CV score at hand is below -7 and LB will be the worst score, -24.7981. <br>\nIs there a similar person? Or do My notebook has some bugs?</p>\n<p>Can I trust CV? I don't know what to do.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1027273,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2020-09-26T02:36:46.550000",
          "content": "<p>I guess, -24.7981 is sample_submission_score.<br>\nThere must be some bugs or mistakes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1027299,
          "author_name": "KFurudate",
          "author_url": "",
          "post_date": "2020-09-26T03:36:27.853000",
          "content": "<p>Wonderful!! <br>\nI was wondering if the CT images couldn't be used for learning because I used CT images for NN when I got this score.</p>\n<p>I've been looking for it for a long time, but I haven't been able to find a bug. <br>\nAnyway, thank you for your kind advice!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1027270,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-26T02:33:52.877000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1026697": "Dear fellow Kagglers,\n\nis your LB score or your CV score the best indication for selecting final submissions?\n\nMy strong feeling is that due to the fact that **only 15% of the hidden dataset** is used for public LB calculation, many teams are currently strongly overfitting to those 15%.\n\nSo we should better trust our CV!\nTo get a proper CV, we need to mitigate any chance of data leakage: e.g. having the same patient in train and in test data. You will most probably have this kind of leakage if you are using (stratified)-K-Fold to split your train data in train and validation set.\n\nThe best way for this competition is a sklearns [GroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html), mitigating that a patient used for training is also part of the validation set and therefore has impact on your CV.\n\nAre you using GroupKFold, is your CV correlated with the LB?\nFor me, it's mostly the case. \nWhat are your results? Are you using GroupKFold? Or even GroupKFold *and* stratification?\n",
    "1026740": "I agree the CV should be more representative of the model generalization than LB. There are 188 patients in the private set, and 15% are used to calculate LB (~ 28 patients). If you train using 5-groupkfold you will validate your model on 5 folds of 35 patients each. \n\nUsing stratified groupkfold have worked well for me. ",
    "1028130": "Initially I did not find any correlation between CV and LB scores. Later, I started calculating the CV based on last 3 weeks only and now the CV and LB scores show a very good correlation. ",
    "1026880": "Yes, I think `GroupKFold` is suitable. The LB is saturated is with high-scoring public kernels with hyperparameter tuning, which is **NOT** suitable for the private LB as a huge shake-up is imminent.\n\nAlways **trust your CV!**",
    "1027062": "I agree with you!\n\nBut group Fold CV score at hand is below -7 and LB will be the worst score, -24.7981. \nIs there a similar person? Or do My notebook has some bugs?\n\nCan I trust CV? I don't know what to do.",
    "1027270": ""
  }
}