{
  "id": 165820,
  "title": "Public LB dramatically higher than local CV",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165820",
  "author_name": "",
  "post_date": "2020-07-11T06:38:06.421881300Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Inspired by Abhishek’s notebooks (<a href=\"https://www.kaggle.com/abhishek/melanoma-detection-with-pytorch\">1</a>,<a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\">2</a>), finally I took the plunge and built my very first pytorch model this week (<a href=\"https://www.kaggle.com/dunklerwald/effnetb1-public-lb-is-higher-than-local-cv\">here</a>). What I can’t explain is why my model stably generates suspiciously high public LB.\nE.g.: \nfolds=3, LR= 0.001, local CV 0.89268, public LB 0.897\nfolds=3, LR= 0.0001, local CV 0.86827, public LB 0.909. </p>\n\n<p>I played out with different number of folds, learning rates and scheduler settings – in all experiments the model demonstrated public LB higher than local CV. I assume what I observe is not LB overfitting but rather some trivial mistake in the code I can’t identify. Would be very grateful for your hints.</p>",
  "messages": [
    {
      "id": "923844",
      "postDate": "07/11/2020 06:38:06",
      "content": "<p>Inspired by Abhishek’s notebooks (<a href=\"https://www.kaggle.com/abhishek/melanoma-detection-with-pytorch\">1</a>,<a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\">2</a>), finally I took the plunge and built my very first pytorch model this week (<a href=\"https://www.kaggle.com/dunklerwald/effnetb1-public-lb-is-higher-than-local-cv\">here</a>). What I can’t explain is why my model stably generates suspiciously high public LB.\nE.g.: \nfolds=3, LR= 0.001, local CV 0.89268, public LB 0.897\nfolds=3, LR= 0.0001, local CV 0.86827, public LB 0.909. </p>\n\n<p>I played out with different number of folds, learning rates and scheduler settings – in all experiments the model demonstrated public LB higher than local CV. I assume what I observe is not LB overfitting but rather some trivial mistake in the code I can’t identify. Would be very grateful for your hints.</p>",
      "rawMarkdown": "Inspired by Abhishek’s notebooks ([1](https://www.kaggle.com/abhishek/melanoma-detection-with-pytorch),[2](https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu)), finally I took the plunge and built my very first pytorch model this week ([here](https://www.kaggle.com/dunklerwald/effnetb1-public-lb-is-higher-than-local-cv)). What I can’t explain is why my model stably generates suspiciously high public LB.\nE.g.: \nfolds=3, LR= 0.001, local CV 0.89268, public LB 0.897\nfolds=3, LR= 0.0001, local CV 0.86827, public LB 0.909. \n\nI played out with different number of folds, learning rates and scheduler settings – in all experiments the model demonstrated public LB higher than local CV. I assume what I observe is not LB overfitting but rather some trivial mistake in the code I can’t identify. Would be very grateful for your hints.",
      "votes": null
    },
    {
      "id": "924787",
      "postDate": "07/11/2020 16:20:34",
      "content": "<p>I was trying a new pipeline and submitted after 1 epoch : my validation score was 0.83 but LB 0.88+</p>\n\n<p>I guess your scores are not \"dramatically higher\" than local CV : try to swtich to 5 fold cv and don't forget that if you submit the average of your folds models then it's expected to have a better LB than CV (since you are ensembling your folds).</p>",
      "rawMarkdown": "I was trying a new pipeline and submitted after 1 epoch : my validation score was 0.83 but LB 0.88+\n\nI guess your scores are not \"dramatically higher\" than local CV : try to swtich to 5 fold cv and don't forget that if you submit the average of your folds models then it's expected to have a better LB than CV (since you are ensembling your folds).",
      "votes": null
    },
    {
      "id": "924812",
      "postDate": "07/11/2020 16:31:43",
      "content": "<p>I would tend to trust your CV though when it comes to making choices for the private leaderboard. It is often the case that following your CV will give you better final results than selecting your best public LB values. This is especially true if you are trying to maximize your public LB scores. A good model will work well with both.</p>",
      "rawMarkdown": "I would tend to trust your CV though when it comes to making choices for the private leaderboard. It is often the case that following your CV will give you better final results than selecting your best public LB values. This is especially true if you are trying to maximize your public LB scores. A good model will work well with both.",
      "votes": null
    },
    {
      "id": "926689",
      "postDate": "07/12/2020 21:58:38",
      "content": "<p>It may be right in my case.\nFor the sake of experiment I changed seed, re-ran the model and calculated the average of fold models on the train data set in the same fashion as I generated submission.</p>\n\n<p>fold0 validation score: 0.8927\nfold1 validation score: 0.8751\nfold2 validation score: 0.8914\nlocal CV: <strong>0.88358</strong></p>\n\n<p>score for average of fold models on train: <strong>0.98600</strong></p>\n\n<p>public LB:<strong>0.907</strong></p>\n\n<p>Still feel very uncertain about all this.\nPS with 5 folds it keeps the same - public LB ~0.02-0.04 higher than local CV. </p>",
      "rawMarkdown": "It may be right in my case.\nFor the sake of experiment I changed seed, re-ran the model and calculated the average of fold models on the train data set in the same fashion as I generated submission.\n\nfold0 validation score: 0.8927\nfold1 validation score: 0.8751\nfold2 validation score: 0.8914\nlocal CV: **0.88358**\n\nscore for average of fold models on train: **0.98600**\n\npublic LB:**0.907**\n\nStill feel very uncertain about all this.\nPS with 5 folds it keeps the same - public LB ~0.02-0.04 higher than local CV.",
      "votes": null
    },
    {
      "id": "926692",
      "postDate": "07/12/2020 22:08:26",
      "content": "<p>And that's the thing. I focus on my CV that results in a suspiciously high public LB. Based on CV\\LB other kaggles shared in this competition, nobody experiences such a gap. Reflecting on this, I start doubting if I can trust my local CV and therefore my model. An infinite loop.</p>",
      "rawMarkdown": "And that's the thing. I focus on my CV that results in a suspiciously high public LB. Based on CV\\LB other kaggles shared in this competition, nobody experiences such a gap. Reflecting on this, I start doubting if I can trust my local CV and therefore my model. An infinite loop.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 924787,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "07/11/2020 16:20:34",
      "content": "<p>I was trying a new pipeline and submitted after 1 epoch : my validation score was 0.83 but LB 0.88+</p>\n\n<p>I guess your scores are not \"dramatically higher\" than local CV : try to swtich to 5 fold cv and don't forget that if you submit the average of your folds models then it's expected to have a better LB than CV (since you are ensembling your folds).</p>",
      "votes": null,
      "replies": [
        {
          "id": 926689,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "07/12/2020 21:58:38",
          "content": "<p>It may be right in my case.\nFor the sake of experiment I changed seed, re-ran the model and calculated the average of fold models on the train data set in the same fashion as I generated submission.</p>\n\n<p>fold0 validation score: 0.8927\nfold1 validation score: 0.8751\nfold2 validation score: 0.8914\nlocal CV: <strong>0.88358</strong></p>\n\n<p>score for average of fold models on train: <strong>0.98600</strong></p>\n\n<p>public LB:<strong>0.907</strong></p>\n\n<p>Still feel very uncertain about all this.\nPS with 5 folds it keeps the same - public LB ~0.02-0.04 higher than local CV. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 924812,
      "author_name": "cbryant",
      "author_url": "",
      "post_date": "07/11/2020 16:31:43",
      "content": "<p>I would tend to trust your CV though when it comes to making choices for the private leaderboard. It is often the case that following your CV will give you better final results than selecting your best public LB values. This is especially true if you are trying to maximize your public LB scores. A good model will work well with both.</p>",
      "votes": null,
      "replies": [
        {
          "id": 926692,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "07/12/2020 22:08:26",
          "content": "<p>And that's the thing. I focus on my CV that results in a suspiciously high public LB. Based on CV\\LB other kaggles shared in this competition, nobody experiences such a gap. Reflecting on this, I start doubting if I can trust my local CV and therefore my model. An infinite loop.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "923844": "Inspired by Abhishek’s notebooks ([1](https://www.kaggle.com/abhishek/melanoma-detection-with-pytorch),[2](https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu)), finally I took the plunge and built my very first pytorch model this week ([here](https://www.kaggle.com/dunklerwald/effnetb1-public-lb-is-higher-than-local-cv)). What I can’t explain is why my model stably generates suspiciously high public LB.\nE.g.: \nfolds=3, LR= 0.001, local CV 0.89268, public LB 0.897\nfolds=3, LR= 0.0001, local CV 0.86827, public LB 0.909. \n\nI played out with different number of folds, learning rates and scheduler settings – in all experiments the model demonstrated public LB higher than local CV. I assume what I observe is not LB overfitting but rather some trivial mistake in the code I can’t identify. Would be very grateful for your hints.",
    "924787": "I was trying a new pipeline and submitted after 1 epoch : my validation score was 0.83 but LB 0.88+\n\nI guess your scores are not \"dramatically higher\" than local CV : try to swtich to 5 fold cv and don't forget that if you submit the average of your folds models then it's expected to have a better LB than CV (since you are ensembling your folds).",
    "924812": "I would tend to trust your CV though when it comes to making choices for the private leaderboard. It is often the case that following your CV will give you better final results than selecting your best public LB values. This is especially true if you are trying to maximize your public LB scores. A good model will work well with both.",
    "926689": "It may be right in my case.\nFor the sake of experiment I changed seed, re-ran the model and calculated the average of fold models on the train data set in the same fashion as I generated submission.\n\nfold0 validation score: 0.8927\nfold1 validation score: 0.8751\nfold2 validation score: 0.8914\nlocal CV: **0.88358**\n\nscore for average of fold models on train: **0.98600**\n\npublic LB:**0.907**\n\nStill feel very uncertain about all this.\nPS with 5 folds it keeps the same - public LB ~0.02-0.04 higher than local CV.",
    "926692": "And that's the thing. I focus on my CV that results in a suspiciously high public LB. Based on CV\\LB other kaggles shared in this competition, nobody experiences such a gap. Reflecting on this, I start doubting if I can trust my local CV and therefore my model. An infinite loop."
  },
  "source": "meta"
}