{
  "id": 199548,
  "title": "Insights of Not Over-fitting to LB",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199548",
  "author_name": "",
  "post_date": "2020-11-26T06:42:00.573659Z",
  "votes": 71,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Dear Kagglers,</p>\n<p>I'm writing this post to share some ideas of not getting overfitted to public LB.</p>\n<p>According to the data description:</p>\n<pre><code>Files\n[train/test]_images the image files. The full set of test images will only be available to your notebook when it is submitted for scoring. Expect to see roughly 15,000 images in the test set.\n</code></pre>\n<p>In other words:<br>\n<strong>Public LB</strong>: 40% -&gt; <strong>6000</strong> images, <br>\n<strong>Private LB</strong>: 60% -&gt; <strong>9000</strong> images<br>\n<strong>Cross Valiation</strong> -&gt; <strong>21397</strong> images to validate our models</p>\n<p>So I would say in most cases, we should probably trust our cv. </p>\n<p><strong>We could also treat lb as an alternative fold</strong>, then it is better to see that every <strong>validation\\test</strong> (every fold) <strong>improves at the same time</strong>., which is also how I used to select models as a final solution. Don't get obsessed with public LB and ignore the cv score; otherwise, we could end up overfitting to public LB.</p>\n<p><strong>Regarding the data size, accuracy might not be a stable indicator</strong>. In public lb, if we have ~0.001 difference in terms of accuracy, which means 6 images of the wrong predictions, I would not say the difference is significant. </p>\n<p>To sum up, tune\\develop your ideas based on cv score. If the cv score improves, you could further test if public lb improves. If the public-cv relationship is roughly proportional, I guess it is good to go. <strong>Another safe way is to do cross-validation with different random seeds (so the folds splitting and model randomness is different), then we conclude its performance by taking an average across different seeds.</strong></p>\n<p>I also see another post with deeper insights from CPMP and the paper mentioned in that post. Here are the links:<br>\n<a href=\"https://arxiv.org/abs/1811.12808\" target=\"_blank\">Paper regarding the model selection</a> <br>\n<a href=\"https://www.kaggle.com/c/lish-moa/discussion/196913\" target=\"_blank\">Post from CPMP</a>.</p>\n<p>Thanks for reading the posts. Hope it helps people to better decide their final models in the future :)</p>",
  "messages": [
    {
      "id": "1091636",
      "postDate": "11/26/2020 06:42:00",
      "content": "<p>Dear Kagglers,</p>\n<p>I'm writing this post to share some ideas of not getting overfitted to public LB.</p>\n<p>According to the data description:</p>\n<pre><code>Files\n[train/test]_images the image files. The full set of test images will only be available to your notebook when it is submitted for scoring. Expect to see roughly 15,000 images in the test set.\n</code></pre>\n<p>In other words:<br>\n<strong>Public LB</strong>: 40% -&gt; <strong>6000</strong> images, <br>\n<strong>Private LB</strong>: 60% -&gt; <strong>9000</strong> images<br>\n<strong>Cross Valiation</strong> -&gt; <strong>21397</strong> images to validate our models</p>\n<p>So I would say in most cases, we should probably trust our cv. </p>\n<p><strong>We could also treat lb as an alternative fold</strong>, then it is better to see that every <strong>validation\\test</strong> (every fold) <strong>improves at the same time</strong>., which is also how I used to select models as a final solution. Don't get obsessed with public LB and ignore the cv score; otherwise, we could end up overfitting to public LB.</p>\n<p><strong>Regarding the data size, accuracy might not be a stable indicator</strong>. In public lb, if we have ~0.001 difference in terms of accuracy, which means 6 images of the wrong predictions, I would not say the difference is significant. </p>\n<p>To sum up, tune\\develop your ideas based on cv score. If the cv score improves, you could further test if public lb improves. If the public-cv relationship is roughly proportional, I guess it is good to go. <strong>Another safe way is to do cross-validation with different random seeds (so the folds splitting and model randomness is different), then we conclude its performance by taking an average across different seeds.</strong></p>\n<p>I also see another post with deeper insights from CPMP and the paper mentioned in that post. Here are the links:<br>\n<a href=\"https://arxiv.org/abs/1811.12808\" target=\"_blank\">Paper regarding the model selection</a> <br>\n<a href=\"https://www.kaggle.com/c/lish-moa/discussion/196913\" target=\"_blank\">Post from CPMP</a>.</p>\n<p>Thanks for reading the posts. Hope it helps people to better decide their final models in the future :)</p>",
      "rawMarkdown": "Dear Kagglers,\n\nI'm writing this post to share some ideas of not getting overfitted to public LB.\n\nAccording to the data description:\n```\nFiles\n[train/test]_images the image files. The full set of test images will only be available to your notebook when it is submitted for scoring. Expect to see roughly 15,000 images in the test set.\n```\nIn other words:\n**Public LB**: 40% -> **6000** images, \n**Private LB**: 60% -> **9000** images\n**Cross Valiation** -> **21397** images to validate our models\n\nSo I would say in most cases, we should probably trust our cv. \n\n**We could also treat lb as an alternative fold**, then it is better to see that every **validation\\test** (every fold) **improves at the same time**., which is also how I used to select models as a final solution. Don't get obsessed with public LB and ignore the cv score; otherwise, we could end up overfitting to public LB.\n\n**Regarding the data size, accuracy might not be a stable indicator**. In public lb, if we have ~0.001 difference in terms of accuracy, which means 6 images of the wrong predictions, I would not say the difference is significant. \n\nTo sum up, tune\\develop your ideas based on cv score. If the cv score improves, you could further test if public lb improves. If the public-cv relationship is roughly proportional, I guess it is good to go. **Another safe way is to do cross-validation with different random seeds (so the folds splitting and model randomness is different), then we conclude its performance by taking an average across different seeds.**\n\nI also see another post with deeper insights from CPMP and the paper mentioned in that post. Here are the links:\n[Paper regarding the model selection](https://arxiv.org/abs/1811.12808) \n[Post from CPMP](https://www.kaggle.com/c/lish-moa/discussion/196913).\n\nThanks for reading the posts. Hope it helps people to better decide their final models in the future :)",
      "votes": null
    },
    {
      "id": "1093243",
      "postDate": "11/27/2020 15:03:39",
      "content": "<p>Pretty useful tips - and a good wake up call! Thanks <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "rawMarkdown": "Pretty useful tips - and a good wake up call! Thanks @khyeh0719",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1093243,
      "author_name": "reighns",
      "author_url": "",
      "post_date": "11/27/2020 15:03:39",
      "content": "<p>Pretty useful tips - and a good wake up call! Thanks <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091636": "Dear Kagglers,\n\nI'm writing this post to share some ideas of not getting overfitted to public LB.\n\nAccording to the data description:\n```\nFiles\n[train/test]_images the image files. The full set of test images will only be available to your notebook when it is submitted for scoring. Expect to see roughly 15,000 images in the test set.\n```\nIn other words:\n**Public LB**: 40% -> **6000** images, \n**Private LB**: 60% -> **9000** images\n**Cross Valiation** -> **21397** images to validate our models\n\nSo I would say in most cases, we should probably trust our cv. \n\n**We could also treat lb as an alternative fold**, then it is better to see that every **validation\\test** (every fold) **improves at the same time**., which is also how I used to select models as a final solution. Don't get obsessed with public LB and ignore the cv score; otherwise, we could end up overfitting to public LB.\n\n**Regarding the data size, accuracy might not be a stable indicator**. In public lb, if we have ~0.001 difference in terms of accuracy, which means 6 images of the wrong predictions, I would not say the difference is significant. \n\nTo sum up, tune\\develop your ideas based on cv score. If the cv score improves, you could further test if public lb improves. If the public-cv relationship is roughly proportional, I guess it is good to go. **Another safe way is to do cross-validation with different random seeds (so the folds splitting and model randomness is different), then we conclude its performance by taking an average across different seeds.**\n\nI also see another post with deeper insights from CPMP and the paper mentioned in that post. Here are the links:\n[Paper regarding the model selection](https://arxiv.org/abs/1811.12808) \n[Post from CPMP](https://www.kaggle.com/c/lish-moa/discussion/196913).\n\nThanks for reading the posts. Hope it helps people to better decide their final models in the future :)",
    "1093243": "Pretty useful tips - and a good wake up call! Thanks @khyeh0719"
  },
  "source": "meta"
}