{
  "id": 420275,
  "title": "How to prevent private LB shaking?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420275",
  "author_name": "",
  "post_date": "2023-06-30T04:43:58.462794100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, I want to thank you Kaggle, the hosts and the Kaggle community for all the sharings. You give us a good chance to learn ML knowledge.<br>\nI'm a new kaggler and this is my first competition. I am 52nd on the public LB, my model is seems overfitting so I only get a bronze medal on the private LB. I find an interesting thing that I have a model which has public score 0.693, but it have private scores 0.698. By contrast, I mix 2 Catboost, 1 XGBoost and 1 LightGBM, they have public score 0.703, but private score 0.697. I think this is unbelievable, Do you have some ideas to prevent the overfitting and how do you choose the final submission model?</p>",
  "messages": [
    {
      "id": "2323615",
      "postDate": "06/30/2023 04:43:58",
      "content": "<p>First of all, I want to thank you Kaggle, the hosts and the Kaggle community for all the sharings. You give us a good chance to learn ML knowledge.<br>\nI'm a new kaggler and this is my first competition. I am 52nd on the public LB, my model is seems overfitting so I only get a bronze medal on the private LB. I find an interesting thing that I have a model which has public score 0.693, but it have private scores 0.698. By contrast, I mix 2 Catboost, 1 XGBoost and 1 LightGBM, they have public score 0.703, but private score 0.697. I think this is unbelievable, Do you have some ideas to prevent the overfitting and how do you choose the final submission model?</p>",
      "rawMarkdown": "First of all, I want to thank you Kaggle, the hosts and the Kaggle community for all the sharings. You give us a good chance to learn ML knowledge.\nI'm a new kaggler and this is my first competition. I am 52nd on the public LB, my model is seems overfitting so I only get a bronze medal on the private LB. I find an interesting thing that I have a model which has public score 0.693, but it have private scores 0.698. By contrast, I mix 2 Catboost, 1 XGBoost and 1 LightGBM, they have public score 0.703, but private score 0.697. I think this is unbelievable, Do you have some ideas to prevent the overfitting and how do you choose the final submission model?",
      "votes": null
    },
    {
      "id": "2323829",
      "postDate": "06/30/2023 07:29:14",
      "content": "<p>From my experience, I often do two things</p>\n<ul>\n<li>Use a CV strategy that is as robust as possible to avoid data leakagae.</li>\n<li>Observe score changes between CV and LB. If you have one submission, it's hard to do this. But if you have several submissions, your score will likely change incrementally (+/- 0.00x after each sub). If there comes a time when you start to see that changes in CV does not align with changes in LB (e.g. CV improves but LB does not or vice versa), it signals some kind if instability (your model is overfitting; data distribution shift between CV and LB; etc.) that should be treated carefully.</li>\n</ul>",
      "rawMarkdown": "From my experience, I often do two things\n- Use a CV strategy that is as robust as possible to avoid data leakagae.\n- Observe score changes between CV and LB. If you have one submission, it's hard to do this. But if you have several submissions, your score will likely change incrementally (+/- 0.00x after each sub). If there comes a time when you start to see that changes in CV does not align with changes in LB (e.g. CV improves but LB does not or vice versa), it signals some kind if instability (your model is overfitting; data distribution shift between CV and LB; etc.) that should be treated carefully.",
      "votes": null
    },
    {
      "id": "2324546",
      "postDate": "06/30/2023 17:10:42",
      "content": "<p>When there are doubts or problems with public/private LB (as it also happened with GoDaddy competition), I believe the public LB score could do more harm than good while you decide which model to choose. In these cases, sticking with a robust CV strategy becomes essential.</p>",
      "rawMarkdown": "When there are doubts or problems with public/private LB (as it also happened with GoDaddy competition), I believe the public LB score could do more harm than good while you decide which model to choose. In these cases, sticking with a robust CV strategy becomes essential.",
      "votes": null
    },
    {
      "id": "2324726",
      "postDate": "06/30/2023 20:29:00",
      "content": "<p>At the core, you need to BOTH control for CV overfitting AND public LB overfitting.</p>\n<p>Each one is individually semi-easy, in fact public LB really helps as an extra holdout to check for CV overfitting.<br>\nBoth together is much more tricky. There's no single recipe, it's more like a full semester course to understand 'why' the top data scientist experts do X, Y, Z, etc, and then apply those lessons in a nuanced way based on the specifics of the competition and your specific experiments and results.</p>\n<p>But the basics for this competition ONCE the train data doubled is that CV was much more trustworthy. It was very clear that there would be much less public LB (and private LB) total data, thus higher chance of overfitting public LB due to random luck.</p>",
      "rawMarkdown": "At the core, you need to BOTH control for CV overfitting AND public LB overfitting.\n\nEach one is individually semi-easy, in fact public LB really helps as an extra holdout to check for CV overfitting.\nBoth together is much more tricky. There's no single recipe, it's more like a full semester course to understand 'why' the top data scientist experts do X, Y, Z, etc, and then apply those lessons in a nuanced way based on the specifics of the competition and your specific experiments and results.\n\nBut the basics for this competition ONCE the train data doubled is that CV was much more trustworthy. It was very clear that there would be much less public LB (and private LB) total data, thus higher chance of overfitting public LB due to random luck.",
      "votes": null
    },
    {
      "id": "2324843",
      "postDate": "06/30/2023 23:11:12",
      "content": "<p>I wrote about it few years ago and I think most of the content is still valid: <a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/196913\" target=\"_blank\">Some tips to avoid overfitting</a></p>",
      "rawMarkdown": "I wrote about it few years ago and I think most of the content is still valid: [Some tips to avoid overfitting](https://www.kaggle.com/competitions/lish-moa/discussion/196913)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2323829,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/30/2023 07:29:14",
      "content": "<p>From my experience, I often do two things</p>\n<ul>\n<li>Use a CV strategy that is as robust as possible to avoid data leakagae.</li>\n<li>Observe score changes between CV and LB. If you have one submission, it's hard to do this. But if you have several submissions, your score will likely change incrementally (+/- 0.00x after each sub). If there comes a time when you start to see that changes in CV does not align with changes in LB (e.g. CV improves but LB does not or vice versa), it signals some kind if instability (your model is overfitting; data distribution shift between CV and LB; etc.) that should be treated carefully.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2324546,
      "author_name": "sametmarasli",
      "author_url": "",
      "post_date": "06/30/2023 17:10:42",
      "content": "<p>When there are doubts or problems with public/private LB (as it also happened with GoDaddy competition), I believe the public LB score could do more harm than good while you decide which model to choose. In these cases, sticking with a robust CV strategy becomes essential.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2324726,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "06/30/2023 20:29:00",
      "content": "<p>At the core, you need to BOTH control for CV overfitting AND public LB overfitting.</p>\n<p>Each one is individually semi-easy, in fact public LB really helps as an extra holdout to check for CV overfitting.<br>\nBoth together is much more tricky. There's no single recipe, it's more like a full semester course to understand 'why' the top data scientist experts do X, Y, Z, etc, and then apply those lessons in a nuanced way based on the specifics of the competition and your specific experiments and results.</p>\n<p>But the basics for this competition ONCE the train data doubled is that CV was much more trustworthy. It was very clear that there would be much less public LB (and private LB) total data, thus higher chance of overfitting public LB due to random luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2324843,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/30/2023 23:11:12",
      "content": "<p>I wrote about it few years ago and I think most of the content is still valid: <a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/196913\" target=\"_blank\">Some tips to avoid overfitting</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2323615": "First of all, I want to thank you Kaggle, the hosts and the Kaggle community for all the sharings. You give us a good chance to learn ML knowledge.\nI'm a new kaggler and this is my first competition. I am 52nd on the public LB, my model is seems overfitting so I only get a bronze medal on the private LB. I find an interesting thing that I have a model which has public score 0.693, but it have private scores 0.698. By contrast, I mix 2 Catboost, 1 XGBoost and 1 LightGBM, they have public score 0.703, but private score 0.697. I think this is unbelievable, Do you have some ideas to prevent the overfitting and how do you choose the final submission model?",
    "2323829": "From my experience, I often do two things\n- Use a CV strategy that is as robust as possible to avoid data leakagae.\n- Observe score changes between CV and LB. If you have one submission, it's hard to do this. But if you have several submissions, your score will likely change incrementally (+/- 0.00x after each sub). If there comes a time when you start to see that changes in CV does not align with changes in LB (e.g. CV improves but LB does not or vice versa), it signals some kind if instability (your model is overfitting; data distribution shift between CV and LB; etc.) that should be treated carefully.",
    "2324546": "When there are doubts or problems with public/private LB (as it also happened with GoDaddy competition), I believe the public LB score could do more harm than good while you decide which model to choose. In these cases, sticking with a robust CV strategy becomes essential.",
    "2324726": "At the core, you need to BOTH control for CV overfitting AND public LB overfitting.\n\nEach one is individually semi-easy, in fact public LB really helps as an extra holdout to check for CV overfitting.\nBoth together is much more tricky. There's no single recipe, it's more like a full semester course to understand 'why' the top data scientist experts do X, Y, Z, etc, and then apply those lessons in a nuanced way based on the specifics of the competition and your specific experiments and results.\n\nBut the basics for this competition ONCE the train data doubled is that CV was much more trustworthy. It was very clear that there would be much less public LB (and private LB) total data, thus higher chance of overfitting public LB due to random luck.",
    "2324843": "I wrote about it few years ago and I think most of the content is still valid: [Some tips to avoid overfitting](https://www.kaggle.com/competitions/lish-moa/discussion/196913)"
  },
  "source": "meta"
}