{
  "id": 159345,
  "title": "overfitting on public lb?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159345",
  "author_name": "",
  "post_date": "2020-06-17T07:09:10.666348100Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>After ensembling my 9 xlm-roberta models(trained on different datasets), my public score has reached to <strong>0.9475</strong>.</p>\n\n<p>How to know whether I am overfitting on the public lb or not? I have used undersampling ( trained with all 7 languages).</p>\n\n<p>Can there be a large difference between public and private lb? How much deviation we can expect?\nAlso, i have the submission.csv generated offline and directly using those in kernel and averaging. This won't cause any issue right?</p>",
  "messages": [
    {
      "id": "889825",
      "postDate": "06/17/2020 07:09:10",
      "content": "<p>After ensembling my 9 xlm-roberta models(trained on different datasets), my public score has reached to <strong>0.9475</strong>.</p>\n\n<p>How to know whether I am overfitting on the public lb or not? I have used undersampling ( trained with all 7 languages).</p>\n\n<p>Can there be a large difference between public and private lb? How much deviation we can expect?\nAlso, i have the submission.csv generated offline and directly using those in kernel and averaging. This won't cause any issue right?</p>",
      "rawMarkdown": "After ensembling my 9 xlm-roberta models(trained on different datasets), my public score has reached to **0.9475**.\n\nHow to know whether I am overfitting on the public lb or not? I have used undersampling ( trained with all 7 languages).\n\nCan there be a large difference between public and private lb? How much deviation we can expect?\nAlso, i have the submission.csv generated offline and directly using those in kernel and averaging. This won't cause any issue right?",
      "votes": null
    },
    {
      "id": "889848",
      "postDate": "06/17/2020 07:25:03",
      "content": "<p><a href=\"/ankitsajwan\">@ankitsajwan</a> It's very probable that we're all overfitting a bit\nBecause the validation is much easier than the public LB, I'd also expect the private lb to be a bit easier, but it could very weel be the opposite ;)\nWe'll all see the deviation in 6 days! lol</p>\n\n<p>ps.: I believe your csv's should be fine\nGood kaggling!</p>",
      "rawMarkdown": "ankitsajwan It's very probable that we're all overfitting a bit\nBecause the validation is much easier than the public LB, I'd also expect the private lb to be a bit easier, but it could very weel be the opposite ;)\nWe'll all see the deviation in 6 days! lol\n\nps.: I believe your csv's should be fine\nGood kaggling!",
      "votes": null
    },
    {
      "id": "890262",
      "postDate": "06/17/2020 12:16:45",
      "content": "<p>we're all overfitting to the public LB in a sense as this is the way Kaggle works, i.e. we have a public standing we try to reach the top of and hope we won't be kicked down when the day X comes with the private LB. that's why there exists a holy rule which states to trust your CV 😏 given you have built it properly</p>\n\n<blockquote>\n  <p>Can there be a large difference between public and private lb? How much deviation we can expect?</p>\n</blockquote>\n\n<p>you have data available to you and thus you can try to infer by looking at distributions and taking your statistics knowledge to test. should we have the right answer for that, would there be an intense competition?</p>",
      "rawMarkdown": "we're all overfitting to the public LB in a sense as this is the way Kaggle works, i.e. we have a public standing we try to reach the top of and hope we won't be kicked down when the day X comes with the private LB. that's why there exists a holy rule which states to trust your CV 😏 given you have built it properly\n\n&gt; Can there be a large difference between public and private lb? How much deviation we can expect?\n\nyou have data available to you and thus you can try to infer by looking at distributions and taking your statistics knowledge to test. should we have the right answer for that, would there be an intense competition?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 889848,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "06/17/2020 07:25:03",
      "content": "<p><a href=\"/ankitsajwan\">@ankitsajwan</a> It's very probable that we're all overfitting a bit\nBecause the validation is much easier than the public LB, I'd also expect the private lb to be a bit easier, but it could very weel be the opposite ;)\nWe'll all see the deviation in 6 days! lol</p>\n\n<p>ps.: I believe your csv's should be fine\nGood kaggling!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 890262,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "06/17/2020 12:16:45",
      "content": "<p>we're all overfitting to the public LB in a sense as this is the way Kaggle works, i.e. we have a public standing we try to reach the top of and hope we won't be kicked down when the day X comes with the private LB. that's why there exists a holy rule which states to trust your CV 😏 given you have built it properly</p>\n\n<blockquote>\n  <p>Can there be a large difference between public and private lb? How much deviation we can expect?</p>\n</blockquote>\n\n<p>you have data available to you and thus you can try to infer by looking at distributions and taking your statistics knowledge to test. should we have the right answer for that, would there be an intense competition?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "889825": "After ensembling my 9 xlm-roberta models(trained on different datasets), my public score has reached to **0.9475**.\n\nHow to know whether I am overfitting on the public lb or not? I have used undersampling ( trained with all 7 languages).\n\nCan there be a large difference between public and private lb? How much deviation we can expect?\nAlso, i have the submission.csv generated offline and directly using those in kernel and averaging. This won't cause any issue right?",
    "889848": "ankitsajwan It's very probable that we're all overfitting a bit\nBecause the validation is much easier than the public LB, I'd also expect the private lb to be a bit easier, but it could very weel be the opposite ;)\nWe'll all see the deviation in 6 days! lol\n\nps.: I believe your csv's should be fine\nGood kaggling!",
    "890262": "we're all overfitting to the public LB in a sense as this is the way Kaggle works, i.e. we have a public standing we try to reach the top of and hope we won't be kicked down when the day X comes with the private LB. that's why there exists a holy rule which states to trust your CV 😏 given you have built it properly\n\n&gt; Can there be a large difference between public and private lb? How much deviation we can expect?\n\nyou have data available to you and thus you can try to infer by looking at distributions and taking your statistics knowledge to test. should we have the right answer for that, would there be an intense competition?"
  },
  "source": "meta"
}