{
  "id": 417261,
  "title": "Most Top Solutions Overfitted the Dataset",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417261",
  "author_name": "",
  "post_date": "2023-06-15T00:50:39.750588200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Now that the private leaderboard is out, I observed several contestants at the top dropped points significantly.  My guess is overfitting, what do you think?</p>",
  "messages": [
    {
      "id": "2302915",
      "postDate": "06/15/2023 00:50:39",
      "content": "<p>Now that the private leaderboard is out, I observed several contestants at the top dropped points significantly.  My guess is overfitting, what do you think?</p>",
      "rawMarkdown": "Now that the private leaderboard is out, I observed several contestants at the top dropped points significantly.  My guess is overfitting, what do you think?",
      "votes": null
    },
    {
      "id": "2303303",
      "postDate": "06/15/2023 08:22:44",
      "content": "<p>Yes, most likely. </p>",
      "rawMarkdown": "Yes, most likely.",
      "votes": null
    },
    {
      "id": "2303328",
      "postDate": "06/15/2023 08:39:33",
      "content": "<p>Yes, a drop of 0.15 of the best score between public and private dataset is quite big. But I think that is mostly due to the little amount of data we had for training and also the test set. The statistics extracted of that are not representative at all. But yes, excessive overfitting and I don't think that the models will be really helpful for the organizers.</p>",
      "rawMarkdown": "Yes, a drop of 0.15 of the best score between public and private dataset is quite big. But I think that is mostly due to the little amount of data we had for training and also the test set. The statistics extracted of that are not representative at all. But yes, excessive overfitting and I don't think that the models will be really helpful for the organizers.",
      "votes": null
    },
    {
      "id": "2303346",
      "postDate": "06/15/2023 08:54:48",
      "content": "<p>I think generally validation just by fragment1 (just holdout) may be dangerous.<br>\nIndeed model trained by fragment2, 3 seems to perform well in public LB, but it doesn't mean the model will perform well in unseen data. It means that the model perform well only to 10% of test data. <br>\nif public test data was randomly sampled, we could trust public LB to some extent, but this is not according to some discussions.<br>\nSo trusting the public score leads to  overfitting. <br>\nHigher ranker who probe test data and get some information don't have to do so, but all most all we need is trust cv, I think.</p>",
      "rawMarkdown": "I think generally validation just by fragment1 (just holdout) may be dangerous.\nIndeed model trained by fragment2, 3 seems to perform well in public LB, but it doesn't mean the model will perform well in unseen data. It means that the model perform well only to 10% of test data. \nif public test data was randomly sampled, we could trust public LB to some extent, but this is not according to some discussions.\nSo trusting the public score leads to  overfitting. \nHigher ranker who probe test data and get some information don't have to do so, but all most all we need is trust cv, I think.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303303,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/15/2023 08:22:44",
      "content": "<p>Yes, most likely. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303328,
      "author_name": "yankir",
      "author_url": "",
      "post_date": "06/15/2023 08:39:33",
      "content": "<p>Yes, a drop of 0.15 of the best score between public and private dataset is quite big. But I think that is mostly due to the little amount of data we had for training and also the test set. The statistics extracted of that are not representative at all. But yes, excessive overfitting and I don't think that the models will be really helpful for the organizers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2303346,
      "author_name": "clearwaterkzk",
      "author_url": "",
      "post_date": "06/15/2023 08:54:48",
      "content": "<p>I think generally validation just by fragment1 (just holdout) may be dangerous.<br>\nIndeed model trained by fragment2, 3 seems to perform well in public LB, but it doesn't mean the model will perform well in unseen data. It means that the model perform well only to 10% of test data. <br>\nif public test data was randomly sampled, we could trust public LB to some extent, but this is not according to some discussions.<br>\nSo trusting the public score leads to  overfitting. <br>\nHigher ranker who probe test data and get some information don't have to do so, but all most all we need is trust cv, I think.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2302915": "Now that the private leaderboard is out, I observed several contestants at the top dropped points significantly.  My guess is overfitting, what do you think?",
    "2303303": "Yes, most likely.",
    "2303328": "Yes, a drop of 0.15 of the best score between public and private dataset is quite big. But I think that is mostly due to the little amount of data we had for training and also the test set. The statistics extracted of that are not representative at all. But yes, excessive overfitting and I don't think that the models will be really helpful for the organizers.",
    "2303346": "I think generally validation just by fragment1 (just holdout) may be dangerous.\nIndeed model trained by fragment2, 3 seems to perform well in public LB, but it doesn't mean the model will perform well in unseen data. It means that the model perform well only to 10% of test data. \nif public test data was randomly sampled, we could trust public LB to some extent, but this is not according to some discussions.\nSo trusting the public score leads to  overfitting. \nHigher ranker who probe test data and get some information don't have to do so, but all most all we need is trust cv, I think."
  },
  "source": "meta"
}