{
  "id": 511590,
  "title": "We should be given 10-20% of private dataset (actual test data) for training",
  "url": "/competitions/birdclef-2024/discussion/511590",
  "author_name": "Daniel Kalicki",
  "post_date": "2024-06-11T11:00:37.705000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I see a lot of people are disappointed with low place, in spite putting a lot of effort (not me, I joined late). Some public notebook flooded silver/bronze part of leaderboard. Some high places are just merges of public notebooks.</p>\n<p>In my <em>opinion</em> it is organizers' mistake. They should share some part of private dataset with us.</p>\n<ul>\n<li>it would decrease randomness of final scores, resulting in more fair competition</li>\n<li>it would improve final scores. People would learn what and why works/doesnt work faster, not by probing LB… It would result in better preprocessing/augmentations/models/etc.</li>\n</ul>\n<p>One could say the way it is we will have more \"robust\" models, less overfit to test data. No, people will still overfit to private set, just by probing it (top 5 teams have about 250-300 entries on average…). </p>",
  "messages": [
    {
      "id": 2866464,
      "postDate": "2024-06-11T11:00:37.707Z",
      "content": "<p>I see a lot of people are disappointed with low place, in spite putting a lot of effort (not me, I joined late). Some public notebook flooded silver/bronze part of leaderboard. Some high places are just merges of public notebooks.</p>\n<p>In my <em>opinion</em> it is organizers' mistake. They should share some part of private dataset with us.</p>\n<ul>\n<li>it would decrease randomness of final scores, resulting in more fair competition</li>\n<li>it would improve final scores. People would learn what and why works/doesnt work faster, not by probing LB… It would result in better preprocessing/augmentations/models/etc.</li>\n</ul>\n<p>One could say the way it is we will have more \"robust\" models, less overfit to test data. No, people will still overfit to private set, just by probing it (top 5 teams have about 250-300 entries on average…). </p>",
      "rawMarkdown": "I see a lot of people are disappointed with low place, in spite putting a lot of effort (not me, I joined late). Some public notebook flooded silver/bronze part of leaderboard. Some high places are just merges of public notebooks.\n\nIn my *opinion* it is organizers' mistake. They should share some part of private dataset with us.\n- it would decrease randomness of final scores, resulting in more fair competition\n- it would improve final scores. People would learn what and why works/doesnt work faster, not by probing LB... It would result in better preprocessing/augmentations/models/etc.\n\nOne could say the way it is we will have more \"robust\" models, less overfit to test data. No, people will still overfit to private set, just by probing it (top 5 teams have about 250-300 entries on average...). ",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2866464": "I see a lot of people are disappointed with low place, in spite putting a lot of effort (not me, I joined late). Some public notebook flooded silver/bronze part of leaderboard. Some high places are just merges of public notebooks.\n\nIn my *opinion* it is organizers' mistake. They should share some part of private dataset with us.\n- it would decrease randomness of final scores, resulting in more fair competition\n- it would improve final scores. People would learn what and why works/doesnt work faster, not by probing LB... It would result in better preprocessing/augmentations/models/etc.\n\nOne could say the way it is we will have more \"robust\" models, less overfit to test data. No, people will still overfit to private set, just by probing it (top 5 teams have about 250-300 entries on average...). "
  }
}