{
  "id": 552545,
  "title": "People dropped the competition actually won!",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552545",
  "author_name": "",
  "post_date": "2024-12-20T08:23:11.130086800Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I noticed that a lot of people who just did few submissions a lot of time ago are really high in the leaderboard!<br>\nMany examples of winning medals with less than 10 submissions and the last more than 1m ago.</p>\n<p>Maybe they were discouraged to see no advances in the public board and dropped the comp, while working on a good solution ?</p>\n<p>🥲 Thoughts ?</p>",
  "messages": [
    {
      "id": "3076746",
      "postDate": "12/20/2024 08:23:11",
      "content": "<p>I noticed that a lot of people who just did few submissions a lot of time ago are really high in the leaderboard!<br>\nMany examples of winning medals with less than 10 submissions and the last more than 1m ago.</p>\n<p>Maybe they were discouraged to see no advances in the public board and dropped the comp, while working on a good solution ?</p>\n<p>🥲 Thoughts ?</p>",
      "rawMarkdown": "I noticed that a lot of people who just did few submissions a lot of time ago are really high in the leaderboard!\nMany examples of winning medals with less than 10 submissions and the last more than 1m ago.\n\nMaybe they were discouraged to see no advances in the public board and dropped the comp, while working on a good solution ?\n\n🥲 Thoughts ?",
      "votes": null
    },
    {
      "id": "3076760",
      "postDate": "12/20/2024 08:34:26",
      "content": "<p>This is usually the scenario with luck driven  competitions<br>\nThis is actually a good trick in such  pandemoniums as optimizing the metric usually causes a downfall <a href=\"https://www.kaggle.com/faibioss\" target=\"_blank\">@faibioss</a> </p>",
      "rawMarkdown": "This is usually the scenario with luck driven ~~lotteries~~ competitions\nThis is actually a good trick in such ~~competitions~~ pandemoniums as optimizing the metric usually causes a downfall @faibioss",
      "votes": null
    },
    {
      "id": "3076779",
      "postDate": "12/20/2024 09:00:43",
      "content": "<p>The real lottery winner (or actually smart) is number 2 who just left at the initial phase of the competition. His last submission was 3 months ago.</p>",
      "rawMarkdown": "The real lottery winner (or actually smart) is number 2 who just left at the initial phase of the competition. His last submission was 3 months ago.",
      "votes": null
    },
    {
      "id": "3076898",
      "postDate": "12/20/2024 11:30:03",
      "content": "<p>How did teams at the top of the public leaderboard were able to overfit on a dataset they couldn't even see? </p>\n<p>Another question : is the split between the public dataset (38% of total I believe) and the private one (remaining 62%) actually random? If so, I don't see why the rankings would be so different even in presence of potential overfitting (to be proven)</p>",
      "rawMarkdown": "How did teams at the top of the public leaderboard were able to overfit on a dataset they couldn't even see? \n\nAnother question : is the split between the public dataset (38% of total I believe) and the private one (remaining 62%) actually random? If so, I don't see why the rankings would be so different even in presence of potential overfitting (to be proven)",
      "votes": null
    },
    {
      "id": "3076907",
      "postDate": "12/20/2024 11:42:55",
      "content": "<p>Overfitting in this case means that they kept on improving their LB score even when the CV score wasn't being stable/not improving therefore fitting the 38% of the data really well but failed miserably on the other 62%</p>",
      "rawMarkdown": "Overfitting in this case means that they kept on improving their LB score even when the CV score wasn't being stable/not improving therefore fitting the 38% of the data really well but failed miserably on the other 62%",
      "votes": null
    },
    {
      "id": "3076922",
      "postDate": "12/20/2024 11:52:58",
      "content": "<p>How do we know that their CV scores weren't stable?</p>\n<p>Also, even if it's the case, why would the LB public score would be so different from the private one? The datasets are supposed to be just 38% and 68% of the same macro dataset split randomly, right?</p>",
      "rawMarkdown": "How do we know that their CV scores weren't stable?\n\nAlso, even if it's the case, why would the LB public score would be so different from the private one? The datasets are supposed to be just 38% and 68% of the same macro dataset split randomly, right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076760,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/20/2024 08:34:26",
      "content": "<p>This is usually the scenario with luck driven  competitions<br>\nThis is actually a good trick in such  pandemoniums as optimizing the metric usually causes a downfall <a href=\"https://www.kaggle.com/faibioss\" target=\"_blank\">@faibioss</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076779,
      "author_name": "nikhil1e9",
      "author_url": "",
      "post_date": "12/20/2024 09:00:43",
      "content": "<p>The real lottery winner (or actually smart) is number 2 who just left at the initial phase of the competition. His last submission was 3 months ago.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076898,
      "author_name": "mezianek",
      "author_url": "",
      "post_date": "12/20/2024 11:30:03",
      "content": "<p>How did teams at the top of the public leaderboard were able to overfit on a dataset they couldn't even see? </p>\n<p>Another question : is the split between the public dataset (38% of total I believe) and the private one (remaining 62%) actually random? If so, I don't see why the rankings would be so different even in presence of potential overfitting (to be proven)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076907,
          "author_name": "nikhil1e9",
          "author_url": "",
          "post_date": "12/20/2024 11:42:55",
          "content": "<p>Overfitting in this case means that they kept on improving their LB score even when the CV score wasn't being stable/not improving therefore fitting the 38% of the data really well but failed miserably on the other 62%</p>",
          "votes": null,
          "replies": [
            {
              "id": 3076922,
              "author_name": "mezianek",
              "author_url": "",
              "post_date": "12/20/2024 11:52:58",
              "content": "<p>How do we know that their CV scores weren't stable?</p>\n<p>Also, even if it's the case, why would the LB public score would be so different from the private one? The datasets are supposed to be just 38% and 68% of the same macro dataset split randomly, right?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3076746": "I noticed that a lot of people who just did few submissions a lot of time ago are really high in the leaderboard!\nMany examples of winning medals with less than 10 submissions and the last more than 1m ago.\n\nMaybe they were discouraged to see no advances in the public board and dropped the comp, while working on a good solution ?\n\n🥲 Thoughts ?",
    "3076760": "This is usually the scenario with luck driven ~~lotteries~~ competitions\nThis is actually a good trick in such ~~competitions~~ pandemoniums as optimizing the metric usually causes a downfall @faibioss",
    "3076779": "The real lottery winner (or actually smart) is number 2 who just left at the initial phase of the competition. His last submission was 3 months ago.",
    "3076898": "How did teams at the top of the public leaderboard were able to overfit on a dataset they couldn't even see? \n\nAnother question : is the split between the public dataset (38% of total I believe) and the private one (remaining 62%) actually random? If so, I don't see why the rankings would be so different even in presence of potential overfitting (to be proven)",
    "3076907": "Overfitting in this case means that they kept on improving their LB score even when the CV score wasn't being stable/not improving therefore fitting the 38% of the data really well but failed miserably on the other 62%",
    "3076922": "How do we know that their CV scores weren't stable?\n\nAlso, even if it's the case, why would the LB public score would be so different from the private one? The datasets are supposed to be just 38% and 68% of the same macro dataset split randomly, right?"
  },
  "source": "meta"
}