{
  "id": 456090,
  "title": "Leakage happening due to test sampling",
  "url": "/competitions/predict-ai-model-runtime/discussion/456090",
  "author_name": "",
  "post_date": "2023-11-18T02:07:23.802708600Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We can know using the prediction results of random as the default results can greatly improve both lb and pb from this discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083</a></p>\n<p>My team also can achieve a very high score with this leakage<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F9c4996f8a1c8e4b5705479554f392598%2Fslow_or_fast_leak.png?generation=1700273145933699&amp;alt=media\" alt=\"\"></p>\n<p>After discussed with other teams, we think the possible reason is from test data sampling. The original test lables(random and default both) are ranked first,then use the same seed for sampling,so they become same lable for this metrics(Kendal Tau Correlation ).</p>\n<p>Some teams maybe used this leakage, some didn't, it's not fair to all teams, we hope can re-calculate the learderboard with same leakage or without leakage. </p>\n<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> </p>",
  "messages": [
    {
      "id": "2529160",
      "postDate": "11/18/2023 02:07:23",
      "content": "<p>We can know using the prediction results of random as the default results can greatly improve both lb and pb from this discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083\" target=\"_blank\">https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083</a></p>\n<p>My team also can achieve a very high score with this leakage<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F9c4996f8a1c8e4b5705479554f392598%2Fslow_or_fast_leak.png?generation=1700273145933699&amp;alt=media\" alt=\"\"></p>\n<p>After discussed with other teams, we think the possible reason is from test data sampling. The original test lables(random and default both) are ranked first,then use the same seed for sampling,so they become same lable for this metrics(Kendal Tau Correlation ).</p>\n<p>Some teams maybe used this leakage, some didn't, it's not fair to all teams, we hope can re-calculate the learderboard with same leakage or without leakage. </p>\n<p><a href=\"https://www.kaggle.com/mangpophothilimthana\" target=\"_blank\">@mangpophothilimthana</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> </p>",
      "rawMarkdown": "We can know using the prediction results of random as the default results can greatly improve both lb and pb from this discussion.\nhttps://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083\n\nMy team also can achieve a very high score with this leakage![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F9c4996f8a1c8e4b5705479554f392598%2Fslow_or_fast_leak.png?generation=1700273145933699&alt=media)\n\nAfter discussed with other teams, we think the possible reason is from test data sampling. The original test lables(random and default both) are ranked first,then use the same seed for sampling,so they become same lable for this metrics(Kendal Tau Correlation ).\n\nSome teams maybe used this leakage, some didn't, it's not fair to all teams, we hope can re-calculate the learderboard with same leakage or without leakage. \n\n@mangpophothilimthana @ashleychow",
      "votes": null
    },
    {
      "id": "2532344",
      "postDate": "11/21/2023 01:24:59",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>! The leaderboard has been finalized. Thanks!</p>",
      "rawMarkdown": "Thanks @senkin13! The leaderboard has been finalized. Thanks!",
      "votes": null
    },
    {
      "id": "2533077",
      "postDate": "11/21/2023 15:16:14",
      "content": "<p>We did some analysis. We see that none of the top 6 winning teams have cheated. Specifically, because all of their scores on \"{nlp, xla}:default\" collection are much worse than their scores on \"{nlp, xla}:random\" collection.</p>\n<p>However, in the ranks of &gt;6 (7 through 30), we find there are two teams with very similar \"default\" and \"random\" scores, which might've used the trick in <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083\" target=\"_blank\">discussion/456083</a></p>",
      "rawMarkdown": "We did some analysis. We see that none of the top 6 winning teams have cheated. Specifically, because all of their scores on \"{nlp, xla}:default\" collection are much worse than their scores on \"{nlp, xla}:random\" collection.\n\nHowever, in the ranks of >6 (7 through 30), we find there are two teams with very similar \"default\" and \"random\" scores, which might've used the trick in [discussion/456083](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2532344,
      "author_name": "ashleychow",
      "author_url": "",
      "post_date": "11/21/2023 01:24:59",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>! The leaderboard has been finalized. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2533077,
      "author_name": "samihaija",
      "author_url": "",
      "post_date": "11/21/2023 15:16:14",
      "content": "<p>We did some analysis. We see that none of the top 6 winning teams have cheated. Specifically, because all of their scores on \"{nlp, xla}:default\" collection are much worse than their scores on \"{nlp, xla}:random\" collection.</p>\n<p>However, in the ranks of &gt;6 (7 through 30), we find there are two teams with very similar \"default\" and \"random\" scores, which might've used the trick in <a href=\"https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083\" target=\"_blank\">discussion/456083</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2529160": "We can know using the prediction results of random as the default results can greatly improve both lb and pb from this discussion.\nhttps://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083\n\nMy team also can achieve a very high score with this leakage![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F9c4996f8a1c8e4b5705479554f392598%2Fslow_or_fast_leak.png?generation=1700273145933699&alt=media)\n\nAfter discussed with other teams, we think the possible reason is from test data sampling. The original test lables(random and default both) are ranked first,then use the same seed for sampling,so they become same lable for this metrics(Kendal Tau Correlation ).\n\nSome teams maybe used this leakage, some didn't, it's not fair to all teams, we hope can re-calculate the learderboard with same leakage or without leakage. \n\n@mangpophothilimthana @ashleychow",
    "2532344": "Thanks @senkin13! The leaderboard has been finalized. Thanks!",
    "2533077": "We did some analysis. We see that none of the top 6 winning teams have cheated. Specifically, because all of their scores on \"{nlp, xla}:default\" collection are much worse than their scores on \"{nlp, xla}:random\" collection.\n\nHowever, in the ranks of >6 (7 through 30), we find there are two teams with very similar \"default\" and \"random\" scores, which might've used the trick in [discussion/456083](https://www.kaggle.com/competitions/predict-ai-model-runtime/discussion/456083)"
  },
  "source": "meta"
}