{
  "id": 580459,
  "title": "[Resolved] Test set label leakage !",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580459",
  "author_name": "Ye_Ai",
  "post_date": "2025-05-24T05:44:22.606000",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>A team has already achieved a perfect score of -1😂<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14634193%2F207813f5f3ce354f3727c901f13c18cf%2Fdrw-leak.png?generation=1748065363453086&amp;alt=media\" alt=\"\"><br>\nupdate: The score of -1 is caused by all-zero submission, the score calculation rule of kaggle system sets the invalid score of nan to -1 to ensure that it is the lowest ranking.</p>\n<p>Thank you to the organizers for actively maintaining the fairness of the competition.</p>",
  "messages": [
    {
      "id": 3208426,
      "postDate": "2025-05-24T05:44:22.607Z",
      "content": "<p>A team has already achieved a perfect score of -1😂<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14634193%2F207813f5f3ce354f3727c901f13c18cf%2Fdrw-leak.png?generation=1748065363453086&amp;alt=media\" alt=\"\"><br>\nupdate: The score of -1 is caused by all-zero submission, the score calculation rule of kaggle system sets the invalid score of nan to -1 to ensure that it is the lowest ranking.</p>\n<p>Thank you to the organizers for actively maintaining the fairness of the competition.</p>",
      "rawMarkdown": "A team has already achieved a perfect score of -1😂\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14634193%2F207813f5f3ce354f3727c901f13c18cf%2Fdrw-leak.png?generation=1748065363453086&alt=media)\nupdate: The score of -1 is caused by all-zero submission, the score calculation rule of kaggle system sets the invalid score of nan to -1 to ensure that it is the lowest ranking.\n\nThank you to the organizers for actively maintaining the fairness of the competition.",
      "votes": 11
    },
    {
      "id": 3208657,
      "postDate": "2025-05-24T12:32:58.493Z",
      "content": "<p>Thank you for raising this. We’ve reviewed the submission, and the “perfect” score of -1 reflects how Kaggle evaluates a naive all-zero submission. At this time, we see no indication of data leakage in the test set. We will continue to monitor and validate the integrity of the competition data.</p>",
      "rawMarkdown": "Thank you for raising this. We’ve reviewed the submission, and the “perfect” score of -1 reflects how Kaggle evaluates a naive all-zero submission. At this time, we see no indication of data leakage in the test set. We will continue to monitor and validate the integrity of the competition data.",
      "votes": 5,
      "replies": [
        {
          "id": 3208672,
          "postDate": "2025-05-24T13:08:28.480Z",
          "content": "<p>Thank you for your clarification and for maintaining the fairness of the competition.</p>",
          "rawMarkdown": "Thank you for your clarification and for maintaining the fairness of the competition."
        },
        {
          "id": 3227170,
          "postDate": "2025-06-18T16:00:59.523Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3227264,
          "postDate": "2025-06-18T18:18:16.923Z",
          "content": "<p>It does seem that test target has been reverse engineered - there are LB scores of -0.83, 0.56 etc. What do you guys think about all this? It would be an interesting competition but winners will clearly be among those who manage to reconstruct the target and \"tune\" their models accordingly. <br>\n <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> </p>",
          "rawMarkdown": "It does seem that test target has been reverse engineered - there are LB scores of -0.83, 0.56 etc. What do you guys think about all this? It would be an interesting competition but winners will clearly be among those who manage to reconstruct the target and \"tune\" their models accordingly. \n @drwtrading ",
          "votes": 4
        },
        {
          "id": 3234689,
          "postDate": "2025-06-28T07:58:20.933Z",
          "content": "<p>We DO need an explanation on the top scores on the LB</p>",
          "rawMarkdown": "We DO need an explanation on the top scores on the LB",
          "votes": 1
        }
      ]
    },
    {
      "id": 3208569,
      "postDate": "2025-05-24T11:15:51.977Z",
      "content": "<p>I guess was pretty inevitable : <a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580239\" target=\"_blank\">https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580239</a></p>\n<p>If target is just BTC close price change 2024-03 onwards, early stopping on test won't be difficult to hide..</p>",
      "rawMarkdown": "I guess was pretty inevitable : https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580239\n\nIf target is just BTC close price change 2024-03 onwards, early stopping on test won't be difficult to hide..",
      "votes": 3
    },
    {
      "id": 3236442,
      "postDate": "2025-06-30T09:52:14.693Z",
      "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> perhaps you can consider swapping out the test data at the end for completely different unseen period to the provided test.parquet ? Pretty clear the competition data integrity is no longer there now</p>",
      "rawMarkdown": "@drwtrading perhaps you can consider swapping out the test data at the end for completely different unseen period to the provided test.parquet ? Pretty clear the competition data integrity is no longer there now",
      "votes": 2
    },
    {
      "id": 3208647,
      "postDate": "2025-05-24T12:21:00.903Z",
      "content": "<p>With just 2 submissions XP</p>",
      "rawMarkdown": "With just 2 submissions XP",
      "votes": 1
    },
    {
      "id": 3208553,
      "postDate": "2025-05-24T10:50:46.727Z",
      "content": "<p>what does it mean minus 1?</p>",
      "rawMarkdown": "what does it mean minus 1?",
      "votes": 1,
      "replies": [
        {
          "id": 3208566,
          "postDate": "2025-05-24T11:11:51.100Z",
          "content": "<p>The evaluation metric for this competition is the Pearson correlation coefficient, where a score of -1 or 1 means a perfect score.<br>\nThis indicates that a team has found the original data source and obtained the label of the private test set for the competition.</p>",
          "rawMarkdown": "The evaluation metric for this competition is the Pearson correlation coefficient, where a score of -1 or 1 means a perfect score.\nThis indicates that a team has found the original data source and obtained the label of the private test set for the competition.",
          "votes": 2,
          "replies": [
            {
              "id": 3209378,
              "postDate": "2025-05-25T17:18:04.077Z",
              "content": "<p>But how did they find the original data?</p>",
              "rawMarkdown": "But how did they find the original data?"
            }
          ]
        }
      ]
    },
    {
      "id": 3208462,
      "postDate": "2025-05-24T07:27:44.243Z",
      "content": "<blockquote>\n  <p>The public leaderboard during the competition will not be scored and serves only for authoring your model submissions using the public testing data. Once the active submission phase ends, we will update the private leaderboard using more recent data, and this will be used to determine the final team rankings.</p>\n</blockquote>\n<p>Does this mean that the competition organizers will change the test set at the end of the competition?​</p>",
      "rawMarkdown": ">The public leaderboard during the competition will not be scored and serves only for authoring your model submissions using the public testing data. Once the active submission phase ends, we will update the private leaderboard using more recent data, and this will be used to determine the final team rankings.\n\nDoes this mean that the competition organizers will change the test set at the end of the competition?​",
      "votes": 1,
      "replies": [
        {
          "id": 3208471,
          "postDate": "2025-05-24T07:34:25.903Z",
          "content": "<p>I think they use 49% of the data for the public score, probably the 49% oldest points and then they will extend with the 51% more recent data. Only a guess.</p>",
          "rawMarkdown": "I think they use 49% of the data for the public score, probably the 49% oldest points and then they will extend with the 51% more recent data. Only a guess.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3208446,
      "postDate": "2025-05-24T06:40:49.023Z",
      "content": "<p>They said this in the overview, so I assume test label leakage wouldn't be a huge problem for this competition? Do correct me if I'm wrong here: <br>\n\"Note: For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.\"</p>",
      "rawMarkdown": "They said this in the overview, so I assume test label leakage wouldn't be a huge problem for this competition? Do correct me if I'm wrong here: \n\"Note: For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.\"",
      "votes": 1,
      "replies": [
        {
          "id": 3208458,
          "postDate": "2025-05-24T07:17:08.577Z",
          "content": "<p>In my opinion, if a team obtains the true labels of the test set, they can optimize feature engineering and model training parameters for the test set.<br>\nsuch as: </p>\n<ol>\n<li>Using features that perform well on both the training and test sets, discarding features that are prone to overfitting. </li>\n<li>Training on the training set, but searching for model parameters based on the test set score, which leads to overfitting the test set. </li>\n</ol>\n<p>Both of these methods can evade code review and achieve significant score improvements on the test set. This is unfair to other teams.</p>",
          "rawMarkdown": "In my opinion, if a team obtains the true labels of the test set, they can optimize feature engineering and model training parameters for the test set.\nsuch as: \n1. Using features that perform well on both the training and test sets, discarding features that are prone to overfitting. \n2. Training on the training set, but searching for model parameters based on the test set score, which leads to overfitting the test set. \n\nBoth of these methods can evade code review and achieve significant score improvements on the test set. This is unfair to other teams.",
          "votes": 5,
          "replies": [
            {
              "id": 3208757,
              "postDate": "2025-05-24T16:10:29.390Z",
              "content": "<p>Got it! Thanks for the clarification!</p>",
              "rawMarkdown": "Got it! Thanks for the clarification!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3242755,
      "postDate": "2025-07-06T10:43:57.473Z",
      "content": "<p>what's the relationship between '-1' &amp; test leak?</p>",
      "rawMarkdown": "what's the relationship between '-1' & test leak?"
    }
  ],
  "comments": [
    {
      "id": 3208657,
      "author_name": "DRW Trading",
      "author_url": "",
      "post_date": "2025-05-24T12:32:58.493000",
      "content": "<p>Thank you for raising this. We’ve reviewed the submission, and the “perfect” score of -1 reflects how Kaggle evaluates a naive all-zero submission. At this time, we see no indication of data leakage in the test set. We will continue to monitor and validate the integrity of the competition data.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3208672,
          "author_name": "Ye_Ai",
          "author_url": "",
          "post_date": "2025-05-24T13:08:28.480000",
          "content": "<p>Thank you for your clarification and for maintaining the fairness of the competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3227170,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-06-18T16:00:59.523000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3227264,
          "author_name": "Etm117",
          "author_url": "",
          "post_date": "2025-06-18T18:18:16.923000",
          "content": "<p>It does seem that test target has been reverse engineered - there are LB scores of -0.83, 0.56 etc. What do you guys think about all this? It would be an interesting competition but winners will clearly be among those who manage to reconstruct the target and \"tune\" their models accordingly. <br>\n <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3234689,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-06-28T07:58:20.933000",
          "content": "<p>We DO need an explanation on the top scores on the LB</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3208569,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2025-05-24T11:15:51.977000",
      "content": "<p>I guess was pretty inevitable : <a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580239\" target=\"_blank\">https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580239</a></p>\n<p>If target is just BTC close price change 2024-03 onwards, early stopping on test won't be difficult to hide..</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3236442,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2025-06-30T09:52:14.693000",
      "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> perhaps you can consider swapping out the test data at the end for completely different unseen period to the provided test.parquet ? Pretty clear the competition data integrity is no longer there now</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3208647,
      "author_name": "AC",
      "author_url": "",
      "post_date": "2025-05-24T12:21:00.903000",
      "content": "<p>With just 2 submissions XP</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3208553,
      "author_name": "Jubayer Hasan",
      "author_url": "",
      "post_date": "2025-05-24T10:50:46.727000",
      "content": "<p>what does it mean minus 1?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3208566,
          "author_name": "Ye_Ai",
          "author_url": "",
          "post_date": "2025-05-24T11:11:51.100000",
          "content": "<p>The evaluation metric for this competition is the Pearson correlation coefficient, where a score of -1 or 1 means a perfect score.<br>\nThis indicates that a team has found the original data source and obtained the label of the private test set for the competition.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3209378,
              "author_name": "Stable Space",
              "author_url": "",
              "post_date": "2025-05-25T17:18:04.077000",
              "content": "<p>But how did they find the original data?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3208462,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2025-05-24T07:27:44.243000",
      "content": "<blockquote>\n  <p>The public leaderboard during the competition will not be scored and serves only for authoring your model submissions using the public testing data. Once the active submission phase ends, we will update the private leaderboard using more recent data, and this will be used to determine the final team rankings.</p>\n</blockquote>\n<p>Does this mean that the competition organizers will change the test set at the end of the competition?​</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3208471,
          "author_name": "Quentin Adatte",
          "author_url": "",
          "post_date": "2025-05-24T07:34:25.903000",
          "content": "<p>I think they use 49% of the data for the public score, probably the 49% oldest points and then they will extend with the 51% more recent data. Only a guess.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3208446,
      "author_name": "hongan",
      "author_url": "",
      "post_date": "2025-05-24T06:40:49.023000",
      "content": "<p>They said this in the overview, so I assume test label leakage wouldn't be a huge problem for this competition? Do correct me if I'm wrong here: <br>\n\"Note: For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.\"</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3208458,
          "author_name": "Ye_Ai",
          "author_url": "",
          "post_date": "2025-05-24T07:17:08.577000",
          "content": "<p>In my opinion, if a team obtains the true labels of the test set, they can optimize feature engineering and model training parameters for the test set.<br>\nsuch as: </p>\n<ol>\n<li>Using features that perform well on both the training and test sets, discarding features that are prone to overfitting. </li>\n<li>Training on the training set, but searching for model parameters based on the test set score, which leads to overfitting the test set. </li>\n</ol>\n<p>Both of these methods can evade code review and achieve significant score improvements on the test set. This is unfair to other teams.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3208757,
              "author_name": "hongan",
              "author_url": "",
              "post_date": "2025-05-24T16:10:29.390000",
              "content": "<p>Got it! Thanks for the clarification!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3242755,
      "author_name": "SHENYUCHEN",
      "author_url": "",
      "post_date": "2025-07-06T10:43:57.473000",
      "content": "<p>what's the relationship between '-1' &amp; test leak?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3208426": "A team has already achieved a perfect score of -1😂\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14634193%2F207813f5f3ce354f3727c901f13c18cf%2Fdrw-leak.png?generation=1748065363453086&alt=media)\nupdate: The score of -1 is caused by all-zero submission, the score calculation rule of kaggle system sets the invalid score of nan to -1 to ensure that it is the lowest ranking.\n\nThank you to the organizers for actively maintaining the fairness of the competition.",
    "3208657": "Thank you for raising this. We’ve reviewed the submission, and the “perfect” score of -1 reflects how Kaggle evaluates a naive all-zero submission. At this time, we see no indication of data leakage in the test set. We will continue to monitor and validate the integrity of the competition data.",
    "3208569": "I guess was pretty inevitable : https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/580239\n\nIf target is just BTC close price change 2024-03 onwards, early stopping on test won't be difficult to hide..",
    "3236442": "@drwtrading perhaps you can consider swapping out the test data at the end for completely different unseen period to the provided test.parquet ? Pretty clear the competition data integrity is no longer there now",
    "3208647": "With just 2 submissions XP",
    "3208553": "what does it mean minus 1?",
    "3208462": ">The public leaderboard during the competition will not be scored and serves only for authoring your model submissions using the public testing data. Once the active submission phase ends, we will update the private leaderboard using more recent data, and this will be used to determine the final team rankings.\n\nDoes this mean that the competition organizers will change the test set at the end of the competition?​",
    "3208446": "They said this in the overview, so I assume test label leakage wouldn't be a huge problem for this competition? Do correct me if I'm wrong here: \n\"Note: For top-performing teams eligible for prizes, submission of a Jupyter Notebook capable of successfully generating the prediction output is mandatory to qualify for the award. We will conduct a code review to verify the reproducibility of results and ensure that no future peeking occurred during predictions. Failure to provide a compliant notebook may result in disqualification from prize eligibility. DRW reserves the right, at its sole discretion, to review submitted code and disqualify any participant found to be engaging in such practices.\"",
    "3242755": "what's the relationship between '-1' & test leak?"
  }
}