{
  "id": 589052,
  "title": "[Important] Leaderboard Reboot with Updated Competition Data",
  "url": "/competitions/drw-crypto-market-prediction/discussion/589052",
  "author_name": "DRW Trading",
  "post_date": "2025-07-09T23:27:44.667000",
  "votes": 12,
  "comment_count": 69,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>First, thank you all for your enthusiasm, effort, and valuable contributions to the competition so far. We understand it may be frustrating to see a data update mid-competition, but this decision was made after careful consideration to ensure fairness and integrity.</p>\n<p>The competition data has been updated with a newly shuffled test set, and as a result, the leaderboard will be rebooted. The previous leaderboard no longer reflects the current data (which now has different true test timestamps), so please <strong>re-download the dataset</strong>, rerun your code, and submit updated predictions.</p>\n<p>We are in contact with Kaggle regarding a leaderboard reboot. Please don’t be concerned if the updated scores are not yet visible.</p>\n<p>As part of this update:</p>\n<ul>\n<li>Test data have been newly shuffled. The rule against future peaking still applies.</li>\n<li>Columns with excessive abnormal values have been removed to streamline data preprocessing. You will now see <strong>X1–X780 instead of X1–X890</strong>.</li>\n<li>The remaining features have been normalized.</li>\n</ul>\n<p>While the proprietary and market features have been revised, they remain accurate and consistent in meaning. We've taken great care to preserve feature semantics and ensure that prior analyses remain valid and minimally impacted by these changes.</p>\n<p>Thank you again for your understanding and continued participation—we're excited to see how your models perform on the refreshed dataset!</p>\n<p>Best regards,<br>\nThe DRW &amp; Cumberland Team</p>\n<hr>\n<p><strong>Update on Jul 11 (UTC): All submissions made before the dataset update have been removed. A very small number may appear rescored if the same participant submitted again around the time of the dataset update, due to technical reasons—but this has no impact on the leaderboard.</strong></p>",
  "messages": [
    {
      "id": 3247577,
      "postDate": "2025-07-13T03:14:29.867Z",
      "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> , can you help us understand what the code review will entail? I still don't fully understand how notebooks will be approved/disproved.</p>\n<p>If you look at the results of running simple linear regression based feature analysis on the train dataset and the test dataset (using pseudo labels through unshuffling the data and using time series to estimate the prediction values), you can extract insights on the behavior of the test dataset. The below is based on a linear regression analysis (for simplicity) but the impact is the same. Anyone who is able to create psuedo labels based on the unshuffling and time-series prediction of the test data is then able to chose features that are most compatible with the train and test dataset - significantly improving scores and avoiding instability, noise, and drift.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fb5be68800f0317246b0846671f70b379%2F__results___1_1.png?generation=1752376298824875&amp;alt=media\" alt=\"\"></p>\n<p>I am not suggesting anyone do this. Nor will I submit a notebook using these forward looking insights. However, I want to emphasize again, that in a submission notebook it would be impossible to tell if someone find a good feature or feature interaction through time series based psuedo label analysis (against the rules) or whether they found the feature through some other analysis that is based off train only.</p>",
      "rawMarkdown": "@drwtrading , can you help us understand what the code review will entail? I still don't fully understand how notebooks will be approved/disproved.\n\nIf you look at the results of running simple linear regression based feature analysis on the train dataset and the test dataset (using pseudo labels through unshuffling the data and using time series to estimate the prediction values), you can extract insights on the behavior of the test dataset. The below is based on a linear regression analysis (for simplicity) but the impact is the same. Anyone who is able to create psuedo labels based on the unshuffling and time-series prediction of the test data is then able to chose features that are most compatible with the train and test dataset - significantly improving scores and avoiding instability, noise, and drift.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fb5be68800f0317246b0846671f70b379%2F__results___1_1.png?generation=1752376298824875&alt=media)\n\nI am not suggesting anyone do this. Nor will I submit a notebook using these forward looking insights. However, I want to emphasize again, that in a submission notebook it would be impossible to tell if someone find a good feature or feature interaction through time series based psuedo label analysis (against the rules) or whether they found the feature through some other analysis that is based off train only.",
      "votes": 16
    },
    {
      "id": 3245833,
      "postDate": "2025-07-09T23:27:44.667Z",
      "content": "<p>Hi Kagglers,</p>\n<p>First, thank you all for your enthusiasm, effort, and valuable contributions to the competition so far. We understand it may be frustrating to see a data update mid-competition, but this decision was made after careful consideration to ensure fairness and integrity.</p>\n<p>The competition data has been updated with a newly shuffled test set, and as a result, the leaderboard will be rebooted. The previous leaderboard no longer reflects the current data (which now has different true test timestamps), so please <strong>re-download the dataset</strong>, rerun your code, and submit updated predictions.</p>\n<p>We are in contact with Kaggle regarding a leaderboard reboot. Please don’t be concerned if the updated scores are not yet visible.</p>\n<p>As part of this update:</p>\n<ul>\n<li>Test data have been newly shuffled. The rule against future peaking still applies.</li>\n<li>Columns with excessive abnormal values have been removed to streamline data preprocessing. You will now see <strong>X1–X780 instead of X1–X890</strong>.</li>\n<li>The remaining features have been normalized.</li>\n</ul>\n<p>While the proprietary and market features have been revised, they remain accurate and consistent in meaning. We've taken great care to preserve feature semantics and ensure that prior analyses remain valid and minimally impacted by these changes.</p>\n<p>Thank you again for your understanding and continued participation—we're excited to see how your models perform on the refreshed dataset!</p>\n<p>Best regards,<br>\nThe DRW &amp; Cumberland Team</p>\n<hr>\n<p><strong>Update on Jul 11 (UTC): All submissions made before the dataset update have been removed. A very small number may appear rescored if the same participant submitted again around the time of the dataset update, due to technical reasons—but this has no impact on the leaderboard.</strong></p>",
      "rawMarkdown": "Hi Kagglers,\n\nFirst, thank you all for your enthusiasm, effort, and valuable contributions to the competition so far. We understand it may be frustrating to see a data update mid-competition, but this decision was made after careful consideration to ensure fairness and integrity.\n\nThe competition data has been updated with a newly shuffled test set, and as a result, the leaderboard will be rebooted. The previous leaderboard no longer reflects the current data (which now has different true test timestamps), so please **re-download the dataset**, rerun your code, and submit updated predictions.\n\nWe are in contact with Kaggle regarding a leaderboard reboot. Please don’t be concerned if the updated scores are not yet visible.\n\nAs part of this update:\n\n- Test data have been newly shuffled. The rule against future peaking still applies.\n- Columns with excessive abnormal values have been removed to streamline data preprocessing. You will now see **X1–X780 instead of X1–X890**.\n- The remaining features have been normalized.\n\nWhile the proprietary and market features have been revised, they remain accurate and consistent in meaning. We've taken great care to preserve feature semantics and ensure that prior analyses remain valid and minimally impacted by these changes.\n\nThank you again for your understanding and continued participation—we're excited to see how your models perform on the refreshed dataset!\n\nBest regards,\nThe DRW & Cumberland Team\n\n------------\n\n**Update on Jul 11 (UTC): All submissions made before the dataset update have been removed. A very small number may appear rescored if the same participant submitted again around the time of the dataset update, due to technical reasons—but this has no impact on the leaderboard.**",
      "votes": 11
    },
    {
      "id": 3250135,
      "postDate": "2025-07-17T18:43:04.133Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a>,</p>\n<p>Could you specify the format of final submission, given lots of abnormal public scores? Is everyone required to submit a Kaggle Notebook, or you will reach out to potential winners for their code solutions</p>",
      "rawMarkdown": "Hi @drwtrading,\n\nCould you specify the format of final submission, given lots of abnormal public scores? Is everyone required to submit a Kaggle Notebook, or you will reach out to potential winners for their code solutions",
      "votes": 5,
      "replies": [
        {
          "id": 3250170,
          "postDate": "2025-07-17T20:08:59.463Z",
          "content": "<p>+1 Same question as him.</p>",
          "rawMarkdown": "+1 Same question as him.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3245903,
      "postDate": "2025-07-10T02:39:46.153Z",
      "content": "<p>Hi DRW,</p>\n<p>What do you mean by the remaining features have been normalized, does that mean you applied a z-score standardizer on the train dataset and transformed the test dataset? Also did all the columns get shuffled?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi DRW,\n\nWhat do you mean by the remaining features have been normalized, does that mean you applied a z-score standardizer on the train dataset and transformed the test dataset? Also did all the columns get shuffled?\n\nThanks",
      "votes": 5,
      "replies": [
        {
          "id": 3245915,
          "postDate": "2025-07-10T03:24:26.383Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3246196,
          "postDate": "2025-07-10T14:30:48.687Z",
          "content": "<p>Did DRW respond to you and then delete the comment? I've seen a few deleted comments, but I'm not sure if they are from DRW, although I think at least one of them was.</p>\n<p>People who saw this comment before it was deleted might have a competitive advantage.</p>",
          "rawMarkdown": "Did DRW respond to you and then delete the comment? I've seen a few deleted comments, but I'm not sure if they are from DRW, although I think at least one of them was.\n\nPeople who saw this comment before it was deleted might have a competitive advantage.",
          "replies": [
            {
              "id": 3246217,
              "postDate": "2025-07-10T14:58:27.490Z",
              "content": "<p>I don’t know. I hope DRW responds because if the columns are shuffled it will affect our models.</p>",
              "rawMarkdown": "I don’t know. I hope DRW responds because if the columns are shuffled it will affect our models."
            },
            {
              "id": 3246220,
              "postDate": "2025-07-10T15:05:45.917Z",
              "content": "<p>I mean, we probably have to redo our feature engineering and selection from scratch anyway right? Some of the features have been wiped and I assume all the features have been significantly transformed to the point of being essentially new features (otherwise competitors could just apply the same transformation to the unscrambled test data and have twice as much data to train on)</p>",
              "rawMarkdown": "I mean, we probably have to redo our feature engineering and selection from scratch anyway right? Some of the features have been wiped and I assume all the features have been significantly transformed to the point of being essentially new features (otherwise competitors could just apply the same transformation to the unscrambled test data and have twice as much data to train on)"
            },
            {
              "id": 3246247,
              "postDate": "2025-07-10T15:22:51.523Z",
              "content": "<p>My initial research suggest that while the features are not the same, there are patterns that can be found between the old data and the new data. Thus people who tuned their models on the unmasked data can run sub-models to figure out relationships between new features and old features in order to incorporate ill-derived learnings from the unmasked data into the new data.</p>\n<p>An oversimplified way of deconstructing this would be to:</p>\n<ol>\n<li>Review the most important old features, especially the old features that were tuned based on the unmasked data.</li>\n<li>Train a model that predicts these old features using the new proprietary features. If you can train a model that accurately predict important old features (again especially old features that were tuned based on the unmasked data) with certain new features, then you can effectively generate proxies for the old features.</li>\n</ol>\n<p>For example, using the traditional XGBoost configuration that got scores of around .13 on the LB, I was able to do the following:</p>\n<p>1.) Use the original features in the XGBoost model.<br>\n2.) Find important features in the unmasked dataset, with consideration for drift, stability, etc.<br>\n3.) Incorporate those important features and/or derivatives thereof into the XGBoost model.<br>\n4.) X856_times_X598 was a feature that showed preliminary promise in the unmasked dataset.<br>\n5.) With the new dataset, I just need to figure out what new features can accurately predict the old features of X856_times_X598, then select those new features themselves or use a sub-model to recreate a proxy variable that would estimate the old X856_times_X598.<br>\n6.) Similar relationships and patterns between the old and new data can likely be found, even if the new features are truly shuffled and transformed, the power of models allows us to find relationships.</p>",
              "rawMarkdown": "My initial research suggest that while the features are not the same, there are patterns that can be found between the old data and the new data. Thus people who tuned their models on the unmasked data can run sub-models to figure out relationships between new features and old features in order to incorporate ill-derived learnings from the unmasked data into the new data.\n\nAn oversimplified way of deconstructing this would be to:\n\n1. Review the most important old features, especially the old features that were tuned based on the unmasked data.\n2. Train a model that predicts these old features using the new proprietary features. If you can train a model that accurately predict important old features (again especially old features that were tuned based on the unmasked data) with certain new features, then you can effectively generate proxies for the old features.\n\nFor example, using the traditional XGBoost configuration that got scores of around .13 on the LB, I was able to do the following:\n\n1.) Use the original features in the XGBoost model.\n2.) Find important features in the unmasked dataset, with consideration for drift, stability, etc.\n3.) Incorporate those important features and/or derivatives thereof into the XGBoost model.\n4.) X856_times_X598 was a feature that showed preliminary promise in the unmasked dataset.\n5.) With the new dataset, I just need to figure out what new features can accurately predict the old features of X856_times_X598, then select those new features themselves or use a sub-model to recreate a proxy variable that would estimate the old X856_times_X598.\n6.) Similar relationships and patterns between the old and new data can likely be found, even if the new features are truly shuffled and transformed, the power of models allows us to find relationships."
            },
            {
              "id": 3246292,
              "postDate": "2025-07-10T16:31:10.057Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/longissachairfay\" target=\"_blank\">@longissachairfay</a> for sharing your valuable insight, but how do you make sure that patterns found before will generalize well in the new private test set</p>",
              "rawMarkdown": "Thanks @longissachairfay for sharing your valuable insight, but how do you make sure that patterns found before will generalize well in the new private test set"
            },
            {
              "id": 3246298,
              "postDate": "2025-07-10T16:43:53.800Z",
              "content": "<p>Tracking train vs CV scores and stability will help with that.</p>\n<p>I'd say that building a model that generalizes well has been the primary challenge in this competition, or any competition for that matter, and this challenge existed even before the dataset was reset.</p>\n<p>Using the previously unmasked test data can help guide feature engineering, model architectures, and other decisions if done correctly, regardless of the dataset being shuffled and transformed again. Even though transformations was made with the new dataset, I believe the following assumptions true:</p>\n<p>1.) The old dataset and new dataset are both crypto datasets and may follow similar patterns.<br>\n2.) Transformations and adjustments in the new dataset can occur, but it is unlikely these transformations are so extreme and obfuscated as to eliminate all possible patterns and relationships between the original and new dataset. If the data was completely obfuscated it would provide provide less value to DRW in real world trading exercises. The new data could be mixed up, adjusted in scale, and transformed using various patterns and functions, but ML models may be able to pickup and estimate these patterns, at least in part.</p>",
              "rawMarkdown": "Tracking train vs CV scores and stability will help with that.\n\nI'd say that building a model that generalizes well has been the primary challenge in this competition, or any competition for that matter, and this challenge existed even before the dataset was reset.\n\nUsing the previously unmasked test data can help guide feature engineering, model architectures, and other decisions if done correctly, regardless of the dataset being shuffled and transformed again. Even though transformations was made with the new dataset, I believe the following assumptions true:\n\n1.) The old dataset and new dataset are both crypto datasets and may follow similar patterns.\n2.) Transformations and adjustments in the new dataset can occur, but it is unlikely these transformations are so extreme and obfuscated as to eliminate all possible patterns and relationships between the original and new dataset. If the data was completely obfuscated it would provide provide less value to DRW in real world trading exercises. The new data could be mixed up, adjusted in scale, and transformed using various patterns and functions, but ML models may be able to pickup and estimate these patterns, at least in part.\n\n"
            },
            {
              "id": 3246358,
              "postDate": "2025-07-10T18:52:37.967Z",
              "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> please tell us how you normalized the data and if the column names were shuffled, the distribution seems to be different.</p>",
              "rawMarkdown": "@drwtrading please tell us how you normalized the data and if the column names were shuffled, the distribution seems to be different."
            }
          ]
        }
      ]
    },
    {
      "id": 3246454,
      "postDate": "2025-07-11T01:34:48.243Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> , I know you are working hard to manage the current situation, but I would suggest that you provide an update on what is happening with scoring and leaderboard changes, preferably before changes are made.</p>\n<p>It is very confusing seeing the leaderboard change, some notebooks removed, other notebooks staying - there needs to be an official announcement on what is happening with more detail.</p>",
      "rawMarkdown": "Dear @drwtrading , I know you are working hard to manage the current situation, but I would suggest that you provide an update on what is happening with scoring and leaderboard changes, preferably before changes are made.\n\nIt is very confusing seeing the leaderboard change, some notebooks removed, other notebooks staying - there needs to be an official announcement on what is happening with more detail.",
      "votes": 6,
      "replies": [
        {
          "id": 3246571,
          "postDate": "2025-07-11T06:54:35.583Z",
          "content": "<p>Dear Kaggler, thank you for your feedback. All submissions made before the dataset update have now been removed or rescored.</p>",
          "rawMarkdown": "Dear Kaggler, thank you for your feedback. All submissions made before the dataset update have now been removed or rescored.",
          "votes": -1,
          "replies": [
            {
              "id": 3246572,
              "postDate": "2025-07-11T07:01:55.730Z",
              "content": "<p>Have you noticed the data has leak again? Someone has figured it out.</p>",
              "rawMarkdown": "Have you noticed the data has leak again? Someone has figured it out.",
              "votes": 2
            },
            {
              "id": 3246739,
              "postDate": "2025-07-11T12:40:21.223Z",
              "content": "<p>How about the \"Code\" section? There are still notebooks showing scores of 0.5+.</p>",
              "rawMarkdown": "How about the \"Code\" section? There are still notebooks showing scores of 0.5+."
            },
            {
              "id": 3246745,
              "postDate": "2025-07-11T12:47:15.777Z",
              "content": "<p>DRW, are you going to take any action to make this competition fair again? As myself and others have discovered, in a very short amount of time, it is entirely possible to use insights from the old dataset and its unshuffled test component to get a competitive edge in the rebooted competition.</p>",
              "rawMarkdown": "DRW, are you going to take any action to make this competition fair again? As myself and others have discovered, in a very short amount of time, it is entirely possible to use insights from the old dataset and its unshuffled test component to get a competitive edge in the rebooted competition.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 3245858,
      "postDate": "2025-07-10T00:45:28.960Z",
      "content": "<p>I'm honestly at a loss here. All of this is very disappointing to be honest, which from your message I think you already understand.</p>\n<p>However, reshuffling the data does not solve the problem because it does not eliminate the insights that people already gathered when analyzing the unshuffled dataset. These insights can be used to intelligently create features for training and test which will give significant advantages to some teams, without showing any visible signs of forward peaking in the notebook.</p>\n<p>Also, can you help me understand what the statement below means:</p>\n<p>\"While the proprietary and market features have been revised, they remain accurate and consistent in meaning.\"</p>\n<p>If they have been revised, how are they consistent in meaning? I hope this wasn't just a scaling adjustment, but that is the only way I could see \"consistent in meaning\" to be accurate.</p>",
      "rawMarkdown": "I'm honestly at a loss here. All of this is very disappointing to be honest, which from your message I think you already understand.\n\nHowever, reshuffling the data does not solve the problem because it does not eliminate the insights that people already gathered when analyzing the unshuffled dataset. These insights can be used to intelligently create features for training and test which will give significant advantages to some teams, without showing any visible signs of forward peaking in the notebook.\n\nAlso, can you help me understand what the statement below means:\n\n\"While the proprietary and market features have been revised, they remain accurate and consistent in meaning.\"\n\nIf they have been revised, how are they consistent in meaning? I hope this wasn't just a scaling adjustment, but that is the only way I could see \"consistent in meaning\" to be accurate.\n\n\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 3245860,
          "postDate": "2025-07-10T00:49:11.340Z",
          "content": "<p>We understand your concerns and appreciate you sharing them openly. You're absolutely right that once certain insights are gained, they can't be “unlearned.” However, we’ve employed several changes—not just reshuffling—to ensure that the current test set remains valid, while also rendering previously hacked labels ineffective. These measures aim to level the playing field as much as possible going forward.</p>\n<p>We’re committed to maintaining fairness and will continue monitoring for potential leakage during the review phase.</p>",
          "rawMarkdown": "We understand your concerns and appreciate you sharing them openly. You're absolutely right that once certain insights are gained, they can't be “unlearned.” However, we’ve employed several changes—not just reshuffling—to ensure that the current test set remains valid, while also rendering previously hacked labels ineffective. These measures aim to level the playing field as much as possible going forward.\n\nWe’re committed to maintaining fairness and will continue monitoring for potential leakage during the review phase.",
          "replies": [
            {
              "id": 3245864,
              "postDate": "2025-07-10T00:53:19.003Z",
              "content": "<p>Thank you for your efforts on this, I do appreciate this.</p>\n<p>I do have my doubts about the ability to render insights gained through unshuffling as ineffective. But I suppose there are ways to highly obfuscate the data in a way that makes it highly unlikely those insights would be useful.</p>\n<p>Also, I do appreciate you cleaning the proprietary dataset. There were some unnecessary columns in the original dataset in my opinion.</p>\n<p>I'm still in the game, just a bit of a speed bump.</p>",
              "rawMarkdown": "Thank you for your efforts on this, I do appreciate this.\n\nI do have my doubts about the ability to render insights gained through unshuffling as ineffective. But I suppose there are ways to highly obfuscate the data in a way that makes it highly unlikely those insights would be useful.\n\nAlso, I do appreciate you cleaning the proprietary dataset. There were some unnecessary columns in the original dataset in my opinion.\n\nI'm still in the game, just a bit of a speed bump."
            }
          ]
        }
      ]
    },
    {
      "id": 3246872,
      "postDate": "2025-07-11T17:09:41.047Z",
      "content": "<p>I found yet another way to reverse the time for the test set. At this point I am not sure if the organizers will be able to keep the competition fair. </p>\n<p>We will only know after code review, I guess.<br>\nThen, again, how would one control for subjective calls about features (some of which can be pure lookahead)</p>",
      "rawMarkdown": "I found yet another way to reverse the time for the test set. At this point I am not sure if the organizers will be able to keep the competition fair. \n\nWe will only know after code review, I guess.\nThen, again, how would one control for subjective calls about features (some of which can be pure lookahead)",
      "votes": 3,
      "replies": [
        {
          "id": 3246875,
          "postDate": "2025-07-11T17:17:54.520Z",
          "content": "<p>In my opinion, it is impossible to keep it fair without a complete reset.</p>\n<p>Once you reverse engineer the order of the test labels, then you use these labels to intelligently augment your training model by selecting features that better align with the test labels.</p>\n<p>There is no way to catch this in code review because there is no way to prove that someone engineered a feature based on the reverse engineered test labels or if they are just good at feature engineering.</p>",
          "rawMarkdown": "In my opinion, it is impossible to keep it fair without a complete reset.\n\nOnce you reverse engineer the order of the test labels, then you use these labels to intelligently augment your training model by selecting features that better align with the test labels.\n\nThere is no way to catch this in code review because there is no way to prove that someone engineered a feature based on the reverse engineered test labels or if they are just good at feature engineering.",
          "votes": 6,
          "replies": [
            {
              "id": 3247058,
              "postDate": "2025-07-12T03:05:45.063Z",
              "content": "<p>Completely agree. It is impossible to quantify to which extent this is an issue. Subjective calls in this case can be completely lookahead. Right now, there are several ways to re-construct the test order and then choose the best features based on the test set.</p>\n<p>The only way out is completely restarting the competitions, with entirely different features and NO test data available for download, with a model-based submission (like most other trading competitions these days).</p>\n<p>This can be summarized by the title of the piece of code posted by quantju: \"ALL U NEED IS GIVE UP\".</p>",
              "rawMarkdown": "Completely agree. It is impossible to quantify to which extent this is an issue. Subjective calls in this case can be completely lookahead. Right now, there are several ways to re-construct the test order and then choose the best features based on the test set.\n\nThe only way out is completely restarting the competitions, with entirely different features and NO test data available for download, with a model-based submission (like most other trading competitions these days).\n\nThis can be summarized by the title of the piece of code posted by quantju: \"ALL U NEED IS GIVE UP\".\n",
              "votes": 2
            },
            {
              "id": 3247066,
              "postDate": "2025-07-12T03:24:47.427Z",
              "content": "<p>All true, I am still learning new things when playing around with the data though, so its all good. I'm not really here hoping I would get a prize.</p>",
              "rawMarkdown": "All true, I am still learning new things when playing around with the data though, so its all good. I'm not really here hoping I would get a prize."
            },
            {
              "id": 3247145,
              "postDate": "2025-07-12T07:28:49.710Z",
              "content": "<p>I think DRW should have a new test set, with half of existing data from the first test set that is for the public leaderboard, and half of new data for the private leaderboard. Future-peeking the public leaderboard has happened in a lot of past trading competitions, what we want is for it not to happen in the private one. So if they give a new test set, one where only the half of the data that is for the public leaderboard has already been seen then the people that do future-peeking would just hurt themselves. I think this is the best way forward for DRW. The public leaderboard will be a mess but again this has always happened in trading competitions.</p>",
              "rawMarkdown": "I think DRW should have a new test set, with half of existing data from the first test set that is for the public leaderboard, and half of new data for the private leaderboard. Future-peeking the public leaderboard has happened in a lot of past trading competitions, what we want is for it not to happen in the private one. So if they give a new test set, one where only the half of the data that is for the public leaderboard has already been seen then the people that do future-peeking would just hurt themselves. I think this is the best way forward for DRW. The public leaderboard will be a mess but again this has always happened in trading competitions.",
              "votes": 2
            }
          ]
        },
        {
          "id": 3246946,
          "postDate": "2025-07-11T20:26:06.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> you might want to recheck the data now with the high LB score </p>",
          "rawMarkdown": "@drwtrading you might want to recheck the data now with the high LB score ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3246003,
      "postDate": "2025-07-10T06:22:08.177Z",
      "content": "<p>Are column names this same?</p>",
      "rawMarkdown": "Are column names this same?",
      "votes": 1,
      "replies": [
        {
          "id": 3246186,
          "postDate": "2025-07-10T14:16:50.887Z",
          "content": "<p>I would also like to know this! Would save time checking things like constant columns, etc.</p>",
          "rawMarkdown": "I would also like to know this! Would save time checking things like constant columns, etc."
        },
        {
          "id": 3246191,
          "postDate": "2025-07-10T14:22:25.480Z",
          "content": "<p>OK just checked and all the constant columns have been removed, i.e. the column names are <em>not</em> the same. I'm assuming some of the very highly-correlated columns have also been removed?</p>",
          "rawMarkdown": "OK just checked and all the constant columns have been removed, i.e. the column names are _not_ the same. I'm assuming some of the very highly-correlated columns have also been removed?",
          "replies": [
            {
              "id": 3246193,
              "postDate": "2025-07-10T14:25:37.737Z",
              "content": "<p>That is what I noticed too</p>",
              "rawMarkdown": "That is what I noticed too"
            }
          ]
        }
      ]
    },
    {
      "id": 3250183,
      "postDate": "2025-07-17T20:42:37.293Z",
      "content": "<p>it doesn't make sense to go through all the CSV files and manually check every code solution. There should be a new test set along with a notebook to run everything efficiently!!!</p>",
      "rawMarkdown": "it doesn't make sense to go through all the CSV files and manually check every code solution. There should be a new test set along with a notebook to run everything efficiently!!!",
      "votes": 2
    },
    {
      "id": 3249373,
      "postDate": "2025-07-16T10:31:46.583Z",
      "content": "<p>Hi, thank you for your effort. For the final submission, do we need to load and process the test data using our model, or is submitting a CSV file with the predictions based on the current test data sufficient?</p>",
      "rawMarkdown": "Hi, thank you for your effort. For the final submission, do we need to load and process the test data using our model, or is submitting a CSV file with the predictions based on the current test data sufficient?",
      "replies": [
        {
          "id": 3249445,
          "postDate": "2025-07-16T13:38:19.843Z",
          "content": "<p>same question here</p>",
          "rawMarkdown": "same question here",
          "votes": 1
        }
      ]
    },
    {
      "id": 3246170,
      "postDate": "2025-07-10T13:22:36.910Z",
      "content": "<p>Dear DRW Hosts (or anyone who can help clarify),</p>\n<p>I'm puzzled by how this solves the issue. Users will already have downloaded and unshuffled test data locally. At the very least this allows them to train their models on the test data period biasing their models to the correct regime? As stated this is still against the rules, but who is going to catch the 79th place entry who boosted their score by a 0.05 by training on the forbidden data?</p>\n<p>I'm a total beginner so possibly I'm missing something obvious but as it stands I'm releuctant to spend the time to adapt to the new data and chasing an extra 0.01 if it'll be impossible to beat someone who just uses a baseline model and cheats for just enough gain to look believable.</p>\n<p>Thanks </p>",
      "rawMarkdown": "Dear DRW Hosts (or anyone who can help clarify),\n\nI'm puzzled by how this solves the issue. Users will already have downloaded and unshuffled test data locally. At the very least this allows them to train their models on the test data period biasing their models to the correct regime? As stated this is still against the rules, but who is going to catch the 79th place entry who boosted their score by a 0.05 by training on the forbidden data?\n\nI'm a total beginner so possibly I'm missing something obvious but as it stands I'm releuctant to spend the time to adapt to the new data and chasing an extra 0.01 if it'll be impossible to beat someone who just uses a baseline model and cheats for just enough gain to look believable.\n\nThanks ",
      "votes": 2,
      "replies": [
        {
          "id": 3246175,
          "postDate": "2025-07-10T13:36:04.110Z",
          "content": "<p>The updated test data includes different true timestamps, making it difficult to map or reuse previous labels. We anticipated this and applied additional manipulations to ensure that prior future-peaking insights are no longer effective.</p>\n<p>We will continuously monitor for suspicious submissions with unusually high scores and conduct thorough code reviews to uphold fairness throughout the competition.</p>",
          "rawMarkdown": "The updated test data includes different true timestamps, making it difficult to map or reuse previous labels. We anticipated this and applied additional manipulations to ensure that prior future-peaking insights are no longer effective.\n\nWe will continuously monitor for suspicious submissions with unusually high scores and conduct thorough code reviews to uphold fairness throughout the competition.",
          "votes": 4,
          "replies": [
            {
              "id": 3246180,
              "postDate": "2025-07-10T13:58:10.120Z",
              "content": "<p>ah ok, the new test data is a completely different time period? Thanks for clarifying.<br>\nActually just one more question for clarity if you don't mind; at the very least those with the old, unshuffled test data have twice as much data to train on, can we assume the new data is sufficiently transformed so that the old test data can't simply be transformed, clipped of the now-redundant features and trained on? Even if its a different period now, having twice as much data to train on is a big advantage.</p>",
              "rawMarkdown": "ah ok, the new test data is a completely different time period? Thanks for clarifying.\nActually just one more question for clarity if you don't mind; at the very least those with the old, unshuffled test data have twice as much data to train on, can we assume the new data is sufficiently transformed so that the old test data can't simply be transformed, clipped of the now-redundant features and trained on? Even if its a different period now, having twice as much data to train on is a big advantage."
            },
            {
              "id": 3249017,
              "postDate": "2025-07-15T14:20:30.490Z",
              "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> what about a mid-range score that uses the leak? Perhaps one could use the leak, drop useless features and then train a model. <br>\nWhat is the intended response in this case?</p>",
              "rawMarkdown": "@drwtrading what about a mid-range score that uses the leak? Perhaps one could use the leak, drop useless features and then train a model. \nWhat is the intended response in this case?",
              "votes": 1
            },
            {
              "id": 3249848,
              "postDate": "2025-07-17T08:13:50.683Z",
              "content": "<p>it seems a waste of time here. not a good setup.</p>",
              "rawMarkdown": "it seems a waste of time here. not a good setup.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3245896,
      "postDate": "2025-07-10T02:25:23.673Z",
      "content": "<p>Hi DRW &amp; Cumberland Team,</p>\n<p>Thanks for your efforts and fast response to keep this match fair. Do you intend to extend the deadline given new dataset used, or the deadline will stay the same? </p>",
      "rawMarkdown": "Hi DRW & Cumberland Team,\n\nThanks for your efforts and fast response to keep this match fair. Do you intend to extend the deadline given new dataset used, or the deadline will stay the same? ",
      "replies": [
        {
          "id": 3245897,
          "postDate": "2025-07-10T02:26:06.553Z",
          "content": "<p>We're still discussing the possibility of an extension and will share an update soon.</p>",
          "rawMarkdown": "We're still discussing the possibility of an extension and will share an update soon.\n",
          "replies": [
            {
              "id": 3245902,
              "postDate": "2025-07-10T02:36:43.983Z",
              "content": "<p>Its already been two months though</p>",
              "rawMarkdown": "Its already been two months though",
              "votes": 1
            },
            {
              "id": 3245918,
              "postDate": "2025-07-10T03:29:07.983Z",
              "content": "<p>True <br>\nThis should have been planned better <br>\nHope kaggle takes necessary measures to prevent effort and time waste of participants <a href=\"https://www.kaggle.com/paperxd\" target=\"_blank\">@paperxd</a> </p>",
              "rawMarkdown": "True \nThis should have been planned better \nHope kaggle takes necessary measures to prevent effort and time waste of participants @paperxd ",
              "votes": 2
            },
            {
              "id": 3246104,
              "postDate": "2025-07-10T10:24:57.750Z",
              "content": "<p>since the feature has been updated, 2-4 weeks extension may be fine. If there is a new leaderboard without hacking, that would be better. Thanks for your organizing</p>",
              "rawMarkdown": "since the feature has been updated, 2-4 weeks extension may be fine. If there is a new leaderboard without hacking, that would be better. Thanks for your organizing",
              "votes": 1
            },
            {
              "id": 3246200,
              "postDate": "2025-07-10T14:38:39.993Z",
              "content": "<p>I don't think an extension will cut it. I've looked at the new dataset along with the old train data and unmasked old test data. In my opinion, there are insights gathered through the unmasking of the old test data that can be applied to the new data to give a competitive advantage.</p>",
              "rawMarkdown": "I don't think an extension will cut it. I've looked at the new dataset along with the old train data and unmasked old test data. In my opinion, there are insights gathered through the unmasking of the old test data that can be applied to the new data to give a competitive advantage."
            },
            {
              "id": 3247002,
              "postDate": "2025-07-12T00:20:57.843Z",
              "content": "<p>We’re looking forward to the final outcome of your discussions. We truly hope that, if an extension is granted, it could be for at least two more weeks — that would be greatly appreciated!</p>",
              "rawMarkdown": "We’re looking forward to the final outcome of your discussions. We truly hope that, if an extension is granted, it could be for at least two more weeks — that would be greatly appreciated!"
            },
            {
              "id": 3247005,
              "postDate": "2025-07-12T00:26:29.740Z",
              "content": "<p>We think there is no point in an extension, if we have an extension more people will just find ways to hack the leaderboard.</p>",
              "rawMarkdown": "We think there is no point in an extension, if we have an extension more people will just find ways to hack the leaderboard."
            },
            {
              "id": 3247008,
              "postDate": "2025-07-12T00:31:19.457Z",
              "content": "<p>They hacked the leaderboard already</p>",
              "rawMarkdown": "They hacked the leaderboard already"
            }
          ]
        }
      ]
    },
    {
      "id": 3253401,
      "postDate": "2025-07-24T15:51:46.837Z",
      "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> I was wondering if it is possible to use different test set for the final submission. </p>",
      "rawMarkdown": "@drwtrading I was wondering if it is possible to use different test set for the final submission. ",
      "replies": [
        {
          "id": 3253403,
          "postDate": "2025-07-24T16:01:14.267Z",
          "content": "<p>I believe a straightforward way to detect overfitted notebook submissions is by reshuffling the test set. This would cause models that rely heavily on timestamp order or lag-based features to perform poorly.</p>",
          "rawMarkdown": "I believe a straightforward way to detect overfitted notebook submissions is by reshuffling the test set. This would cause models that rely heavily on timestamp order or lag-based features to perform poorly.\n",
          "replies": [
            {
              "id": 3253404,
              "postDate": "2025-07-24T16:02:55.533Z",
              "content": "<p>I agree with you. I think another option is using different set of test data.</p>",
              "rawMarkdown": "I agree with you. I think another option is using different set of test data."
            },
            {
              "id": 3253411,
              "postDate": "2025-07-24T16:15:50.780Z",
              "content": "<p>Using a completely different test set would make the competition winners largely a matter of random luck. To be honest, no model can generalize perfectly to any arbitrary dataset — it's like trying to use one key to open every door, which is impossible. A better way to eliminate notebooks that rely on future-peeking is by reshuffling the test set during the evaluation phase.</p>",
              "rawMarkdown": "Using a completely different test set would make the competition winners largely a matter of random luck. To be honest, no model can generalize perfectly to any arbitrary dataset — it's like trying to use one key to open every door, which is impossible. A better way to eliminate notebooks that rely on future-peeking is by reshuffling the test set during the evaluation phase."
            }
          ]
        }
      ]
    },
    {
      "id": 3245865,
      "postDate": "2025-07-10T00:55:20.103Z",
      "content": "<p>Would it be possible to clarify if all prior submissions will be removed? Otherwise, the leaderboard is misleading. Additionally, will the competition time be extended?</p>",
      "rawMarkdown": "Would it be possible to clarify if all prior submissions will be removed? Otherwise, the leaderboard is misleading. Additionally, will the competition time be extended?",
      "replies": [
        {
          "id": 3245870,
          "postDate": "2025-07-10T00:58:21.600Z",
          "content": "<p>All prior submissions will be considered invalid since the test data has changed. The leaderboard scores will be updated accordingly to reflect submissions made on the new dataset.</p>\n<p>As for the competition timeline, we’re still discussing whether an extension will be granted and will share an update soon.</p>",
          "rawMarkdown": "All prior submissions will be considered invalid since the test data has changed. The leaderboard scores will be updated accordingly to reflect submissions made on the new dataset.\n\nAs for the competition timeline, we’re still discussing whether an extension will be granted and will share an update soon.",
          "votes": 1,
          "replies": [
            {
              "id": 3246197,
              "postDate": "2025-07-10T14:32:12.220Z",
              "content": "<p>In my opinion, the current competition should be shutdown, an apology issued. Then a reboot competition started, and you may want to increase the number of award slots to quell angst and discontent. Just my 2 cents.</p>\n<p>There is no way to ensure that any of the future peeking doesn't taint the results with unfair advantages.</p>",
              "rawMarkdown": "In my opinion, the current competition should be shutdown, an apology issued. Then a reboot competition started, and you may want to increase the number of award slots to quell angst and discontent. Just my 2 cents.\n\nThere is no way to ensure that any of the future peeking doesn't taint the results with unfair advantages.",
              "votes": 3
            },
            {
              "id": 3246460,
              "postDate": "2025-07-11T01:46:04.553Z",
              "content": "<p>This seems like the only way forward. Even with adjusted features, teams/individuals can use insights gained from the previous unshuffled test dataset to get an unfair advantage. </p>",
              "rawMarkdown": "This seems like the only way forward. Even with adjusted features, teams/individuals can use insights gained from the previous unshuffled test dataset to get an unfair advantage. "
            },
            {
              "id": 3246762,
              "postDate": "2025-07-11T13:46:18.893Z",
              "content": "<p>As expected, someone has likely already reversed engineered the new dataset, possibly using the old dataset as well for guidance.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fe527868ac7281b3a70f7858f8e302d51%2F2025-07-11_09-45.png?generation=1752241573177090&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "As expected, someone has likely already reversed engineered the new dataset, possibly using the old dataset as well for guidance.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fe527868ac7281b3a70f7858f8e302d51%2F2025-07-11_09-45.png?generation=1752241573177090&alt=media)\n",
              "votes": 2
            },
            {
              "id": 3246853,
              "postDate": "2025-07-11T16:35:16.570Z",
              "content": "<p>This might have been a glitch that was using the old LB</p>",
              "rawMarkdown": "This might have been a glitch that was using the old LB"
            },
            {
              "id": 3246868,
              "postDate": "2025-07-11T16:53:56.990Z",
              "content": "<p>Unfortunately, it is not. The new order within the test dataset has been decoded. The obfuscation and reordering done by DRW was insufficient. I was able to score 0.65 just now.</p>",
              "rawMarkdown": "Unfortunately, it is not. The new order within the test dataset has been decoded. The obfuscation and reordering done by DRW was insufficient. I was able to score 0.65 just now.",
              "votes": 2
            },
            {
              "id": 3246877,
              "postDate": "2025-07-11T17:27:39.963Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/taylorsamarel\" target=\"_blank\">@taylorsamarel</a> , are you sure that you reorder the entire test set right or only the public score part?</p>",
              "rawMarkdown": "Hi @taylorsamarel , are you sure that you reorder the entire test set right or only the public score part?",
              "votes": 1
            },
            {
              "id": 3246884,
              "postDate": "2025-07-11T17:35:58.857Z",
              "content": "<p>It may not be perfect, but it is pretty close. It wasn't all of my work, I was able to build off of other notebooks.</p>\n<p>It seems that the administrators of this competition are discounting notebooks that score greater than ~0.6, but I didn't see this in an announcement anywhere. I understand forward looking is against the competition rules, but I think there is learning to be had when experimenting with notebooks, as long as participants - like myself - understand that these 'peaking' notebooks won't be considered for a prize.</p>\n<p>I think the fundamental issue with this competition is that somewhat arbitrary and non-transparent actions are being taken in an attempt to 'level the playing field'. However, the data continues to provide 'leaked' insights for those that can decompose the new dataset.</p>\n<p>The issue comes from the fact that this 'decomposed insights' can be used to intelligently tune a model, that, at its face, does not involve rule breaking forward peaking.</p>",
              "rawMarkdown": "It may not be perfect, but it is pretty close. It wasn't all of my work, I was able to build off of other notebooks.\n\nIt seems that the administrators of this competition are discounting notebooks that score greater than ~0.6, but I didn't see this in an announcement anywhere. I understand forward looking is against the competition rules, but I think there is learning to be had when experimenting with notebooks, as long as participants - like myself - understand that these 'peaking' notebooks won't be considered for a prize.\n\nI think the fundamental issue with this competition is that somewhat arbitrary and non-transparent actions are being taken in an attempt to 'level the playing field'. However, the data continues to provide 'leaked' insights for those that can decompose the new dataset.\n\nThe issue comes from the fact that this 'decomposed insights' can be used to intelligently tune a model, that, at its face, does not involve rule breaking forward peaking.",
              "votes": 1
            },
            {
              "id": 3246915,
              "postDate": "2025-07-11T19:11:54.910Z",
              "content": "<p>DRW will conduct code reviews though</p>",
              "rawMarkdown": "DRW will conduct code reviews though"
            },
            {
              "id": 3246923,
              "postDate": "2025-07-11T19:30:24.260Z",
              "content": "<p>Have you deleted your submission yourself ? Or has it been moderated</p>",
              "rawMarkdown": "Have you deleted your submission yourself ? Or has it been moderated"
            },
            {
              "id": 3246929,
              "postDate": "2025-07-11T19:49:22.433Z",
              "content": "<p>Seems like it was administratively removed, not exactly sure what that means. Anything above 0.6 didn't stay on the board, but I still see the .599 score on the code section if you sort by score (not on the leaderboard though).</p>",
              "rawMarkdown": "Seems like it was administratively removed, not exactly sure what that means. Anything above 0.6 didn't stay on the board, but I still see the .599 score on the code section if you sort by score (not on the leaderboard though)."
            },
            {
              "id": 3246930,
              "postDate": "2025-07-11T19:50:19.257Z",
              "content": "<p>There is no way to determine if an engineered feature was a good guess, or if it was derived through future peaking. Someone could derive the perfect features to engineer using the shuffled label data and then just hard code those features into a separate notebook.</p>",
              "rawMarkdown": "There is no way to determine if an engineered feature was a good guess, or if it was derived through future peaking. Someone could derive the perfect features to engineer using the shuffled label data and then just hard code those features into a separate notebook."
            }
          ]
        }
      ]
    },
    {
      "id": 3246069,
      "postDate": "2025-07-10T08:37:53.613Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3246011,
      "postDate": "2025-07-10T06:47:00.427Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3245877,
      "postDate": "2025-07-10T01:25:21.040Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3245879,
          "postDate": "2025-07-10T01:43:28.370Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3245838,
      "postDate": "2025-07-09T23:41:56.770Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3247577,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-07-13T03:14:29.867000",
      "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> , can you help us understand what the code review will entail? I still don't fully understand how notebooks will be approved/disproved.</p>\n<p>If you look at the results of running simple linear regression based feature analysis on the train dataset and the test dataset (using pseudo labels through unshuffling the data and using time series to estimate the prediction values), you can extract insights on the behavior of the test dataset. The below is based on a linear regression analysis (for simplicity) but the impact is the same. Anyone who is able to create psuedo labels based on the unshuffling and time-series prediction of the test data is then able to chose features that are most compatible with the train and test dataset - significantly improving scores and avoiding instability, noise, and drift.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fb5be68800f0317246b0846671f70b379%2F__results___1_1.png?generation=1752376298824875&amp;alt=media\" alt=\"\"></p>\n<p>I am not suggesting anyone do this. Nor will I submit a notebook using these forward looking insights. However, I want to emphasize again, that in a submission notebook it would be impossible to tell if someone find a good feature or feature interaction through time series based psuedo label analysis (against the rules) or whether they found the feature through some other analysis that is based off train only.</p>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 3250135,
      "author_name": "Alex Zhongs",
      "author_url": "",
      "post_date": "2025-07-17T18:43:04.133000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a>,</p>\n<p>Could you specify the format of final submission, given lots of abnormal public scores? Is everyone required to submit a Kaggle Notebook, or you will reach out to potential winners for their code solutions</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3250170,
          "author_name": "paperxd",
          "author_url": "",
          "post_date": "2025-07-17T20:08:59.463000",
          "content": "<p>+1 Same question as him.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3245903,
      "author_name": "paperxd",
      "author_url": "",
      "post_date": "2025-07-10T02:39:46.153000",
      "content": "<p>Hi DRW,</p>\n<p>What do you mean by the remaining features have been normalized, does that mean you applied a z-score standardizer on the train dataset and transformed the test dataset? Also did all the columns get shuffled?</p>\n<p>Thanks</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3245915,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-10T03:24:26.383000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3246196,
          "author_name": "Taylor S. Amarel",
          "author_url": "",
          "post_date": "2025-07-10T14:30:48.687000",
          "content": "<p>Did DRW respond to you and then delete the comment? I've seen a few deleted comments, but I'm not sure if they are from DRW, although I think at least one of them was.</p>\n<p>People who saw this comment before it was deleted might have a competitive advantage.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3246217,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "2025-07-10T14:58:27.490000",
              "content": "<p>I don’t know. I hope DRW responds because if the columns are shuffled it will affect our models.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246220,
              "author_name": "GreatestCutie",
              "author_url": "",
              "post_date": "2025-07-10T15:05:45.917000",
              "content": "<p>I mean, we probably have to redo our feature engineering and selection from scratch anyway right? Some of the features have been wiped and I assume all the features have been significantly transformed to the point of being essentially new features (otherwise competitors could just apply the same transformation to the unscrambled test data and have twice as much data to train on)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246247,
              "author_name": "Long Issac Hair Fay",
              "author_url": "",
              "post_date": "2025-07-10T15:22:51.523000",
              "content": "<p>My initial research suggest that while the features are not the same, there are patterns that can be found between the old data and the new data. Thus people who tuned their models on the unmasked data can run sub-models to figure out relationships between new features and old features in order to incorporate ill-derived learnings from the unmasked data into the new data.</p>\n<p>An oversimplified way of deconstructing this would be to:</p>\n<ol>\n<li>Review the most important old features, especially the old features that were tuned based on the unmasked data.</li>\n<li>Train a model that predicts these old features using the new proprietary features. If you can train a model that accurately predict important old features (again especially old features that were tuned based on the unmasked data) with certain new features, then you can effectively generate proxies for the old features.</li>\n</ol>\n<p>For example, using the traditional XGBoost configuration that got scores of around .13 on the LB, I was able to do the following:</p>\n<p>1.) Use the original features in the XGBoost model.<br>\n2.) Find important features in the unmasked dataset, with consideration for drift, stability, etc.<br>\n3.) Incorporate those important features and/or derivatives thereof into the XGBoost model.<br>\n4.) X856_times_X598 was a feature that showed preliminary promise in the unmasked dataset.<br>\n5.) With the new dataset, I just need to figure out what new features can accurately predict the old features of X856_times_X598, then select those new features themselves or use a sub-model to recreate a proxy variable that would estimate the old X856_times_X598.<br>\n6.) Similar relationships and patterns between the old and new data can likely be found, even if the new features are truly shuffled and transformed, the power of models allows us to find relationships.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246292,
              "author_name": "Alex Zhongs",
              "author_url": "",
              "post_date": "2025-07-10T16:31:10.057000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/longissachairfay\" target=\"_blank\">@longissachairfay</a> for sharing your valuable insight, but how do you make sure that patterns found before will generalize well in the new private test set</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246298,
              "author_name": "Long Issac Hair Fay",
              "author_url": "",
              "post_date": "2025-07-10T16:43:53.800000",
              "content": "<p>Tracking train vs CV scores and stability will help with that.</p>\n<p>I'd say that building a model that generalizes well has been the primary challenge in this competition, or any competition for that matter, and this challenge existed even before the dataset was reset.</p>\n<p>Using the previously unmasked test data can help guide feature engineering, model architectures, and other decisions if done correctly, regardless of the dataset being shuffled and transformed again. Even though transformations was made with the new dataset, I believe the following assumptions true:</p>\n<p>1.) The old dataset and new dataset are both crypto datasets and may follow similar patterns.<br>\n2.) Transformations and adjustments in the new dataset can occur, but it is unlikely these transformations are so extreme and obfuscated as to eliminate all possible patterns and relationships between the original and new dataset. If the data was completely obfuscated it would provide provide less value to DRW in real world trading exercises. The new data could be mixed up, adjusted in scale, and transformed using various patterns and functions, but ML models may be able to pickup and estimate these patterns, at least in part.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246358,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "2025-07-10T18:52:37.967000",
              "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> please tell us how you normalized the data and if the column names were shuffled, the distribution seems to be different.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3246454,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-07-11T01:34:48.243000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> , I know you are working hard to manage the current situation, but I would suggest that you provide an update on what is happening with scoring and leaderboard changes, preferably before changes are made.</p>\n<p>It is very confusing seeing the leaderboard change, some notebooks removed, other notebooks staying - there needs to be an official announcement on what is happening with more detail.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3246571,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-11T06:54:35.583000",
          "content": "<p>Dear Kaggler, thank you for your feedback. All submissions made before the dataset update have now been removed or rescored.</p>",
          "votes": -1,
          "replies": [
            {
              "id": 3246572,
              "author_name": "wangliang9274",
              "author_url": "",
              "post_date": "2025-07-11T07:01:55.730000",
              "content": "<p>Have you noticed the data has leak again? Someone has figured it out.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3246739,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-11T12:40:21.223000",
              "content": "<p>How about the \"Code\" section? There are still notebooks showing scores of 0.5+.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246745,
              "author_name": "Long Issac Hair Fay",
              "author_url": "",
              "post_date": "2025-07-11T12:47:15.777000",
              "content": "<p>DRW, are you going to take any action to make this competition fair again? As myself and others have discovered, in a very short amount of time, it is entirely possible to use insights from the old dataset and its unshuffled test component to get a competitive edge in the rebooted competition.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3245858,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-07-10T00:45:28.960000",
      "content": "<p>I'm honestly at a loss here. All of this is very disappointing to be honest, which from your message I think you already understand.</p>\n<p>However, reshuffling the data does not solve the problem because it does not eliminate the insights that people already gathered when analyzing the unshuffled dataset. These insights can be used to intelligently create features for training and test which will give significant advantages to some teams, without showing any visible signs of forward peaking in the notebook.</p>\n<p>Also, can you help me understand what the statement below means:</p>\n<p>\"While the proprietary and market features have been revised, they remain accurate and consistent in meaning.\"</p>\n<p>If they have been revised, how are they consistent in meaning? I hope this wasn't just a scaling adjustment, but that is the only way I could see \"consistent in meaning\" to be accurate.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3245860,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-10T00:49:11.340000",
          "content": "<p>We understand your concerns and appreciate you sharing them openly. You're absolutely right that once certain insights are gained, they can't be “unlearned.” However, we’ve employed several changes—not just reshuffling—to ensure that the current test set remains valid, while also rendering previously hacked labels ineffective. These measures aim to level the playing field as much as possible going forward.</p>\n<p>We’re committed to maintaining fairness and will continue monitoring for potential leakage during the review phase.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3245864,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-10T00:53:19.003000",
              "content": "<p>Thank you for your efforts on this, I do appreciate this.</p>\n<p>I do have my doubts about the ability to render insights gained through unshuffling as ineffective. But I suppose there are ways to highly obfuscate the data in a way that makes it highly unlikely those insights would be useful.</p>\n<p>Also, I do appreciate you cleaning the proprietary dataset. There were some unnecessary columns in the original dataset in my opinion.</p>\n<p>I'm still in the game, just a bit of a speed bump.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3246872,
      "author_name": "Vladimir Khismatullin",
      "author_url": "",
      "post_date": "2025-07-11T17:09:41.047000",
      "content": "<p>I found yet another way to reverse the time for the test set. At this point I am not sure if the organizers will be able to keep the competition fair. </p>\n<p>We will only know after code review, I guess.<br>\nThen, again, how would one control for subjective calls about features (some of which can be pure lookahead)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3246875,
          "author_name": "Taylor S. Amarel",
          "author_url": "",
          "post_date": "2025-07-11T17:17:54.520000",
          "content": "<p>In my opinion, it is impossible to keep it fair without a complete reset.</p>\n<p>Once you reverse engineer the order of the test labels, then you use these labels to intelligently augment your training model by selecting features that better align with the test labels.</p>\n<p>There is no way to catch this in code review because there is no way to prove that someone engineered a feature based on the reverse engineered test labels or if they are just good at feature engineering.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 3247058,
              "author_name": "Vladimir Khismatullin",
              "author_url": "",
              "post_date": "2025-07-12T03:05:45.063000",
              "content": "<p>Completely agree. It is impossible to quantify to which extent this is an issue. Subjective calls in this case can be completely lookahead. Right now, there are several ways to re-construct the test order and then choose the best features based on the test set.</p>\n<p>The only way out is completely restarting the competitions, with entirely different features and NO test data available for download, with a model-based submission (like most other trading competitions these days).</p>\n<p>This can be summarized by the title of the piece of code posted by quantju: \"ALL U NEED IS GIVE UP\".</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3247066,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-12T03:24:47.427000",
              "content": "<p>All true, I am still learning new things when playing around with the data though, so its all good. I'm not really here hoping I would get a prize.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3247145,
              "author_name": "NN",
              "author_url": "",
              "post_date": "2025-07-12T07:28:49.710000",
              "content": "<p>I think DRW should have a new test set, with half of existing data from the first test set that is for the public leaderboard, and half of new data for the private leaderboard. Future-peeking the public leaderboard has happened in a lot of past trading competitions, what we want is for it not to happen in the private one. So if they give a new test set, one where only the half of the data that is for the public leaderboard has already been seen then the people that do future-peeking would just hurt themselves. I think this is the best way forward for DRW. The public leaderboard will be a mess but again this has always happened in trading competitions.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3246946,
          "author_name": "Anh Quang Phan",
          "author_url": "",
          "post_date": "2025-07-11T20:26:06.793000",
          "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> you might want to recheck the data now with the high LB score </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3246003,
      "author_name": "Rafał Pawłowski",
      "author_url": "",
      "post_date": "2025-07-10T06:22:08.177000",
      "content": "<p>Are column names this same?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3246186,
          "author_name": "Sarah Jeffreson",
          "author_url": "",
          "post_date": "2025-07-10T14:16:50.887000",
          "content": "<p>I would also like to know this! Would save time checking things like constant columns, etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3246191,
          "author_name": "Sarah Jeffreson",
          "author_url": "",
          "post_date": "2025-07-10T14:22:25.480000",
          "content": "<p>OK just checked and all the constant columns have been removed, i.e. the column names are <em>not</em> the same. I'm assuming some of the very highly-correlated columns have also been removed?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3246193,
              "author_name": "byunjins",
              "author_url": "",
              "post_date": "2025-07-10T14:25:37.737000",
              "content": "<p>That is what I noticed too</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3250183,
      "author_name": "Cyrus",
      "author_url": "",
      "post_date": "2025-07-17T20:42:37.293000",
      "content": "<p>it doesn't make sense to go through all the CSV files and manually check every code solution. There should be a new test set along with a notebook to run everything efficiently!!!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3249373,
      "author_name": "Cyrus",
      "author_url": "",
      "post_date": "2025-07-16T10:31:46.583000",
      "content": "<p>Hi, thank you for your effort. For the final submission, do we need to load and process the test data using our model, or is submitting a CSV file with the predictions based on the current test data sufficient?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3249445,
          "author_name": "Alex Zhongs",
          "author_url": "",
          "post_date": "2025-07-16T13:38:19.843000",
          "content": "<p>same question here</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3246170,
      "author_name": "GreatestCutie",
      "author_url": "",
      "post_date": "2025-07-10T13:22:36.910000",
      "content": "<p>Dear DRW Hosts (or anyone who can help clarify),</p>\n<p>I'm puzzled by how this solves the issue. Users will already have downloaded and unshuffled test data locally. At the very least this allows them to train their models on the test data period biasing their models to the correct regime? As stated this is still against the rules, but who is going to catch the 79th place entry who boosted their score by a 0.05 by training on the forbidden data?</p>\n<p>I'm a total beginner so possibly I'm missing something obvious but as it stands I'm releuctant to spend the time to adapt to the new data and chasing an extra 0.01 if it'll be impossible to beat someone who just uses a baseline model and cheats for just enough gain to look believable.</p>\n<p>Thanks </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3246175,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-10T13:36:04.110000",
          "content": "<p>The updated test data includes different true timestamps, making it difficult to map or reuse previous labels. We anticipated this and applied additional manipulations to ensure that prior future-peaking insights are no longer effective.</p>\n<p>We will continuously monitor for suspicious submissions with unusually high scores and conduct thorough code reviews to uphold fairness throughout the competition.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3246180,
              "author_name": "GreatestCutie",
              "author_url": "",
              "post_date": "2025-07-10T13:58:10.120000",
              "content": "<p>ah ok, the new test data is a completely different time period? Thanks for clarifying.<br>\nActually just one more question for clarity if you don't mind; at the very least those with the old, unshuffled test data have twice as much data to train on, can we assume the new data is sufficiently transformed so that the old test data can't simply be transformed, clipped of the now-redundant features and trained on? Even if its a different period now, having twice as much data to train on is a big advantage.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3249017,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2025-07-15T14:20:30.490000",
              "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> what about a mid-range score that uses the leak? Perhaps one could use the leak, drop useless features and then train a model. <br>\nWhat is the intended response in this case?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3249848,
              "author_name": "Cyrus",
              "author_url": "",
              "post_date": "2025-07-17T08:13:50.683000",
              "content": "<p>it seems a waste of time here. not a good setup.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3245896,
      "author_name": "Alex Zhongs",
      "author_url": "",
      "post_date": "2025-07-10T02:25:23.673000",
      "content": "<p>Hi DRW &amp; Cumberland Team,</p>\n<p>Thanks for your efforts and fast response to keep this match fair. Do you intend to extend the deadline given new dataset used, or the deadline will stay the same? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3245897,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-10T02:26:06.553000",
          "content": "<p>We're still discussing the possibility of an extension and will share an update soon.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3245902,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "2025-07-10T02:36:43.983000",
              "content": "<p>Its already been two months though</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3245918,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2025-07-10T03:29:07.983000",
              "content": "<p>True <br>\nThis should have been planned better <br>\nHope kaggle takes necessary measures to prevent effort and time waste of participants <a href=\"https://www.kaggle.com/paperxd\" target=\"_blank\">@paperxd</a> </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3246104,
              "author_name": "Jackeylov3",
              "author_url": "",
              "post_date": "2025-07-10T10:24:57.750000",
              "content": "<p>since the feature has been updated, 2-4 weeks extension may be fine. If there is a new leaderboard without hacking, that would be better. Thanks for your organizing</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3246200,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-10T14:38:39.993000",
              "content": "<p>I don't think an extension will cut it. I've looked at the new dataset along with the old train data and unmasked old test data. In my opinion, there are insights gathered through the unmasking of the old test data that can be applied to the new data to give a competitive advantage.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3247002,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-12T00:20:57.843000",
              "content": "<p>We’re looking forward to the final outcome of your discussions. We truly hope that, if an extension is granted, it could be for at least two more weeks — that would be greatly appreciated!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3247005,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "2025-07-12T00:26:29.740000",
              "content": "<p>We think there is no point in an extension, if we have an extension more people will just find ways to hack the leaderboard.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3247008,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-12T00:31:19.457000",
              "content": "<p>They hacked the leaderboard already</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3253401,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "2025-07-24T15:51:46.837000",
      "content": "<p><a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> I was wondering if it is possible to use different test set for the final submission. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3253403,
          "author_name": "EL Younes",
          "author_url": "",
          "post_date": "2025-07-24T16:01:14.267000",
          "content": "<p>I believe a straightforward way to detect overfitted notebook submissions is by reshuffling the test set. This would cause models that rely heavily on timestamp order or lag-based features to perform poorly.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3253404,
              "author_name": "byunjins",
              "author_url": "",
              "post_date": "2025-07-24T16:02:55.533000",
              "content": "<p>I agree with you. I think another option is using different set of test data.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3253411,
              "author_name": "EL Younes",
              "author_url": "",
              "post_date": "2025-07-24T16:15:50.780000",
              "content": "<p>Using a completely different test set would make the competition winners largely a matter of random luck. To be honest, no model can generalize perfectly to any arbitrary dataset — it's like trying to use one key to open every door, which is impossible. A better way to eliminate notebooks that rely on future-peeking is by reshuffling the test set during the evaluation phase.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3245865,
      "author_name": "KIT",
      "author_url": "",
      "post_date": "2025-07-10T00:55:20.103000",
      "content": "<p>Would it be possible to clarify if all prior submissions will be removed? Otherwise, the leaderboard is misleading. Additionally, will the competition time be extended?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3245870,
          "author_name": "DRW Trading",
          "author_url": "",
          "post_date": "2025-07-10T00:58:21.600000",
          "content": "<p>All prior submissions will be considered invalid since the test data has changed. The leaderboard scores will be updated accordingly to reflect submissions made on the new dataset.</p>\n<p>As for the competition timeline, we’re still discussing whether an extension will be granted and will share an update soon.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3246197,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-10T14:32:12.220000",
              "content": "<p>In my opinion, the current competition should be shutdown, an apology issued. Then a reboot competition started, and you may want to increase the number of award slots to quell angst and discontent. Just my 2 cents.</p>\n<p>There is no way to ensure that any of the future peeking doesn't taint the results with unfair advantages.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3246460,
              "author_name": "Long Issac Hair Fay",
              "author_url": "",
              "post_date": "2025-07-11T01:46:04.553000",
              "content": "<p>This seems like the only way forward. Even with adjusted features, teams/individuals can use insights gained from the previous unshuffled test dataset to get an unfair advantage. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246762,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-11T13:46:18.893000",
              "content": "<p>As expected, someone has likely already reversed engineered the new dataset, possibly using the old dataset as well for guidance.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fe527868ac7281b3a70f7858f8e302d51%2F2025-07-11_09-45.png?generation=1752241573177090&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3246853,
              "author_name": "Deepak Saldanha",
              "author_url": "",
              "post_date": "2025-07-11T16:35:16.570000",
              "content": "<p>This might have been a glitch that was using the old LB</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246868,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-11T16:53:56.990000",
              "content": "<p>Unfortunately, it is not. The new order within the test dataset has been decoded. The obfuscation and reordering done by DRW was insufficient. I was able to score 0.65 just now.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3246877,
              "author_name": "Alex Zhongs",
              "author_url": "",
              "post_date": "2025-07-11T17:27:39.963000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/taylorsamarel\" target=\"_blank\">@taylorsamarel</a> , are you sure that you reorder the entire test set right or only the public score part?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3246884,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-11T17:35:58.857000",
              "content": "<p>It may not be perfect, but it is pretty close. It wasn't all of my work, I was able to build off of other notebooks.</p>\n<p>It seems that the administrators of this competition are discounting notebooks that score greater than ~0.6, but I didn't see this in an announcement anywhere. I understand forward looking is against the competition rules, but I think there is learning to be had when experimenting with notebooks, as long as participants - like myself - understand that these 'peaking' notebooks won't be considered for a prize.</p>\n<p>I think the fundamental issue with this competition is that somewhat arbitrary and non-transparent actions are being taken in an attempt to 'level the playing field'. However, the data continues to provide 'leaked' insights for those that can decompose the new dataset.</p>\n<p>The issue comes from the fact that this 'decomposed insights' can be used to intelligently tune a model, that, at its face, does not involve rule breaking forward peaking.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3246915,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "2025-07-11T19:11:54.910000",
              "content": "<p>DRW will conduct code reviews though</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246923,
              "author_name": "YannFb",
              "author_url": "",
              "post_date": "2025-07-11T19:30:24.260000",
              "content": "<p>Have you deleted your submission yourself ? Or has it been moderated</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246929,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-11T19:49:22.433000",
              "content": "<p>Seems like it was administratively removed, not exactly sure what that means. Anything above 0.6 didn't stay on the board, but I still see the .599 score on the code section if you sort by score (not on the leaderboard though).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3246930,
              "author_name": "Taylor S. Amarel",
              "author_url": "",
              "post_date": "2025-07-11T19:50:19.257000",
              "content": "<p>There is no way to determine if an engineered feature was a good guess, or if it was derived through future peaking. Someone could derive the perfect features to engineer using the shuffled label data and then just hard code those features into a separate notebook.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3246069,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-10T08:37:53.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3246011,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-10T06:47:00.427000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3245877,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-10T01:25:21.040000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3245879,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-10T01:43:28.370000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3245838,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-09T23:41:56.770000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3247577": "@drwtrading , can you help us understand what the code review will entail? I still don't fully understand how notebooks will be approved/disproved.\n\nIf you look at the results of running simple linear regression based feature analysis on the train dataset and the test dataset (using pseudo labels through unshuffling the data and using time series to estimate the prediction values), you can extract insights on the behavior of the test dataset. The below is based on a linear regression analysis (for simplicity) but the impact is the same. Anyone who is able to create psuedo labels based on the unshuffling and time-series prediction of the test data is then able to chose features that are most compatible with the train and test dataset - significantly improving scores and avoiding instability, noise, and drift.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10414474%2Fb5be68800f0317246b0846671f70b379%2F__results___1_1.png?generation=1752376298824875&alt=media)\n\nI am not suggesting anyone do this. Nor will I submit a notebook using these forward looking insights. However, I want to emphasize again, that in a submission notebook it would be impossible to tell if someone find a good feature or feature interaction through time series based psuedo label analysis (against the rules) or whether they found the feature through some other analysis that is based off train only.",
    "3245833": "Hi Kagglers,\n\nFirst, thank you all for your enthusiasm, effort, and valuable contributions to the competition so far. We understand it may be frustrating to see a data update mid-competition, but this decision was made after careful consideration to ensure fairness and integrity.\n\nThe competition data has been updated with a newly shuffled test set, and as a result, the leaderboard will be rebooted. The previous leaderboard no longer reflects the current data (which now has different true test timestamps), so please **re-download the dataset**, rerun your code, and submit updated predictions.\n\nWe are in contact with Kaggle regarding a leaderboard reboot. Please don’t be concerned if the updated scores are not yet visible.\n\nAs part of this update:\n\n- Test data have been newly shuffled. The rule against future peaking still applies.\n- Columns with excessive abnormal values have been removed to streamline data preprocessing. You will now see **X1–X780 instead of X1–X890**.\n- The remaining features have been normalized.\n\nWhile the proprietary and market features have been revised, they remain accurate and consistent in meaning. We've taken great care to preserve feature semantics and ensure that prior analyses remain valid and minimally impacted by these changes.\n\nThank you again for your understanding and continued participation—we're excited to see how your models perform on the refreshed dataset!\n\nBest regards,\nThe DRW & Cumberland Team\n\n------------\n\n**Update on Jul 11 (UTC): All submissions made before the dataset update have been removed. A very small number may appear rescored if the same participant submitted again around the time of the dataset update, due to technical reasons—but this has no impact on the leaderboard.**",
    "3250135": "Hi @drwtrading,\n\nCould you specify the format of final submission, given lots of abnormal public scores? Is everyone required to submit a Kaggle Notebook, or you will reach out to potential winners for their code solutions",
    "3245903": "Hi DRW,\n\nWhat do you mean by the remaining features have been normalized, does that mean you applied a z-score standardizer on the train dataset and transformed the test dataset? Also did all the columns get shuffled?\n\nThanks",
    "3246454": "Dear @drwtrading , I know you are working hard to manage the current situation, but I would suggest that you provide an update on what is happening with scoring and leaderboard changes, preferably before changes are made.\n\nIt is very confusing seeing the leaderboard change, some notebooks removed, other notebooks staying - there needs to be an official announcement on what is happening with more detail.",
    "3245858": "I'm honestly at a loss here. All of this is very disappointing to be honest, which from your message I think you already understand.\n\nHowever, reshuffling the data does not solve the problem because it does not eliminate the insights that people already gathered when analyzing the unshuffled dataset. These insights can be used to intelligently create features for training and test which will give significant advantages to some teams, without showing any visible signs of forward peaking in the notebook.\n\nAlso, can you help me understand what the statement below means:\n\n\"While the proprietary and market features have been revised, they remain accurate and consistent in meaning.\"\n\nIf they have been revised, how are they consistent in meaning? I hope this wasn't just a scaling adjustment, but that is the only way I could see \"consistent in meaning\" to be accurate.\n\n\n\n",
    "3246872": "I found yet another way to reverse the time for the test set. At this point I am not sure if the organizers will be able to keep the competition fair. \n\nWe will only know after code review, I guess.\nThen, again, how would one control for subjective calls about features (some of which can be pure lookahead)",
    "3246003": "Are column names this same?",
    "3250183": "it doesn't make sense to go through all the CSV files and manually check every code solution. There should be a new test set along with a notebook to run everything efficiently!!!",
    "3249373": "Hi, thank you for your effort. For the final submission, do we need to load and process the test data using our model, or is submitting a CSV file with the predictions based on the current test data sufficient?",
    "3246170": "Dear DRW Hosts (or anyone who can help clarify),\n\nI'm puzzled by how this solves the issue. Users will already have downloaded and unshuffled test data locally. At the very least this allows them to train their models on the test data period biasing their models to the correct regime? As stated this is still against the rules, but who is going to catch the 79th place entry who boosted their score by a 0.05 by training on the forbidden data?\n\nI'm a total beginner so possibly I'm missing something obvious but as it stands I'm releuctant to spend the time to adapt to the new data and chasing an extra 0.01 if it'll be impossible to beat someone who just uses a baseline model and cheats for just enough gain to look believable.\n\nThanks ",
    "3245896": "Hi DRW & Cumberland Team,\n\nThanks for your efforts and fast response to keep this match fair. Do you intend to extend the deadline given new dataset used, or the deadline will stay the same? ",
    "3253401": "@drwtrading I was wondering if it is possible to use different test set for the final submission. ",
    "3245865": "Would it be possible to clarify if all prior submissions will be removed? Otherwise, the leaderboard is misleading. Additionally, will the competition time be extended?",
    "3246069": "",
    "3246011": "",
    "3245877": "",
    "3245838": ""
  }
}