{
  "id": 580062,
  "title": "Masking timestamps in the test dataset",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580062",
  "author_name": "",
  "post_date": "2025-05-22T07:56:38.578347200Z",
  "votes": 22,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I understand why timestamps in the test set were masked. </p>\n<p>But this prevents the most relevant feature engineering techniques - capturing trends in prices, volume, bid-ask spreads.</p>\n<p>Typically in any time-series forecasting, the strongest features are the ones related to lagged target features - where the lag window depends on the timing of data availability during inference. But we cannot use them here. </p>\n<p>the only proper way to do that is via time series API, but i guess that is only available in featured competitions</p>",
  "messages": [
    {
      "id": "3207081",
      "postDate": "05/22/2025 07:56:38",
      "content": "<p>I understand why timestamps in the test set were masked. </p>\n<p>But this prevents the most relevant feature engineering techniques - capturing trends in prices, volume, bid-ask spreads.</p>\n<p>Typically in any time-series forecasting, the strongest features are the ones related to lagged target features - where the lag window depends on the timing of data availability during inference. But we cannot use them here. </p>\n<p>the only proper way to do that is via time series API, but i guess that is only available in featured competitions</p>",
      "rawMarkdown": "I understand why timestamps in the test set were masked. \n\nBut this prevents the most relevant feature engineering techniques - capturing trends in prices, volume, bid-ask spreads.\n\nTypically in any time-series forecasting, the strongest features are the ones related to lagged target features - where the lag window depends on the timing of data availability during inference. But we cannot use them here. \n\nthe only proper way to do that is via time series API, but i guess that is only available in featured competitions",
      "votes": null
    },
    {
      "id": "3207103",
      "postDate": "05/22/2025 08:33:26",
      "content": "<blockquote>\n  <p>To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.</p>\n</blockquote>\n<p>This means we cannot utilize any rolling features or time-series models. 🤔</p>",
      "rawMarkdown": "> To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.\n\nThis means we cannot utilize any rolling features or time-series models. 🤔",
      "votes": null
    },
    {
      "id": "3207107",
      "postDate": "05/22/2025 08:39:41",
      "content": "<p>That is exactly the point</p>",
      "rawMarkdown": "That is exactly the point",
      "votes": null
    },
    {
      "id": "3207131",
      "postDate": "05/22/2025 09:26:40",
      "content": "<p>Ironically, some people would probably try to reverse engineer timestamps and they might even succeed too</p>",
      "rawMarkdown": "Ironically, some people would probably try to reverse engineer timestamps and they might even succeed too",
      "votes": null
    },
    {
      "id": "3207139",
      "postDate": "05/22/2025 09:36:38",
      "content": "<blockquote>\n  <p>Just keep in mind that, similar to the note about using future data, any external dataset containing future information will be considered a rule violation. So make sure everything you use would have been available at the time of prediction!</p>\n</blockquote>\n<p>Reverse engineering timestamps without access to future data is challenging, yet achievable with future data. However, I'm uncertain how the organizer can ascertain whether the model has been trained on and memorized this future data.</p>",
      "rawMarkdown": "> Just keep in mind that, similar to the note about using future data, any external dataset containing future information will be considered a rule violation. So make sure everything you use would have been available at the time of prediction!\n\nReverse engineering timestamps without access to future data is challenging, yet achievable with future data. However, I'm uncertain how the organizer can ascertain whether the model has been trained on and memorized this future data.",
      "votes": null
    },
    {
      "id": "3207147",
      "postDate": "05/22/2025 09:42:54",
      "content": "<p>Does the ID in the test file maintain the chronological order?</p>",
      "rawMarkdown": "Does the ID in the test file maintain the chronological order?",
      "votes": null
    },
    {
      "id": "3207148",
      "postDate": "05/22/2025 09:44:43",
      "content": "<blockquote>\n  <p>timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.</p>\n</blockquote>",
      "rawMarkdown": ">timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.",
      "votes": null
    },
    {
      "id": "3207210",
      "postDate": "05/22/2025 12:07:32",
      "content": "<p>This is exactly what is going to happen</p>",
      "rawMarkdown": "This is exactly what is going to happen",
      "votes": null
    },
    {
      "id": "3207300",
      "postDate": "05/22/2025 14:50:32",
      "content": "<p>I was thinking the same, so basically here, we simply have 800 features that needs to be independently mapped to a value between 0 and 1. Is that correct?</p>",
      "rawMarkdown": "I was thinking the same, so basically here, we simply have 800 features that needs to be independently mapped to a value between 0 and 1. Is that correct?",
      "votes": null
    },
    {
      "id": "3207324",
      "postDate": "05/22/2025 15:28:09",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> sir , is this more likely a regression problem rather than a time series competition ? I unnderstand they masked test timestamps , but even tho they are not sorted I guess unlike train data . so for me it sounds a regression problem , how to handle things , any guidence plaase ? I am new at time series . How to properly approach and handle it ?</p>",
      "rawMarkdown": "narsil sir , is this more likely a regression problem rather than a time series competition ? I unnderstand they masked test timestamps , but even tho they are not sorted I guess unlike train data . so for me it sounds a regression problem , how to handle things , any guidence plaase ? I am new at time series . How to properly approach and handle it ?",
      "votes": null
    },
    {
      "id": "3207410",
      "postDate": "05/22/2025 17:28:17",
      "content": "<p>maybe it's about working with the instantaneous market state, a different kind of challenge </p>",
      "rawMarkdown": "maybe it's about working with the instantaneous market state, a different kind of challenge",
      "votes": null
    },
    {
      "id": "3208315",
      "postDate": "05/23/2025 21:28:57",
      "content": "<p>I don't think we have a constraint of 0 to 1; we can map them on any scale. They are checked using Pearson correlation, so the range doesn't matter.</p>",
      "rawMarkdown": "I don't think we have a constraint of 0 to 1; we can map them on any scale. They are checked using Pearson correlation, so the range doesn't matter.",
      "votes": null
    },
    {
      "id": "3208341",
      "postDate": "05/23/2025 23:33:19",
      "content": "<p>Yeah I confirm we dont, I was mislead by one of my plots, sorry!</p>",
      "rawMarkdown": "Yeah I confirm we dont, I was mislead by one of my plots, sorry!",
      "votes": null
    },
    {
      "id": "3208666",
      "postDate": "05/24/2025 12:56:13",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> thank you for this thread. first thing that came into mind was typical time series FE, I was wondering how to get started when timestamps are masked in test data.</p>",
      "rawMarkdown": "narsil thank you for this thread. first thing that came into mind was typical time series FE, I was wondering how to get started when timestamps are masked in test data.",
      "votes": null
    },
    {
      "id": "3208738",
      "postDate": "05/24/2025 15:40:39",
      "content": "<p>I guess this is a <strong>regression problem</strong>!</p>",
      "rawMarkdown": "I guess this is a **regression problem**!",
      "votes": null
    },
    {
      "id": "3208903",
      "postDate": "05/24/2025 22:31:05",
      "content": "<p>if that's actually possible, is that allowed according to competition rules?</p>",
      "rawMarkdown": "if that's actually possible, is that allowed according to competition rules?",
      "votes": null
    },
    {
      "id": "3210837",
      "postDate": "05/27/2025 18:07:17",
      "content": "<p>I think yes.<br>\nTime can be predicted using a model and nothing prevents us from training more models <a href=\"https://www.kaggle.com/meetbabariya\" target=\"_blank\">@meetbabariya</a> </p>",
      "rawMarkdown": "I think yes.\nTime can be predicted using a model and nothing prevents us from training more models @meetbabariya",
      "votes": null
    },
    {
      "id": "3219347",
      "postDate": "06/07/2025 14:52:17",
      "content": "<p>I agree with your opinions.<br>\nI wrote the discussion to suggest for test data structure.<br>\n<a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/583522\" target=\"_blank\">https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/583522</a></p>",
      "rawMarkdown": "I agree with your opinions.\nI wrote the discussion to suggest for test data structure.\nhttps://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/583522",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3207103,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "05/22/2025 08:33:26",
      "content": "<blockquote>\n  <p>To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.</p>\n</blockquote>\n<p>This means we cannot utilize any rolling features or time-series models. 🤔</p>",
      "votes": null,
      "replies": [
        {
          "id": 3207107,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "05/22/2025 08:39:41",
          "content": "<p>That is exactly the point</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219347,
              "author_name": "motono0223",
              "author_url": "",
              "post_date": "06/07/2025 14:52:17",
              "content": "<p>I agree with your opinions.<br>\nI wrote the discussion to suggest for test data structure.<br>\n<a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/583522\" target=\"_blank\">https://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/583522</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3207131,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "05/22/2025 09:26:40",
      "content": "<p>Ironically, some people would probably try to reverse engineer timestamps and they might even succeed too</p>",
      "votes": null,
      "replies": [
        {
          "id": 3207139,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "05/22/2025 09:36:38",
          "content": "<blockquote>\n  <p>Just keep in mind that, similar to the note about using future data, any external dataset containing future information will be considered a rule violation. So make sure everything you use would have been available at the time of prediction!</p>\n</blockquote>\n<p>Reverse engineering timestamps without access to future data is challenging, yet achievable with future data. However, I'm uncertain how the organizer can ascertain whether the model has been trained on and memorized this future data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3207210,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "05/22/2025 12:07:32",
          "content": "<p>This is exactly what is going to happen</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3208903,
          "author_name": "meetbabariya",
          "author_url": "",
          "post_date": "05/24/2025 22:31:05",
          "content": "<p>if that's actually possible, is that allowed according to competition rules?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3210837,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "05/27/2025 18:07:17",
              "content": "<p>I think yes.<br>\nTime can be predicted using a model and nothing prevents us from training more models <a href=\"https://www.kaggle.com/meetbabariya\" target=\"_blank\">@meetbabariya</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3207147,
      "author_name": "chenxin1991",
      "author_url": "",
      "post_date": "05/22/2025 09:42:54",
      "content": "<p>Does the ID in the test file maintain the chronological order?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3207148,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "05/22/2025 09:44:43",
          "content": "<blockquote>\n  <p>timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3207300,
      "author_name": "quentinadatte1307",
      "author_url": "",
      "post_date": "05/22/2025 14:50:32",
      "content": "<p>I was thinking the same, so basically here, we simply have 800 features that needs to be independently mapped to a value between 0 and 1. Is that correct?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3208315,
          "author_name": "akashpatel001",
          "author_url": "",
          "post_date": "05/23/2025 21:28:57",
          "content": "<p>I don't think we have a constraint of 0 to 1; we can map them on any scale. They are checked using Pearson correlation, so the range doesn't matter.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3208341,
              "author_name": "quentinadatte1307",
              "author_url": "",
              "post_date": "05/23/2025 23:33:19",
              "content": "<p>Yeah I confirm we dont, I was mislead by one of my plots, sorry!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3207324,
      "author_name": "ayushkhaire",
      "author_url": "",
      "post_date": "05/22/2025 15:28:09",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> sir , is this more likely a regression problem rather than a time series competition ? I unnderstand they masked test timestamps , but even tho they are not sorted I guess unlike train data . so for me it sounds a regression problem , how to handle things , any guidence plaase ? I am new at time series . How to properly approach and handle it ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3208738,
          "author_name": "andrewguanzc",
          "author_url": "",
          "post_date": "05/24/2025 15:40:39",
          "content": "<p>I guess this is a <strong>regression problem</strong>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3207410,
      "author_name": "jaejohn",
      "author_url": "",
      "post_date": "05/22/2025 17:28:17",
      "content": "<p>maybe it's about working with the instantaneous market state, a different kind of challenge </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3208666,
      "author_name": "snehalkhandewale",
      "author_url": "",
      "post_date": "05/24/2025 12:56:13",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> thank you for this thread. first thing that came into mind was typical time series FE, I was wondering how to get started when timestamps are masked in test data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3207081": "I understand why timestamps in the test set were masked. \n\nBut this prevents the most relevant feature engineering techniques - capturing trends in prices, volume, bid-ask spreads.\n\nTypically in any time-series forecasting, the strongest features are the ones related to lagged target features - where the lag window depends on the timing of data availability during inference. But we cannot use them here. \n\nthe only proper way to do that is via time series API, but i guess that is only available in featured competitions",
    "3207103": "> To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.\n\nThis means we cannot utilize any rolling features or time-series models. 🤔",
    "3207107": "That is exactly the point",
    "3207131": "Ironically, some people would probably try to reverse engineer timestamps and they might even succeed too",
    "3207139": "> Just keep in mind that, similar to the note about using future data, any external dataset containing future information will be considered a rule violation. So make sure everything you use would have been available at the time of prediction!\n\nReverse engineering timestamps without access to future data is challenging, yet achievable with future data. However, I'm uncertain how the organizer can ascertain whether the model has been trained on and memorized this future data.",
    "3207147": "Does the ID in the test file maintain the chronological order?",
    "3207148": ">timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.",
    "3207210": "This is exactly what is going to happen",
    "3207300": "I was thinking the same, so basically here, we simply have 800 features that needs to be independently mapped to a value between 0 and 1. Is that correct?",
    "3207324": "narsil sir , is this more likely a regression problem rather than a time series competition ? I unnderstand they masked test timestamps , but even tho they are not sorted I guess unlike train data . so for me it sounds a regression problem , how to handle things , any guidence plaase ? I am new at time series . How to properly approach and handle it ?",
    "3207410": "maybe it's about working with the instantaneous market state, a different kind of challenge",
    "3208315": "I don't think we have a constraint of 0 to 1; we can map them on any scale. They are checked using Pearson correlation, so the range doesn't matter.",
    "3208341": "Yeah I confirm we dont, I was mislead by one of my plots, sorry!",
    "3208666": "narsil thank you for this thread. first thing that came into mind was typical time series FE, I was wondering how to get started when timestamps are masked in test data.",
    "3208738": "I guess this is a **regression problem**!",
    "3208903": "if that's actually possible, is that allowed according to competition rules?",
    "3210837": "I think yes.\nTime can be predicted using a model and nothing prevents us from training more models @meetbabariya",
    "3219347": "I agree with your opinions.\nI wrote the discussion to suggest for test data structure.\nhttps://www.kaggle.com/competitions/drw-crypto-market-prediction/discussion/583522"
  },
  "source": "meta"
}