{
  "id": 588445,
  "title": "Is this e regression or a forecasting competition?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588445",
  "author_name": "",
  "post_date": "2025-07-06T13:43:42.756348900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was wondering if this competition can be considered a forecasting or regression competition. We know that train data has timestamps and  is temporally consecutive with a range of 1 minute between one observation and the next. If you only look at train dataset you could consider this as a forecasting problem where you have to predict a 1 minute price movement. But if you look at test dataset there is no timestamp and there is only an ID column. This wouldn't be a problem if the rows were consecutive from a temporal point of view even if the timestamp is masked because you could also find in test dataset the patters identified  in training dataset. But data in test dataset not only have an ID column instead of timestemp but rows are also shuffled.  In this way you lose any temporal pattern in the test dataset. You could try to find out time patterns in rows regardless of timestamps for example if you try to figure out if  among the more than 800 features columns there are some patterns related to time, maybe some columns are just lagged version of other columns and you can find time patters anyway. Another consideration is that test dataset has more then 500000 rows so the forecasting period would be  huge. But I think that you have to predict price movement on a much shorter time horizon.  So I think that if there are some time patterns maybe these patterns are inside the more than 800 features of the dataset but it's not easy to exploit them; it's about finding the right combinations of features of each rows that represent some time pattern in the data that could be exploited in test dataset.</p>",
  "messages": [
    {
      "id": "3242852",
      "postDate": "07/06/2025 13:43:42",
      "content": "<p>I was wondering if this competition can be considered a forecasting or regression competition. We know that train data has timestamps and  is temporally consecutive with a range of 1 minute between one observation and the next. If you only look at train dataset you could consider this as a forecasting problem where you have to predict a 1 minute price movement. But if you look at test dataset there is no timestamp and there is only an ID column. This wouldn't be a problem if the rows were consecutive from a temporal point of view even if the timestamp is masked because you could also find in test dataset the patters identified  in training dataset. But data in test dataset not only have an ID column instead of timestemp but rows are also shuffled.  In this way you lose any temporal pattern in the test dataset. You could try to find out time patterns in rows regardless of timestamps for example if you try to figure out if  among the more than 800 features columns there are some patterns related to time, maybe some columns are just lagged version of other columns and you can find time patters anyway. Another consideration is that test dataset has more then 500000 rows so the forecasting period would be  huge. But I think that you have to predict price movement on a much shorter time horizon.  So I think that if there are some time patterns maybe these patterns are inside the more than 800 features of the dataset but it's not easy to exploit them; it's about finding the right combinations of features of each rows that represent some time pattern in the data that could be exploited in test dataset.</p>",
      "rawMarkdown": "I was wondering if this competition can be considered a forecasting or regression competition. We know that train data has timestamps and  is temporally consecutive with a range of 1 minute between one observation and the next. If you only look at train dataset you could consider this as a forecasting problem where you have to predict a 1 minute price movement. But if you look at test dataset there is no timestamp and there is only an ID column. This wouldn't be a problem if the rows were consecutive from a temporal point of view even if the timestamp is masked because you could also find in test dataset the patters identified  in training dataset. But data in test dataset not only have an ID column instead of timestemp but rows are also shuffled.  In this way you lose any temporal pattern in the test dataset. You could try to find out time patterns in rows regardless of timestamps for example if you try to figure out if  among the more than 800 features columns there are some patterns related to time, maybe some columns are just lagged version of other columns and you can find time patters anyway. Another consideration is that test dataset has more then 500000 rows so the forecasting period would be  huge. But I think that you have to predict price movement on a much shorter time horizon.  So I think that if there are some time patterns maybe these patterns are inside the more than 800 features of the dataset but it's not easy to exploit them; it's about finding the right combinations of features of each rows that represent some time pattern in the data that could be exploited in test dataset.",
      "votes": null
    },
    {
      "id": "3242954",
      "postDate": "07/06/2025 15:28:30",
      "content": "<p>I have read in many discussion threads that the test dataset in shuffled but there isn't any mention of that being the case in the description of the competition. From where does this idea come ?</p>",
      "rawMarkdown": "I have read in many discussion threads that the test dataset in shuffled but there isn't any mention of that being the case in the description of the competition. From where does this idea come ?",
      "votes": null
    },
    {
      "id": "3242979",
      "postDate": "07/06/2025 16:00:06",
      "content": "<p>It's written in Data tab-&gt; timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.</p>",
      "rawMarkdown": "It's written in Data tab-> timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.",
      "votes": null
    },
    {
      "id": "3242994",
      "postDate": "07/06/2025 16:18:30",
      "content": "<p>This is a regression task, not a time series task.</p>",
      "rawMarkdown": "This is a regression task, not a time series task.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3242954,
      "author_name": "yannfb",
      "author_url": "",
      "post_date": "07/06/2025 15:28:30",
      "content": "<p>I have read in many discussion threads that the test dataset in shuffled but there isn't any mention of that being the case in the description of the competition. From where does this idea come ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3242979,
          "author_name": "obiaf88",
          "author_url": "",
          "post_date": "07/06/2025 16:00:06",
          "content": "<p>It's written in Data tab-&gt; timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3242994,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/06/2025 16:18:30",
      "content": "<p>This is a regression task, not a time series task.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3242852": "I was wondering if this competition can be considered a forecasting or regression competition. We know that train data has timestamps and  is temporally consecutive with a range of 1 minute between one observation and the next. If you only look at train dataset you could consider this as a forecasting problem where you have to predict a 1 minute price movement. But if you look at test dataset there is no timestamp and there is only an ID column. This wouldn't be a problem if the rows were consecutive from a temporal point of view even if the timestamp is masked because you could also find in test dataset the patters identified  in training dataset. But data in test dataset not only have an ID column instead of timestemp but rows are also shuffled.  In this way you lose any temporal pattern in the test dataset. You could try to find out time patterns in rows regardless of timestamps for example if you try to figure out if  among the more than 800 features columns there are some patterns related to time, maybe some columns are just lagged version of other columns and you can find time patters anyway. Another consideration is that test dataset has more then 500000 rows so the forecasting period would be  huge. But I think that you have to predict price movement on a much shorter time horizon.  So I think that if there are some time patterns maybe these patterns are inside the more than 800 features of the dataset but it's not easy to exploit them; it's about finding the right combinations of features of each rows that represent some time pattern in the data that could be exploited in test dataset.",
    "3242954": "I have read in many discussion threads that the test dataset in shuffled but there isn't any mention of that being the case in the description of the competition. From where does this idea come ?",
    "3242979": "It's written in Data tab-> timestamp: To prevent future peeking, all timestamps are masked, shuffled, and replaced with a unique ID.",
    "3242994": "This is a regression task, not a time series task."
  },
  "source": "meta"
}