{
  "id": 581126,
  "title": "Reverse Engineer the timestamps in Test?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/581126",
  "author_name": "Giovanni Marco Dall'Olio",
  "post_date": "2025-05-28T13:01:43.482000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In this competition, the timestamps in the test data have been anonymized and shuffled. This is a shame, because it means we cannot use any time series approach.</p>\n<p>Do you think it would be possible to reverse-engineer the original order of the rows? </p>\n<p>We don't know the name and significance of most of the columns in the train dataset, but maybe, some of these features contain aggregated information, such as averages over sliding windows or percentages of changes since the previous step. Is there a way to figure this out? </p>\n<p>Another approach is to use the train dataset to predict month, season, and other dates that may be important for the price of bitcoin. I would train a model using all features first, then select the ones that have higher importance, and train again. Does this make sense?</p>",
  "messages": [
    {
      "id": 3211452,
      "postDate": "2025-05-28T13:01:43.483Z",
      "content": "<p>In this competition, the timestamps in the test data have been anonymized and shuffled. This is a shame, because it means we cannot use any time series approach.</p>\n<p>Do you think it would be possible to reverse-engineer the original order of the rows? </p>\n<p>We don't know the name and significance of most of the columns in the train dataset, but maybe, some of these features contain aggregated information, such as averages over sliding windows or percentages of changes since the previous step. Is there a way to figure this out? </p>\n<p>Another approach is to use the train dataset to predict month, season, and other dates that may be important for the price of bitcoin. I would train a model using all features first, then select the ones that have higher importance, and train again. Does this make sense?</p>",
      "rawMarkdown": "In this competition, the timestamps in the test data have been anonymized and shuffled. This is a shame, because it means we cannot use any time series approach.\n\nDo you think it would be possible to reverse-engineer the original order of the rows? \n\nWe don't know the name and significance of most of the columns in the train dataset, but maybe, some of these features contain aggregated information, such as averages over sliding windows or percentages of changes since the previous step. Is there a way to figure this out? \n\nAnother approach is to use the train dataset to predict month, season, and other dates that may be important for the price of bitcoin. I would train a model using all features first, then select the ones that have higher importance, and train again. Does this make sense?",
      "votes": 7
    },
    {
      "id": 3215835,
      "postDate": "2025-06-02T17:39:00.923Z",
      "content": "<p>I tried to use the LSTM architecture, then I selected 1440 batch files, but none of this gave any result. Data without timestamps, we will not be able to reduce them to time series, alas(</p>",
      "rawMarkdown": "I tried to use the LSTM architecture, then I selected 1440 batch files, but none of this gave any result. Data without timestamps, we will not be able to reduce them to time series, alas(",
      "votes": 1
    },
    {
      "id": 3213677,
      "postDate": "2025-05-30T10:14:05.137Z",
      "content": "<p>I think, out of those 890 anonymous features, we have to consider some important ones only. The question is how do we select the best features.<br>\nand the timestamps could be handled by weight sampling.</p>",
      "rawMarkdown": "I think, out of those 890 anonymous features, we have to consider some important ones only. The question is how do we select the best features.\nand the timestamps could be handled by weight sampling.",
      "votes": 1
    },
    {
      "id": 3214956,
      "postDate": "2025-06-01T10:29:33.483Z",
      "content": "<p>If we make a mistake in the order of the rows and predict the target based on cumulative features, we will probably capture information from the future. So there is no point in restoring the order of the rows.</p>",
      "rawMarkdown": "If we make a mistake in the order of the rows and predict the target based on cumulative features, we will probably capture information from the future. So there is no point in restoring the order of the rows.",
      "votes": 2
    },
    {
      "id": 3212241,
      "postDate": "2025-05-29T14:37:05.113Z",
      "content": "<p>This is a good idea. May be you can make it. But this does not make sense. Importantly is not helpful for us to understand the market.</p>",
      "rawMarkdown": "This is a good idea. May be you can make it. But this does not make sense. Importantly is not helpful for us to understand the market.",
      "replies": [
        {
          "id": 3213746,
          "postDate": "2025-05-30T11:33:34.333Z",
          "content": "<p>I'm trying that here: <a href=\"https://www.kaggle.com/code/dalloliogm/rev-engineering-the-original-order-of-test-rows\" target=\"_blank\">https://www.kaggle.com/code/dalloliogm/rev-engineering-the-original-order-of-test-rows</a><br>\nI agree it is very impractical, but I wanted to try anyways :-)</p>",
          "rawMarkdown": "I'm trying that here: https://www.kaggle.com/code/dalloliogm/rev-engineering-the-original-order-of-test-rows\nI agree it is very impractical, but I wanted to try anyways :-)",
          "votes": -2
        }
      ]
    },
    {
      "id": 3213615,
      "postDate": "2025-05-30T08:52:02.740Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3215835,
      "author_name": "Dmitry Kiryukhin",
      "author_url": "",
      "post_date": "2025-06-02T17:39:00.923000",
      "content": "<p>I tried to use the LSTM architecture, then I selected 1440 batch files, but none of this gave any result. Data without timestamps, we will not be able to reduce them to time series, alas(</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3213677,
      "author_name": "Vishal Painjane",
      "author_url": "",
      "post_date": "2025-05-30T10:14:05.137000",
      "content": "<p>I think, out of those 890 anonymous features, we have to consider some important ones only. The question is how do we select the best features.<br>\nand the timestamps could be handled by weight sampling.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3214956,
      "author_name": "AntonLKaggler",
      "author_url": "",
      "post_date": "2025-06-01T10:29:33.483000",
      "content": "<p>If we make a mistake in the order of the rows and predict the target based on cumulative features, we will probably capture information from the future. So there is no point in restoring the order of the rows.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3212241,
      "author_name": "AndrewGuanzc",
      "author_url": "",
      "post_date": "2025-05-29T14:37:05.113000",
      "content": "<p>This is a good idea. May be you can make it. But this does not make sense. Importantly is not helpful for us to understand the market.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3213746,
          "author_name": "Giovanni Marco Dall'Olio",
          "author_url": "",
          "post_date": "2025-05-30T11:33:34.333000",
          "content": "<p>I'm trying that here: <a href=\"https://www.kaggle.com/code/dalloliogm/rev-engineering-the-original-order-of-test-rows\" target=\"_blank\">https://www.kaggle.com/code/dalloliogm/rev-engineering-the-original-order-of-test-rows</a><br>\nI agree it is very impractical, but I wanted to try anyways :-)</p>",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 3213615,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-30T08:52:02.740000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3211452": "In this competition, the timestamps in the test data have been anonymized and shuffled. This is a shame, because it means we cannot use any time series approach.\n\nDo you think it would be possible to reverse-engineer the original order of the rows? \n\nWe don't know the name and significance of most of the columns in the train dataset, but maybe, some of these features contain aggregated information, such as averages over sliding windows or percentages of changes since the previous step. Is there a way to figure this out? \n\nAnother approach is to use the train dataset to predict month, season, and other dates that may be important for the price of bitcoin. I would train a model using all features first, then select the ones that have higher importance, and train again. Does this make sense?",
    "3215835": "I tried to use the LSTM architecture, then I selected 1440 batch files, but none of this gave any result. Data without timestamps, we will not be able to reduce them to time series, alas(",
    "3213677": "I think, out of those 890 anonymous features, we have to consider some important ones only. The question is how do we select the best features.\nand the timestamps could be handled by weight sampling.",
    "3214956": "If we make a mistake in the order of the rows and predict the target based on cumulative features, we will probably capture information from the future. So there is no point in restoring the order of the rows.",
    "3212241": "This is a good idea. May be you can make it. But this does not make sense. Importantly is not helpful for us to understand the market.",
    "3213615": ""
  }
}