{
  "id": 585082,
  "title": "What is wrong with shuffled test.csv",
  "url": "/competitions/drw-crypto-market-prediction/discussion/585082",
  "author_name": "",
  "post_date": "2025-06-17T21:18:59.905083900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>XGBoost will always predict the same way no matter the order of the data of test.csv. Why are people complaining about shuffled test.csv?</p>",
  "messages": [
    {
      "id": "3226597",
      "postDate": "06/17/2025 21:18:59",
      "content": "<p>XGBoost will always predict the same way no matter the order of the data of test.csv. Why are people complaining about shuffled test.csv?</p>",
      "rawMarkdown": "XGBoost will always predict the same way no matter the order of the data of test.csv. Why are people complaining about shuffled test.csv?",
      "votes": null
    },
    {
      "id": "3228207",
      "postDate": "06/19/2025 22:04:30",
      "content": "<p>Masking the timestamp prevents one from leveraging temporal dependencies, which are often critical in time-dependent tasks. Statistics like rolling averages or rolling standard deviations become meaningless without knowing the order. Time-series models can’t be applied either. I read some top solutions from similar competitions, and models like RNNs and Transformers were popular choices!</p>",
      "rawMarkdown": "Masking the timestamp prevents one from leveraging temporal dependencies, which are often critical in time-dependent tasks. Statistics like rolling averages or rolling standard deviations become meaningless without knowing the order. Time-series models can’t be applied either. I read some top solutions from similar competitions, and models like RNNs and Transformers were popular choices!",
      "votes": null
    },
    {
      "id": "3229471",
      "postDate": "06/21/2025 14:33:19",
      "content": "<p>but that becomes useless during inference, right?</p>",
      "rawMarkdown": "but that becomes useless during inference, right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3228207,
      "author_name": "ccyhui",
      "author_url": "",
      "post_date": "06/19/2025 22:04:30",
      "content": "<p>Masking the timestamp prevents one from leveraging temporal dependencies, which are often critical in time-dependent tasks. Statistics like rolling averages or rolling standard deviations become meaningless without knowing the order. Time-series models can’t be applied either. I read some top solutions from similar competitions, and models like RNNs and Transformers were popular choices!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3229471,
          "author_name": "naitikkariwal",
          "author_url": "",
          "post_date": "06/21/2025 14:33:19",
          "content": "<p>but that becomes useless during inference, right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3226597": "XGBoost will always predict the same way no matter the order of the data of test.csv. Why are people complaining about shuffled test.csv?",
    "3228207": "Masking the timestamp prevents one from leveraging temporal dependencies, which are often critical in time-dependent tasks. Statistics like rolling averages or rolling standard deviations become meaningless without knowing the order. Time-series models can’t be applied either. I read some top solutions from similar competitions, and models like RNNs and Transformers were popular choices!",
    "3229471": "but that becomes useless during inference, right?"
  },
  "source": "meta"
}