{
  "id": 590237,
  "title": "Should we continue with non time series approaches and stuck at 300+ on LB or use time series?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/590237",
  "author_name": "",
  "post_date": "2025-07-18T23:59:30.742984Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>What is your approach? Will they delete hacked solutions?</p>",
  "messages": [
    {
      "id": "3250702",
      "postDate": "07/18/2025 23:59:30",
      "content": "<p>What is your approach? Will they delete hacked solutions?</p>",
      "rawMarkdown": "What is your approach? Will they delete hacked solutions?",
      "votes": null
    },
    {
      "id": "3250949",
      "postDate": "07/19/2025 14:58:39",
      "content": "<p>Hey, I am not unshuffling the testing set so I guess no temporal features/assumptions can be used in the modelling, just trying to rebuild the best distribution of test y given the test features at any point. So no lagged features, nothing rolling etc, no arima, no LSTM etc. Thats a choice, just a matter of interpretation of the rule. Adding temporality in the training for feature selection/importance/stability/interaction might be a good idea though. can't wait to see where this challenge is going, although my scoring will certainly be miles away from top scorer who peak into the test</p>",
      "rawMarkdown": "Hey, I am not unshuffling the testing set so I guess no temporal features/assumptions can be used in the modelling, just trying to rebuild the best distribution of test y given the test features at any point. So no lagged features, nothing rolling etc, no arima, no LSTM etc. Thats a choice, just a matter of interpretation of the rule. Adding temporality in the training for feature selection/importance/stability/interaction might be a good idea though. can't wait to see where this challenge is going, although my scoring will certainly be miles away from top scorer who peak into the test",
      "votes": null
    },
    {
      "id": "3250953",
      "postDate": "07/19/2025 15:05:48",
      "content": "<p>And not saying Im not using temporal features, LSTM  and other sequential concepts, simply I am doing that to extract meaningful features etc from the training, Im not using it to directly predict on the test. </p>",
      "rawMarkdown": "And not saying Im not using temporal features, LSTM  and other sequential concepts, simply I am doing that to extract meaningful features etc from the training, Im not using it to directly predict on the test.",
      "votes": null
    },
    {
      "id": "3250973",
      "postDate": "07/19/2025 15:51:55",
      "content": "<p>Who knows. Host has been silent since the timestamps were re-hacked. I don't see that there's point in doing either. There is zero reason to do time series other than a personal challenge. But there's probably zero reason to stick to \"fair\" approaches too since so much has been deduced about the test set by those who bothered to reverse engineer it that it's no doubt easy enough to boost your score a little into the 0.15-0.2 range by using light, undetectable cheating (e.g. choosing features most stable into the test set). Your choices are; waste time doing time series, waste time doing it fairly or cheat a little to even the playing field (not endorsing that ofc).</p>",
      "rawMarkdown": "Who knows. Host has been silent since the timestamps were re-hacked. I don't see that there's point in doing either. There is zero reason to do time series other than a personal challenge. But there's probably zero reason to stick to \"fair\" approaches too since so much has been deduced about the test set by those who bothered to reverse engineer it that it's no doubt easy enough to boost your score a little into the 0.15-0.2 range by using light, undetectable cheating (e.g. choosing features most stable into the test set). Your choices are; waste time doing time series, waste time doing it fairly or cheat a little to even the playing field (not endorsing that ofc).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3250949,
      "author_name": "eliasroubache",
      "author_url": "",
      "post_date": "07/19/2025 14:58:39",
      "content": "<p>Hey, I am not unshuffling the testing set so I guess no temporal features/assumptions can be used in the modelling, just trying to rebuild the best distribution of test y given the test features at any point. So no lagged features, nothing rolling etc, no arima, no LSTM etc. Thats a choice, just a matter of interpretation of the rule. Adding temporality in the training for feature selection/importance/stability/interaction might be a good idea though. can't wait to see where this challenge is going, although my scoring will certainly be miles away from top scorer who peak into the test</p>",
      "votes": null,
      "replies": [
        {
          "id": 3250953,
          "author_name": "eliasroubache",
          "author_url": "",
          "post_date": "07/19/2025 15:05:48",
          "content": "<p>And not saying Im not using temporal features, LSTM  and other sequential concepts, simply I am doing that to extract meaningful features etc from the training, Im not using it to directly predict on the test. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3250973,
      "author_name": "greatestcutie",
      "author_url": "",
      "post_date": "07/19/2025 15:51:55",
      "content": "<p>Who knows. Host has been silent since the timestamps were re-hacked. I don't see that there's point in doing either. There is zero reason to do time series other than a personal challenge. But there's probably zero reason to stick to \"fair\" approaches too since so much has been deduced about the test set by those who bothered to reverse engineer it that it's no doubt easy enough to boost your score a little into the 0.15-0.2 range by using light, undetectable cheating (e.g. choosing features most stable into the test set). Your choices are; waste time doing time series, waste time doing it fairly or cheat a little to even the playing field (not endorsing that ofc).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3250702": "What is your approach? Will they delete hacked solutions?",
    "3250949": "Hey, I am not unshuffling the testing set so I guess no temporal features/assumptions can be used in the modelling, just trying to rebuild the best distribution of test y given the test features at any point. So no lagged features, nothing rolling etc, no arima, no LSTM etc. Thats a choice, just a matter of interpretation of the rule. Adding temporality in the training for feature selection/importance/stability/interaction might be a good idea though. can't wait to see where this challenge is going, although my scoring will certainly be miles away from top scorer who peak into the test",
    "3250953": "And not saying Im not using temporal features, LSTM  and other sequential concepts, simply I am doing that to extract meaningful features etc from the training, Im not using it to directly predict on the test.",
    "3250973": "Who knows. Host has been silent since the timestamps were re-hacked. I don't see that there's point in doing either. There is zero reason to do time series other than a personal challenge. But there's probably zero reason to stick to \"fair\" approaches too since so much has been deduced about the test set by those who bothered to reverse engineer it that it's no doubt easy enough to boost your score a little into the 0.15-0.2 range by using light, undetectable cheating (e.g. choosing features most stable into the test set). Your choices are; waste time doing time series, waste time doing it fairly or cheat a little to even the playing field (not endorsing that ofc)."
  },
  "source": "meta"
}