{
  "id": 598865,
  "title": "Top 13 solution",
  "url": "/competitions/drw-crypto-market-prediction/writeups/top-13-solution",
  "author_name": "",
  "post_date": "2025-08-13T06:30:53.513Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First I would like to thank DRW for a great experience. About my solution, I have to admit that I got pretty lucky in order to have such a high rank in the private leaderboard. You can checkout my solution at my <a href=\"https://github.com/justpqa/drw-crypto-market-prediction\" target=\"_blank\">Repo</a> (you could give it a star if you found it useful 😉), but here are a few details on what I've done:</p>\n<ul>\n<li>Features: Starting from top 100 features that is more correlated with target, I used SHAP on GBDT models to choose top features. I experimented with both anonymized and market variables (including extra market variables created from original market variables), and I found that market variables work well before the reboot but worse after, so my model use mainly anonymous features</li>\n<li>CV: 4-fold \"walk-forward\" CV with 4 months train - 1 month gap - 4 months validation, which gives me a decent correlation &amp; CV score is somewhat close to LB score</li>\n<li>Model: XGBoost + MLP (2 layers) on chosen features, ensembled on 3 different seeds, and also ensembled on multiple time window (I found that training multiple models on different time windows of training data really help prediction to be more robust)<br>\nOverall, it was a great experience building and learning from everyone, and I don't think I could go this far without everyone sharing their understanding. Hope everyone also have a good time.</li>\n</ul>",
  "messages": [
    {
      "id": "3268569",
      "postDate": "08/13/2025 06:30:45",
      "content": "<p>First I would like to thank DRW for a great experience. About my solution, I have to admit that I got pretty lucky in order to have such a high rank in the private leaderboard. You can checkout my solution at my <a href=\"https://github.com/justpqa/drw-crypto-market-prediction\" target=\"_blank\">Repo</a> (you could give it a star if you found it useful 😉), but here are a few details on what I've done:</p>\n<ul>\n<li>Features: Starting from top 100 features that is more correlated with target, I used SHAP on GBDT models to choose top features. I experimented with both anonymized and market variables (including extra market variables created from original market variables), and I found that market variables work well before the reboot but worse after, so my model use mainly anonymous features</li>\n<li>CV: 4-fold \"walk-forward\" CV with 4 months train - 1 month gap - 4 months validation, which gives me a decent correlation &amp; CV score is somewhat close to LB score</li>\n<li>Model: XGBoost + MLP (2 layers) on chosen features, ensembled on 3 different seeds, and also ensembled on multiple time window (I found that training multiple models on different time windows of training data really help prediction to be more robust)<br>\nOverall, it was a great experience building and learning from everyone, and I don't think I could go this far without everyone sharing their understanding. Hope everyone also have a good time.</li>\n</ul>",
      "rawMarkdown": "First I would like to thank DRW for a great experience. About my solution, I have to admit that I got pretty lucky in order to have such a high rank in the private leaderboard. You can checkout my solution at my [Repo](https://github.com/justpqa/drw-crypto-market-prediction) (you could give it a star if you found it useful 😉), but here are a few details on what I've done:\n- Features: Starting from top 100 features that is more correlated with target, I used SHAP on GBDT models to choose top features. I experimented with both anonymized and market variables (including extra market variables created from original market variables), and I found that market variables work well before the reboot but worse after, so my model use mainly anonymous features\n- CV: 4-fold \"walk-forward\" CV with 4 months train - 1 month gap - 4 months validation, which gives me a decent correlation & CV score is somewhat close to LB score\n- Model: XGBoost + MLP (2 layers) on chosen features, ensembled on 3 different seeds, and also ensembled on multiple time window (I found that training multiple models on different time windows of training data really help prediction to be more robust)\nOverall, it was a great experience building and learning from everyone, and I don't think I could go this far without everyone sharing their understanding. Hope everyone also have a good time.",
      "votes": null
    },
    {
      "id": "3273990",
      "postDate": "08/23/2025 18:25:39",
      "content": "<p>Another great MLP solution!</p>",
      "rawMarkdown": "Another great MLP solution!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3273990,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "08/23/2025 18:25:39",
      "content": "<p>Another great MLP solution!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3268569": "First I would like to thank DRW for a great experience. About my solution, I have to admit that I got pretty lucky in order to have such a high rank in the private leaderboard. You can checkout my solution at my [Repo](https://github.com/justpqa/drw-crypto-market-prediction) (you could give it a star if you found it useful 😉), but here are a few details on what I've done:\n- Features: Starting from top 100 features that is more correlated with target, I used SHAP on GBDT models to choose top features. I experimented with both anonymized and market variables (including extra market variables created from original market variables), and I found that market variables work well before the reboot but worse after, so my model use mainly anonymous features\n- CV: 4-fold \"walk-forward\" CV with 4 months train - 1 month gap - 4 months validation, which gives me a decent correlation & CV score is somewhat close to LB score\n- Model: XGBoost + MLP (2 layers) on chosen features, ensembled on 3 different seeds, and also ensembled on multiple time window (I found that training multiple models on different time windows of training data really help prediction to be more robust)\nOverall, it was a great experience building and learning from everyone, and I don't think I could go this far without everyone sharing their understanding. Hope everyone also have a good time.",
    "3273990": "Another great MLP solution!"
  },
  "source": "meta"
}