{
  "id": 591074,
  "title": "Lost in space solution",
  "url": "/competitions/drw-crypto-market-prediction/writeups/rafa-paw-owski-lost-in-space-solution",
  "author_name": "",
  "post_date": "2025-07-25T06:41:42.222584Z",
  "votes": 14,
  "comment_count": 11,
  "views": 0,
  "content": "<p>My solution bases on computing power and heavy ensemble learning. I use autogluon with feature selection + training on dataset subsets. For example, the first solution uses an entire train_df, the second one is a mean of solutions for even and odd rows. In general, a nth solution is a mean of solutions for subsets train_df[::n], train_df[1::n], …, train_df[n-1::n]. I apply n equal to 100. The problem seems to have high variance, so ensemble learning is very efficient. This score booster can be applied to every solution. The solution can be easily extended by adding more diversity. We can modify sample weights, a number of folds, take only a recent or oldest part of data or select different subset of features. The computing power is a main obstacle. I have provided to host a simplified solution (1 of 100) due to time limitation. Parallel computing is recommended in this case or being very patient. <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> </p>",
  "messages": [
    {
      "id": "3253723",
      "postDate": "07/25/2025 06:41:42",
      "content": "<p>My solution bases on computing power and heavy ensemble learning. I use autogluon with feature selection + training on dataset subsets. For example, the first solution uses an entire train_df, the second one is a mean of solutions for even and odd rows. In general, a nth solution is a mean of solutions for subsets train_df[::n], train_df[1::n], …, train_df[n-1::n]. I apply n equal to 100. The problem seems to have high variance, so ensemble learning is very efficient. This score booster can be applied to every solution. The solution can be easily extended by adding more diversity. We can modify sample weights, a number of folds, take only a recent or oldest part of data or select different subset of features. The computing power is a main obstacle. I have provided to host a simplified solution (1 of 100) due to time limitation. Parallel computing is recommended in this case or being very patient. <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> </p>",
      "rawMarkdown": "My solution bases on computing power and heavy ensemble learning. I use autogluon with feature selection + training on dataset subsets. For example, the first solution uses an entire train_df, the second one is a mean of solutions for even and odd rows. In general, a nth solution is a mean of solutions for subsets train_df[::n], train_df[1::n], ..., train_df[n-1::n]. I apply n equal to 100. The problem seems to have high variance, so ensemble learning is very efficient. This score booster can be applied to every solution. The solution can be easily extended by adding more diversity. We can modify sample weights, a number of folds, take only a recent or oldest part of data or select different subset of features. The computing power is a main obstacle. I have provided to host a simplified solution (1 of 100) due to time limitation. Parallel computing is recommended in this case or being very patient. @drwtrading",
      "votes": null
    },
    {
      "id": "3253737",
      "postDate": "07/25/2025 07:05:05",
      "content": "<p>How did you submit your solution to DRW?</p>",
      "rawMarkdown": "How did you submit your solution to DRW?",
      "votes": null
    },
    {
      "id": "3253745",
      "postDate": "07/25/2025 07:15:01",
      "content": "<p>You can share notebook only with the host. See a share button near to save version one. </p>",
      "rawMarkdown": "You can share notebook only with the host. See a share button near to save version one.",
      "votes": null
    },
    {
      "id": "3253763",
      "postDate": "07/25/2025 08:05:37",
      "content": "<p>Has the host explicitly asked the participants to share their solutions with them as you described? I can't see anything in the discussion section regarding this, and I haven't received any emails about it either.</p>",
      "rawMarkdown": "Has the host explicitly asked the participants to share their solutions with them as you described? I can't see anything in the discussion section regarding this, and I haven't received any emails about it either.",
      "votes": null
    },
    {
      "id": "3253865",
      "postDate": "07/25/2025 12:41:38",
      "content": "<p>Thank you for sharing your innovative ensembling strategy. I learned a lot from this post. I have a couple of quick questions to better understand your feature handling:</p>\n<ul>\n<li><p>Did you create any new features from the original dataset (e.g., interaction terms, aggregations), or did you rely solely on the provided features?</p></li>\n<li><p>To clarify, was the feature selection you mentioned handled automatically within AutoGluon, or was it a separate, manual step you performed before training?</p></li>\n</ul>",
      "rawMarkdown": "Thank you for sharing your innovative ensembling strategy. I learned a lot from this post. I have a couple of quick questions to better understand your feature handling:\n\n- Did you create any new features from the original dataset (e.g., interaction terms, aggregations), or did you rely solely on the provided features?\n\n- To clarify, was the feature selection you mentioned handled automatically within AutoGluon, or was it a separate, manual step you performed before training?",
      "votes": null
    },
    {
      "id": "3253872",
      "postDate": "07/25/2025 12:50:37",
      "content": "<ol>\n<li>Yes, I have created extra features. </li>\n<li>First a manual, then autogluon selection. Autogluon was not able to find good features from an entire dataset. </li>\n</ol>",
      "rawMarkdown": "1. Yes, I have created extra features. \n2. First a manual, then autogluon selection. Autogluon was not able to find good features from an entire dataset.",
      "votes": null
    },
    {
      "id": "3253876",
      "postDate": "07/25/2025 12:53:41",
      "content": "<p>Thanks! I agree with you on #2</p>",
      "rawMarkdown": "Thanks! I agree with you on #2",
      "votes": null
    },
    {
      "id": "3253964",
      "postDate": "07/25/2025 15:12:54",
      "content": "<p>It's an really interested way to do it! Thanks for sharing!  Is it why you have a huge jump from 0.14 to 0.16 because of changing strategy? <br>\nFor me, we have already optimized the models and we used multiple models to achieve this goal. I have checked the range of my prediction. The prediction range is really diverse and this results in the jump of our scores.</p>\n<p>Also, one question. Which model did you choose to do this strategy? Or did you use multiple models?</p>",
      "rawMarkdown": "It's an really interested way to do it! Thanks for sharing!  Is it why you have a huge jump from 0.14 to 0.16 because of changing strategy? \nFor me, we have already optimized the models and we used multiple models to achieve this goal. I have checked the range of my prediction. The prediction range is really diverse and this results in the jump of our scores.\n\nAlso, one question. Which model did you choose to do this strategy? Or did you use multiple models?",
      "votes": null
    },
    {
      "id": "3254046",
      "postDate": "07/25/2025 18:09:24",
      "content": "<p>Thanks so much for sharing! I also ensembled over many sets of manually-selected and engineered features, but I'm glad to learn about Autogluon, and your heavy-ensembling strategy. Is there a particular reason (periodicity?) you chose to sample regular intervals for the different models in the ensemble, or did you experiment with a few strategies?</p>",
      "rawMarkdown": "Thanks so much for sharing! I also ensembled over many sets of manually-selected and engineered features, but I'm glad to learn about Autogluon, and your heavy-ensembling strategy. Is there a particular reason (periodicity?) you chose to sample regular intervals for the different models in the ensemble, or did you experiment with a few strategies?",
      "votes": null
    },
    {
      "id": "3254975",
      "postDate": "07/27/2025 16:31:30",
      "content": "<p>I have noticed that regular big intervals make neighboring rows less similar, and then this problem can be more treated as a regression task. I try now with irregular subsets. It means, I select only labels which are close to smoothed labels (with usage of an Exponential Moving Average or a Savitzky–Golay filter). </p>",
      "rawMarkdown": "I have noticed that regular big intervals make neighboring rows less similar, and then this problem can be more treated as a regression task. I try now with irregular subsets. It means, I select only labels which are close to smoothed labels (with usage of an Exponential Moving Average or a Savitzky–Golay filter).",
      "votes": null
    },
    {
      "id": "3258051",
      "postDate": "07/30/2025 02:13:36",
      "content": "<p>What an interesting solution. Did you customize any settings such as folds/hyperparameters or just use default settings for your models?</p>",
      "rawMarkdown": "What an interesting solution. Did you customize any settings such as folds/hyperparameters or just use default settings for your models?",
      "votes": null
    },
    {
      "id": "3258861",
      "postDate": "07/31/2025 11:02:29",
      "content": "<p>A simple linear regression could win it.</p>",
      "rawMarkdown": "A simple linear regression could win it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3253737,
      "author_name": "nnwithsf",
      "author_url": "",
      "post_date": "07/25/2025 07:05:05",
      "content": "<p>How did you submit your solution to DRW?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3253745,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "07/25/2025 07:15:01",
          "content": "<p>You can share notebook only with the host. See a share button near to save version one. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3253763,
              "author_name": "ravaghi",
              "author_url": "",
              "post_date": "07/25/2025 08:05:37",
              "content": "<p>Has the host explicitly asked the participants to share their solutions with them as you described? I can't see anything in the discussion section regarding this, and I haven't received any emails about it either.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3253865,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "07/25/2025 12:41:38",
      "content": "<p>Thank you for sharing your innovative ensembling strategy. I learned a lot from this post. I have a couple of quick questions to better understand your feature handling:</p>\n<ul>\n<li><p>Did you create any new features from the original dataset (e.g., interaction terms, aggregations), or did you rely solely on the provided features?</p></li>\n<li><p>To clarify, was the feature selection you mentioned handled automatically within AutoGluon, or was it a separate, manual step you performed before training?</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3253872,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "07/25/2025 12:50:37",
          "content": "<ol>\n<li>Yes, I have created extra features. </li>\n<li>First a manual, then autogluon selection. Autogluon was not able to find good features from an entire dataset. </li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 3253876,
              "author_name": "byunjins",
              "author_url": "",
              "post_date": "07/25/2025 12:53:41",
              "content": "<p>Thanks! I agree with you on #2</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3253964,
      "author_name": "dingyangwang",
      "author_url": "",
      "post_date": "07/25/2025 15:12:54",
      "content": "<p>It's an really interested way to do it! Thanks for sharing!  Is it why you have a huge jump from 0.14 to 0.16 because of changing strategy? <br>\nFor me, we have already optimized the models and we used multiple models to achieve this goal. I have checked the range of my prediction. The prediction range is really diverse and this results in the jump of our scores.</p>\n<p>Also, one question. Which model did you choose to do this strategy? Or did you use multiple models?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3254046,
      "author_name": "sarahjeffreson",
      "author_url": "",
      "post_date": "07/25/2025 18:09:24",
      "content": "<p>Thanks so much for sharing! I also ensembled over many sets of manually-selected and engineered features, but I'm glad to learn about Autogluon, and your heavy-ensembling strategy. Is there a particular reason (periodicity?) you chose to sample regular intervals for the different models in the ensemble, or did you experiment with a few strategies?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3254975,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "07/27/2025 16:31:30",
          "content": "<p>I have noticed that regular big intervals make neighboring rows less similar, and then this problem can be more treated as a regression task. I try now with irregular subsets. It means, I select only labels which are close to smoothed labels (with usage of an Exponential Moving Average or a Savitzky–Golay filter). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3258051,
      "author_name": "byunjuns",
      "author_url": "",
      "post_date": "07/30/2025 02:13:36",
      "content": "<p>What an interesting solution. Did you customize any settings such as folds/hyperparameters or just use default settings for your models?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3258861,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "07/31/2025 11:02:29",
      "content": "<p>A simple linear regression could win it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3253723": "My solution bases on computing power and heavy ensemble learning. I use autogluon with feature selection + training on dataset subsets. For example, the first solution uses an entire train_df, the second one is a mean of solutions for even and odd rows. In general, a nth solution is a mean of solutions for subsets train_df[::n], train_df[1::n], ..., train_df[n-1::n]. I apply n equal to 100. The problem seems to have high variance, so ensemble learning is very efficient. This score booster can be applied to every solution. The solution can be easily extended by adding more diversity. We can modify sample weights, a number of folds, take only a recent or oldest part of data or select different subset of features. The computing power is a main obstacle. I have provided to host a simplified solution (1 of 100) due to time limitation. Parallel computing is recommended in this case or being very patient. @drwtrading",
    "3253737": "How did you submit your solution to DRW?",
    "3253745": "You can share notebook only with the host. See a share button near to save version one.",
    "3253763": "Has the host explicitly asked the participants to share their solutions with them as you described? I can't see anything in the discussion section regarding this, and I haven't received any emails about it either.",
    "3253865": "Thank you for sharing your innovative ensembling strategy. I learned a lot from this post. I have a couple of quick questions to better understand your feature handling:\n\n- Did you create any new features from the original dataset (e.g., interaction terms, aggregations), or did you rely solely on the provided features?\n\n- To clarify, was the feature selection you mentioned handled automatically within AutoGluon, or was it a separate, manual step you performed before training?",
    "3253872": "1. Yes, I have created extra features. \n2. First a manual, then autogluon selection. Autogluon was not able to find good features from an entire dataset.",
    "3253876": "Thanks! I agree with you on #2",
    "3253964": "It's an really interested way to do it! Thanks for sharing!  Is it why you have a huge jump from 0.14 to 0.16 because of changing strategy? \nFor me, we have already optimized the models and we used multiple models to achieve this goal. I have checked the range of my prediction. The prediction range is really diverse and this results in the jump of our scores.\n\nAlso, one question. Which model did you choose to do this strategy? Or did you use multiple models?",
    "3254046": "Thanks so much for sharing! I also ensembled over many sets of manually-selected and engineered features, but I'm glad to learn about Autogluon, and your heavy-ensembling strategy. Is there a particular reason (periodicity?) you chose to sample regular intervals for the different models in the ensemble, or did you experiment with a few strategies?",
    "3254975": "I have noticed that regular big intervals make neighboring rows less similar, and then this problem can be more treated as a regression task. I try now with irregular subsets. It means, I select only labels which are close to smoothed labels (with usage of an Exponential Moving Average or a Savitzky–Golay filter).",
    "3258051": "What an interesting solution. Did you customize any settings such as folds/hyperparameters or just use default settings for your models?",
    "3258861": "A simple linear regression could win it."
  },
  "source": "meta"
}