{
  "id": 589724,
  "title": "[Non Peeking] Stuck on sub 0.05 scores",
  "url": "/competitions/drw-crypto-market-prediction/discussion/589724",
  "author_name": "",
  "post_date": "2025-07-15T04:24:21.800960900Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm stuck on sub 0.05 scores and I feel like I've tried everything. I initially had a simple LGBM which scored near 0.07, but then after the reboot it went down to 0.05. Now, I've tried all sorts of techniques like using CatBoost, feature engineering based on importance, feature engineering based on stabillity, combining my \"working\" model with a ridge regressor, a LGBM feature selection + Neural Network approach, and more. I'm out of submissions for today. I'm thinking of trying a very simple ridge regressor tomorrow, since I think I'm getting too complex. </p>\n<p>Additionally, does anyone have a good solid way to evaluate your model without submitting? It sucks having to use up a submission on a model that does horribly. </p>\n<p>This is my first competition and I was hoping someone could give me ideas on what to try. I don't think theres a point in doing the test set reverse engineering since its against the rules and will get pruned anyways. If anyone has any suggestions for what I can try, please let me know.</p>",
  "messages": [
    {
      "id": "3248711",
      "postDate": "07/15/2025 04:24:21",
      "content": "<p>I'm stuck on sub 0.05 scores and I feel like I've tried everything. I initially had a simple LGBM which scored near 0.07, but then after the reboot it went down to 0.05. Now, I've tried all sorts of techniques like using CatBoost, feature engineering based on importance, feature engineering based on stabillity, combining my \"working\" model with a ridge regressor, a LGBM feature selection + Neural Network approach, and more. I'm out of submissions for today. I'm thinking of trying a very simple ridge regressor tomorrow, since I think I'm getting too complex. </p>\n<p>Additionally, does anyone have a good solid way to evaluate your model without submitting? It sucks having to use up a submission on a model that does horribly. </p>\n<p>This is my first competition and I was hoping someone could give me ideas on what to try. I don't think theres a point in doing the test set reverse engineering since its against the rules and will get pruned anyways. If anyone has any suggestions for what I can try, please let me know.</p>",
      "rawMarkdown": "I'm stuck on sub 0.05 scores and I feel like I've tried everything. I initially had a simple LGBM which scored near 0.07, but then after the reboot it went down to 0.05. Now, I've tried all sorts of techniques like using CatBoost, feature engineering based on importance, feature engineering based on stabillity, combining my \"working\" model with a ridge regressor, a LGBM feature selection + Neural Network approach, and more. I'm out of submissions for today. I'm thinking of trying a very simple ridge regressor tomorrow, since I think I'm getting too complex. \n\nAdditionally, does anyone have a good solid way to evaluate your model without submitting? It sucks having to use up a submission on a model that does horribly. \n\nThis is my first competition and I was hoping someone could give me ideas on what to try. I don't think theres a point in doing the test set reverse engineering since its against the rules and will get pruned anyways. If anyone has any suggestions for what I can try, please let me know.",
      "votes": null
    },
    {
      "id": "3248776",
      "postDate": "07/15/2025 07:04:29",
      "content": "<p>Please reduce your factors to a range of 40. You can choose the factors you think are important. I believe this will enhance your scores.</p>",
      "rawMarkdown": "Please reduce your factors to a range of 40. You can choose the factors you think are important. I believe this will enhance your scores.",
      "votes": null
    },
    {
      "id": "3248947",
      "postDate": "07/15/2025 12:49:47",
      "content": "<p>The following notebook may help: <a href=\"https://www.kaggle.com/code/taylorsamarel/reboot-fire-and-ice\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/reboot-fire-and-ice</a></p>\n<p>As <a href=\"https://www.kaggle.com/visterbai\" target=\"_blank\">@visterbai</a> mentioned, try not to add too many features too fast. I have found that in many cases you can create feature interactions from 2-4 features that are more stable that the individual features themselves.</p>",
      "rawMarkdown": "The following notebook may help: https://www.kaggle.com/code/taylorsamarel/reboot-fire-and-ice\n\nAs @visterbai mentioned, try not to add too many features too fast. I have found that in many cases you can create feature interactions from 2-4 features that are more stable that the individual features themselves.",
      "votes": null
    },
    {
      "id": "3249039",
      "postDate": "07/15/2025 14:55:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/visterbai\" target=\"_blank\">@visterbai</a> , </p>\n<p>Have you found a way to justify whether the features set does well because it did well in certain time period or it is simply the best set for BTC? I am curious if there is way to find the optimal feature set for a single symbol (BTC)</p>",
      "rawMarkdown": "Hi @visterbai , \n\nHave you found a way to justify whether the features set does well because it did well in certain time period or it is simply the best set for BTC? I am curious if there is way to find the optimal feature set for a single symbol (BTC)",
      "votes": null
    },
    {
      "id": "3249402",
      "postDate": "07/16/2025 12:10:26",
      "content": "<p>With a linear regression, you can score 0.075.</p>",
      "rawMarkdown": "With a linear regression, you can score 0.075.",
      "votes": null
    },
    {
      "id": "3249675",
      "postDate": "07/16/2025 22:27:30",
      "content": "<p>You can submit just the values of the feature called X752 (the most correlated feature with the target). You will get ~0.07</p>",
      "rawMarkdown": "You can submit just the values of the feature called X752 (the most correlated feature with the target). You will get ~0.07",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3248776,
      "author_name": "visterbai",
      "author_url": "",
      "post_date": "07/15/2025 07:04:29",
      "content": "<p>Please reduce your factors to a range of 40. You can choose the factors you think are important. I believe this will enhance your scores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3249039,
          "author_name": "alexzhongs",
          "author_url": "",
          "post_date": "07/15/2025 14:55:17",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/visterbai\" target=\"_blank\">@visterbai</a> , </p>\n<p>Have you found a way to justify whether the features set does well because it did well in certain time period or it is simply the best set for BTC? I am curious if there is way to find the optimal feature set for a single symbol (BTC)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3248947,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/15/2025 12:49:47",
      "content": "<p>The following notebook may help: <a href=\"https://www.kaggle.com/code/taylorsamarel/reboot-fire-and-ice\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/reboot-fire-and-ice</a></p>\n<p>As <a href=\"https://www.kaggle.com/visterbai\" target=\"_blank\">@visterbai</a> mentioned, try not to add too many features too fast. I have found that in many cases you can create feature interactions from 2-4 features that are more stable that the individual features themselves.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3249402,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "07/16/2025 12:10:26",
      "content": "<p>With a linear regression, you can score 0.075.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3249675,
      "author_name": "gromml",
      "author_url": "",
      "post_date": "07/16/2025 22:27:30",
      "content": "<p>You can submit just the values of the feature called X752 (the most correlated feature with the target). You will get ~0.07</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3248711": "I'm stuck on sub 0.05 scores and I feel like I've tried everything. I initially had a simple LGBM which scored near 0.07, but then after the reboot it went down to 0.05. Now, I've tried all sorts of techniques like using CatBoost, feature engineering based on importance, feature engineering based on stabillity, combining my \"working\" model with a ridge regressor, a LGBM feature selection + Neural Network approach, and more. I'm out of submissions for today. I'm thinking of trying a very simple ridge regressor tomorrow, since I think I'm getting too complex. \n\nAdditionally, does anyone have a good solid way to evaluate your model without submitting? It sucks having to use up a submission on a model that does horribly. \n\nThis is my first competition and I was hoping someone could give me ideas on what to try. I don't think theres a point in doing the test set reverse engineering since its against the rules and will get pruned anyways. If anyone has any suggestions for what I can try, please let me know.",
    "3248776": "Please reduce your factors to a range of 40. You can choose the factors you think are important. I believe this will enhance your scores.",
    "3248947": "The following notebook may help: https://www.kaggle.com/code/taylorsamarel/reboot-fire-and-ice\n\nAs @visterbai mentioned, try not to add too many features too fast. I have found that in many cases you can create feature interactions from 2-4 features that are more stable that the individual features themselves.",
    "3249039": "Hi @visterbai , \n\nHave you found a way to justify whether the features set does well because it did well in certain time period or it is simply the best set for BTC? I am curious if there is way to find the optimal feature set for a single symbol (BTC)",
    "3249402": "With a linear regression, you can score 0.075.",
    "3249675": "You can submit just the values of the feature called X752 (the most correlated feature with the target). You will get ~0.07"
  },
  "source": "meta"
}