{
  "id": 583526,
  "title": "DRW Crypto Market Prediction — Model Results, Ensembling Struggles, and Meta-Learner Help?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/583526",
  "author_name": "",
  "post_date": "2025-06-07T14:55:39.268853400Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,<br>\nI’m still a beginner and actively learning, so I’d really appreciate any guidance or help from the community. I've been actively participating in this competition and have hit a bit of a wall. After trying out various models, advanced feature engineering, and multiple ensembling strategies, I still can't push my public score beyond 0.0965. Below is a detailed breakdown of my experiments so far. I'd love your feedback, ideas, or pointers if you've faced similar bottlenecks!</p>\n<p>✅ Individual Model Scores (Public LB)<br>\nModel            Score<br>\nLightGBM    0.10489<br>\nCatBoost            0.06139<br>\nXGBoost            0.10001<br>\nHistGBM            0.08649<br>\nRidge            0.08280<br>\nMLP                    0.00996 (worst)</p>\n<p>-&gt; <strong>What I’ve Tried So Far</strong></p>\n<ol>\n<li>Advanced Feature Engineering<br>\nBuilt a pipeline generating over 900+ features, including:</li>\n</ol>\n<p>Realized volatility proxies<br>\nPressure × spread interactions<br>\nLiquidity imbalance<br>\nRolling stats, z-scores, etc.<br>\nFinal dataset size is ~13 GB+.</p>\n<ol>\n<li><p>Weighted Ensembling with Optuna<br>\nUsed holdout predictions from best-performing models.<br>\nTried Optuna to find best weighted combination (maximizing Pearson).<br>\nResult: Score decreased to ~0.10176 🤯<br>\n(Worse than my best individual model — LGBM)</p></li>\n<li><p>Holdout Strategy with Time Awareness<br>\nCreated time-aware holdout sets (15%)<br>\nMulti-seed training (3 seeds × 5 folds)<br>\nValidation scores were okay, but leaderboard didn’t improve.</p></li>\n<li><p>LightGBM Focused Tuning<br>\nTried multiple configs — including dart, goss, and high-capacity GBDT with GPU.<br>\nBest LightGBM params:<br>\n'num_leaves': 2048, 'learning_rate': 0.005, 'reg_alpha': 0.1, <br>\n'reg_lambda': 0.1, 'bagging_fraction': 0.8, 'feature_fraction': 0.8,<br>\n'n_estimators': 5000, 'early_stopping_rounds': 300, 'device': 'gpu'</p></li>\n</ol>\n<p>❌ Still Stuck — Main Issues:<br>\nWhy is ensembling degrading performance?<br>\nMy top models perform decently on their own — but Optuna-weighted predictions underperform even basic LightGBM!</p>\n<p>Does meta-modeling (stacking) make sense here?<br>\nI was thinking of training a meta-learner (e.g., Ridge/XGB/LGB) on out-of-fold preds from my base models. Has anyone had success with that in this competition?</p>\n<p>Pearson Optimization ≠ RMSE Optimization?<br>\nMost models are minimizing RMSE — but the leaderboard uses Pearson correlation. Could this mismatch be hurting performance?</p>\n<p>Any tips for final push beyond 0.0965?<br>\nShould I reduce feature space using SHAP?<br>\nTry different ensembling (GBRank, voting, stacking)?<br>\nAny post-processing tips that helped you?</p>\n<p>🧠 What I'm Looking For<br>\nIf you've:<br>\nManaged to boost your LB score beyond 0.10<br>\nGotten value from stacking or meta-modeling<br>\nFound Pearson-specific tricks that helped<br>\nUsed clever time-series validation</p>\n<p>I’d love to hear how you approached it!<br>\nI’m happy to share more code or notebook logic if needed — thanks in advance for helping me.</p>\n<p>Happy modeling!</p>",
  "messages": [
    {
      "id": "3219350",
      "postDate": "06/07/2025 14:55:39",
      "content": "<p>Hi everyone,<br>\nI’m still a beginner and actively learning, so I’d really appreciate any guidance or help from the community. I've been actively participating in this competition and have hit a bit of a wall. After trying out various models, advanced feature engineering, and multiple ensembling strategies, I still can't push my public score beyond 0.0965. Below is a detailed breakdown of my experiments so far. I'd love your feedback, ideas, or pointers if you've faced similar bottlenecks!</p>\n<p>✅ Individual Model Scores (Public LB)<br>\nModel            Score<br>\nLightGBM    0.10489<br>\nCatBoost            0.06139<br>\nXGBoost            0.10001<br>\nHistGBM            0.08649<br>\nRidge            0.08280<br>\nMLP                    0.00996 (worst)</p>\n<p>-&gt; <strong>What I’ve Tried So Far</strong></p>\n<ol>\n<li>Advanced Feature Engineering<br>\nBuilt a pipeline generating over 900+ features, including:</li>\n</ol>\n<p>Realized volatility proxies<br>\nPressure × spread interactions<br>\nLiquidity imbalance<br>\nRolling stats, z-scores, etc.<br>\nFinal dataset size is ~13 GB+.</p>\n<ol>\n<li><p>Weighted Ensembling with Optuna<br>\nUsed holdout predictions from best-performing models.<br>\nTried Optuna to find best weighted combination (maximizing Pearson).<br>\nResult: Score decreased to ~0.10176 🤯<br>\n(Worse than my best individual model — LGBM)</p></li>\n<li><p>Holdout Strategy with Time Awareness<br>\nCreated time-aware holdout sets (15%)<br>\nMulti-seed training (3 seeds × 5 folds)<br>\nValidation scores were okay, but leaderboard didn’t improve.</p></li>\n<li><p>LightGBM Focused Tuning<br>\nTried multiple configs — including dart, goss, and high-capacity GBDT with GPU.<br>\nBest LightGBM params:<br>\n'num_leaves': 2048, 'learning_rate': 0.005, 'reg_alpha': 0.1, <br>\n'reg_lambda': 0.1, 'bagging_fraction': 0.8, 'feature_fraction': 0.8,<br>\n'n_estimators': 5000, 'early_stopping_rounds': 300, 'device': 'gpu'</p></li>\n</ol>\n<p>❌ Still Stuck — Main Issues:<br>\nWhy is ensembling degrading performance?<br>\nMy top models perform decently on their own — but Optuna-weighted predictions underperform even basic LightGBM!</p>\n<p>Does meta-modeling (stacking) make sense here?<br>\nI was thinking of training a meta-learner (e.g., Ridge/XGB/LGB) on out-of-fold preds from my base models. Has anyone had success with that in this competition?</p>\n<p>Pearson Optimization ≠ RMSE Optimization?<br>\nMost models are minimizing RMSE — but the leaderboard uses Pearson correlation. Could this mismatch be hurting performance?</p>\n<p>Any tips for final push beyond 0.0965?<br>\nShould I reduce feature space using SHAP?<br>\nTry different ensembling (GBRank, voting, stacking)?<br>\nAny post-processing tips that helped you?</p>\n<p>🧠 What I'm Looking For<br>\nIf you've:<br>\nManaged to boost your LB score beyond 0.10<br>\nGotten value from stacking or meta-modeling<br>\nFound Pearson-specific tricks that helped<br>\nUsed clever time-series validation</p>\n<p>I’d love to hear how you approached it!<br>\nI’m happy to share more code or notebook logic if needed — thanks in advance for helping me.</p>\n<p>Happy modeling!</p>",
      "rawMarkdown": "Hi everyone,\nI’m still a beginner and actively learning, so I’d really appreciate any guidance or help from the community. I've been actively participating in this competition and have hit a bit of a wall. After trying out various models, advanced feature engineering, and multiple ensembling strategies, I still can't push my public score beyond 0.0965. Below is a detailed breakdown of my experiments so far. I'd love your feedback, ideas, or pointers if you've faced similar bottlenecks!\n\n✅ Individual Model Scores (Public LB)\nModel\t        Score\nLightGBM\t0.10489\nCatBoost\t        0.06139\nXGBoost\t        0.10001\nHistGBM\t        0.08649\nRidge\t        0.08280\nMLP\t                0.00996 (worst)\n\n-> **What I’ve Tried So Far**\n1. Advanced Feature Engineering\nBuilt a pipeline generating over 900+ features, including:\n\nRealized volatility proxies\nPressure × spread interactions\nLiquidity imbalance\nRolling stats, z-scores, etc.\nFinal dataset size is ~13 GB+.\n\n2. Weighted Ensembling with Optuna\nUsed holdout predictions from best-performing models.\nTried Optuna to find best weighted combination (maximizing Pearson).\nResult: Score decreased to ~0.10176 🤯\n(Worse than my best individual model — LGBM)\n\n3. Holdout Strategy with Time Awareness\nCreated time-aware holdout sets (15%)\nMulti-seed training (3 seeds × 5 folds)\nValidation scores were okay, but leaderboard didn’t improve.\n\n4. LightGBM Focused Tuning\nTried multiple configs — including dart, goss, and high-capacity GBDT with GPU.\nBest LightGBM params:\n'num_leaves': 2048, 'learning_rate': 0.005, 'reg_alpha': 0.1, \n'reg_lambda': 0.1, 'bagging_fraction': 0.8, 'feature_fraction': 0.8,\n'n_estimators': 5000, 'early_stopping_rounds': 300, 'device': 'gpu'\n\n\n\n❌ Still Stuck — Main Issues:\nWhy is ensembling degrading performance?\nMy top models perform decently on their own — but Optuna-weighted predictions underperform even basic LightGBM!\n\nDoes meta-modeling (stacking) make sense here?\nI was thinking of training a meta-learner (e.g., Ridge/XGB/LGB) on out-of-fold preds from my base models. Has anyone had success with that in this competition?\n\nPearson Optimization ≠ RMSE Optimization?\nMost models are minimizing RMSE — but the leaderboard uses Pearson correlation. Could this mismatch be hurting performance?\n\nAny tips for final push beyond 0.0965?\nShould I reduce feature space using SHAP?\nTry different ensembling (GBRank, voting, stacking)?\nAny post-processing tips that helped you?\n\n🧠 What I'm Looking For\nIf you've:\nManaged to boost your LB score beyond 0.10\nGotten value from stacking or meta-modeling\nFound Pearson-specific tricks that helped\nUsed clever time-series validation\n\nI’d love to hear how you approached it!\nI’m happy to share more code or notebook logic if needed — thanks in advance for helping me.\n\nHappy modeling!",
      "votes": null
    },
    {
      "id": "3219498",
      "postDate": "06/07/2025 20:09:31",
      "content": "<p>I guess why ensemble method here degrades final performance could be: there are too many sub-optimal models in the list: Since the data here is quite noisy, the performance metrics will be noisy as well, as a result the optimized ensemble weight may fit a lot of noise</p>",
      "rawMarkdown": "I guess why ensemble method here degrades final performance could be: there are too many sub-optimal models in the list: Since the data here is quite noisy, the performance metrics will be noisy as well, as a result the optimized ensemble weight may fit a lot of noise",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3219498,
      "author_name": "alexzhongs",
      "author_url": "",
      "post_date": "06/07/2025 20:09:31",
      "content": "<p>I guess why ensemble method here degrades final performance could be: there are too many sub-optimal models in the list: Since the data here is quite noisy, the performance metrics will be noisy as well, as a result the optimized ensemble weight may fit a lot of noise</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3219350": "Hi everyone,\nI’m still a beginner and actively learning, so I’d really appreciate any guidance or help from the community. I've been actively participating in this competition and have hit a bit of a wall. After trying out various models, advanced feature engineering, and multiple ensembling strategies, I still can't push my public score beyond 0.0965. Below is a detailed breakdown of my experiments so far. I'd love your feedback, ideas, or pointers if you've faced similar bottlenecks!\n\n✅ Individual Model Scores (Public LB)\nModel\t        Score\nLightGBM\t0.10489\nCatBoost\t        0.06139\nXGBoost\t        0.10001\nHistGBM\t        0.08649\nRidge\t        0.08280\nMLP\t                0.00996 (worst)\n\n-> **What I’ve Tried So Far**\n1. Advanced Feature Engineering\nBuilt a pipeline generating over 900+ features, including:\n\nRealized volatility proxies\nPressure × spread interactions\nLiquidity imbalance\nRolling stats, z-scores, etc.\nFinal dataset size is ~13 GB+.\n\n2. Weighted Ensembling with Optuna\nUsed holdout predictions from best-performing models.\nTried Optuna to find best weighted combination (maximizing Pearson).\nResult: Score decreased to ~0.10176 🤯\n(Worse than my best individual model — LGBM)\n\n3. Holdout Strategy with Time Awareness\nCreated time-aware holdout sets (15%)\nMulti-seed training (3 seeds × 5 folds)\nValidation scores were okay, but leaderboard didn’t improve.\n\n4. LightGBM Focused Tuning\nTried multiple configs — including dart, goss, and high-capacity GBDT with GPU.\nBest LightGBM params:\n'num_leaves': 2048, 'learning_rate': 0.005, 'reg_alpha': 0.1, \n'reg_lambda': 0.1, 'bagging_fraction': 0.8, 'feature_fraction': 0.8,\n'n_estimators': 5000, 'early_stopping_rounds': 300, 'device': 'gpu'\n\n\n\n❌ Still Stuck — Main Issues:\nWhy is ensembling degrading performance?\nMy top models perform decently on their own — but Optuna-weighted predictions underperform even basic LightGBM!\n\nDoes meta-modeling (stacking) make sense here?\nI was thinking of training a meta-learner (e.g., Ridge/XGB/LGB) on out-of-fold preds from my base models. Has anyone had success with that in this competition?\n\nPearson Optimization ≠ RMSE Optimization?\nMost models are minimizing RMSE — but the leaderboard uses Pearson correlation. Could this mismatch be hurting performance?\n\nAny tips for final push beyond 0.0965?\nShould I reduce feature space using SHAP?\nTry different ensembling (GBRank, voting, stacking)?\nAny post-processing tips that helped you?\n\n🧠 What I'm Looking For\nIf you've:\nManaged to boost your LB score beyond 0.10\nGotten value from stacking or meta-modeling\nFound Pearson-specific tricks that helped\nUsed clever time-series validation\n\nI’d love to hear how you approached it!\nI’m happy to share more code or notebook logic if needed — thanks in advance for helping me.\n\nHappy modeling!",
    "3219498": "I guess why ensemble method here degrades final performance could be: there are too many sub-optimal models in the list: Since the data here is quite noisy, the performance metrics will be noisy as well, as a result the optimized ensemble weight may fit a lot of noise"
  },
  "source": "meta"
}