{
  "id": 585986,
  "title": "Why More Features ≠ Better: Lessons from MLPs in Noisy Crypto Data",
  "url": "/competitions/drw-crypto-market-prediction/discussion/585986",
  "author_name": "",
  "post_date": "2025-06-24T12:30:19.106222700Z",
  "votes": 7,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hey everyone ,</p>\n<p>I’ve been working on the DRW Crypto Market Prediction competition and wanted to share some unexpected findings and hard-won lessons from trying to model short-term price movements using an MLPRegressor on the full training set.<br>\nThe Setup<br>\nDataset: ~525,000 rows with ~890 features<br>\nTarget: label (appears to represent short-term price movement)<br>\nMetric: Pearson correlation<br>\nModel: MLPRegressor (sklearn), incremental feature selection loop from top-k F-scores<br>\nGoal: Find the optimal number of features (in steps of 5) that yield the best Pearson correlation<br>\nKey Observations</p>\n<ol>\n<li>More features ≠ better performance<br>\nI naively assumed adding more features would boost performance, but RMSE and Pearson correlation actually worsened as I increased feature count in some cases. The model became less stable, and generalization dropped — even after F-test feature filtering.<br>\nTop 5 features: Pearson = 0.0296<br>\nTop 25 features: Pearson = 0.0006<br>\nTop 30 features: Pearson = 0.0565<br>\nTop 40 features: Pearson = 0.0146</li>\n<li>MLPRegressor struggled to converge<br>\nEven with max_iter=300, I received convergence warnings at nearly every step. As feature count increased, training time ballooned (some steps took over 20 minutes), but results didn’t improve proportionally.</li>\n<li>Label distribution is wide-tailed but centered<br>\nThe label column has:<br>\nMean ~0.03<br>\nStd ~1.01<br>\nRange: –24 to +20<br>\nThe label isn’t flat, which is great for modeling, but the outliers massively skew RMSE and make convergence harder for neural nets.</li>\n</ol>\n<p>Has anyone seen consistent gains using MLPs or Keras models?<br>\nHas feature engineering helped over brute force top-k methods?<br>\nAnyone tried building hybrid models using both proprietary and public volume-based features?<br>\nLet’s discuss below — I’d love to hear your experiments, failures, and wins 🙌</p>",
  "messages": [
    {
      "id": "3231474",
      "postDate": "06/24/2025 12:30:19",
      "content": "<p>Hey everyone ,</p>\n<p>I’ve been working on the DRW Crypto Market Prediction competition and wanted to share some unexpected findings and hard-won lessons from trying to model short-term price movements using an MLPRegressor on the full training set.<br>\nThe Setup<br>\nDataset: ~525,000 rows with ~890 features<br>\nTarget: label (appears to represent short-term price movement)<br>\nMetric: Pearson correlation<br>\nModel: MLPRegressor (sklearn), incremental feature selection loop from top-k F-scores<br>\nGoal: Find the optimal number of features (in steps of 5) that yield the best Pearson correlation<br>\nKey Observations</p>\n<ol>\n<li>More features ≠ better performance<br>\nI naively assumed adding more features would boost performance, but RMSE and Pearson correlation actually worsened as I increased feature count in some cases. The model became less stable, and generalization dropped — even after F-test feature filtering.<br>\nTop 5 features: Pearson = 0.0296<br>\nTop 25 features: Pearson = 0.0006<br>\nTop 30 features: Pearson = 0.0565<br>\nTop 40 features: Pearson = 0.0146</li>\n<li>MLPRegressor struggled to converge<br>\nEven with max_iter=300, I received convergence warnings at nearly every step. As feature count increased, training time ballooned (some steps took over 20 minutes), but results didn’t improve proportionally.</li>\n<li>Label distribution is wide-tailed but centered<br>\nThe label column has:<br>\nMean ~0.03<br>\nStd ~1.01<br>\nRange: –24 to +20<br>\nThe label isn’t flat, which is great for modeling, but the outliers massively skew RMSE and make convergence harder for neural nets.</li>\n</ol>\n<p>Has anyone seen consistent gains using MLPs or Keras models?<br>\nHas feature engineering helped over brute force top-k methods?<br>\nAnyone tried building hybrid models using both proprietary and public volume-based features?<br>\nLet’s discuss below — I’d love to hear your experiments, failures, and wins 🙌</p>",
      "rawMarkdown": "Hey everyone ,\n\nI’ve been working on the DRW Crypto Market Prediction competition and wanted to share some unexpected findings and hard-won lessons from trying to model short-term price movements using an MLPRegressor on the full training set.\nThe Setup\nDataset: ~525,000 rows with ~890 features\nTarget: label (appears to represent short-term price movement)\nMetric: Pearson correlation\nModel: MLPRegressor (sklearn), incremental feature selection loop from top-k F-scores\nGoal: Find the optimal number of features (in steps of 5) that yield the best Pearson correlation\nKey Observations\n1. More features ≠ better performance\nI naively assumed adding more features would boost performance, but RMSE and Pearson correlation actually worsened as I increased feature count in some cases. The model became less stable, and generalization dropped — even after F-test feature filtering.\nTop 5 features: Pearson = 0.0296\nTop 25 features: Pearson = 0.0006\nTop 30 features: Pearson = 0.0565\nTop 40 features: Pearson = 0.0146\n2. MLPRegressor struggled to converge\nEven with max_iter=300, I received convergence warnings at nearly every step. As feature count increased, training time ballooned (some steps took over 20 minutes), but results didn’t improve proportionally.\n3. Label distribution is wide-tailed but centered\nThe label column has:\nMean ~0.03\nStd ~1.01\nRange: –24 to +20\nThe label isn’t flat, which is great for modeling, but the outliers massively skew RMSE and make convergence harder for neural nets.\n\n\n\n\nHas anyone seen consistent gains using MLPs or Keras models?\nHas feature engineering helped over brute force top-k methods?\nAnyone tried building hybrid models using both proprietary and public volume-based features?\nLet’s discuss below — I’d love to hear your experiments, failures, and wins 🙌",
      "votes": null
    },
    {
      "id": "3231479",
      "postDate": "06/24/2025 12:39:10",
      "content": "<p>Top 5 features: Pearson = 0.0296<br>\nTop 25 features: Pearson = 0.0006<br>\nTop 30 features: Pearson = 0.0565<br>\nTop 40 features: Pearson = 0.0146</p>\n<p>Are these CV scores or LB scores? What CV are you using?</p>",
      "rawMarkdown": "Top 5 features: Pearson = 0.0296\nTop 25 features: Pearson = 0.0006\nTop 30 features: Pearson = 0.0565\nTop 40 features: Pearson = 0.0146\n\nAre these CV scores or LB scores? What CV are you using?",
      "votes": null
    },
    {
      "id": "3231489",
      "postDate": "06/24/2025 13:00:15",
      "content": "<p>Hi byunjins,<br>\nThanks for your question! Those Pearson scores are from the validation set (CV) during incremental feature selection — specifically a time-based split where the last 20% of the training data is used as validation to avoid future data leakage.<br>\nThey are not leaderboard (LB) scores — I haven’t yet submitted those predictions.<br>\nI’m using a simple train_test_split with shuffle=False to preserve the temporal order, mimicking a realistic scenario for the crypto price prediction task.<br>\nLet me know if you want me to share the exact CV code or results on the LB!</p>",
      "rawMarkdown": "Hi byunjins,\nThanks for your question! Those Pearson scores are from the validation set (CV) during incremental feature selection — specifically a time-based split where the last 20% of the training data is used as validation to avoid future data leakage.\nThey are not leaderboard (LB) scores — I haven’t yet submitted those predictions.\nI’m using a simple train_test_split with shuffle=False to preserve the temporal order, mimicking a realistic scenario for the crypto price prediction task.\nLet me know if you want me to share the exact CV code or results on the LB!",
      "votes": null
    },
    {
      "id": "3232082",
      "postDate": "06/25/2025 10:51:16",
      "content": "<p>How are you getting such low scores(assuming they reflect actual LB like scores) on CV on pearson? Am I missing something? Because when I try pearson I get high scores. So I thought that my test data is from different distribution. I am not even able to debug if there is something else causing a dataleak.</p>",
      "rawMarkdown": "How are you getting such low scores(assuming they reflect actual LB like scores) on CV on pearson? Am I missing something? Because when I try pearson I get high scores. So I thought that my test data is from different distribution. I am not even able to debug if there is something else causing a dataleak.",
      "votes": null
    },
    {
      "id": "3232326",
      "postDate": "06/25/2025 15:22:11",
      "content": "<p>I used a BP neural network model with 2 hidden layers, AdamW optimizer, and CosineAnnealingLR scheduler. I also used 5 - fold cross - validation. The highest LB score is 0.10459, but the model may be unstable.</p>",
      "rawMarkdown": "I used a BP neural network model with 2 hidden layers, AdamW optimizer, and CosineAnnealingLR scheduler. I also used 5 - fold cross - validation. The highest LB score is 0.10459, but the model may be unstable.",
      "votes": null
    },
    {
      "id": "3232470",
      "postDate": "06/25/2025 18:59:09",
      "content": "<p>Really insightful post — thanks for sharing your findings so transparently!</p>\n<p>Your experience with MLPRegressor resonates a lot. In noisy financial data like crypto, I’ve noticed that increasing feature count often adds more noise than signal, especially with models like MLPs that don’t inherently handle feature importance or sparsity well. Regularization can help a bit, but it’s often not enough.</p>\n<p>A few quick thoughts and follow-up questions:</p>\n<p>Have you tried dimensionality reduction methods like PCA or autoencoders before feeding into the MLP? Sometimes they help stabilize convergence.</p>\n<p>For the label distribution, did you experiment with scaling or clipping extreme values? That can reduce the influence of outliers during training.</p>\n<p>I’ve seen some gains using hybrid models — combining MLPs with tree-based models or feature selectors from LGBM/XGBoost. The idea is to let the trees handle the feature noise and then pass refined features to a neural net.</p>\n<p>Also curious if anyone has tried contrastive learning or any self-supervised approaches on crypto time series — might be a better way to learn useful representations in such a volatile domain.</p>",
      "rawMarkdown": "Really insightful post — thanks for sharing your findings so transparently!\n\nYour experience with MLPRegressor resonates a lot. In noisy financial data like crypto, I’ve noticed that increasing feature count often adds more noise than signal, especially with models like MLPs that don’t inherently handle feature importance or sparsity well. Regularization can help a bit, but it’s often not enough.\n\nA few quick thoughts and follow-up questions:\n\nHave you tried dimensionality reduction methods like PCA or autoencoders before feeding into the MLP? Sometimes they help stabilize convergence.\n\nFor the label distribution, did you experiment with scaling or clipping extreme values? That can reduce the influence of outliers during training.\n\nI’ve seen some gains using hybrid models — combining MLPs with tree-based models or feature selectors from LGBM/XGBoost. The idea is to let the trees handle the feature noise and then pass refined features to a neural net.\n\n\nAlso curious if anyone has tried contrastive learning or any self-supervised approaches on crypto time series — might be a better way to learn useful representations in such a volatile domain.",
      "votes": null
    },
    {
      "id": "3232765",
      "postDate": "06/26/2025 07:08:59",
      "content": "<p>Spot on. That \"more features = worse performance\" issue is a classic.</p>\n<p>I've found it often comes down to two things:</p>\n<p>Data Leakage: Some features might have look-ahead bias baked into them.<br>\nFactor Decay: Some features look great on the training set but their predictive power just dies off on new data.<br>\nI had the same experience—I removed a single feature, \"X612\", and my model's performance got a noticeable boost. It really is quality over quantity.</p>",
      "rawMarkdown": "Spot on. That \"more features = worse performance\" issue is a classic.\n\nI've found it often comes down to two things:\n\nData Leakage: Some features might have look-ahead bias baked into them.\nFactor Decay: Some features look great on the training set but their predictive power just dies off on new data.\nI had the same experience—I removed a single feature, \"X612\", and my model's performance got a noticeable boost. It really is quality over quantity.",
      "votes": null
    },
    {
      "id": "3233317",
      "postDate": "06/26/2025 17:05:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/visterbai\" target=\"_blank\">@visterbai</a> ,</p>\n<p>Thanks for your thought-provoking insight! I am curious how you detect data leakage in a systematic way? I did find some features seem to be filled nan with future info (if I remember correctly</p>",
      "rawMarkdown": "Hi @visterbai ,\n\nThanks for your thought-provoking insight! I am curious how you detect data leakage in a systematic way? I did find some features seem to be filled nan with future info (if I remember correctly",
      "votes": null
    },
    {
      "id": "3233398",
      "postDate": "06/26/2025 18:29:55",
      "content": "<p>I ran a few simple tests and extracted some features that seemed important based on SHAP values, then submitted them individually as submissions. I noticed that some features have a much higher IC (information coefficient) on the test set than on the training set, while others show the opposite. If a feature performs similarly on both the training and test sets with an IC above 0.02🌟, it likely has good generalization and tends to yield solid results.</p>",
      "rawMarkdown": "I ran a few simple tests and extracted some features that seemed important based on SHAP values, then submitted them individually as submissions. I noticed that some features have a much higher IC (information coefficient) on the test set than on the training set, while others show the opposite. If a feature performs similarly on both the training and test sets with an IC above 0.02🌟, it likely has good generalization and tends to yield solid results.",
      "votes": null
    },
    {
      "id": "3233499",
      "postDate": "06/26/2025 21:43:40",
      "content": "<p>Absolutely, totally agree.<br>\nI’ve run into the same issues — especially factor decay — where certain features shine during training but completely fall apart in live or test data.<br>\nAnd yeah, data leakage is sneaky. One seemingly harmless feature can wreck your entire validation if you're not careful.<br>\nRemoving \"X612\" and seeing a boost is a perfect example — sometimes less is more. I’ve started adopting a \"trust but verify\" approach for every feature now.</p>",
      "rawMarkdown": "Absolutely, totally agree.\nI’ve run into the same issues — especially factor decay — where certain features shine during training but completely fall apart in live or test data.\nAnd yeah, data leakage is sneaky. One seemingly harmless feature can wreck your entire validation if you're not careful.\n\nRemoving \"X612\" and seeing a boost is a perfect example — sometimes less is more. I’ve started adopting a \"trust but verify\" approach for every feature now.",
      "votes": null
    },
    {
      "id": "3235441",
      "postDate": "06/29/2025 08:52:13",
      "content": "<p>I also use a neural network, but my model is deeper, I use 4 to 8 layers of hidden layers. My best score is only around 0.03</p>",
      "rawMarkdown": "I also use a neural network, but my model is deeper, I use 4 to 8 layers of hidden layers. My best score is only around 0.03",
      "votes": null
    },
    {
      "id": "3237397",
      "postDate": "07/01/2025 04:34:17",
      "content": "<p>Try injecting noise and using high drop out levels.</p>",
      "rawMarkdown": "Try injecting noise and using high drop out levels.",
      "votes": null
    },
    {
      "id": "3243188",
      "postDate": "07/06/2025 21:09:21",
      "content": "<p>Very interesting post and questions, thank you for sharing!!<br>\nLooking forword to see other people's responses and experiences about this. <br>\nThank you all !!</p>",
      "rawMarkdown": "Very interesting post and questions, thank you for sharing!!\nLooking forword to see other people's responses and experiences about this. \nThank you all !!",
      "votes": null
    },
    {
      "id": "3243242",
      "postDate": "07/06/2025 23:03:07",
      "content": "<p>I've been able to get MLP models and GANDALF models to train with significant stability to around 105-150 features. However, adding more features is becoming difficult again!</p>",
      "rawMarkdown": "I've been able to get MLP models and GANDALF models to train with significant stability to around 105-150 features. However, adding more features is becoming difficult again!",
      "votes": null
    },
    {
      "id": "3243438",
      "postDate": "07/07/2025 06:28:12",
      "content": "<p>Have you tried ensemble of models by training them on few set of features on each in combination? If so could you share was it useful doing it or not?</p>",
      "rawMarkdown": "Have you tried ensemble of models by training them on few set of features on each in combination? If so could you share was it useful doing it or not?",
      "votes": null
    },
    {
      "id": "3243647",
      "postDate": "07/07/2025 11:43:29",
      "content": "<p>This is possible, and is something I am trying at the moment - with marginal success. However, the task of adding new features becomes slower and slower and I find myself only getting useful results using a greedy approach of manually trying different ensemble combinations and feature sets. Everytime I try to do an advanced algorithm to optimize everything at once I either get extreme overfitting, instability, or run out of memory/compute time.</p>",
      "rawMarkdown": "This is possible, and is something I am trying at the moment - with marginal success. However, the task of adding new features becomes slower and slower and I find myself only getting useful results using a greedy approach of manually trying different ensemble combinations and feature sets. Everytime I try to do an advanced algorithm to optimize everything at once I either get extreme overfitting, instability, or run out of memory/compute time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3231479,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "06/24/2025 12:39:10",
      "content": "<p>Top 5 features: Pearson = 0.0296<br>\nTop 25 features: Pearson = 0.0006<br>\nTop 30 features: Pearson = 0.0565<br>\nTop 40 features: Pearson = 0.0146</p>\n<p>Are these CV scores or LB scores? What CV are you using?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3231489,
          "author_name": "sanketpai",
          "author_url": "",
          "post_date": "06/24/2025 13:00:15",
          "content": "<p>Hi byunjins,<br>\nThanks for your question! Those Pearson scores are from the validation set (CV) during incremental feature selection — specifically a time-based split where the last 20% of the training data is used as validation to avoid future data leakage.<br>\nThey are not leaderboard (LB) scores — I haven’t yet submitted those predictions.<br>\nI’m using a simple train_test_split with shuffle=False to preserve the temporal order, mimicking a realistic scenario for the crypto price prediction task.<br>\nLet me know if you want me to share the exact CV code or results on the LB!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3232082,
              "author_name": "electroknight",
              "author_url": "",
              "post_date": "06/25/2025 10:51:16",
              "content": "<p>How are you getting such low scores(assuming they reflect actual LB like scores) on CV on pearson? Am I missing something? Because when I try pearson I get high scores. So I thought that my test data is from different distribution. I am not even able to debug if there is something else causing a dataleak.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3232326,
      "author_name": "jiaoyouzhang",
      "author_url": "",
      "post_date": "06/25/2025 15:22:11",
      "content": "<p>I used a BP neural network model with 2 hidden layers, AdamW optimizer, and CosineAnnealingLR scheduler. I also used 5 - fold cross - validation. The highest LB score is 0.10459, but the model may be unstable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3235441,
          "author_name": "jinghongliu",
          "author_url": "",
          "post_date": "06/29/2025 08:52:13",
          "content": "<p>I also use a neural network, but my model is deeper, I use 4 to 8 layers of hidden layers. My best score is only around 0.03</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3232470,
      "author_name": "mohamedmahmoud111",
      "author_url": "",
      "post_date": "06/25/2025 18:59:09",
      "content": "<p>Really insightful post — thanks for sharing your findings so transparently!</p>\n<p>Your experience with MLPRegressor resonates a lot. In noisy financial data like crypto, I’ve noticed that increasing feature count often adds more noise than signal, especially with models like MLPs that don’t inherently handle feature importance or sparsity well. Regularization can help a bit, but it’s often not enough.</p>\n<p>A few quick thoughts and follow-up questions:</p>\n<p>Have you tried dimensionality reduction methods like PCA or autoencoders before feeding into the MLP? Sometimes they help stabilize convergence.</p>\n<p>For the label distribution, did you experiment with scaling or clipping extreme values? That can reduce the influence of outliers during training.</p>\n<p>I’ve seen some gains using hybrid models — combining MLPs with tree-based models or feature selectors from LGBM/XGBoost. The idea is to let the trees handle the feature noise and then pass refined features to a neural net.</p>\n<p>Also curious if anyone has tried contrastive learning or any self-supervised approaches on crypto time series — might be a better way to learn useful representations in such a volatile domain.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3232765,
      "author_name": "visterbai",
      "author_url": "",
      "post_date": "06/26/2025 07:08:59",
      "content": "<p>Spot on. That \"more features = worse performance\" issue is a classic.</p>\n<p>I've found it often comes down to two things:</p>\n<p>Data Leakage: Some features might have look-ahead bias baked into them.<br>\nFactor Decay: Some features look great on the training set but their predictive power just dies off on new data.<br>\nI had the same experience—I removed a single feature, \"X612\", and my model's performance got a noticeable boost. It really is quality over quantity.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3233317,
          "author_name": "alexzhongs",
          "author_url": "",
          "post_date": "06/26/2025 17:05:50",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/visterbai\" target=\"_blank\">@visterbai</a> ,</p>\n<p>Thanks for your thought-provoking insight! I am curious how you detect data leakage in a systematic way? I did find some features seem to be filled nan with future info (if I remember correctly</p>",
          "votes": null,
          "replies": [
            {
              "id": 3233398,
              "author_name": "visterbai",
              "author_url": "",
              "post_date": "06/26/2025 18:29:55",
              "content": "<p>I ran a few simple tests and extracted some features that seemed important based on SHAP values, then submitted them individually as submissions. I noticed that some features have a much higher IC (information coefficient) on the test set than on the training set, while others show the opposite. If a feature performs similarly on both the training and test sets with an IC above 0.02🌟, it likely has good generalization and tends to yield solid results.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3233499,
          "author_name": "mohamedmahmoud111",
          "author_url": "",
          "post_date": "06/26/2025 21:43:40",
          "content": "<p>Absolutely, totally agree.<br>\nI’ve run into the same issues — especially factor decay — where certain features shine during training but completely fall apart in live or test data.<br>\nAnd yeah, data leakage is sneaky. One seemingly harmless feature can wreck your entire validation if you're not careful.<br>\nRemoving \"X612\" and seeing a boost is a perfect example — sometimes less is more. I’ve started adopting a \"trust but verify\" approach for every feature now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237397,
      "author_name": "carlosmendozaii",
      "author_url": "",
      "post_date": "07/01/2025 04:34:17",
      "content": "<p>Try injecting noise and using high drop out levels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3243188,
      "author_name": "abdelhakouanzougui2",
      "author_url": "",
      "post_date": "07/06/2025 21:09:21",
      "content": "<p>Very interesting post and questions, thank you for sharing!!<br>\nLooking forword to see other people's responses and experiences about this. <br>\nThank you all !!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3243242,
          "author_name": "taylorsamarel",
          "author_url": "",
          "post_date": "07/06/2025 23:03:07",
          "content": "<p>I've been able to get MLP models and GANDALF models to train with significant stability to around 105-150 features. However, adding more features is becoming difficult again!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3243438,
              "author_name": "electroknight",
              "author_url": "",
              "post_date": "07/07/2025 06:28:12",
              "content": "<p>Have you tried ensemble of models by training them on few set of features on each in combination? If so could you share was it useful doing it or not?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3243647,
                  "author_name": "taylorsamarel",
                  "author_url": "",
                  "post_date": "07/07/2025 11:43:29",
                  "content": "<p>This is possible, and is something I am trying at the moment - with marginal success. However, the task of adding new features becomes slower and slower and I find myself only getting useful results using a greedy approach of manually trying different ensemble combinations and feature sets. Everytime I try to do an advanced algorithm to optimize everything at once I either get extreme overfitting, instability, or run out of memory/compute time.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3231474": "Hey everyone ,\n\nI’ve been working on the DRW Crypto Market Prediction competition and wanted to share some unexpected findings and hard-won lessons from trying to model short-term price movements using an MLPRegressor on the full training set.\nThe Setup\nDataset: ~525,000 rows with ~890 features\nTarget: label (appears to represent short-term price movement)\nMetric: Pearson correlation\nModel: MLPRegressor (sklearn), incremental feature selection loop from top-k F-scores\nGoal: Find the optimal number of features (in steps of 5) that yield the best Pearson correlation\nKey Observations\n1. More features ≠ better performance\nI naively assumed adding more features would boost performance, but RMSE and Pearson correlation actually worsened as I increased feature count in some cases. The model became less stable, and generalization dropped — even after F-test feature filtering.\nTop 5 features: Pearson = 0.0296\nTop 25 features: Pearson = 0.0006\nTop 30 features: Pearson = 0.0565\nTop 40 features: Pearson = 0.0146\n2. MLPRegressor struggled to converge\nEven with max_iter=300, I received convergence warnings at nearly every step. As feature count increased, training time ballooned (some steps took over 20 minutes), but results didn’t improve proportionally.\n3. Label distribution is wide-tailed but centered\nThe label column has:\nMean ~0.03\nStd ~1.01\nRange: –24 to +20\nThe label isn’t flat, which is great for modeling, but the outliers massively skew RMSE and make convergence harder for neural nets.\n\n\n\n\nHas anyone seen consistent gains using MLPs or Keras models?\nHas feature engineering helped over brute force top-k methods?\nAnyone tried building hybrid models using both proprietary and public volume-based features?\nLet’s discuss below — I’d love to hear your experiments, failures, and wins 🙌",
    "3231479": "Top 5 features: Pearson = 0.0296\nTop 25 features: Pearson = 0.0006\nTop 30 features: Pearson = 0.0565\nTop 40 features: Pearson = 0.0146\n\nAre these CV scores or LB scores? What CV are you using?",
    "3231489": "Hi byunjins,\nThanks for your question! Those Pearson scores are from the validation set (CV) during incremental feature selection — specifically a time-based split where the last 20% of the training data is used as validation to avoid future data leakage.\nThey are not leaderboard (LB) scores — I haven’t yet submitted those predictions.\nI’m using a simple train_test_split with shuffle=False to preserve the temporal order, mimicking a realistic scenario for the crypto price prediction task.\nLet me know if you want me to share the exact CV code or results on the LB!",
    "3232082": "How are you getting such low scores(assuming they reflect actual LB like scores) on CV on pearson? Am I missing something? Because when I try pearson I get high scores. So I thought that my test data is from different distribution. I am not even able to debug if there is something else causing a dataleak.",
    "3232326": "I used a BP neural network model with 2 hidden layers, AdamW optimizer, and CosineAnnealingLR scheduler. I also used 5 - fold cross - validation. The highest LB score is 0.10459, but the model may be unstable.",
    "3232470": "Really insightful post — thanks for sharing your findings so transparently!\n\nYour experience with MLPRegressor resonates a lot. In noisy financial data like crypto, I’ve noticed that increasing feature count often adds more noise than signal, especially with models like MLPs that don’t inherently handle feature importance or sparsity well. Regularization can help a bit, but it’s often not enough.\n\nA few quick thoughts and follow-up questions:\n\nHave you tried dimensionality reduction methods like PCA or autoencoders before feeding into the MLP? Sometimes they help stabilize convergence.\n\nFor the label distribution, did you experiment with scaling or clipping extreme values? That can reduce the influence of outliers during training.\n\nI’ve seen some gains using hybrid models — combining MLPs with tree-based models or feature selectors from LGBM/XGBoost. The idea is to let the trees handle the feature noise and then pass refined features to a neural net.\n\n\nAlso curious if anyone has tried contrastive learning or any self-supervised approaches on crypto time series — might be a better way to learn useful representations in such a volatile domain.",
    "3232765": "Spot on. That \"more features = worse performance\" issue is a classic.\n\nI've found it often comes down to two things:\n\nData Leakage: Some features might have look-ahead bias baked into them.\nFactor Decay: Some features look great on the training set but their predictive power just dies off on new data.\nI had the same experience—I removed a single feature, \"X612\", and my model's performance got a noticeable boost. It really is quality over quantity.",
    "3233317": "Hi @visterbai ,\n\nThanks for your thought-provoking insight! I am curious how you detect data leakage in a systematic way? I did find some features seem to be filled nan with future info (if I remember correctly",
    "3233398": "I ran a few simple tests and extracted some features that seemed important based on SHAP values, then submitted them individually as submissions. I noticed that some features have a much higher IC (information coefficient) on the test set than on the training set, while others show the opposite. If a feature performs similarly on both the training and test sets with an IC above 0.02🌟, it likely has good generalization and tends to yield solid results.",
    "3233499": "Absolutely, totally agree.\nI’ve run into the same issues — especially factor decay — where certain features shine during training but completely fall apart in live or test data.\nAnd yeah, data leakage is sneaky. One seemingly harmless feature can wreck your entire validation if you're not careful.\n\nRemoving \"X612\" and seeing a boost is a perfect example — sometimes less is more. I’ve started adopting a \"trust but verify\" approach for every feature now.",
    "3235441": "I also use a neural network, but my model is deeper, I use 4 to 8 layers of hidden layers. My best score is only around 0.03",
    "3237397": "Try injecting noise and using high drop out levels.",
    "3243188": "Very interesting post and questions, thank you for sharing!!\nLooking forword to see other people's responses and experiences about this. \nThank you all !!",
    "3243242": "I've been able to get MLP models and GANDALF models to train with significant stability to around 105-150 features. However, adding more features is becoming difficult again!",
    "3243438": "Have you tried ensemble of models by training them on few set of features on each in combination? If so could you share was it useful doing it or not?",
    "3243647": "This is possible, and is something I am trying at the moment - with marginal success. However, the task of adding new features becomes slower and slower and I find myself only getting useful results using a greedy approach of manually trying different ensemble combinations and feature sets. Everytime I try to do an advanced algorithm to optimize everything at once I either get extreme overfitting, instability, or run out of memory/compute time."
  },
  "source": "meta"
}