{
  "id": 588823,
  "title": "Does anyone know why specific features like X137, X168, X174, etc. work so well?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588823",
  "author_name": "",
  "post_date": "2025-07-08T13:24:02.186053900Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I've noticed that some high-scoring public notebooks use a specific set of features like:</p>\n<p>[\"X137\", \"X168\", \"X174\", \"X178\", \"X292\", \"X302\", \"X333\", \"X344\", \"X345\", \"X385\", \"X415\", \"X421\", \"X532\", \"X586\", \"X598\", \"X603\", \"X612\", \"X674\", \"X817\", \"X852\", \"X855\", \"X856\", \"X860\", \"X862\", \"X863\", \"X888\"]</p>\n<p>However, I haven't found any clear explanation in the shared code or discussions about why these features were selected. I'm wondering how the authors chose them.</p>\n<p>So far, I've tried several approaches:</p>\n<p>Selecting features based on their Pearson correlation with the target</p>\n<p>Reducing dimensionality using PCA (e.g., keeping components explaining 90% of variance)</p>\n<p>Grouping and pruning highly correlated features (e.g., removing one from each pair with correlation &gt; 0.98, or averaging within clusters)</p>\n<p>Unfortunately, none of these approaches have led to significant improvement — my scores are still stuck below 0.05.</p>\n<p>If anyone has any insights or suggestions on how to effectively select or identify good features in this competition, I would really appreciate your advice.</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "3244724",
      "postDate": "07/08/2025 13:24:02",
      "content": "<p>Hi everyone,</p>\n<p>I've noticed that some high-scoring public notebooks use a specific set of features like:</p>\n<p>[\"X137\", \"X168\", \"X174\", \"X178\", \"X292\", \"X302\", \"X333\", \"X344\", \"X345\", \"X385\", \"X415\", \"X421\", \"X532\", \"X586\", \"X598\", \"X603\", \"X612\", \"X674\", \"X817\", \"X852\", \"X855\", \"X856\", \"X860\", \"X862\", \"X863\", \"X888\"]</p>\n<p>However, I haven't found any clear explanation in the shared code or discussions about why these features were selected. I'm wondering how the authors chose them.</p>\n<p>So far, I've tried several approaches:</p>\n<p>Selecting features based on their Pearson correlation with the target</p>\n<p>Reducing dimensionality using PCA (e.g., keeping components explaining 90% of variance)</p>\n<p>Grouping and pruning highly correlated features (e.g., removing one from each pair with correlation &gt; 0.98, or averaging within clusters)</p>\n<p>Unfortunately, none of these approaches have led to significant improvement — my scores are still stuck below 0.05.</p>\n<p>If anyone has any insights or suggestions on how to effectively select or identify good features in this competition, I would really appreciate your advice.</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi everyone,\n\nI've noticed that some high-scoring public notebooks use a specific set of features like:\n\n[\"X137\", \"X168\", \"X174\", \"X178\", \"X292\", \"X302\", \"X333\", \"X344\", \"X345\", \"X385\", \"X415\", \"X421\", \"X532\", \"X586\", \"X598\", \"X603\", \"X612\", \"X674\", \"X817\", \"X852\", \"X855\", \"X856\", \"X860\", \"X862\", \"X863\", \"X888\"]\n\nHowever, I haven't found any clear explanation in the shared code or discussions about why these features were selected. I'm wondering how the authors chose them.\n\nSo far, I've tried several approaches:\n\nSelecting features based on their Pearson correlation with the target\n\nReducing dimensionality using PCA (e.g., keeping components explaining 90% of variance)\n\nGrouping and pruning highly correlated features (e.g., removing one from each pair with correlation > 0.98, or averaging within clusters)\n\n\nUnfortunately, none of these approaches have led to significant improvement — my scores are still stuck below 0.05.\n\nIf anyone has any insights or suggestions on how to effectively select or identify good features in this competition, I would really appreciate your advice.\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "3245028",
      "postDate": "07/08/2025 19:05:22",
      "content": "<p>Have the same doubt here.</p>\n<p>I also found those features not significantly linearly correlated with target. That say, fitting linear structure may not help public/private score.</p>",
      "rawMarkdown": "Have the same doubt here.\n\nI also found those features not significantly linearly correlated with target. That say, fitting linear structure may not help public/private score.",
      "votes": null
    },
    {
      "id": "3245082",
      "postDate": "07/08/2025 22:07:20",
      "content": "<p>Thanks for the reply — good to know I’m not the only one thinking this way.<br>\nYeah, I also found that most of those features don’t show strong linear correlation with the target, so I agree that linear structures (like PCA or Pearson-based selection) may not be effective.</p>\n<p>Do you happen to know any good methods to find useful features beyond linear approaches?<br>\nI’ve tried things like SHAP values and feature grouping based on inter-feature correlation, but haven’t seen much improvement yet. Any ideas would be really appreciated!</p>",
      "rawMarkdown": "Thanks for the reply — good to know I’m not the only one thinking this way.\nYeah, I also found that most of those features don’t show strong linear correlation with the target, so I agree that linear structures (like PCA or Pearson-based selection) may not be effective.\n\nDo you happen to know any good methods to find useful features beyond linear approaches?\nI’ve tried things like SHAP values and feature grouping based on inter-feature correlation, but haven’t seen much improvement yet. Any ideas would be really appreciated!",
      "votes": null
    },
    {
      "id": "3245332",
      "postDate": "07/09/2025 07:59:03",
      "content": "<p>In the beginning, I also tried to look for useful features among the 800+ anonymous variables. However, no matter which model I used, whether I looked for linear or nonlinear correlations, the features I identified didn’t perform well. In real-world trading scenarios, these anonymized features might actually be “factors” derived from order book data—possibly generated by genetic programming or similar approaches. Unfortunately, the organizers didn’t provide any detailed explanation of these features, which makes feature engineering quite challenging.</p>",
      "rawMarkdown": "In the beginning, I also tried to look for useful features among the 800+ anonymous variables. However, no matter which model I used, whether I looked for linear or nonlinear correlations, the features I identified didn’t perform well. In real-world trading scenarios, these anonymized features might actually be “factors” derived from order book data—possibly generated by genetic programming or similar approaches. Unfortunately, the organizers didn’t provide any detailed explanation of these features, which makes feature engineering quite challenging.",
      "votes": null
    },
    {
      "id": "3245347",
      "postDate": "07/09/2025 08:16:56",
      "content": "<p>good question! </p>",
      "rawMarkdown": "good question!",
      "votes": null
    },
    {
      "id": "3245493",
      "postDate": "07/09/2025 12:45:40",
      "content": "<p>Thanks for sharing your thoughts — I totally agree, it's really tough to work with so many anonymous variables without any context or documentation.</p>\n<p>I also thought they might be some kind of derived factors (e.g., order book signals or engineered features), possibly from feature generators or even genetic programming, as you mentioned.</p>\n<p>That said, have you found any practical methods or heuristics that helped you identify promising features in your own experiments — even without knowing their meaning?</p>\n<p>For example, did you find SHAP values, time-based analysis, or clustering methods helpful?<br>\nI'm still trying to figure out an effective strategy, so any tips would be much appreciated.</p>",
      "rawMarkdown": "Thanks for sharing your thoughts — I totally agree, it's really tough to work with so many anonymous variables without any context or documentation.\n\nI also thought they might be some kind of derived factors (e.g., order book signals or engineered features), possibly from feature generators or even genetic programming, as you mentioned.\n\nThat said, have you found any practical methods or heuristics that helped you identify promising features in your own experiments — even without knowing their meaning?\n\nFor example, did you find SHAP values, time-based analysis, or clustering methods helpful?\nI'm still trying to figure out an effective strategy, so any tips would be much appreciated.",
      "votes": null
    },
    {
      "id": "3245537",
      "postDate": "07/09/2025 14:24:00",
      "content": "<p>It frankly feels more like an investment problem to me than a hard science problem -- the goal is to find a robust solution instead of greedily optimizing performance in the training set. </p>\n<p>Given that few people (in public discussion) find consistent local CV with leaderboard score and limited effective training set data points are provided, it is pretty hard to tell whether optimizing local performance is chasing some short-term phenomena or uncovering true latent structure</p>",
      "rawMarkdown": "It frankly feels more like an investment problem to me than a hard science problem -- the goal is to find a robust solution instead of greedily optimizing performance in the training set. \n\nGiven that few people (in public discussion) find consistent local CV with leaderboard score and limited effective training set data points are provided, it is pretty hard to tell whether optimizing local performance is chasing some short-term phenomena or uncovering true latent structure",
      "votes": null
    },
    {
      "id": "3245786",
      "postDate": "07/09/2025 21:06:15",
      "content": "<p>Yes, I think you're absolutely right — that’s exactly how it feels to me too.</p>\n<p>Honestly, I’ve been struggling with this for quite a while.<br>\nLately, I’ve been trying various quick fixes just to push up my public LB score, but I’ve reached a point where I’m not even sure what I’m really doing or why it works (or doesn’t).</p>\n<p>It's frustrating not knowing whether I'm chasing short-term noise or building something truly robust.<br>\nThanks for putting it into words — it really resonates with me.</p>",
      "rawMarkdown": "Yes, I think you're absolutely right — that’s exactly how it feels to me too.\n\nHonestly, I’ve been struggling with this for quite a while.\nLately, I’ve been trying various quick fixes just to push up my public LB score, but I’ve reached a point where I’m not even sure what I’m really doing or why it works (or doesn’t).\n\nIt's frustrating not knowing whether I'm chasing short-term noise or building something truly robust.\nThanks for putting it into words — it really resonates with me.",
      "votes": null
    },
    {
      "id": "3248634",
      "postDate": "07/15/2025 00:17:15",
      "content": "<p>These features have generally high stability in the train and test datasets.</p>",
      "rawMarkdown": "These features have generally high stability in the train and test datasets.",
      "votes": null
    },
    {
      "id": "3249162",
      "postDate": "07/15/2025 21:08:06",
      "content": "<p>That makes sense — so you mean those features maintain similar distributions and behavior across both train and test sets?<br>\nI guess I should start checking feature stability more carefully. Thanks for the insight!</p>",
      "rawMarkdown": "That makes sense — so you mean those features maintain similar distributions and behavior across both train and test sets?\nI guess I should start checking feature stability more carefully. Thanks for the insight!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3245028,
      "author_name": "alexzhongs",
      "author_url": "",
      "post_date": "07/08/2025 19:05:22",
      "content": "<p>Have the same doubt here.</p>\n<p>I also found those features not significantly linearly correlated with target. That say, fitting linear structure may not help public/private score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3245082,
          "author_name": "japmunk",
          "author_url": "",
          "post_date": "07/08/2025 22:07:20",
          "content": "<p>Thanks for the reply — good to know I’m not the only one thinking this way.<br>\nYeah, I also found that most of those features don’t show strong linear correlation with the target, so I agree that linear structures (like PCA or Pearson-based selection) may not be effective.</p>\n<p>Do you happen to know any good methods to find useful features beyond linear approaches?<br>\nI’ve tried things like SHAP values and feature grouping based on inter-feature correlation, but haven’t seen much improvement yet. Any ideas would be really appreciated!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3245537,
              "author_name": "alexzhongs",
              "author_url": "",
              "post_date": "07/09/2025 14:24:00",
              "content": "<p>It frankly feels more like an investment problem to me than a hard science problem -- the goal is to find a robust solution instead of greedily optimizing performance in the training set. </p>\n<p>Given that few people (in public discussion) find consistent local CV with leaderboard score and limited effective training set data points are provided, it is pretty hard to tell whether optimizing local performance is chasing some short-term phenomena or uncovering true latent structure</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3245786,
                  "author_name": "japmunk",
                  "author_url": "",
                  "post_date": "07/09/2025 21:06:15",
                  "content": "<p>Yes, I think you're absolutely right — that’s exactly how it feels to me too.</p>\n<p>Honestly, I’ve been struggling with this for quite a while.<br>\nLately, I’ve been trying various quick fixes just to push up my public LB score, but I’ve reached a point where I’m not even sure what I’m really doing or why it works (or doesn’t).</p>\n<p>It's frustrating not knowing whether I'm chasing short-term noise or building something truly robust.<br>\nThanks for putting it into words — it really resonates with me.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3245332,
      "author_name": "liujianingcfec",
      "author_url": "",
      "post_date": "07/09/2025 07:59:03",
      "content": "<p>In the beginning, I also tried to look for useful features among the 800+ anonymous variables. However, no matter which model I used, whether I looked for linear or nonlinear correlations, the features I identified didn’t perform well. In real-world trading scenarios, these anonymized features might actually be “factors” derived from order book data—possibly generated by genetic programming or similar approaches. Unfortunately, the organizers didn’t provide any detailed explanation of these features, which makes feature engineering quite challenging.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3245493,
          "author_name": "japmunk",
          "author_url": "",
          "post_date": "07/09/2025 12:45:40",
          "content": "<p>Thanks for sharing your thoughts — I totally agree, it's really tough to work with so many anonymous variables without any context or documentation.</p>\n<p>I also thought they might be some kind of derived factors (e.g., order book signals or engineered features), possibly from feature generators or even genetic programming, as you mentioned.</p>\n<p>That said, have you found any practical methods or heuristics that helped you identify promising features in your own experiments — even without knowing their meaning?</p>\n<p>For example, did you find SHAP values, time-based analysis, or clustering methods helpful?<br>\nI'm still trying to figure out an effective strategy, so any tips would be much appreciated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3245347,
      "author_name": "kingui",
      "author_url": "",
      "post_date": "07/09/2025 08:16:56",
      "content": "<p>good question! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3248634,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/15/2025 00:17:15",
      "content": "<p>These features have generally high stability in the train and test datasets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3249162,
          "author_name": "japmunk",
          "author_url": "",
          "post_date": "07/15/2025 21:08:06",
          "content": "<p>That makes sense — so you mean those features maintain similar distributions and behavior across both train and test sets?<br>\nI guess I should start checking feature stability more carefully. Thanks for the insight!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3244724": "Hi everyone,\n\nI've noticed that some high-scoring public notebooks use a specific set of features like:\n\n[\"X137\", \"X168\", \"X174\", \"X178\", \"X292\", \"X302\", \"X333\", \"X344\", \"X345\", \"X385\", \"X415\", \"X421\", \"X532\", \"X586\", \"X598\", \"X603\", \"X612\", \"X674\", \"X817\", \"X852\", \"X855\", \"X856\", \"X860\", \"X862\", \"X863\", \"X888\"]\n\nHowever, I haven't found any clear explanation in the shared code or discussions about why these features were selected. I'm wondering how the authors chose them.\n\nSo far, I've tried several approaches:\n\nSelecting features based on their Pearson correlation with the target\n\nReducing dimensionality using PCA (e.g., keeping components explaining 90% of variance)\n\nGrouping and pruning highly correlated features (e.g., removing one from each pair with correlation > 0.98, or averaging within clusters)\n\n\nUnfortunately, none of these approaches have led to significant improvement — my scores are still stuck below 0.05.\n\nIf anyone has any insights or suggestions on how to effectively select or identify good features in this competition, I would really appreciate your advice.\n\nThanks in advance!",
    "3245028": "Have the same doubt here.\n\nI also found those features not significantly linearly correlated with target. That say, fitting linear structure may not help public/private score.",
    "3245082": "Thanks for the reply — good to know I’m not the only one thinking this way.\nYeah, I also found that most of those features don’t show strong linear correlation with the target, so I agree that linear structures (like PCA or Pearson-based selection) may not be effective.\n\nDo you happen to know any good methods to find useful features beyond linear approaches?\nI’ve tried things like SHAP values and feature grouping based on inter-feature correlation, but haven’t seen much improvement yet. Any ideas would be really appreciated!",
    "3245332": "In the beginning, I also tried to look for useful features among the 800+ anonymous variables. However, no matter which model I used, whether I looked for linear or nonlinear correlations, the features I identified didn’t perform well. In real-world trading scenarios, these anonymized features might actually be “factors” derived from order book data—possibly generated by genetic programming or similar approaches. Unfortunately, the organizers didn’t provide any detailed explanation of these features, which makes feature engineering quite challenging.",
    "3245347": "good question!",
    "3245493": "Thanks for sharing your thoughts — I totally agree, it's really tough to work with so many anonymous variables without any context or documentation.\n\nI also thought they might be some kind of derived factors (e.g., order book signals or engineered features), possibly from feature generators or even genetic programming, as you mentioned.\n\nThat said, have you found any practical methods or heuristics that helped you identify promising features in your own experiments — even without knowing their meaning?\n\nFor example, did you find SHAP values, time-based analysis, or clustering methods helpful?\nI'm still trying to figure out an effective strategy, so any tips would be much appreciated.",
    "3245537": "It frankly feels more like an investment problem to me than a hard science problem -- the goal is to find a robust solution instead of greedily optimizing performance in the training set. \n\nGiven that few people (in public discussion) find consistent local CV with leaderboard score and limited effective training set data points are provided, it is pretty hard to tell whether optimizing local performance is chasing some short-term phenomena or uncovering true latent structure",
    "3245786": "Yes, I think you're absolutely right — that’s exactly how it feels to me too.\n\nHonestly, I’ve been struggling with this for quite a while.\nLately, I’ve been trying various quick fixes just to push up my public LB score, but I’ve reached a point where I’m not even sure what I’m really doing or why it works (or doesn’t).\n\nIt's frustrating not knowing whether I'm chasing short-term noise or building something truly robust.\nThanks for putting it into words — it really resonates with me.",
    "3248634": "These features have generally high stability in the train and test datasets.",
    "3249162": "That makes sense — so you mean those features maintain similar distributions and behavior across both train and test sets?\nI guess I should start checking feature stability more carefully. Thanks for the insight!"
  },
  "source": "meta"
}