{
  "id": 589819,
  "title": "Methods to \"diversify\" features' effect within the model",
  "url": "/competitions/drw-crypto-market-prediction/discussion/589819",
  "author_name": "Alex Zhongs",
  "post_date": "2025-07-15T15:46:55.789000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Many practitioners prefer a lean and effective feature set, assuming selected features will generalize well to the test set. However, this assumption often breaks when certain features perform well in training but shift significantly in the test set—leading to inconsistent local CV vs. public leaderboard results (as frequently observed in discussions).</p>\n<p>To reduce over-reliance on a few dominant features and build more robust models—especially for real-world deployment—it's helpful to diversify the feature contribution. Below are a couple of techniques I found useful that have consistently provided marginal improvements in both local CV and public scores:</p>\n<p>For linear models: Group features by the magnitude of their coefficients (e.g., place features with |coef| ≈ 0.1 in one bucket, those with |coef| ≈ 0.01 in another). This mitigates the risk of the model blowing up if high-weight features shift, while still preserving signal from more stable, lower-weighted features.</p>\n<p>For tree-based models: Use techniques that encourage inclusion of less important in-sample (IS) features. This prevents the model from overfitting to a narrow set of dominant features, especially in cases where the training period is short (only 1 year).</p>",
  "messages": [
    {
      "id": 3249055,
      "postDate": "2025-07-15T15:46:55.790Z",
      "content": "<p>Many practitioners prefer a lean and effective feature set, assuming selected features will generalize well to the test set. However, this assumption often breaks when certain features perform well in training but shift significantly in the test set—leading to inconsistent local CV vs. public leaderboard results (as frequently observed in discussions).</p>\n<p>To reduce over-reliance on a few dominant features and build more robust models—especially for real-world deployment—it's helpful to diversify the feature contribution. Below are a couple of techniques I found useful that have consistently provided marginal improvements in both local CV and public scores:</p>\n<p>For linear models: Group features by the magnitude of their coefficients (e.g., place features with |coef| ≈ 0.1 in one bucket, those with |coef| ≈ 0.01 in another). This mitigates the risk of the model blowing up if high-weight features shift, while still preserving signal from more stable, lower-weighted features.</p>\n<p>For tree-based models: Use techniques that encourage inclusion of less important in-sample (IS) features. This prevents the model from overfitting to a narrow set of dominant features, especially in cases where the training period is short (only 1 year).</p>",
      "rawMarkdown": "Many practitioners prefer a lean and effective feature set, assuming selected features will generalize well to the test set. However, this assumption often breaks when certain features perform well in training but shift significantly in the test set—leading to inconsistent local CV vs. public leaderboard results (as frequently observed in discussions).\n\nTo reduce over-reliance on a few dominant features and build more robust models—especially for real-world deployment—it's helpful to diversify the feature contribution. Below are a couple of techniques I found useful that have consistently provided marginal improvements in both local CV and public scores:\n\nFor linear models: Group features by the magnitude of their coefficients (e.g., place features with |coef| ≈ 0.1 in one bucket, those with |coef| ≈ 0.01 in another). This mitigates the risk of the model blowing up if high-weight features shift, while still preserving signal from more stable, lower-weighted features.\n\nFor tree-based models: Use techniques that encourage inclusion of less important in-sample (IS) features. This prevents the model from overfitting to a narrow set of dominant features, especially in cases where the training period is short (only 1 year).",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3249055": "Many practitioners prefer a lean and effective feature set, assuming selected features will generalize well to the test set. However, this assumption often breaks when certain features perform well in training but shift significantly in the test set—leading to inconsistent local CV vs. public leaderboard results (as frequently observed in discussions).\n\nTo reduce over-reliance on a few dominant features and build more robust models—especially for real-world deployment—it's helpful to diversify the feature contribution. Below are a couple of techniques I found useful that have consistently provided marginal improvements in both local CV and public scores:\n\nFor linear models: Group features by the magnitude of their coefficients (e.g., place features with |coef| ≈ 0.1 in one bucket, those with |coef| ≈ 0.01 in another). This mitigates the risk of the model blowing up if high-weight features shift, while still preserving signal from more stable, lower-weighted features.\n\nFor tree-based models: Use techniques that encourage inclusion of less important in-sample (IS) features. This prevents the model from overfitting to a narrow set of dominant features, especially in cases where the training period is short (only 1 year)."
  }
}