{
  "id": 584777,
  "title": "May I ask if anyone can explain the data processing flow of this competition?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584777",
  "author_name": "",
  "post_date": "2025-06-16T02:18:53.717959300Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/sadettinamilverdil/yat-r-m-tavsiyesi-de-ildir\">Here</a> is the notebook.</p>\n<p>I have two questions about this code, one is how to perform feature selection, and the other is how to adjust the parameters of xgboost. Is there anyone willing to give me some guidance?</p>\n<p>I know that the data for this competition has a very low signal-to-noise ratio, so I can understand that the value of the regularization parameter is very large.</p>",
  "messages": [
    {
      "id": "3225110",
      "postDate": "06/16/2025 02:18:53",
      "content": "<p><a href=\"https://www.kaggle.com/code/sadettinamilverdil/yat-r-m-tavsiyesi-de-ildir\">Here</a> is the notebook.</p>\n<p>I have two questions about this code, one is how to perform feature selection, and the other is how to adjust the parameters of xgboost. Is there anyone willing to give me some guidance?</p>\n<p>I know that the data for this competition has a very low signal-to-noise ratio, so I can understand that the value of the regularization parameter is very large.</p>",
      "rawMarkdown": "<a href=\"https://www.kaggle.com/code/sadettinamilverdil/yat-r-m-tavsiyesi-de-ildir\">Here</a> is the notebook.\n\nI have two questions about this code, one is how to perform feature selection, and the other is how to adjust the parameters of xgboost. Is there anyone willing to give me some guidance?\n\nI know that the data for this competition has a very low signal-to-noise ratio, so I can understand that the value of the regularization parameter is very large.",
      "votes": null
    },
    {
      "id": "3225191",
      "postDate": "06/16/2025 04:53:28",
      "content": "<ol>\n<li><p>Feature selection via SHAP, it like lgbm, xgb importance value. You can train some model then select columns that in top ft importance on almost model.</p></li>\n<li><p>I think author run simple hyperparameters search, maybe use kflod without shuffle or traintestsplit.</p></li>\n</ol>\n<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "rawMarkdown": "1. Feature selection via SHAP, it like lgbm, xgb importance value. You can train some model then select columns that in top ft importance on almost model.\n\n2. I think author run simple hyperparameters search, maybe use kflod without shuffle or traintestsplit.\n\n@yunsuxiaozi",
      "votes": null
    },
    {
      "id": "3225974",
      "postDate": "06/17/2025 05:23:53",
      "content": "<p>My dp is that for feature selection, good performance in local cv doesn't necessarily guarantee good public score, which I guess is related to distributional shift. For that I haven't found a solid solution.</p>",
      "rawMarkdown": "My dp is that for feature selection, good performance in local cv doesn't necessarily guarantee good public score, which I guess is related to distributional shift. For that I haven't found a solid solution.",
      "votes": null
    },
    {
      "id": "3226801",
      "postDate": "06/18/2025 05:28:25",
      "content": "<p>Data Processing Pipeline (LightGBM Example)<br>\nHere’s a high-level view of my preprocessing workflow:<br>\nDrop highly correlated features (&gt;0.98 correlation threshold)<br>\nRemove low-variance columns and NaN-heavy columns<br>\nFeature Selection using SelectKBest + f_regression<br>\nStandard scaling (mostly to assist with interpretability and consistency in model input)<br>\nModel Training (LightGBM with GPU)<br>\nIncrementally trained models with 5 to 65 top features to analyze performance gains.</p>\n<p>for XGBoost you can try using k fold cv + grid search to find optimal features but Im not sure if theres any set protocol that you could follow to achieve great results apart from this.</p>\n<p>from some of my findings it seems that there is a major cutoff in signal strength after a set number of features. So maybe focus on narrowing down to fewer more important features. </p>\n<p>You could try using PCA to reduce the dimensionality of the data. Start there and see where you can go.</p>\n<p>this worked fine for me but im sure theres better ways.<br>\nhope it helps.</p>",
      "rawMarkdown": "Data Processing Pipeline (LightGBM Example)\nHere’s a high-level view of my preprocessing workflow:\nDrop highly correlated features (>0.98 correlation threshold)\nRemove low-variance columns and NaN-heavy columns\nFeature Selection using SelectKBest + f_regression\nStandard scaling (mostly to assist with interpretability and consistency in model input)\nModel Training (LightGBM with GPU)\nIncrementally trained models with 5 to 65 top features to analyze performance gains.\n\n\nfor XGBoost you can try using k fold cv + grid search to find optimal features but Im not sure if theres any set protocol that you could follow to achieve great results apart from this.\n\nfrom some of my findings it seems that there is a major cutoff in signal strength after a set number of features. So maybe focus on narrowing down to fewer more important features. \n\nYou could try using PCA to reduce the dimensionality of the data. Start there and see where you can go.\n\n\nthis worked fine for me but im sure theres better ways.\nhope it helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3225191,
      "author_name": "nguyennguyen599",
      "author_url": "",
      "post_date": "06/16/2025 04:53:28",
      "content": "<ol>\n<li><p>Feature selection via SHAP, it like lgbm, xgb importance value. You can train some model then select columns that in top ft importance on almost model.</p></li>\n<li><p>I think author run simple hyperparameters search, maybe use kflod without shuffle or traintestsplit.</p></li>\n</ol>\n<p><a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3225974,
      "author_name": "alexzhongs",
      "author_url": "",
      "post_date": "06/17/2025 05:23:53",
      "content": "<p>My dp is that for feature selection, good performance in local cv doesn't necessarily guarantee good public score, which I guess is related to distributional shift. For that I haven't found a solid solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3226801,
      "author_name": "sanketpai",
      "author_url": "",
      "post_date": "06/18/2025 05:28:25",
      "content": "<p>Data Processing Pipeline (LightGBM Example)<br>\nHere’s a high-level view of my preprocessing workflow:<br>\nDrop highly correlated features (&gt;0.98 correlation threshold)<br>\nRemove low-variance columns and NaN-heavy columns<br>\nFeature Selection using SelectKBest + f_regression<br>\nStandard scaling (mostly to assist with interpretability and consistency in model input)<br>\nModel Training (LightGBM with GPU)<br>\nIncrementally trained models with 5 to 65 top features to analyze performance gains.</p>\n<p>for XGBoost you can try using k fold cv + grid search to find optimal features but Im not sure if theres any set protocol that you could follow to achieve great results apart from this.</p>\n<p>from some of my findings it seems that there is a major cutoff in signal strength after a set number of features. So maybe focus on narrowing down to fewer more important features. </p>\n<p>You could try using PCA to reduce the dimensionality of the data. Start there and see where you can go.</p>\n<p>this worked fine for me but im sure theres better ways.<br>\nhope it helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3225110": "<a href=\"https://www.kaggle.com/code/sadettinamilverdil/yat-r-m-tavsiyesi-de-ildir\">Here</a> is the notebook.\n\nI have two questions about this code, one is how to perform feature selection, and the other is how to adjust the parameters of xgboost. Is there anyone willing to give me some guidance?\n\nI know that the data for this competition has a very low signal-to-noise ratio, so I can understand that the value of the regularization parameter is very large.",
    "3225191": "1. Feature selection via SHAP, it like lgbm, xgb importance value. You can train some model then select columns that in top ft importance on almost model.\n\n2. I think author run simple hyperparameters search, maybe use kflod without shuffle or traintestsplit.\n\n@yunsuxiaozi",
    "3225974": "My dp is that for feature selection, good performance in local cv doesn't necessarily guarantee good public score, which I guess is related to distributional shift. For that I haven't found a solid solution.",
    "3226801": "Data Processing Pipeline (LightGBM Example)\nHere’s a high-level view of my preprocessing workflow:\nDrop highly correlated features (>0.98 correlation threshold)\nRemove low-variance columns and NaN-heavy columns\nFeature Selection using SelectKBest + f_regression\nStandard scaling (mostly to assist with interpretability and consistency in model input)\nModel Training (LightGBM with GPU)\nIncrementally trained models with 5 to 65 top features to analyze performance gains.\n\n\nfor XGBoost you can try using k fold cv + grid search to find optimal features but Im not sure if theres any set protocol that you could follow to achieve great results apart from this.\n\nfrom some of my findings it seems that there is a major cutoff in signal strength after a set number of features. So maybe focus on narrowing down to fewer more important features. \n\nYou could try using PCA to reduce the dimensionality of the data. Start there and see where you can go.\n\n\nthis worked fine for me but im sure theres better ways.\nhope it helps."
  },
  "source": "meta"
}