{
  "id": 587510,
  "title": "Feature Drift between training and test data",
  "url": "/competitions/drw-crypto-market-prediction/discussion/587510",
  "author_name": "",
  "post_date": "2025-07-01T10:06:15.400970900Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>When I trained an XGBoost Regressor on the training data, I was getting pretty good correlation scores on my validation set (around 0.6 to 0.8). But when I submitted the predictions for the test set, the leaderboard score was much lower — around 0.04 to 0.08.</p>\n<p>I tried to understand what was going wrong and checked the distributions of the features in the training and test sets. I noticed that while some features looked similar in both datasets, a lot of them had very different distributions.</p>\n<p>So I’m thinking maybe the model didn’t generalize well because it was relying on features that behave differently in the test set.</p>\n<p>My question is: Is it fair/ethical to train the model using only those features that have similar distributions in the train and test data? I feel like this might help the model generalize better and reduce overfitting to patterns that don’t hold up in the test set.</p>\n<p>Has anyone else dealt with something like this before? Would love to know if this is a common approach or if there are better ways to handle this kind of issue.</p>",
  "messages": [
    {
      "id": "3237794",
      "postDate": "07/01/2025 10:06:15",
      "content": "<p>When I trained an XGBoost Regressor on the training data, I was getting pretty good correlation scores on my validation set (around 0.6 to 0.8). But when I submitted the predictions for the test set, the leaderboard score was much lower — around 0.04 to 0.08.</p>\n<p>I tried to understand what was going wrong and checked the distributions of the features in the training and test sets. I noticed that while some features looked similar in both datasets, a lot of them had very different distributions.</p>\n<p>So I’m thinking maybe the model didn’t generalize well because it was relying on features that behave differently in the test set.</p>\n<p>My question is: Is it fair/ethical to train the model using only those features that have similar distributions in the train and test data? I feel like this might help the model generalize better and reduce overfitting to patterns that don’t hold up in the test set.</p>\n<p>Has anyone else dealt with something like this before? Would love to know if this is a common approach or if there are better ways to handle this kind of issue.</p>",
      "rawMarkdown": "When I trained an XGBoost Regressor on the training data, I was getting pretty good correlation scores on my validation set (around 0.6 to 0.8). But when I submitted the predictions for the test set, the leaderboard score was much lower — around 0.04 to 0.08.\n\nI tried to understand what was going wrong and checked the distributions of the features in the training and test sets. I noticed that while some features looked similar in both datasets, a lot of them had very different distributions.\n\nSo I’m thinking maybe the model didn’t generalize well because it was relying on features that behave differently in the test set.\n\nMy question is: Is it fair/ethical to train the model using only those features that have similar distributions in the train and test data? I feel like this might help the model generalize better and reduce overfitting to patterns that don’t hold up in the test set.\n\nHas anyone else dealt with something like this before? Would love to know if this is a common approach or if there are better ways to handle this kind of issue.",
      "votes": null
    },
    {
      "id": "3238015",
      "postDate": "07/01/2025 13:13:23",
      "content": "<p>same here the test data distribution is very different and i think lot of people are getting high scores just by using test data distribution in training as well </p>",
      "rawMarkdown": "same here the test data distribution is very different and i think lot of people are getting high scores just by using test data distribution in training as well",
      "votes": null
    },
    {
      "id": "3238042",
      "postDate": "07/01/2025 13:48:56",
      "content": "<p>I've had this problem in some noisy datasets and was able to make the features drift-robust by binning or ranking the non-stable features. This often works well with noisy financial data where there is low signal to noise.</p>",
      "rawMarkdown": "I've had this problem in some noisy datasets and was able to make the features drift-robust by binning or ranking the non-stable features. This often works well with noisy financial data where there is low signal to noise.",
      "votes": null
    },
    {
      "id": "3238056",
      "postDate": "07/01/2025 13:56:47",
      "content": "<p>According to my experience, using the test set to gain any insights would be wrong.</p>",
      "rawMarkdown": "According to my experience, using the test set to gain any insights would be wrong.",
      "votes": null
    },
    {
      "id": "3238057",
      "postDate": "07/01/2025 13:57:43",
      "content": "<p>Thanks for your insights. Binning the non-stable features would be a good idea. But I want to know if it would make sense to check if there is a feature drift between the training and test set as it would mean that I am manipulating the training data based on future information.</p>",
      "rawMarkdown": "Thanks for your insights. Binning the non-stable features would be a good idea. But I want to know if it would make sense to check if there is a feature drift between the training and test set as it would mean that I am manipulating the training data based on future information.",
      "votes": null
    },
    {
      "id": "3238084",
      "postDate": "07/01/2025 14:26:37",
      "content": "<p>I don't think you could compare it to the test set. But you could make an assumption that variables that are stable throughout the train dataset are more likely to be stable within the test set. Or more complex assumptions like variables that becomes stable at the end of the train dataset may or may not continue to be stable throughout the test set.</p>",
      "rawMarkdown": "I don't think you could compare it to the test set. But you could make an assumption that variables that are stable throughout the train dataset are more likely to be stable within the test set. Or more complex assumptions like variables that becomes stable at the end of the train dataset may or may not continue to be stable throughout the test set.",
      "votes": null
    },
    {
      "id": "3238091",
      "postDate": "07/01/2025 14:32:04",
      "content": "<p>Have u thought about features attention mechanism like train a model to give different attention to each feature maybe this can handle correlation drift </p>",
      "rawMarkdown": "Have u thought about features attention mechanism like train a model to give different attention to each feature maybe this can handle correlation drift",
      "votes": null
    },
    {
      "id": "3238369",
      "postDate": "07/01/2025 19:24:19",
      "content": "<p>It feels like not only distributional drift but also concept drift</p>",
      "rawMarkdown": "It feels like not only distributional drift but also concept drift",
      "votes": null
    },
    {
      "id": "3238391",
      "postDate": "07/01/2025 19:49:35",
      "content": "<p>There are so many different ways to tackle this, but most of them involve making assumptions about the stability of variables in the future. A tricky proposition. A model could put general guard rails on how features should perform in the test dataset, but even that could be be dis-proven. Even looking at the train dataset, some features are really important for certain market regimes, but those features aren't very important for other market regimes.</p>",
      "rawMarkdown": "There are so many different ways to tackle this, but most of them involve making assumptions about the stability of variables in the future. A tricky proposition. A model could put general guard rails on how features should perform in the test dataset, but even that could be be dis-proven. Even looking at the train dataset, some features are really important for certain market regimes, but those features aren't very important for other market regimes.",
      "votes": null
    },
    {
      "id": "3239300",
      "postDate": "07/02/2025 17:08:01",
      "content": "<p>Could you explain what do you mean by concept drift?</p>",
      "rawMarkdown": "Could you explain what do you mean by concept drift?",
      "votes": null
    },
    {
      "id": "3239307",
      "postDate": "07/02/2025 17:24:05",
      "content": "<p>Some relationships break or even flip side (it is only my feelings</p>",
      "rawMarkdown": "Some relationships break or even flip side (it is only my feelings",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3238015,
      "author_name": "aaryanpathak",
      "author_url": "",
      "post_date": "07/01/2025 13:13:23",
      "content": "<p>same here the test data distribution is very different and i think lot of people are getting high scores just by using test data distribution in training as well </p>",
      "votes": null,
      "replies": [
        {
          "id": 3238056,
          "author_name": "fakhruddinhussain",
          "author_url": "",
          "post_date": "07/01/2025 13:56:47",
          "content": "<p>According to my experience, using the test set to gain any insights would be wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3238042,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/01/2025 13:48:56",
      "content": "<p>I've had this problem in some noisy datasets and was able to make the features drift-robust by binning or ranking the non-stable features. This often works well with noisy financial data where there is low signal to noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238057,
          "author_name": "fakhruddinhussain",
          "author_url": "",
          "post_date": "07/01/2025 13:57:43",
          "content": "<p>Thanks for your insights. Binning the non-stable features would be a good idea. But I want to know if it would make sense to check if there is a feature drift between the training and test set as it would mean that I am manipulating the training data based on future information.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3238084,
              "author_name": "taylorsamarel",
              "author_url": "",
              "post_date": "07/01/2025 14:26:37",
              "content": "<p>I don't think you could compare it to the test set. But you could make an assumption that variables that are stable throughout the train dataset are more likely to be stable within the test set. Or more complex assumptions like variables that becomes stable at the end of the train dataset may or may not continue to be stable throughout the test set.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3238091,
      "author_name": "aaryanpathak",
      "author_url": "",
      "post_date": "07/01/2025 14:32:04",
      "content": "<p>Have u thought about features attention mechanism like train a model to give different attention to each feature maybe this can handle correlation drift </p>",
      "votes": null,
      "replies": [
        {
          "id": 3238391,
          "author_name": "taylorsamarel",
          "author_url": "",
          "post_date": "07/01/2025 19:49:35",
          "content": "<p>There are so many different ways to tackle this, but most of them involve making assumptions about the stability of variables in the future. A tricky proposition. A model could put general guard rails on how features should perform in the test dataset, but even that could be be dis-proven. Even looking at the train dataset, some features are really important for certain market regimes, but those features aren't very important for other market regimes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3238369,
      "author_name": "alexzhongs",
      "author_url": "",
      "post_date": "07/01/2025 19:24:19",
      "content": "<p>It feels like not only distributional drift but also concept drift</p>",
      "votes": null,
      "replies": [
        {
          "id": 3239300,
          "author_name": "fakhruddinhussain",
          "author_url": "",
          "post_date": "07/02/2025 17:08:01",
          "content": "<p>Could you explain what do you mean by concept drift?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3239307,
              "author_name": "alexzhongs",
              "author_url": "",
              "post_date": "07/02/2025 17:24:05",
              "content": "<p>Some relationships break or even flip side (it is only my feelings</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237794": "When I trained an XGBoost Regressor on the training data, I was getting pretty good correlation scores on my validation set (around 0.6 to 0.8). But when I submitted the predictions for the test set, the leaderboard score was much lower — around 0.04 to 0.08.\n\nI tried to understand what was going wrong and checked the distributions of the features in the training and test sets. I noticed that while some features looked similar in both datasets, a lot of them had very different distributions.\n\nSo I’m thinking maybe the model didn’t generalize well because it was relying on features that behave differently in the test set.\n\nMy question is: Is it fair/ethical to train the model using only those features that have similar distributions in the train and test data? I feel like this might help the model generalize better and reduce overfitting to patterns that don’t hold up in the test set.\n\nHas anyone else dealt with something like this before? Would love to know if this is a common approach or if there are better ways to handle this kind of issue.",
    "3238015": "same here the test data distribution is very different and i think lot of people are getting high scores just by using test data distribution in training as well",
    "3238042": "I've had this problem in some noisy datasets and was able to make the features drift-robust by binning or ranking the non-stable features. This often works well with noisy financial data where there is low signal to noise.",
    "3238056": "According to my experience, using the test set to gain any insights would be wrong.",
    "3238057": "Thanks for your insights. Binning the non-stable features would be a good idea. But I want to know if it would make sense to check if there is a feature drift between the training and test set as it would mean that I am manipulating the training data based on future information.",
    "3238084": "I don't think you could compare it to the test set. But you could make an assumption that variables that are stable throughout the train dataset are more likely to be stable within the test set. Or more complex assumptions like variables that becomes stable at the end of the train dataset may or may not continue to be stable throughout the test set.",
    "3238091": "Have u thought about features attention mechanism like train a model to give different attention to each feature maybe this can handle correlation drift",
    "3238369": "It feels like not only distributional drift but also concept drift",
    "3238391": "There are so many different ways to tackle this, but most of them involve making assumptions about the stability of variables in the future. A tricky proposition. A model could put general guard rails on how features should perform in the test dataset, but even that could be be dis-proven. Even looking at the train dataset, some features are really important for certain market regimes, but those features aren't very important for other market regimes.",
    "3239300": "Could you explain what do you mean by concept drift?",
    "3239307": "Some relationships break or even flip side (it is only my feelings"
  },
  "source": "meta"
}