{
  "id": 589645,
  "title": "For those that did not peek into the future..",
  "url": "/competitions/drw-crypto-market-prediction/discussion/589645",
  "author_name": "",
  "post_date": "2025-07-14T11:50:33.526935Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>For those that did not peek into the future, when this competition is over, could you share your methodology to overcome dataset drift? What are the criteria you use to select the features as well as features engineering? </p>\n<p>I tried various preprocessing including binning, but even if my cross validation score on my train set was good and with a tiny standard deviation, the final score in the test set just way off. Consistence result in cross validation doesn't mean consistency in test set.</p>\n<p>I'm interested to learn from others the real world approach in this kind of dataset.</p>",
  "messages": [
    {
      "id": "3248343",
      "postDate": "07/14/2025 11:50:33",
      "content": "<p>For those that did not peek into the future, when this competition is over, could you share your methodology to overcome dataset drift? What are the criteria you use to select the features as well as features engineering? </p>\n<p>I tried various preprocessing including binning, but even if my cross validation score on my train set was good and with a tiny standard deviation, the final score in the test set just way off. Consistence result in cross validation doesn't mean consistency in test set.</p>\n<p>I'm interested to learn from others the real world approach in this kind of dataset.</p>",
      "rawMarkdown": "For those that did not peek into the future, when this competition is over, could you share your methodology to overcome dataset drift? What are the criteria you use to select the features as well as features engineering? \n\nI tried various preprocessing including binning, but even if my cross validation score on my train set was good and with a tiny standard deviation, the final score in the test set just way off. Consistence result in cross validation doesn't mean consistency in test set.\n\nI'm interested to learn from others the real world approach in this kind of dataset.",
      "votes": null
    },
    {
      "id": "3248362",
      "postDate": "07/14/2025 12:37:28",
      "content": "<p>Have you tried looking at the stability of variables via bootstrapping?</p>",
      "rawMarkdown": "Have you tried looking at the stability of variables via bootstrapping?",
      "votes": null
    },
    {
      "id": "3248403",
      "postDate": "07/14/2025 13:55:49",
      "content": "<p>Well, you never truly know if something is bound to drift, unless you see the test data. </p>\n<p>If you want to completely avoid lookahead (no looking into the test data, no reverse-engineering of the features), you can test the drift of the model by varying the train set across time. </p>\n<p>But first, based on your post, there is one thing I have to suggest: using classic CV (i.e., choose indexes for train and val randomly) is a very bad idea in our case, as it induces significant lookahead. <br>\nThink about it: if you split the folds randomly across the data, then for each validation point you are bound to have a train datapoint that is very close in time, so your model will always capture feature drift using this validation setup (meaning the model will always know how features progress over the whole period of time -- that's pure lookahead which you want to avoid).</p>\n<p>So try using a simple split like oldest 80% for training and latest 20% for validation. In this case it's a much better estimate than classic CV. Just make sure your train data never looks ahead of any validation data.</p>\n<p>The next step would be to train models on different subsamples of the data (say 0-50% of dates, 20-70% dates, 30-80% of dates in terms of time progression from oldest to newest) and filtering out features that decrease stability. This is similar to bootstrapping, except it accounts for the lookahead you are bound to get with classic randomization.</p>",
      "rawMarkdown": "Well, you never truly know if something is bound to drift, unless you see the test data. \n\nIf you want to completely avoid lookahead (no looking into the test data, no reverse-engineering of the features), you can test the drift of the model by varying the train set across time. \n\nBut first, based on your post, there is one thing I have to suggest: using classic CV (i.e., choose indexes for train and val randomly) is a very bad idea in our case, as it induces significant lookahead. \nThink about it: if you split the folds randomly across the data, then for each validation point you are bound to have a train datapoint that is very close in time, so your model will always capture feature drift using this validation setup (meaning the model will always know how features progress over the whole period of time -- that's pure lookahead which you want to avoid).\n\nSo try using a simple split like oldest 80% for training and latest 20% for validation. In this case it's a much better estimate than classic CV. Just make sure your train data never looks ahead of any validation data.\n\n\nThe next step would be to train models on different subsamples of the data (say 0-50% of dates, 20-70% dates, 30-80% of dates in terms of time progression from oldest to newest) and filtering out features that decrease stability. This is similar to bootstrapping, except it accounts for the lookahead you are bound to get with classic randomization.",
      "votes": null
    },
    {
      "id": "3248408",
      "postDate": "07/14/2025 14:03:08",
      "content": "<p>I see what you are saying here, and think your last paragraph is great.</p>\n<p>However, simple split like oldest 80% and 20% validation has its own deficiencies. For example, that last 20% for validation could be occurring in a season/regime that is different than the rest of the data. There is likely some seasonality to crypto markets and it can be a bit misleading to use a specific 'season' for validation, when the entire test dataset is much more than just 1 'season'.</p>",
      "rawMarkdown": "I see what you are saying here, and think your last paragraph is great.\n\nHowever, simple split like oldest 80% and 20% validation has its own deficiencies. For example, that last 20% for validation could be occurring in a season/regime that is different than the rest of the data. There is likely some seasonality to crypto markets and it can be a bit misleading to use a specific 'season' for validation, when the entire test dataset is much more than just 1 'season'.",
      "votes": null
    },
    {
      "id": "3248444",
      "postDate": "07/14/2025 15:42:42",
      "content": "<p>That is a valid concern, of course. I don't think there is a single simple solution to this problem. One common way to account for the regime switch is through averaging models trained over different periods of time.</p>\n<p>As for this competition, the scores I got on the last 20% of train are actually very close to what I see on PLB. </p>",
      "rawMarkdown": "That is a valid concern, of course. I don't think there is a single simple solution to this problem. One common way to account for the regime switch is through averaging models trained over different periods of time.\n\nAs for this competition, the scores I got on the last 20% of train are actually very close to what I see on PLB.",
      "votes": null
    },
    {
      "id": "3249024",
      "postDate": "07/15/2025 14:28:46",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/taylorsamarel\" target=\"_blank\">@taylorsamarel</a> </p>\n<p>I assume bootstrapping here means running 100 round of resampled X, Y (with replacement) and train a model like LassoCV on it. Using the results' model coefficient to rank the features. Then nope. I was thinking I could more or less achieve similar result by just referring to their 10 fold cv scores. Ie. features with with high standard deviation scores typically show instability.  </p>",
      "rawMarkdown": "Thanks @taylorsamarel \n\nI assume bootstrapping here means running 100 round of resampled X, Y (with replacement) and train a model like LassoCV on it. Using the results' model coefficient to rank the features. Then nope. I was thinking I could more or less achieve similar result by just referring to their 10 fold cv scores. Ie. features with with high standard deviation scores typically show instability.",
      "votes": null
    },
    {
      "id": "3249033",
      "postDate": "07/15/2025 14:39:44",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/vladimirkhismatullin\" target=\"_blank\">@vladimirkhismatullin</a> for your suggestions.</p>\n<p>Well I did use KFold cv but I did not shuffled. Although this does not completely eliminate lookahead biases as in this competition, the features and label may still overlapped a bit in the border between train and validation set, but I argued that the leaks are minor.</p>",
      "rawMarkdown": "Thank you @vladimirkhismatullin for your suggestions.\n\nWell I did use KFold cv but I did not shuffled. Although this does not completely eliminate lookahead biases as in this competition, the features and label may still overlapped a bit in the border between train and validation set, but I argued that the leaks are minor.",
      "votes": null
    },
    {
      "id": "3249083",
      "postDate": "07/15/2025 16:50:20",
      "content": "<p>I find it so hard.I worked for a long time but the score is only 0.09.I don't know how the Kagglers who have reached a score of 0.9 did it</p>",
      "rawMarkdown": "I find it so hard.I worked for a long time but the score is only 0.09.I don't know how the Kagglers who have reached a score of 0.9 did it",
      "votes": null
    },
    {
      "id": "3249100",
      "postDate": "07/15/2025 17:41:35",
      "content": "<p>Ordered testset is available, then you can use timeseries features…</p>",
      "rawMarkdown": "Ordered testset is available, then you can use timeseries features...",
      "votes": null
    },
    {
      "id": "3249120",
      "postDate": "07/15/2025 18:28:42",
      "content": "<p>No offense -- just to be constructive, from what I understand, you use CV and get an estimate very different from PLB. Isn't that enough evidence to say that the leaks are not minor? </p>\n<p>You might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV is that you induce significant lookahead in the other features (unless you drop them beforehand), which do drift over time (this drift is hard to know for the test data, hence the difference).</p>\n<p>Try doing the following: split train 80-20 by time into new train and fake test. Repeat your CV procedure on the new train and then measure the performance once on fake test. You will likely get a near zero score on fake test.<br>\nThen do the same with features from <a href=\"https://www.kaggle.com/migrantworkerdatahub\" target=\"_blank\">@migrantworkerdatahub</a> 'small improvements' notebook - and there will be close to no difference between scores on CV and fake test.</p>",
      "rawMarkdown": "No offense -- just to be constructive, from what I understand, you use CV and get an estimate very different from PLB. Isn't that enough evidence to say that the leaks are not minor? \n\nYou might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV is that you induce significant lookahead in the other features (unless you drop them beforehand), which do drift over time (this drift is hard to know for the test data, hence the difference).\n\nTry doing the following: split train 80-20 by time into new train and fake test. Repeat your CV procedure on the new train and then measure the performance once on fake test. You will likely get a near zero score on fake test.\nThen do the same with features from @migrantworkerdatahub 'small improvements' notebook - and there will be close to no difference between scores on CV and fake test.",
      "votes": null
    },
    {
      "id": "3249232",
      "postDate": "07/16/2025 04:15:48",
      "content": "<p>The test set is actually out of order，if use timeseries index(train set index is a timeseries) ,I don't know if the model will have the same effect on the test set if it learns based on time.because test set is messy,My idea is to find the data pattern of each row</p>",
      "rawMarkdown": "The test set is actually out of order，if use timeseries index(train set index is a timeseries) ,I don't know if the model will have the same effect on the test set if it learns based on time.because test set is messy,My idea is to find the data pattern of each row",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3248362,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/14/2025 12:37:28",
      "content": "<p>Have you tried looking at the stability of variables via bootstrapping?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3249024,
          "author_name": "hewhew",
          "author_url": "",
          "post_date": "07/15/2025 14:28:46",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/taylorsamarel\" target=\"_blank\">@taylorsamarel</a> </p>\n<p>I assume bootstrapping here means running 100 round of resampled X, Y (with replacement) and train a model like LassoCV on it. Using the results' model coefficient to rank the features. Then nope. I was thinking I could more or less achieve similar result by just referring to their 10 fold cv scores. Ie. features with with high standard deviation scores typically show instability.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3248403,
      "author_name": "vladimirkhismatullin",
      "author_url": "",
      "post_date": "07/14/2025 13:55:49",
      "content": "<p>Well, you never truly know if something is bound to drift, unless you see the test data. </p>\n<p>If you want to completely avoid lookahead (no looking into the test data, no reverse-engineering of the features), you can test the drift of the model by varying the train set across time. </p>\n<p>But first, based on your post, there is one thing I have to suggest: using classic CV (i.e., choose indexes for train and val randomly) is a very bad idea in our case, as it induces significant lookahead. <br>\nThink about it: if you split the folds randomly across the data, then for each validation point you are bound to have a train datapoint that is very close in time, so your model will always capture feature drift using this validation setup (meaning the model will always know how features progress over the whole period of time -- that's pure lookahead which you want to avoid).</p>\n<p>So try using a simple split like oldest 80% for training and latest 20% for validation. In this case it's a much better estimate than classic CV. Just make sure your train data never looks ahead of any validation data.</p>\n<p>The next step would be to train models on different subsamples of the data (say 0-50% of dates, 20-70% dates, 30-80% of dates in terms of time progression from oldest to newest) and filtering out features that decrease stability. This is similar to bootstrapping, except it accounts for the lookahead you are bound to get with classic randomization.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3248408,
          "author_name": "taylorsamarel",
          "author_url": "",
          "post_date": "07/14/2025 14:03:08",
          "content": "<p>I see what you are saying here, and think your last paragraph is great.</p>\n<p>However, simple split like oldest 80% and 20% validation has its own deficiencies. For example, that last 20% for validation could be occurring in a season/regime that is different than the rest of the data. There is likely some seasonality to crypto markets and it can be a bit misleading to use a specific 'season' for validation, when the entire test dataset is much more than just 1 'season'.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3248444,
              "author_name": "vladimirkhismatullin",
              "author_url": "",
              "post_date": "07/14/2025 15:42:42",
              "content": "<p>That is a valid concern, of course. I don't think there is a single simple solution to this problem. One common way to account for the regime switch is through averaging models trained over different periods of time.</p>\n<p>As for this competition, the scores I got on the last 20% of train are actually very close to what I see on PLB. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3249033,
          "author_name": "hewhew",
          "author_url": "",
          "post_date": "07/15/2025 14:39:44",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/vladimirkhismatullin\" target=\"_blank\">@vladimirkhismatullin</a> for your suggestions.</p>\n<p>Well I did use KFold cv but I did not shuffled. Although this does not completely eliminate lookahead biases as in this competition, the features and label may still overlapped a bit in the border between train and validation set, but I argued that the leaks are minor.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3249120,
              "author_name": "vladimirkhismatullin",
              "author_url": "",
              "post_date": "07/15/2025 18:28:42",
              "content": "<p>No offense -- just to be constructive, from what I understand, you use CV and get an estimate very different from PLB. Isn't that enough evidence to say that the leaks are not minor? </p>\n<p>You might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV is that you induce significant lookahead in the other features (unless you drop them beforehand), which do drift over time (this drift is hard to know for the test data, hence the difference).</p>\n<p>Try doing the following: split train 80-20 by time into new train and fake test. Repeat your CV procedure on the new train and then measure the performance once on fake test. You will likely get a near zero score on fake test.<br>\nThen do the same with features from <a href=\"https://www.kaggle.com/migrantworkerdatahub\" target=\"_blank\">@migrantworkerdatahub</a> 'small improvements' notebook - and there will be close to no difference between scores on CV and fake test.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3249083,
      "author_name": "seoyonyip",
      "author_url": "",
      "post_date": "07/15/2025 16:50:20",
      "content": "<p>I find it so hard.I worked for a long time but the score is only 0.09.I don't know how the Kagglers who have reached a score of 0.9 did it</p>",
      "votes": null,
      "replies": [
        {
          "id": 3249100,
          "author_name": "drshredz",
          "author_url": "",
          "post_date": "07/15/2025 17:41:35",
          "content": "<p>Ordered testset is available, then you can use timeseries features…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3249232,
              "author_name": "seoyonyip",
              "author_url": "",
              "post_date": "07/16/2025 04:15:48",
              "content": "<p>The test set is actually out of order，if use timeseries index(train set index is a timeseries) ,I don't know if the model will have the same effect on the test set if it learns based on time.because test set is messy,My idea is to find the data pattern of each row</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3248343": "For those that did not peek into the future, when this competition is over, could you share your methodology to overcome dataset drift? What are the criteria you use to select the features as well as features engineering? \n\nI tried various preprocessing including binning, but even if my cross validation score on my train set was good and with a tiny standard deviation, the final score in the test set just way off. Consistence result in cross validation doesn't mean consistency in test set.\n\nI'm interested to learn from others the real world approach in this kind of dataset.",
    "3248362": "Have you tried looking at the stability of variables via bootstrapping?",
    "3248403": "Well, you never truly know if something is bound to drift, unless you see the test data. \n\nIf you want to completely avoid lookahead (no looking into the test data, no reverse-engineering of the features), you can test the drift of the model by varying the train set across time. \n\nBut first, based on your post, there is one thing I have to suggest: using classic CV (i.e., choose indexes for train and val randomly) is a very bad idea in our case, as it induces significant lookahead. \nThink about it: if you split the folds randomly across the data, then for each validation point you are bound to have a train datapoint that is very close in time, so your model will always capture feature drift using this validation setup (meaning the model will always know how features progress over the whole period of time -- that's pure lookahead which you want to avoid).\n\nSo try using a simple split like oldest 80% for training and latest 20% for validation. In this case it's a much better estimate than classic CV. Just make sure your train data never looks ahead of any validation data.\n\n\nThe next step would be to train models on different subsamples of the data (say 0-50% of dates, 20-70% dates, 30-80% of dates in terms of time progression from oldest to newest) and filtering out features that decrease stability. This is similar to bootstrapping, except it accounts for the lookahead you are bound to get with classic randomization.",
    "3248408": "I see what you are saying here, and think your last paragraph is great.\n\nHowever, simple split like oldest 80% and 20% validation has its own deficiencies. For example, that last 20% for validation could be occurring in a season/regime that is different than the rest of the data. There is likely some seasonality to crypto markets and it can be a bit misleading to use a specific 'season' for validation, when the entire test dataset is much more than just 1 'season'.",
    "3248444": "That is a valid concern, of course. I don't think there is a single simple solution to this problem. One common way to account for the regime switch is through averaging models trained over different periods of time.\n\nAs for this competition, the scores I got on the last 20% of train are actually very close to what I see on PLB.",
    "3249024": "Thanks @taylorsamarel \n\nI assume bootstrapping here means running 100 round of resampled X, Y (with replacement) and train a model like LassoCV on it. Using the results' model coefficient to rank the features. Then nope. I was thinking I could more or less achieve similar result by just referring to their 10 fold cv scores. Ie. features with with high standard deviation scores typically show instability.",
    "3249033": "Thank you @vladimirkhismatullin for your suggestions.\n\nWell I did use KFold cv but I did not shuffled. Although this does not completely eliminate lookahead biases as in this competition, the features and label may still overlapped a bit in the border between train and validation set, but I argued that the leaks are minor.",
    "3249083": "I find it so hard.I worked for a long time but the score is only 0.09.I don't know how the Kagglers who have reached a score of 0.9 did it",
    "3249100": "Ordered testset is available, then you can use timeseries features...",
    "3249120": "No offense -- just to be constructive, from what I understand, you use CV and get an estimate very different from PLB. Isn't that enough evidence to say that the leaks are not minor? \n\nYou might have seen some public codes that achieves good public LB results and use k-fold. BUT, these solutions already selected the 'correct' features that do not drift over time. Once again, the problem with CV is that you induce significant lookahead in the other features (unless you drop them beforehand), which do drift over time (this drift is hard to know for the test data, hence the difference).\n\nTry doing the following: split train 80-20 by time into new train and fake test. Repeat your CV procedure on the new train and then measure the performance once on fake test. You will likely get a near zero score on fake test.\nThen do the same with features from @migrantworkerdatahub 'small improvements' notebook - and there will be close to no difference between scores on CV and fake test.",
    "3249232": "The test set is actually out of order，if use timeseries index(train set index is a timeseries) ,I don't know if the model will have the same effect on the test set if it learns based on time.because test set is messy,My idea is to find the data pattern of each row"
  },
  "source": "meta"
}