{
  "id": 581193,
  "title": "Sample Weights To Bias Model",
  "url": "/competitions/drw-crypto-market-prediction/discussion/581193",
  "author_name": "",
  "post_date": "2025-05-29T00:57:52.843607200Z",
  "votes": 20,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I think a few others have discovered this but although this is not a time series problem, I have discovered that using sample weights can help bias the model to put more weight on the records that occur at the end of the train dataset.</p>\n<p>The theory is that the test dataset would logically come after the train dataset and market dynamics may have slightly changed between the beginning of the train dataset and the end. The end of the train dataset might be more representative of the test data.</p>\n<p>I know this probably isn't ground breaking but I've been able to implement it to improve a few different model architectures with success.</p>",
  "messages": [
    {
      "id": "3211845",
      "postDate": "05/29/2025 00:57:52",
      "content": "<p>I think a few others have discovered this but although this is not a time series problem, I have discovered that using sample weights can help bias the model to put more weight on the records that occur at the end of the train dataset.</p>\n<p>The theory is that the test dataset would logically come after the train dataset and market dynamics may have slightly changed between the beginning of the train dataset and the end. The end of the train dataset might be more representative of the test data.</p>\n<p>I know this probably isn't ground breaking but I've been able to implement it to improve a few different model architectures with success.</p>",
      "rawMarkdown": "I think a few others have discovered this but although this is not a time series problem, I have discovered that using sample weights can help bias the model to put more weight on the records that occur at the end of the train dataset.\n\nThe theory is that the test dataset would logically come after the train dataset and market dynamics may have slightly changed between the beginning of the train dataset and the end. The end of the train dataset might be more representative of the test data.\n\nI know this probably isn't ground breaking but I've been able to implement it to improve a few different model architectures with success.",
      "votes": null
    },
    {
      "id": "3211871",
      "postDate": "05/29/2025 02:22:43",
      "content": "<p>You mentioned that the end of the train data might better represent the test data, but since it's 1 year test, it might predict the first few months of the test better while performing worse for later test dates. So while the public score improves, the private score will drop.<br>\n this creates a classic overfitting scenario to the public leaderboard</p>",
      "rawMarkdown": "You mentioned that the end of the train data might better represent the test data, but since it's 1 year test, it might predict the first few months of the test better while performing worse for later test dates. So while the public score improves, the private score will drop.\n this creates a classic overfitting scenario to the public leaderboard",
      "votes": null
    },
    {
      "id": "3211880",
      "postDate": "05/29/2025 02:44:56",
      "content": "<p>It is possible that it could lead to over fitting depending upon the strength and curve of sample weight decay. I tried multiple methods of decaying weights along different curves and found it to be relatively successful when used in moderation. Using appropriate decay on the sample weights and not completely disregarding the value of the first rows provides a method to avoid over fitting.</p>\n<p>Since we don't have the chronological order of the test data, it becomes a game of striking a balance. If I can improve the first for the first 8 months of test data while slightly increase the errors on the remaining 4 months, on balance it still can produce better results. </p>\n<p>Outside of competitions on Kaggle I've seen dynamic sample weights be used successfully in real trading situations that are not time series oriented.</p>\n<p>Maybe I'm missing something here and would love to hear more thoughts and strategies.</p>",
      "rawMarkdown": "It is possible that it could lead to over fitting depending upon the strength and curve of sample weight decay. I tried multiple methods of decaying weights along different curves and found it to be relatively successful when used in moderation. Using appropriate decay on the sample weights and not completely disregarding the value of the first rows provides a method to avoid over fitting.\n\nSince we don't have the chronological order of the test data, it becomes a game of striking a balance. If I can improve the first for the first 8 months of test data while slightly increase the errors on the remaining 4 months, on balance it still can produce better results. \n\nOutside of competitions on Kaggle I've seen dynamic sample weights be used successfully in real trading situations that are not time series oriented.\n\nMaybe I'm missing something here and would love to hear more thoughts and strategies.",
      "votes": null
    },
    {
      "id": "3211886",
      "postDate": "05/29/2025 02:57:06",
      "content": "<p>If the test data were 1-2 month, I would agree with you :)</p>",
      "rawMarkdown": "If the test data were 1-2 month, I would agree with you :)",
      "votes": null
    },
    {
      "id": "3211903",
      "postDate": "05/29/2025 03:29:06",
      "content": "<p>The great thing about Kaggle is we can just test it.</p>\n<p>I'm trying to put a comparison up here but currently have GPU tasks ongoing: <a href=\"https://www.kaggle.com/code/taylorsamarel/sample-weight-decay-normal-ensemble-comparison/\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/sample-weight-decay-normal-ensemble-comparison/</a></p>\n<p>The notebook is setup to compare 3 different variations (EDIT: I ran the notebook and added scores after each strategy):</p>\n<ol>\n<li>No sample weights; [Submission Score <strong>0.09852</strong>]</li>\n<li>With sample weights including a weight floor and decay rate; (it does admittedly get a little tricky with how weights are integrated into folds); [Submission Score <strong>0.09927</strong>]</li>\n<li>Ensemble based recency weighting by training two models, the first on 100% of the data, the second on the most recent 75% of the data and then combining those to get additional emphasis on more recent data. [Submission Score <strong>0.09991</strong>]</li>\n</ol>\n<p>From my submissions these methods seem to improve the scores on the test data.</p>\n<p>Also, just thinking real world here, near the second half of the train data a new crypto exchange could have launched, or liquidity algorithms change, or an exchange drops out and no longer operates. This creates a new market 'regime' that didn't exist in the beginning of the train set but may continue to exist in the test time period. Granted there are many assumptions in this theory, but it appears to work.</p>\n<p>One of the many anonymized features could capture things that better model market regimes but there is a lot of noise to sort through. </p>",
      "rawMarkdown": "The great thing about Kaggle is we can just test it.\n\nI'm trying to put a comparison up here but currently have GPU tasks ongoing: https://www.kaggle.com/code/taylorsamarel/sample-weight-decay-normal-ensemble-comparison/\n\nThe notebook is setup to compare 3 different variations (EDIT: I ran the notebook and added scores after each strategy):\n\n1. No sample weights; [Submission Score **0.09852**]\n2. With sample weights including a weight floor and decay rate; (it does admittedly get a little tricky with how weights are integrated into folds); [Submission Score **0.09927**]\n3. Ensemble based recency weighting by training two models, the first on 100% of the data, the second on the most recent 75% of the data and then combining those to get additional emphasis on more recent data. [Submission Score **0.09991**]\n\nFrom my submissions these methods seem to improve the scores on the test data.\n\nAlso, just thinking real world here, near the second half of the train data a new crypto exchange could have launched, or liquidity algorithms change, or an exchange drops out and no longer operates. This creates a new market 'regime' that didn't exist in the beginning of the train set but may continue to exist in the test time period. Granted there are many assumptions in this theory, but it appears to work.\n\nOne of the many anonymized features could capture things that better model market regimes but there is a lot of noise to sort through.",
      "votes": null
    },
    {
      "id": "3211923",
      "postDate": "05/29/2025 04:02:53",
      "content": "<p>If the 1-year test dataset had been randomly split into public and private portions, I would consider the score improvement reliable. However, from what I understand, the public data consists of test data from the first 6 months, while the private data contains test data from the last 6 months. This means the private dataset represents a more recent time period.</p>\n<p>The goal is to achieve good performance on the private dataset, but this temporal split raises concerns about whether improvements on the public data will actually translate to success on the private data, since they come from different time periods rather than being randomly distributed</p>",
      "rawMarkdown": "If the 1-year test dataset had been randomly split into public and private portions, I would consider the score improvement reliable. However, from what I understand, the public data consists of test data from the first 6 months, while the private data contains test data from the last 6 months. This means the private dataset represents a more recent time period.\n\nThe goal is to achieve good performance on the private dataset, but this temporal split raises concerns about whether improvements on the public data will actually translate to success on the private data, since they come from different time periods rather than being randomly distributed",
      "votes": null
    },
    {
      "id": "3211937",
      "postDate": "05/29/2025 04:39:14",
      "content": "<p>Ahh, this is an excellent point. I did not consider how the public and private test data is split. In fact, I had to go through the discussion and see your comments with <a href=\"https://www.kaggle.com/yw2735\" target=\"_blank\">@yw2735</a>.</p>\n<p>Considering that the private test data (51%) come after the public test data, the improvements I've been seeing may diminish, but there could also be situations where it performs better than expected, especially if that 51% deviates significantly from the records at the very beginning.</p>\n<p>Its an interesting situation, I wish we were given public test data that spans a larger time frame. Once the 51% comes into play scoring of notebooks will fluctuate a lot.</p>",
      "rawMarkdown": "Ahh, this is an excellent point. I did not consider how the public and private test data is split. In fact, I had to go through the discussion and see your comments with @yw2735.\n\nConsidering that the private test data (51%) come after the public test data, the improvements I've been seeing may diminish, but there could also be situations where it performs better than expected, especially if that 51% deviates significantly from the records at the very beginning.\n\nIts an interesting situation, I wish we were given public test data that spans a larger time frame. Once the 51% comes into play scoring of notebooks will fluctuate a lot.",
      "votes": null
    },
    {
      "id": "3212249",
      "postDate": "05/29/2025 14:51:01",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/onurkoc83\" target=\"_blank\">@onurkoc83</a> that this might cause overfitting in lb. But since we can choose two submissions, I may consider submitting one version with sample weights and one without.</p>",
      "rawMarkdown": "I agree with @onurkoc83 that this might cause overfitting in lb. But since we can choose two submissions, I may consider submitting one version with sample weights and one without.",
      "votes": null
    },
    {
      "id": "3212263",
      "postDate": "05/29/2025 15:13:29",
      "content": "<p>Probably a good strategy to try both for submission.</p>\n<p>I think the argument for overfitting goes both ways. Leaving earlier records with the same weight as the most recent records risks overfitting to the earlier records.</p>\n<p>And to be clear, the sample weight adjustments that I've played around with are very mild. For example, the most recent record having a weight of 1, and the first record having a weight of 0.9. Which comes out to a decay rate of nearly 0.0000002 per record - which some would say is very small.</p>",
      "rawMarkdown": "Probably a good strategy to try both for submission.\n\nI think the argument for overfitting goes both ways. Leaving earlier records with the same weight as the most recent records risks overfitting to the earlier records.\n\nAnd to be clear, the sample weight adjustments that I've played around with are very mild. For example, the most recent record having a weight of 1, and the first record having a weight of 0.9. Which comes out to a decay rate of nearly 0.0000002 per record - which some would say is very small.",
      "votes": null
    },
    {
      "id": "3227849",
      "postDate": "06/19/2025 11:05:42",
      "content": "<p>May I ask how the time weight is assigned to the test data? According to the data tab, the test data has been shuffled and masked, so the order is unknown.</p>",
      "rawMarkdown": "May I ask how the time weight is assigned to the test data? According to the data tab, the test data has been shuffled and masked, so the order is unknown.",
      "votes": null
    },
    {
      "id": "3227964",
      "postDate": "06/19/2025 13:50:11",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ccyhui\" target=\"_blank\">@ccyhui</a>, you're right - we can't assign time-based weights to the test data for exactly the reasons you mentioned.<br>\nSample weighting in the training data could potentially help, but only if:</p>\n<p>Higher-weighted samples better represent the test data distribution<br>\nIn this case, I suspect recent training data (near the end of the dataset) is more representative of test data patterns than earlier records and therefore deserves a higher sample weight when training.</p>\n<p>The trade-off is that this approach carries risks - it might overfit to recent patterns or miss important information from regime changes or structural shifts in the market.</p>",
      "rawMarkdown": "Hi @ccyhui, you're right - we can't assign time-based weights to the test data for exactly the reasons you mentioned.\nSample weighting in the training data could potentially help, but only if:\n\nHigher-weighted samples better represent the test data distribution\nIn this case, I suspect recent training data (near the end of the dataset) is more representative of test data patterns than earlier records and therefore deserves a higher sample weight when training.\n\nThe trade-off is that this approach carries risks - it might overfit to recent patterns or miss important information from regime changes or structural shifts in the market.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3211871,
      "author_name": "onurkoc83",
      "author_url": "",
      "post_date": "05/29/2025 02:22:43",
      "content": "<p>You mentioned that the end of the train data might better represent the test data, but since it's 1 year test, it might predict the first few months of the test better while performing worse for later test dates. So while the public score improves, the private score will drop.<br>\n this creates a classic overfitting scenario to the public leaderboard</p>",
      "votes": null,
      "replies": [
        {
          "id": 3211880,
          "author_name": "taylorsamarel",
          "author_url": "",
          "post_date": "05/29/2025 02:44:56",
          "content": "<p>It is possible that it could lead to over fitting depending upon the strength and curve of sample weight decay. I tried multiple methods of decaying weights along different curves and found it to be relatively successful when used in moderation. Using appropriate decay on the sample weights and not completely disregarding the value of the first rows provides a method to avoid over fitting.</p>\n<p>Since we don't have the chronological order of the test data, it becomes a game of striking a balance. If I can improve the first for the first 8 months of test data while slightly increase the errors on the remaining 4 months, on balance it still can produce better results. </p>\n<p>Outside of competitions on Kaggle I've seen dynamic sample weights be used successfully in real trading situations that are not time series oriented.</p>\n<p>Maybe I'm missing something here and would love to hear more thoughts and strategies.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3211886,
              "author_name": "onurkoc83",
              "author_url": "",
              "post_date": "05/29/2025 02:57:06",
              "content": "<p>If the test data were 1-2 month, I would agree with you :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3211903,
                  "author_name": "taylorsamarel",
                  "author_url": "",
                  "post_date": "05/29/2025 03:29:06",
                  "content": "<p>The great thing about Kaggle is we can just test it.</p>\n<p>I'm trying to put a comparison up here but currently have GPU tasks ongoing: <a href=\"https://www.kaggle.com/code/taylorsamarel/sample-weight-decay-normal-ensemble-comparison/\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/sample-weight-decay-normal-ensemble-comparison/</a></p>\n<p>The notebook is setup to compare 3 different variations (EDIT: I ran the notebook and added scores after each strategy):</p>\n<ol>\n<li>No sample weights; [Submission Score <strong>0.09852</strong>]</li>\n<li>With sample weights including a weight floor and decay rate; (it does admittedly get a little tricky with how weights are integrated into folds); [Submission Score <strong>0.09927</strong>]</li>\n<li>Ensemble based recency weighting by training two models, the first on 100% of the data, the second on the most recent 75% of the data and then combining those to get additional emphasis on more recent data. [Submission Score <strong>0.09991</strong>]</li>\n</ol>\n<p>From my submissions these methods seem to improve the scores on the test data.</p>\n<p>Also, just thinking real world here, near the second half of the train data a new crypto exchange could have launched, or liquidity algorithms change, or an exchange drops out and no longer operates. This creates a new market 'regime' that didn't exist in the beginning of the train set but may continue to exist in the test time period. Granted there are many assumptions in this theory, but it appears to work.</p>\n<p>One of the many anonymized features could capture things that better model market regimes but there is a lot of noise to sort through. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3211923,
                      "author_name": "onurkoc83",
                      "author_url": "",
                      "post_date": "05/29/2025 04:02:53",
                      "content": "<p>If the 1-year test dataset had been randomly split into public and private portions, I would consider the score improvement reliable. However, from what I understand, the public data consists of test data from the first 6 months, while the private data contains test data from the last 6 months. This means the private dataset represents a more recent time period.</p>\n<p>The goal is to achieve good performance on the private dataset, but this temporal split raises concerns about whether improvements on the public data will actually translate to success on the private data, since they come from different time periods rather than being randomly distributed</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3211937,
                          "author_name": "taylorsamarel",
                          "author_url": "",
                          "post_date": "05/29/2025 04:39:14",
                          "content": "<p>Ahh, this is an excellent point. I did not consider how the public and private test data is split. In fact, I had to go through the discussion and see your comments with <a href=\"https://www.kaggle.com/yw2735\" target=\"_blank\">@yw2735</a>.</p>\n<p>Considering that the private test data (51%) come after the public test data, the improvements I've been seeing may diminish, but there could also be situations where it performs better than expected, especially if that 51% deviates significantly from the records at the very beginning.</p>\n<p>Its an interesting situation, I wish we were given public test data that spans a larger time frame. Once the 51% comes into play scoring of notebooks will fluctuate a lot.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3212249,
                              "author_name": "yw2735",
                              "author_url": "",
                              "post_date": "05/29/2025 14:51:01",
                              "content": "<p>I agree with <a href=\"https://www.kaggle.com/onurkoc83\" target=\"_blank\">@onurkoc83</a> that this might cause overfitting in lb. But since we can choose two submissions, I may consider submitting one version with sample weights and one without.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3212263,
                                  "author_name": "taylorsamarel",
                                  "author_url": "",
                                  "post_date": "05/29/2025 15:13:29",
                                  "content": "<p>Probably a good strategy to try both for submission.</p>\n<p>I think the argument for overfitting goes both ways. Leaving earlier records with the same weight as the most recent records risks overfitting to the earlier records.</p>\n<p>And to be clear, the sample weight adjustments that I've played around with are very mild. For example, the most recent record having a weight of 1, and the first record having a weight of 0.9. Which comes out to a decay rate of nearly 0.0000002 per record - which some would say is very small.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3227849,
      "author_name": "ccyhui",
      "author_url": "",
      "post_date": "06/19/2025 11:05:42",
      "content": "<p>May I ask how the time weight is assigned to the test data? According to the data tab, the test data has been shuffled and masked, so the order is unknown.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3227964,
          "author_name": "taylorsamarel",
          "author_url": "",
          "post_date": "06/19/2025 13:50:11",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ccyhui\" target=\"_blank\">@ccyhui</a>, you're right - we can't assign time-based weights to the test data for exactly the reasons you mentioned.<br>\nSample weighting in the training data could potentially help, but only if:</p>\n<p>Higher-weighted samples better represent the test data distribution<br>\nIn this case, I suspect recent training data (near the end of the dataset) is more representative of test data patterns than earlier records and therefore deserves a higher sample weight when training.</p>\n<p>The trade-off is that this approach carries risks - it might overfit to recent patterns or miss important information from regime changes or structural shifts in the market.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3211845": "I think a few others have discovered this but although this is not a time series problem, I have discovered that using sample weights can help bias the model to put more weight on the records that occur at the end of the train dataset.\n\nThe theory is that the test dataset would logically come after the train dataset and market dynamics may have slightly changed between the beginning of the train dataset and the end. The end of the train dataset might be more representative of the test data.\n\nI know this probably isn't ground breaking but I've been able to implement it to improve a few different model architectures with success.",
    "3211871": "You mentioned that the end of the train data might better represent the test data, but since it's 1 year test, it might predict the first few months of the test better while performing worse for later test dates. So while the public score improves, the private score will drop.\n this creates a classic overfitting scenario to the public leaderboard",
    "3211880": "It is possible that it could lead to over fitting depending upon the strength and curve of sample weight decay. I tried multiple methods of decaying weights along different curves and found it to be relatively successful when used in moderation. Using appropriate decay on the sample weights and not completely disregarding the value of the first rows provides a method to avoid over fitting.\n\nSince we don't have the chronological order of the test data, it becomes a game of striking a balance. If I can improve the first for the first 8 months of test data while slightly increase the errors on the remaining 4 months, on balance it still can produce better results. \n\nOutside of competitions on Kaggle I've seen dynamic sample weights be used successfully in real trading situations that are not time series oriented.\n\nMaybe I'm missing something here and would love to hear more thoughts and strategies.",
    "3211886": "If the test data were 1-2 month, I would agree with you :)",
    "3211903": "The great thing about Kaggle is we can just test it.\n\nI'm trying to put a comparison up here but currently have GPU tasks ongoing: https://www.kaggle.com/code/taylorsamarel/sample-weight-decay-normal-ensemble-comparison/\n\nThe notebook is setup to compare 3 different variations (EDIT: I ran the notebook and added scores after each strategy):\n\n1. No sample weights; [Submission Score **0.09852**]\n2. With sample weights including a weight floor and decay rate; (it does admittedly get a little tricky with how weights are integrated into folds); [Submission Score **0.09927**]\n3. Ensemble based recency weighting by training two models, the first on 100% of the data, the second on the most recent 75% of the data and then combining those to get additional emphasis on more recent data. [Submission Score **0.09991**]\n\nFrom my submissions these methods seem to improve the scores on the test data.\n\nAlso, just thinking real world here, near the second half of the train data a new crypto exchange could have launched, or liquidity algorithms change, or an exchange drops out and no longer operates. This creates a new market 'regime' that didn't exist in the beginning of the train set but may continue to exist in the test time period. Granted there are many assumptions in this theory, but it appears to work.\n\nOne of the many anonymized features could capture things that better model market regimes but there is a lot of noise to sort through.",
    "3211923": "If the 1-year test dataset had been randomly split into public and private portions, I would consider the score improvement reliable. However, from what I understand, the public data consists of test data from the first 6 months, while the private data contains test data from the last 6 months. This means the private dataset represents a more recent time period.\n\nThe goal is to achieve good performance on the private dataset, but this temporal split raises concerns about whether improvements on the public data will actually translate to success on the private data, since they come from different time periods rather than being randomly distributed",
    "3211937": "Ahh, this is an excellent point. I did not consider how the public and private test data is split. In fact, I had to go through the discussion and see your comments with @yw2735.\n\nConsidering that the private test data (51%) come after the public test data, the improvements I've been seeing may diminish, but there could also be situations where it performs better than expected, especially if that 51% deviates significantly from the records at the very beginning.\n\nIts an interesting situation, I wish we were given public test data that spans a larger time frame. Once the 51% comes into play scoring of notebooks will fluctuate a lot.",
    "3212249": "I agree with @onurkoc83 that this might cause overfitting in lb. But since we can choose two submissions, I may consider submitting one version with sample weights and one without.",
    "3212263": "Probably a good strategy to try both for submission.\n\nI think the argument for overfitting goes both ways. Leaving earlier records with the same weight as the most recent records risks overfitting to the earlier records.\n\nAnd to be clear, the sample weight adjustments that I've played around with are very mild. For example, the most recent record having a weight of 1, and the first record having a weight of 0.9. Which comes out to a decay rate of nearly 0.0000002 per record - which some would say is very small.",
    "3227849": "May I ask how the time weight is assigned to the test data? According to the data tab, the test data has been shuffled and masked, so the order is unknown.",
    "3227964": "Hi @ccyhui, you're right - we can't assign time-based weights to the test data for exactly the reasons you mentioned.\nSample weighting in the training data could potentially help, but only if:\n\nHigher-weighted samples better represent the test data distribution\nIn this case, I suspect recent training data (near the end of the dataset) is more representative of test data patterns than earlier records and therefore deserves a higher sample weight when training.\n\nThe trade-off is that this approach carries risks - it might overfit to recent patterns or miss important information from regime changes or structural shifts in the market."
  },
  "source": "meta"
}