{
  "id": 554730,
  "title": "NN Model had best performance when validated using K fold split: Why?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/554730",
  "author_name": "rulan",
  "post_date": "2025-01-03T01:14:17.434000",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>When I first began building my NN, I naively used Group K Fold to validate it. I took the best performing 3 models from 3 folds and ensembled them as a submission. The 3 NN ensemble scored a 0.0067 on LB.</p>\n<p>I realized that this was causing data leakage and switched over to using the purged time split. My 3 NN ensemble immediately became 0.0059 on the LB. After adding some features and optimizing some hyperparameters a bit, my 5 NN ensemble (each trained on a unique random seed) is 0.0065 LB.</p>\n<p>My only question is: why? Why did my group k fold NNs do the best on the LB despite that it is obviously validated with data leakage, compared to the properly validated NNs?</p>",
  "messages": [
    {
      "id": 3087198,
      "postDate": "2025-01-03T08:01:31.113Z",
      "content": "<p>That is because in this comp more recent data is much more important for the future data prediction (not surprising in non-stationary time series preidction, right?) - you can easily verify this. That is the reason why you have to leverage the online learning in this comp if you want to have a good final rank, because the gap between solutions with online learning and without online-learning will be even bigger in private leaderboard. In group-kfold setup, each of your models will be exposed to the more recent dataset, while in your purged time split that may not happen. So it's not surprising you will have better scores using group-kfold. But it also has its own problem, which is another topic. </p>",
      "rawMarkdown": "That is because in this comp more recent data is much more important for the future data prediction (not surprising in non-stationary time series preidction, right?) - you can easily verify this. That is the reason why you have to leverage the online learning in this comp if you want to have a good final rank, because the gap between solutions with online learning and without online-learning will be even bigger in private leaderboard. In group-kfold setup, each of your models will be exposed to the more recent dataset, while in your purged time split that may not happen. So it's not surprising you will have better scores using group-kfold. But it also has its own problem, which is another topic. ",
      "votes": 12,
      "replies": [
        {
          "id": 3087498,
          "postDate": "2025-01-03T14:42:11.987Z",
          "content": "<p>This makes sense, thank you. I did not expect that losing out on a million or so rows would result in such a large drop in LB score. </p>",
          "rawMarkdown": "This makes sense, thank you. I did not expect that losing out on a million or so rows would result in such a large drop in LB score. "
        },
        {
          "id": 3087762,
          "postDate": "2025-01-03T20:42:29.213Z",
          "content": "<p>Thank you for your insight. I was wondering can online learning make a big difference? We will have only the ground truth responder for the last timestamp of the day (i.e. the lagged values of the next day) so the amount of data for retraining is really small. I didn't get very high score. </p>",
          "rawMarkdown": "Thank you for your insight. I was wondering can online learning make a big difference? We will have only the ground truth responder for the last timestamp of the day (i.e. the lagged values of the next day) so the amount of data for retraining is really small. I didn't get very high score. ",
          "replies": [
            {
              "id": 3087776,
              "postDate": "2025-01-03T21:31:15.730Z",
              "content": "<p>OK. Here is the thing. From one of my early experiments, say if you train a model using data before 1339, without online learning, the performance scores for data vintages 1339-1458, 1459-1578, 1579-1698 are 0.0204, 0.0172, 0.004, while with online learning, the scores will be improved to 0.0213 (up 0.0009), 0.0233(up 0.0061), 0.0128(up 0.0088). Big or not, you decide.</p>",
              "rawMarkdown": "OK. Here is the thing. From one of my early experiments, say if you train a model using data before 1339, without online learning, the performance scores for data vintages 1339-1458, 1459-1578, 1579-1698 are 0.0204, 0.0172, 0.004, while with online learning, the scores will be improved to 0.0213 (up 0.0009), 0.0233(up 0.0061), 0.0128(up 0.0088). Big or not, you decide.",
              "votes": 7
            },
            {
              "id": 3087781,
              "postDate": "2025-01-03T21:43:40.683Z",
              "content": "<blockquote>\n  <p>We will have only the ground truth responder for the last timestamp of the day</p>\n</blockquote>\n<p>No, we have the true responder for <strong>all timestamps</strong> of the previous day. </p>",
              "rawMarkdown": ">We will have only the ground truth responder for the last timestamp of the day\n\nNo, we have the true responder for **all timestamps** of the previous day. ",
              "votes": 3
            },
            {
              "id": 3087839,
              "postDate": "2025-01-04T00:14:36.627Z",
              "content": "<p>Thank you <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>, I don't know if we're on the same line. Anyway, by looking to your PB score your point of view should be the correct one. On my side I got that we're provided with the responders of the previous day from the last timestamp but I think I misunderstood this concept. From your answer it looks that the responders we receive at the start of each day are their exact values for each timestamp of the previous day, this means that the whole batch of data is usable to retrain.</p>",
              "rawMarkdown": "Thank you @shiyili, I don't know if we're on the same line. Anyway, by looking to your PB score your point of view should be the correct one. On my side I got that we're provided with the responders of the previous day from the last timestamp but I think I misunderstood this concept. From your answer it looks that the responders we receive at the start of each day are their exact values for each timestamp of the previous day, this means that the whole batch of data is usable to retrain."
            },
            {
              "id": 3087845,
              "postDate": "2025-01-04T00:21:02.950Z",
              "content": "<p>If you have doubts about the lags, you can check these two kernel:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p><a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\" target=\"_blank\">https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api</a></p>\n<p>You can also check the source code of the api module, which is very straightforward.</p>",
              "rawMarkdown": "If you have doubts about the lags, you can check these two kernel:\n\nhttps://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\n\nhttps://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\n\nYou can also check the source code of the api module, which is very straightforward.",
              "votes": 1
            },
            {
              "id": 3088388,
              "postDate": "2025-01-04T15:53:04.610Z",
              "content": "<p>Thank you, great notebooks. The API of course can provide some insights but actually the content of the provided lags was not really clear and is not defined in the API. The data provided by the competition host for the lags contain a single time_id. The best would have been to provide a whole day with the exact structure of the data.</p>",
              "rawMarkdown": "Thank you, great notebooks. The API of course can provide some insights but actually the content of the provided lags was not really clear and is not defined in the API. The data provided by the competition host for the lags contain a single time_id. The best would have been to provide a whole day with the exact structure of the data."
            },
            {
              "id": 3088647,
              "postDate": "2025-01-05T00:39:24.430Z",
              "content": "<p>In addition I'm note sure about the second notebook shared: <a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\" target=\"_blank\">https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api</a></p>\n<p>There in particular it looks that the test and lags are passed to the function prediction for day N and N-1 but this should be incorrect.  Day and lags are passed at the start of each day for the same day for that day.</p>",
              "rawMarkdown": "In addition I'm note sure about the second notebook shared: https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\n\nThere in particular it looks that the test and lags are passed to the function prediction for day N and N-1 but this should be incorrect.  Day and lags are passed at the start of each day for the same day for that day."
            },
            {
              "id": 3089538,
              "postDate": "2025-01-06T07:26:12.903Z",
              "content": "<p>in online retrain, should we reponder_6_lag_1 as both feature and target? <br>\nor should we make use a delay method, that using lag 2 day as feature and lag 1 day as target? <br>\ncould you share some insight on this please? thank you</p>",
              "rawMarkdown": "in online retrain, should we reponder_6_lag_1 as both feature and target? \nor should we make use a delay method, that using lag 2 day as feature and lag 1 day as target? \ncould you share some insight on this please? thank you"
            },
            {
              "id": 3089607,
              "postDate": "2025-01-06T09:48:17.670Z",
              "content": "<p>r6-lag-1 should only be used as target. If you want to include r6 of the previous day as feature, you should include r6-lag-2 in your re-train dataset. </p>\n<p>Basically, during inference, you accumulate features for each time-id at date T until the end of the day. Then on the beginning of T+1, true targets of all time-ids for day T gets revealed ( i.e. r6-lag-1). Now you have features of day T in your accumulated cache, and the corresponding targets, you can train the model using data of day T. </p>\n<p>If you want to use the target value of the previous day as features, you should use targets of T-1 for the cache of day T, which is r6-lag-2. </p>",
              "rawMarkdown": "r6-lag-1 should only be used as target. If you want to include r6 of the previous day as feature, you should include r6-lag-2 in your re-train dataset. \n\nBasically, during inference, you accumulate features for each time-id at date T until the end of the day. Then on the beginning of T+1, true targets of all time-ids for day T gets revealed ( i.e. r6-lag-1). Now you have features of day T in your accumulated cache, and the corresponding targets, you can train the model using data of day T. \n\nIf you want to use the target value of the previous day as features, you should use targets of T-1 for the cache of day T, which is r6-lag-2. ",
              "votes": 2
            },
            {
              "id": 3089693,
              "postDate": "2025-01-06T12:20:34.050Z",
              "content": "<p>I don't know if this could help but I shared this notebook <a href=\"https://www.kaggle.com/code/simonedegasperis/online-retrain-poc\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/online-retrain-poc</a> for online retrain. In particular I'm creating a cache with the features and another one with the lags.<br>\nThen when i want to retrain I shift the lags 1 day back because we know that they are the ground truth of the previous day and I merge with features cache. I finally use this data to retrain a model. The shifted lag are going to be the labels.<br>\nThis simple example just wanted to be an easy POC of online retraining because I didn't find much on the public notebooks.</p>",
              "rawMarkdown": "I don't know if this could help but I shared this notebook https://www.kaggle.com/code/simonedegasperis/online-retrain-poc for online retrain. In particular I'm creating a cache with the features and another one with the lags.\nThen when i want to retrain I shift the lags 1 day back because we know that they are the ground truth of the previous day and I merge with features cache. I finally use this data to retrain a model. The shifted lag are going to be the labels.\nThis simple example just wanted to be an easy POC of online retraining because I didn't find much on the public notebooks.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3086998,
      "postDate": "2025-01-03T01:14:17.433Z",
      "content": "<p>When I first began building my NN, I naively used Group K Fold to validate it. I took the best performing 3 models from 3 folds and ensembled them as a submission. The 3 NN ensemble scored a 0.0067 on LB.</p>\n<p>I realized that this was causing data leakage and switched over to using the purged time split. My 3 NN ensemble immediately became 0.0059 on the LB. After adding some features and optimizing some hyperparameters a bit, my 5 NN ensemble (each trained on a unique random seed) is 0.0065 LB.</p>\n<p>My only question is: why? Why did my group k fold NNs do the best on the LB despite that it is obviously validated with data leakage, compared to the properly validated NNs?</p>",
      "rawMarkdown": "When I first began building my NN, I naively used Group K Fold to validate it. I took the best performing 3 models from 3 folds and ensembled them as a submission. The 3 NN ensemble scored a 0.0067 on LB.\n\nI realized that this was causing data leakage and switched over to using the purged time split. My 3 NN ensemble immediately became 0.0059 on the LB. After adding some features and optimizing some hyperparameters a bit, my 5 NN ensemble (each trained on a unique random seed) is 0.0065 LB.\n\nMy only question is: why? Why did my group k fold NNs do the best on the LB despite that it is obviously validated with data leakage, compared to the properly validated NNs?",
      "votes": 3
    },
    {
      "id": 3087187,
      "postDate": "2025-01-03T07:44:37.367Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3087198,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2025-01-03T08:01:31.113000",
      "content": "<p>That is because in this comp more recent data is much more important for the future data prediction (not surprising in non-stationary time series preidction, right?) - you can easily verify this. That is the reason why you have to leverage the online learning in this comp if you want to have a good final rank, because the gap between solutions with online learning and without online-learning will be even bigger in private leaderboard. In group-kfold setup, each of your models will be exposed to the more recent dataset, while in your purged time split that may not happen. So it's not surprising you will have better scores using group-kfold. But it also has its own problem, which is another topic. </p>",
      "votes": 12,
      "replies": [
        {
          "id": 3087498,
          "author_name": "rulan",
          "author_url": "",
          "post_date": "2025-01-03T14:42:11.987000",
          "content": "<p>This makes sense, thank you. I did not expect that losing out on a million or so rows would result in such a large drop in LB score. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3087762,
          "author_name": "Simone De Gasperis",
          "author_url": "",
          "post_date": "2025-01-03T20:42:29.213000",
          "content": "<p>Thank you for your insight. I was wondering can online learning make a big difference? We will have only the ground truth responder for the last timestamp of the day (i.e. the lagged values of the next day) so the amount of data for retraining is really small. I didn't get very high score. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3087776,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-03T21:31:15.730000",
              "content": "<p>OK. Here is the thing. From one of my early experiments, say if you train a model using data before 1339, without online learning, the performance scores for data vintages 1339-1458, 1459-1578, 1579-1698 are 0.0204, 0.0172, 0.004, while with online learning, the scores will be improved to 0.0213 (up 0.0009), 0.0233(up 0.0061), 0.0128(up 0.0088). Big or not, you decide.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3087781,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-03T21:43:40.683000",
              "content": "<blockquote>\n  <p>We will have only the ground truth responder for the last timestamp of the day</p>\n</blockquote>\n<p>No, we have the true responder for <strong>all timestamps</strong> of the previous day. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3087839,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-04T00:14:36.627000",
              "content": "<p>Thank you <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>, I don't know if we're on the same line. Anyway, by looking to your PB score your point of view should be the correct one. On my side I got that we're provided with the responders of the previous day from the last timestamp but I think I misunderstood this concept. From your answer it looks that the responders we receive at the start of each day are their exact values for each timestamp of the previous day, this means that the whole batch of data is usable to retrain.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3087845,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-04T00:21:02.950000",
              "content": "<p>If you have doubts about the lags, you can check these two kernel:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p><a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\" target=\"_blank\">https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api</a></p>\n<p>You can also check the source code of the api module, which is very straightforward.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3088388,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-04T15:53:04.610000",
              "content": "<p>Thank you, great notebooks. The API of course can provide some insights but actually the content of the provided lags was not really clear and is not defined in the API. The data provided by the competition host for the lags contain a single time_id. The best would have been to provide a whole day with the exact structure of the data.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3088647,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-05T00:39:24.430000",
              "content": "<p>In addition I'm note sure about the second notebook shared: <a href=\"https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api\" target=\"_blank\">https://www.kaggle.com/code/chumajin/janestreet-easy-to-understand-new-time-series-api</a></p>\n<p>There in particular it looks that the test and lags are passed to the function prediction for day N and N-1 but this should be incorrect.  Day and lags are passed at the start of each day for the same day for that day.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3089538,
              "author_name": "ZT",
              "author_url": "",
              "post_date": "2025-01-06T07:26:12.903000",
              "content": "<p>in online retrain, should we reponder_6_lag_1 as both feature and target? <br>\nor should we make use a delay method, that using lag 2 day as feature and lag 1 day as target? <br>\ncould you share some insight on this please? thank you</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3089607,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-06T09:48:17.670000",
              "content": "<p>r6-lag-1 should only be used as target. If you want to include r6 of the previous day as feature, you should include r6-lag-2 in your re-train dataset. </p>\n<p>Basically, during inference, you accumulate features for each time-id at date T until the end of the day. Then on the beginning of T+1, true targets of all time-ids for day T gets revealed ( i.e. r6-lag-1). Now you have features of day T in your accumulated cache, and the corresponding targets, you can train the model using data of day T. </p>\n<p>If you want to use the target value of the previous day as features, you should use targets of T-1 for the cache of day T, which is r6-lag-2. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3089693,
              "author_name": "Simone De Gasperis",
              "author_url": "",
              "post_date": "2025-01-06T12:20:34.050000",
              "content": "<p>I don't know if this could help but I shared this notebook <a href=\"https://www.kaggle.com/code/simonedegasperis/online-retrain-poc\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/online-retrain-poc</a> for online retrain. In particular I'm creating a cache with the features and another one with the lags.<br>\nThen when i want to retrain I shift the lags 1 day back because we know that they are the ground truth of the previous day and I merge with features cache. I finally use this data to retrain a model. The shifted lag are going to be the labels.<br>\nThis simple example just wanted to be an easy POC of online retraining because I didn't find much on the public notebooks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3087187,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-03T07:44:37.367000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3087198": "That is because in this comp more recent data is much more important for the future data prediction (not surprising in non-stationary time series preidction, right?) - you can easily verify this. That is the reason why you have to leverage the online learning in this comp if you want to have a good final rank, because the gap between solutions with online learning and without online-learning will be even bigger in private leaderboard. In group-kfold setup, each of your models will be exposed to the more recent dataset, while in your purged time split that may not happen. So it's not surprising you will have better scores using group-kfold. But it also has its own problem, which is another topic. ",
    "3086998": "When I first began building my NN, I naively used Group K Fold to validate it. I took the best performing 3 models from 3 folds and ensembled them as a submission. The 3 NN ensemble scored a 0.0067 on LB.\n\nI realized that this was causing data leakage and switched over to using the purged time split. My 3 NN ensemble immediately became 0.0059 on the LB. After adding some features and optimizing some hyperparameters a bit, my 5 NN ensemble (each trained on a unique random seed) is 0.0065 LB.\n\nMy only question is: why? Why did my group k fold NNs do the best on the LB despite that it is obviously validated with data leakage, compared to the properly validated NNs?",
    "3087187": ""
  }
}