{
  "id": 363939,
  "title": "Yes, We Can Use Test Data Leakage.",
  "url": "/competitions/otto-recommender-system/discussion/363939",
  "author_name": "",
  "post_date": "2022-11-03T19:39:41.162697400Z",
  "votes": 51,
  "comment_count": 17,
  "views": 0,
  "content": "<p>In this competition we are predicting test events occurring within a one week period. </p>\n<p>When we predict events occurring in first day of test data, we can use events occurring in the next 6 days of test data which are from the future to make better predictions. Original post asked \"Can we train our models using test data from the future?\" and the host replied below \"yes we can\".</p>\n<p>Using test data will boost our CV LB as shown in this notebook <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">here</a> where original notebook had LB 0.539 and adding test data to train achieves LB 0.541. </p>\n<p>Below is a diagram of train and test sessions from OTTO's GitHub <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">here</a>. And in the comments below, i post histograms of test session start times and test session end times.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/leak.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2016122",
      "postDate": "11/03/2022 19:39:41",
      "content": "<p>In this competition we are predicting test events occurring within a one week period. </p>\n<p>When we predict events occurring in first day of test data, we can use events occurring in the next 6 days of test data which are from the future to make better predictions. Original post asked \"Can we train our models using test data from the future?\" and the host replied below \"yes we can\".</p>\n<p>Using test data will boost our CV LB as shown in this notebook <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">here</a> where original notebook had LB 0.539 and adding test data to train achieves LB 0.541. </p>\n<p>Below is a diagram of train and test sessions from OTTO's GitHub <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">here</a>. And in the comments below, i post histograms of test session start times and test session end times.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/leak.png\" alt=\"\"></p>",
      "rawMarkdown": "In this competition we are predicting test events occurring within a one week period. \n\nWhen we predict events occurring in first day of test data, we can use events occurring in the next 6 days of test data which are from the future to make better predictions. Original post asked \"Can we train our models using test data from the future?\" and the host replied below \"yes we can\".\n\nUsing test data will boost our CV LB as shown in this notebook [here][1] where original notebook had LB 0.539 and adding test data to train achieves LB 0.541. \n\nBelow is a diagram of train and test sessions from OTTO's GitHub [here][2]. And in the comments below, i post histograms of test session start times and test session end times.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/leak.png)\n\n[1]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[2]: https://github.com/otto-de/recsys-dataset",
      "votes": null
    },
    {
      "id": "2016227",
      "postDate": "11/03/2022 21:01:09",
      "content": "<p>I don't see how this could be disallowed to be honest 🙂 Tracking this in solutions would be a nightmare, to make sure people didn't train on the test set (essentially, impossible). Plus this is not very different from pseudo-labeling, etc.</p>\n<p>That is a great question though <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Would be great if there could be an official confirmation that this is okay 🙂</p>",
      "rawMarkdown": "I don't see how this could be disallowed to be honest 🙂 Tracking this in solutions would be a nightmare, to make sure people didn't train on the test set (essentially, impossible). Plus this is not very different from pseudo-labeling, etc.\n\nThat is a great question though @cdeotte! Would be great if there could be an official confirmation that this is okay 🙂",
      "votes": null
    },
    {
      "id": "2016230",
      "postDate": "11/03/2022 21:05:54",
      "content": "<p>BTW I am not sure your phrasing is correct? It doesn't line up with my understanding -- maybe I am getting something wrong here.</p>\n<p>We are never predicting the first day of the test set for any of the events. We are always predicting the events that come after the truncated sessions in the test set (always day 7+) AFAIU.</p>\n<p>Would be great if someone could verify if that reasoning is right -- conceptually the situation is different from what you describe.</p>",
      "rawMarkdown": "BTW I am not sure your phrasing is correct? It doesn't line up with my understanding -- maybe I am getting something wrong here.\n\nWe are never predicting the first day of the test set for any of the events. We are always predicting the events that come after the truncated sessions in the test set (always day 7+) AFAIU.\n\nWould be great if someone could verify if that reasoning is right -- conceptually the situation is different from what you describe.",
      "votes": null
    },
    {
      "id": "2016265",
      "postDate": "11/03/2022 22:01:25",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for your question and congrats on the current 1st place! In practice, we would not be able to train our models on the truncated futures sessions, as you suggest. However, as <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> mentioned, it is impractical to prevent this for the scope of this competition, and you are allowed to use all the data we provided.</p>",
      "rawMarkdown": "Hey @cdeotte, thanks for your question and congrats on the current 1st place! In practice, we would not be able to train our models on the truncated futures sessions, as you suggest. However, as @radek1 mentioned, it is impractical to prevent this for the scope of this competition, and you are allowed to use all the data we provided.",
      "votes": null
    },
    {
      "id": "2016280",
      "postDate": "11/03/2022 22:31:15",
      "content": "<p>I see why you consider it impractical, I feel that this question was brought up because of developements in the <a href=\"http://www.recsyschallenge.com/2022/\" target=\"_blank\">RecSys Challenge 2022</a> where a strict rule on usage of test set was added: </p>\n<blockquote>\n  <p>When predicting, treat each test session independently of all other test sessions (i.e., when predicting for test session B, the model should not have any knowledge of test session A. Even if that came before it in terms of time-stamp)</p>\n</blockquote>\n<p>I had the same concern in a recently closed community competition where I posted a <a href=\"https://www.kaggle.com/competitions/what-card-should-i-select-next/discussion/353776\" target=\"_blank\">similar thread</a>.<br>\nThose kind of rules usually get in place so that the code is more realistically close to a real-world recommender system. </p>\n<p>But, for sure, enforcing this kind of rules requires a lot of work and in addition to that it is difficult to detect future data leakage (one could just use future data for hyperparameter tuning and no one could know).<br>\n( In recsys challenge 2022 i think around top 20-30 teams were required to submit \"reproducible code\" for their solutions, someone had a lot of work to perform for sure. Just going through  my own solution was a nightmare, I can't imagine how an external reader that validated the code must have felt. )</p>\n<p>Luckly we already have confirmation that it is possible to use future data. </p>\n<p>I just wanted to add some context to the discussion for future readers reference.</p>",
      "rawMarkdown": "I see why you consider it impractical, I feel that this question was brought up because of developements in the [RecSys Challenge 2022](http://www.recsyschallenge.com/2022/) where a strict rule on usage of test set was added: \n\n> When predicting, treat each test session independently of all other test sessions (i.e., when predicting for test session B, the model should not have any knowledge of test session A. Even if that came before it in terms of time-stamp)\n\nI had the same concern in a recently closed community competition where I posted a [similar thread](https://www.kaggle.com/competitions/what-card-should-i-select-next/discussion/353776).\nThose kind of rules usually get in place so that the code is more realistically close to a real-world recommender system. \n\nBut, for sure, enforcing this kind of rules requires a lot of work and in addition to that it is difficult to detect future data leakage (one could just use future data for hyperparameter tuning and no one could know).\n( In recsys challenge 2022 i think around top 20-30 teams were required to submit \"reproducible code\" for their solutions, someone had a lot of work to perform for sure. Just going through  my own solution was a nightmare, I can't imagine how an external reader that validated the code must have felt. )\n\nLuckly we already have confirmation that it is possible to use future data. \n\nI just wanted to add some context to the discussion for future readers reference.",
      "votes": null
    },
    {
      "id": "2016413",
      "postDate": "11/04/2022 01:37:09",
      "content": "<p>I added a diagram to my original post to illustrate. And below are histograms of the 1,671,803 test session end times and start times. We see that roughly 1/7 of the sessions end on Tuesday Aug 30th 2022 and then 6/7 of the test sessions start on the following 6 days. The median session length is 30 seconds.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist2.png\" alt=\"\"></p>",
      "rawMarkdown": "I added a diagram to my original post to illustrate. And below are histograms of the 1,671,803 test session end times and start times. We see that roughly 1/7 of the sessions end on Tuesday Aug 30th 2022 and then 6/7 of the test sessions start on the following 6 days. The median session length is 30 seconds.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist2.png)",
      "votes": null
    },
    {
      "id": "2016771",
      "postDate": "11/04/2022 08:51:30",
      "content": "<p>I really like your question <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I really don't know its exact answer. I have got some answer from comment section. I just wanted you to share the conclusive answer if possible. Thanks for sharing.</p>",
      "rawMarkdown": "I really like your question @cdeotte, I really don't know its exact answer. I have got some answer from comment section. I just wanted you to share the conclusive answer if possible. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2017211",
      "postDate": "11/04/2022 15:53:20",
      "content": "<p>Os dados para teste precisam ser publicados para que mais Cientistas e Analistas consigam realizar treinamentos e se tornarem profissionais melhores.</p>",
      "rawMarkdown": "Os dados para teste precisam ser publicados para que mais Cientistas e Analistas consigam realizar treinamentos e se tornarem profissionais melhores.",
      "votes": null
    },
    {
      "id": "2017432",
      "postDate": "11/04/2022 19:54:49",
      "content": "<p>Nice visualization, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 👌 But I agree with <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> that the phrasings regarding the evaluation are somewhat misleading. The task is neither to predict the 7+ day nor the next day of a truncated session. Instead, your task is to predict the next events of the randomly truncated sessions, which can range over the whole test period but never exceed it. I hope this clears up any misconceptions 😌</p>",
      "rawMarkdown": "Nice visualization, @cdeotte 👌 But I agree with @radek1 that the phrasings regarding the evaluation are somewhat misleading. The task is neither to predict the 7+ day nor the next day of a truncated session. Instead, your task is to predict the next events of the randomly truncated sessions, which can range over the whole test period but never exceed it. I hope this clears up any misconceptions 😌",
      "votes": null
    },
    {
      "id": "2017433",
      "postDate": "11/04/2022 19:55:18",
      "content": "<p>The entire dataset, including the untruncated test sessions, will be published once the competition is finalized.</p>",
      "rawMarkdown": "The entire dataset, including the untruncated test sessions, will be published once the competition is finalized.",
      "votes": null
    },
    {
      "id": "2017465",
      "postDate": "11/04/2022 20:31:48",
      "content": "<p>The host Philipp has said </p>\n<blockquote>\n  <p>you are allowed to use all the data we provided.</p>\n</blockquote>\n<p>in another comment. So, yes, we may train with test data.</p>",
      "rawMarkdown": "The host Philipp has said \n> you are allowed to use all the data we provided.\n\nin another comment. So, yes, we may train with test data.",
      "votes": null
    },
    {
      "id": "2017485",
      "postDate": "11/04/2022 20:49:20",
      "content": "<p>Thanks for clarification <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> . Good point, we are predicting events (not days). I updated my original post to use the word \"events\". Hopefully this together with your explanation makes evaluation clearer.</p>",
      "rawMarkdown": "Thanks for clarification @pnormann . Good point, we are predicting events (not days). I updated my original post to use the word \"events\". Hopefully this together with your explanation makes evaluation clearer.",
      "votes": null
    },
    {
      "id": "2072682",
      "postDate": "12/22/2022 10:46:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> <br>\nIs there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?</p>\n<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nWhy would you call your finding a data \"leakage\"? When training a model on a train set, and you want to make a prediction at a random datetime (still within train set's date range), there are going to be train sessions whose last actions took place X days before this reference datetime, and therefore you can create features that look at what other items were clicked/carted/ordered after the session's last action and before the reference datetime. This way you can incorporate information from the \"future\" of a session, but without cheating - it's not leakage. For example, if I made my last action as a customer 2 days ago, and I went on your website today, your product recommender could also look at what other items have been popular since 2 days ago.<br>\nUnless I have misunderstood on how you're using the Test data.</p>",
      "rawMarkdown": "Hi @pnormann \nIs there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?\n\nHi @cdeotte\nWhy would you call your finding a data \"leakage\"? When training a model on a train set, and you want to make a prediction at a random datetime (still within train set's date range), there are going to be train sessions whose last actions took place X days before this reference datetime, and therefore you can create features that look at what other items were clicked/carted/ordered after the session's last action and before the reference datetime. This way you can incorporate information from the \"future\" of a session, but without cheating - it's not leakage. For example, if I made my last action as a customer 2 days ago, and I went on your website today, your product recommender could also look at what other items have been popular since 2 days ago.\nUnless I have misunderstood on how you're using the Test data.",
      "votes": null
    },
    {
      "id": "2072890",
      "postDate": "12/22/2022 13:52:22",
      "content": "<p><a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a> </p>\n<blockquote>\n  <p>Is there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?</p>\n</blockquote>\n<p>There is no historical data for test users. The users in test data are the users from the week 5 who have no history data. They are not truncated. The train users are truncated. Any user with history data before week 5 is added to train data and their activity is truncated to weeks 1 thru 4.</p>\n<blockquote>\n  <p>Why would you call your finding a data \"leakage\"?</p>\n</blockquote>\n<p>Imagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.</p>",
      "rawMarkdown": "ikogias \n>Is there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?\n\nThere is no historical data for test users. The users in test data are the users from the week 5 who have no history data. They are not truncated. The train users are truncated. Any user with history data before week 5 is added to train data and their activity is truncated to weeks 1 thru 4.\n\n>Why would you call your finding a data \"leakage\"?\n\nImagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.",
      "votes": null
    },
    {
      "id": "2072936",
      "postDate": "12/22/2022 14:29:32",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>There is no historical data for test users. </p>\n</blockquote>\n<p>Didn't realise that, thanks.</p>\n<blockquote>\n  <p>Imagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.</p>\n</blockquote>\n<p>I agree with your interpretation given that OTTO has truncated the sessions. What I had in mind was that test sessions were intact and the \"time of prediction\" was the very end of the test set. Had they done it like this, there wouldn't be any leakage.. </p>",
      "rawMarkdown": "cdeotte \n\n>There is no historical data for test users. \n\nDidn't realise that, thanks.\n\n>Imagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.\n\nI agree with your interpretation given that OTTO has truncated the sessions. What I had in mind was that test sessions were intact and the \"time of prediction\" was the very end of the test set. Had they done it like this, there wouldn't be any leakage..",
      "votes": null
    },
    {
      "id": "2072955",
      "postDate": "12/22/2022 14:43:51",
      "content": "<p><a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a> The leakage is about using <code>test User B</code> to predict <code>test User A</code>. I agree with you if we only use information from <code>User A</code> to predict <code>User A</code> then there is no leakage. But since Kaggle gave us all the test data, we have info from 1.8 million test users. If I use info from the 1,799,999 test users other than the user i am predicting then it is leakage because in real life when we are predicting <code>User A</code> we do not have future information from <code>User B</code>, <code>User C</code> (whom have activity in the future), etc</p>",
      "rawMarkdown": "ikogias The leakage is about using `test User B` to predict `test User A`. I agree with you if we only use information from `User A` to predict `User A` then there is no leakage. But since Kaggle gave us all the test data, we have info from 1.8 million test users. If I use info from the 1,799,999 test users other than the user i am predicting then it is leakage because in real life when we are predicting `User A` we do not have future information from `User B`, `User C` (whom have activity in the future), etc",
      "votes": null
    },
    {
      "id": "2072959",
      "postDate": "12/22/2022 14:46:20",
      "content": "<p>Note we are not predicting events that occur in week 6 which take place after all test user sessions. We are predicting events taking place during week 5. So for User A we have what they did Monday and Tuesday and we need to predict Wednesday, Thursday, Friday.</p>\n<p>At the same time, we have User B with events from Monday, Tuesday, Wednesday, Thursday. And we need to predict what User B does on Friday.</p>\n<p>The leakage is using User B events from Wednesday and Thursday to predict User A events on Wednesday, Thursday, Friday.</p>",
      "rawMarkdown": "Note we are not predicting events that occur in week 6 which take place after all test user sessions. We are predicting events taking place during week 5. So for User A we have what they did Monday and Tuesday and we need to predict Wednesday, Thursday, Friday.\n\nAt the same time, we have User B with events from Monday, Tuesday, Wednesday, Thursday. And we need to predict what User B does on Friday.\n\nThe leakage is using User B events from Wednesday and Thursday to predict User A events on Wednesday, Thursday, Friday.",
      "votes": null
    },
    {
      "id": "2072982",
      "postDate": "12/22/2022 15:01:49",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Agreed. All I was saying was, had OTTO asked us to predict events that occur in week 6 (with week 5 test sessions intact) then there wouldn't be any possibility of data leakage.</p>",
      "rawMarkdown": "cdeotte Agreed. All I was saying was, had OTTO asked us to predict events that occur in week 6 (with week 5 test sessions intact) then there wouldn't be any possibility of data leakage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2016227,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/03/2022 21:01:09",
      "content": "<p>I don't see how this could be disallowed to be honest 🙂 Tracking this in solutions would be a nightmare, to make sure people didn't train on the test set (essentially, impossible). Plus this is not very different from pseudo-labeling, etc.</p>\n<p>That is a great question though <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Would be great if there could be an official confirmation that this is okay 🙂</p>",
      "votes": null,
      "replies": [
        {
          "id": 2016280,
          "author_name": "pietromaldini1",
          "author_url": "",
          "post_date": "11/03/2022 22:31:15",
          "content": "<p>I see why you consider it impractical, I feel that this question was brought up because of developements in the <a href=\"http://www.recsyschallenge.com/2022/\" target=\"_blank\">RecSys Challenge 2022</a> where a strict rule on usage of test set was added: </p>\n<blockquote>\n  <p>When predicting, treat each test session independently of all other test sessions (i.e., when predicting for test session B, the model should not have any knowledge of test session A. Even if that came before it in terms of time-stamp)</p>\n</blockquote>\n<p>I had the same concern in a recently closed community competition where I posted a <a href=\"https://www.kaggle.com/competitions/what-card-should-i-select-next/discussion/353776\" target=\"_blank\">similar thread</a>.<br>\nThose kind of rules usually get in place so that the code is more realistically close to a real-world recommender system. </p>\n<p>But, for sure, enforcing this kind of rules requires a lot of work and in addition to that it is difficult to detect future data leakage (one could just use future data for hyperparameter tuning and no one could know).<br>\n( In recsys challenge 2022 i think around top 20-30 teams were required to submit \"reproducible code\" for their solutions, someone had a lot of work to perform for sure. Just going through  my own solution was a nightmare, I can't imagine how an external reader that validated the code must have felt. )</p>\n<p>Luckly we already have confirmation that it is possible to use future data. </p>\n<p>I just wanted to add some context to the discussion for future readers reference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2016230,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/03/2022 21:05:54",
      "content": "<p>BTW I am not sure your phrasing is correct? It doesn't line up with my understanding -- maybe I am getting something wrong here.</p>\n<p>We are never predicting the first day of the test set for any of the events. We are always predicting the events that come after the truncated sessions in the test set (always day 7+) AFAIU.</p>\n<p>Would be great if someone could verify if that reasoning is right -- conceptually the situation is different from what you describe.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2016413,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/04/2022 01:37:09",
          "content": "<p>I added a diagram to my original post to illustrate. And below are histograms of the 1,671,803 test session end times and start times. We see that roughly 1/7 of the sessions end on Tuesday Aug 30th 2022 and then 6/7 of the test sessions start on the following 6 days. The median session length is 30 seconds.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist2.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017432,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/04/2022 19:54:49",
          "content": "<p>Nice visualization, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 👌 But I agree with <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> that the phrasings regarding the evaluation are somewhat misleading. The task is neither to predict the 7+ day nor the next day of a truncated session. Instead, your task is to predict the next events of the randomly truncated sessions, which can range over the whole test period but never exceed it. I hope this clears up any misconceptions 😌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017485,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/04/2022 20:49:20",
          "content": "<p>Thanks for clarification <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> . Good point, we are predicting events (not days). I updated my original post to use the word \"events\". Hopefully this together with your explanation makes evaluation clearer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2016265,
      "author_name": "pnormann",
      "author_url": "",
      "post_date": "11/03/2022 22:01:25",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for your question and congrats on the current 1st place! In practice, we would not be able to train our models on the truncated futures sessions, as you suggest. However, as <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> mentioned, it is impractical to prevent this for the scope of this competition, and you are allowed to use all the data we provided.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2016771,
      "author_name": "rahul029",
      "author_url": "",
      "post_date": "11/04/2022 08:51:30",
      "content": "<p>I really like your question <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I really don't know its exact answer. I have got some answer from comment section. I just wanted you to share the conclusive answer if possible. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2017465,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/04/2022 20:31:48",
          "content": "<p>The host Philipp has said </p>\n<blockquote>\n  <p>you are allowed to use all the data we provided.</p>\n</blockquote>\n<p>in another comment. So, yes, we may train with test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2017211,
      "author_name": "leandrocarrinho",
      "author_url": "",
      "post_date": "11/04/2022 15:53:20",
      "content": "<p>Os dados para teste precisam ser publicados para que mais Cientistas e Analistas consigam realizar treinamentos e se tornarem profissionais melhores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2017433,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/04/2022 19:55:18",
          "content": "<p>The entire dataset, including the untruncated test sessions, will be published once the competition is finalized.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2072682,
      "author_name": "ikogias",
      "author_url": "",
      "post_date": "12/22/2022 10:46:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> <br>\nIs there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?</p>\n<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nWhy would you call your finding a data \"leakage\"? When training a model on a train set, and you want to make a prediction at a random datetime (still within train set's date range), there are going to be train sessions whose last actions took place X days before this reference datetime, and therefore you can create features that look at what other items were clicked/carted/ordered after the session's last action and before the reference datetime. This way you can incorporate information from the \"future\" of a session, but without cheating - it's not leakage. For example, if I made my last action as a customer 2 days ago, and I went on your website today, your product recommender could also look at what other items have been popular since 2 days ago.<br>\nUnless I have misunderstood on how you're using the Test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2072890,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "12/22/2022 13:52:22",
          "content": "<p><a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a> </p>\n<blockquote>\n  <p>Is there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?</p>\n</blockquote>\n<p>There is no historical data for test users. The users in test data are the users from the week 5 who have no history data. They are not truncated. The train users are truncated. Any user with history data before week 5 is added to train data and their activity is truncated to weeks 1 thru 4.</p>\n<blockquote>\n  <p>Why would you call your finding a data \"leakage\"?</p>\n</blockquote>\n<p>Imagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2072936,
              "author_name": "ikogias",
              "author_url": "",
              "post_date": "12/22/2022 14:29:32",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>There is no historical data for test users. </p>\n</blockquote>\n<p>Didn't realise that, thanks.</p>\n<blockquote>\n  <p>Imagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.</p>\n</blockquote>\n<p>I agree with your interpretation given that OTTO has truncated the sessions. What I had in mind was that test sessions were intact and the \"time of prediction\" was the very end of the test set. Had they done it like this, there wouldn't be any leakage.. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2072955,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "12/22/2022 14:43:51",
                  "content": "<p><a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">@ikogias</a> The leakage is about using <code>test User B</code> to predict <code>test User A</code>. I agree with you if we only use information from <code>User A</code> to predict <code>User A</code> then there is no leakage. But since Kaggle gave us all the test data, we have info from 1.8 million test users. If I use info from the 1,799,999 test users other than the user i am predicting then it is leakage because in real life when we are predicting <code>User A</code> we do not have future information from <code>User B</code>, <code>User C</code> (whom have activity in the future), etc</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2072959,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "12/22/2022 14:46:20",
                  "content": "<p>Note we are not predicting events that occur in week 6 which take place after all test user sessions. We are predicting events taking place during week 5. So for User A we have what they did Monday and Tuesday and we need to predict Wednesday, Thursday, Friday.</p>\n<p>At the same time, we have User B with events from Monday, Tuesday, Wednesday, Thursday. And we need to predict what User B does on Friday.</p>\n<p>The leakage is using User B events from Wednesday and Thursday to predict User A events on Wednesday, Thursday, Friday.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2072982,
                      "author_name": "ikogias",
                      "author_url": "",
                      "post_date": "12/22/2022 15:01:49",
                      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Agreed. All I was saying was, had OTTO asked us to predict events that occur in week 6 (with week 5 test sessions intact) then there wouldn't be any possibility of data leakage.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2016122": "In this competition we are predicting test events occurring within a one week period. \n\nWhen we predict events occurring in first day of test data, we can use events occurring in the next 6 days of test data which are from the future to make better predictions. Original post asked \"Can we train our models using test data from the future?\" and the host replied below \"yes we can\".\n\nUsing test data will boost our CV LB as shown in this notebook [here][1] where original notebook had LB 0.539 and adding test data to train achieves LB 0.541. \n\nBelow is a diagram of train and test sessions from OTTO's GitHub [here][2]. And in the comments below, i post histograms of test session start times and test session end times.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/leak.png)\n\n[1]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[2]: https://github.com/otto-de/recsys-dataset",
    "2016227": "I don't see how this could be disallowed to be honest 🙂 Tracking this in solutions would be a nightmare, to make sure people didn't train on the test set (essentially, impossible). Plus this is not very different from pseudo-labeling, etc.\n\nThat is a great question though @cdeotte! Would be great if there could be an official confirmation that this is okay 🙂",
    "2016230": "BTW I am not sure your phrasing is correct? It doesn't line up with my understanding -- maybe I am getting something wrong here.\n\nWe are never predicting the first day of the test set for any of the events. We are always predicting the events that come after the truncated sessions in the test set (always day 7+) AFAIU.\n\nWould be great if someone could verify if that reasoning is right -- conceptually the situation is different from what you describe.",
    "2016265": "Hey @cdeotte, thanks for your question and congrats on the current 1st place! In practice, we would not be able to train our models on the truncated futures sessions, as you suggest. However, as @radek1 mentioned, it is impractical to prevent this for the scope of this competition, and you are allowed to use all the data we provided.",
    "2016280": "I see why you consider it impractical, I feel that this question was brought up because of developements in the [RecSys Challenge 2022](http://www.recsyschallenge.com/2022/) where a strict rule on usage of test set was added: \n\n> When predicting, treat each test session independently of all other test sessions (i.e., when predicting for test session B, the model should not have any knowledge of test session A. Even if that came before it in terms of time-stamp)\n\nI had the same concern in a recently closed community competition where I posted a [similar thread](https://www.kaggle.com/competitions/what-card-should-i-select-next/discussion/353776).\nThose kind of rules usually get in place so that the code is more realistically close to a real-world recommender system. \n\nBut, for sure, enforcing this kind of rules requires a lot of work and in addition to that it is difficult to detect future data leakage (one could just use future data for hyperparameter tuning and no one could know).\n( In recsys challenge 2022 i think around top 20-30 teams were required to submit \"reproducible code\" for their solutions, someone had a lot of work to perform for sure. Just going through  my own solution was a nightmare, I can't imagine how an external reader that validated the code must have felt. )\n\nLuckly we already have confirmation that it is possible to use future data. \n\nI just wanted to add some context to the discussion for future readers reference.",
    "2016413": "I added a diagram to my original post to illustrate. And below are histograms of the 1,671,803 test session end times and start times. We see that roughly 1/7 of the sessions end on Tuesday Aug 30th 2022 and then 6/7 of the test sessions start on the following 6 days. The median session length is 30 seconds.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/date_hist2.png)",
    "2016771": "I really like your question @cdeotte, I really don't know its exact answer. I have got some answer from comment section. I just wanted you to share the conclusive answer if possible. Thanks for sharing.",
    "2017211": "Os dados para teste precisam ser publicados para que mais Cientistas e Analistas consigam realizar treinamentos e se tornarem profissionais melhores.",
    "2017432": "Nice visualization, @cdeotte 👌 But I agree with @radek1 that the phrasings regarding the evaluation are somewhat misleading. The task is neither to predict the 7+ day nor the next day of a truncated session. Instead, your task is to predict the next events of the randomly truncated sessions, which can range over the whole test period but never exceed it. I hope this clears up any misconceptions 😌",
    "2017433": "The entire dataset, including the untruncated test sessions, will be published once the competition is finalized.",
    "2017465": "The host Philipp has said \n> you are allowed to use all the data we provided.\n\nin another comment. So, yes, we may train with test data.",
    "2017485": "Thanks for clarification @pnormann . Good point, we are predicting events (not days). I updated my original post to use the word \"events\". Hopefully this together with your explanation makes evaluation clearer.",
    "2072682": "Hi @pnormann \nIs there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?\n\nHi @cdeotte\nWhy would you call your finding a data \"leakage\"? When training a model on a train set, and you want to make a prediction at a random datetime (still within train set's date range), there are going to be train sessions whose last actions took place X days before this reference datetime, and therefore you can create features that look at what other items were clicked/carted/ordered after the session's last action and before the reference datetime. This way you can incorporate information from the \"future\" of a session, but without cheating - it's not leakage. For example, if I made my last action as a customer 2 days ago, and I went on your website today, your product recommender could also look at what other items have been popular since 2 days ago.\nUnless I have misunderstood on how you're using the Test data.",
    "2072890": "ikogias \n>Is there a reason why you haven't made available the historical data of the test sessions for dates covering the train set? Conscious that in a real use-case, we would have access to that data for inference, so what was your motivation for truncating the test sessions like this?\n\nThere is no historical data for test users. The users in test data are the users from the week 5 who have no history data. They are not truncated. The train users are truncated. Any user with history data before week 5 is added to train data and their activity is truncated to weeks 1 thru 4.\n\n>Why would you call your finding a data \"leakage\"?\n\nImagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.",
    "2072936": "cdeotte \n\n>There is no historical data for test users. \n\nDidn't realise that, thanks.\n\n>Imagine today is Wednesday and I need to make a recommendation for you. It is true that i can view your activity from Monday and Tuesday. In this competition, i can also view the future of Thursday and Friday (which has not occurred yet) and make you a recommendation of the most popular item from the future dates of Thursday and Friday.\n\nI agree with your interpretation given that OTTO has truncated the sessions. What I had in mind was that test sessions were intact and the \"time of prediction\" was the very end of the test set. Had they done it like this, there wouldn't be any leakage..",
    "2072955": "ikogias The leakage is about using `test User B` to predict `test User A`. I agree with you if we only use information from `User A` to predict `User A` then there is no leakage. But since Kaggle gave us all the test data, we have info from 1.8 million test users. If I use info from the 1,799,999 test users other than the user i am predicting then it is leakage because in real life when we are predicting `User A` we do not have future information from `User B`, `User C` (whom have activity in the future), etc",
    "2072959": "Note we are not predicting events that occur in week 6 which take place after all test user sessions. We are predicting events taking place during week 5. So for User A we have what they did Monday and Tuesday and we need to predict Wednesday, Thursday, Friday.\n\nAt the same time, we have User B with events from Monday, Tuesday, Wednesday, Thursday. And we need to predict what User B does on Friday.\n\nThe leakage is using User B events from Wednesday and Thursday to predict User A events on Wednesday, Thursday, Friday.",
    "2072982": "cdeotte Agreed. All I was saying was, had OTTO asked us to predict events that occur in week 6 (with week 5 test sessions intact) then there wouldn't be any possibility of data leakage."
  },
  "source": "meta"
}