{
  "id": 388479,
  "title": "[FIXED] - Is Test Data Leak Intentional or Accidental?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388479",
  "author_name": "",
  "post_date": "2023-02-17T17:24:07.638804700Z",
  "votes": 85,
  "comment_count": 44,
  "views": 0,
  "content": "<p>In Kaggle's API example notebook, all of <code>level_group = '0-4'</code> comes <strong>before</strong> <code>level_group = '5-12'</code> which comes <strong>before</strong> <code>level_group = '13-22'</code>. This is the <strong>correct</strong> ordering in time.</p>\n<p>During submission, I have noticed that the Kaggle API gives us the levels <strong>backward</strong>. First we see all <code>level_group = '13-22'</code>, then we see all <code>level_group = '5-12'</code>, then we see all <code>level_group = '0-4'</code>. This is <strong>incorrect</strong> ordering in time. This means we can build features from the future before predicting the present.</p>\n<p>Sohier, Maggie, Natalie, Alex, is this <strong>test data leak</strong> intentional or accidental?<br>\n<a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a></p>\n<h1>UPDATE</h1>\n<p>This time leak has been fixed. Thank you Kaggle staff for correction</p>",
  "messages": [
    {
      "id": "2148806",
      "postDate": "02/17/2023 17:24:07",
      "content": "<p>In Kaggle's API example notebook, all of <code>level_group = '0-4'</code> comes <strong>before</strong> <code>level_group = '5-12'</code> which comes <strong>before</strong> <code>level_group = '13-22'</code>. This is the <strong>correct</strong> ordering in time.</p>\n<p>During submission, I have noticed that the Kaggle API gives us the levels <strong>backward</strong>. First we see all <code>level_group = '13-22'</code>, then we see all <code>level_group = '5-12'</code>, then we see all <code>level_group = '0-4'</code>. This is <strong>incorrect</strong> ordering in time. This means we can build features from the future before predicting the present.</p>\n<p>Sohier, Maggie, Natalie, Alex, is this <strong>test data leak</strong> intentional or accidental?<br>\n<a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a></p>\n<h1>UPDATE</h1>\n<p>This time leak has been fixed. Thank you Kaggle staff for correction</p>",
      "rawMarkdown": "In Kaggle's API example notebook, all of `level_group = '0-4'` comes **before** `level_group = '5-12'` which comes **before** `level_group = '13-22'`. This is the **correct** ordering in time.\n\nDuring submission, I have noticed that the Kaggle API gives us the levels **backward**. First we see all `level_group = '13-22'`, then we see all `level_group = '5-12'`, then we see all `level_group = '0-4'`. This is **incorrect** ordering in time. This means we can build features from the future before predicting the present.\n\nSohier, Maggie, Natalie, Alex, is this **test data leak** intentional or accidental?\n@nrambis @alexmlfranklin @sohier @maggiemd\n\n# UPDATE\nThis time leak has been fixed. Thank you Kaggle staff for correction",
      "votes": null
    },
    {
      "id": "2148847",
      "postDate": "02/17/2023 17:58:15",
      "content": "<p>Wow this is new! How do you know that the Kaggle API gives data in such order? I have a lot of questions about the the Kaggle API's data fetching process (batch size, run-time, etc.) but don't know where to find the details. </p>",
      "rawMarkdown": "Wow this is new! How do you know that the Kaggle API gives data in such order? I have a lot of questions about the the Kaggle API's data fetching process (batch size, run-time, etc.) but don't know where to find the details.",
      "votes": null
    },
    {
      "id": "2148855",
      "postDate": "02/17/2023 18:03:59",
      "content": "<p>Thanks Chris! We'll take a look.</p>",
      "rawMarkdown": "Thanks Chris! We'll take a look.",
      "votes": null
    },
    {
      "id": "2148988",
      "postDate": "02/17/2023 20:09:04",
      "content": "<p>Thanks for pointing it out Chris. I am curious, is your best score using the leak ?</p>",
      "rawMarkdown": "Thanks for pointing it out Chris. I am curious, is your best score using the leak ?",
      "votes": null
    },
    {
      "id": "2148992",
      "postDate": "02/17/2023 20:12:32",
      "content": "<p>Yes. My current LB of 703 has CV 704 and uses information from the future.</p>\n<p>My best CV without leak is 695. My best CV with leak is 707.  I have not submitted these two models to the LB yet. (Still debugging some submit errors).</p>",
      "rawMarkdown": "Yes. My current LB of 703 has CV 704 and uses information from the future.\n\nMy best CV without leak is 695. My best CV with leak is 707.  I have not submitted these two models to the LB yet. (Still debugging some submit errors).",
      "votes": null
    },
    {
      "id": "2149018",
      "postDate": "02/17/2023 20:51:53",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> We've checked the data and API inputs and we're having trouble seeing what could be causing this. Would you mind sharing an example? Must be something we're missing.</p>",
      "rawMarkdown": "cdeotte We've checked the data and API inputs and we're having trouble seeing what could be causing this. Would you mind sharing an example? Must be something we're missing.",
      "votes": null
    },
    {
      "id": "2149020",
      "postDate": "02/17/2023 21:00:41",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> The simplest example is <a href=\"https://www.kaggle.com/code/cdeotte/lb-leak-probe\" target=\"_blank\">here</a>. This submission notebook by default will submit all zeros and score LB 0.226. This submission remembers whenever it sees <code>level_group = '5-12'</code> and <code>level_group = '13-22'</code>. The first time it observes one leak (i.e. seeing a user <code>level_group = '0-4'</code> <strong>AFTER</strong> it sees either <code>5-12</code> or <code>13-22</code>), it starts predicting 1s thereafter. We see that this notebook scores LB 0.414 which proves that is <strong>observes at least 1 leak</strong> on public LB. (because LB 0.414 is the LB score of submitting all 1s. Furthermore we know the leak occurs very early since we achieved full LB 0.414)</p>\n<p>The reason i know that all the (public LB) data is leaked is because I have a trained an XGB model which uses features from the future. The features from the future change both my CV by +0.012 and LB by +0.012. This indicates that every single user is leaking. (because if only some users were leaking then LB would not boost the same as CV. Note in this competition my CV and LB are always exactly the same).</p>",
      "rawMarkdown": "philculliton The simplest example is [here][1]. This submission notebook by default will submit all zeros and score LB 0.226. This submission remembers whenever it sees `level_group = '5-12'` and `level_group = '13-22'`. The first time it observes one leak (i.e. seeing a user `level_group = '0-4'` **AFTER** it sees either `5-12` or `13-22`), it starts predicting 1s thereafter. We see that this notebook scores LB 0.414 which proves that is **observes at least 1 leak** on public LB. (because LB 0.414 is the LB score of submitting all 1s. Furthermore we know the leak occurs very early since we achieved full LB 0.414)\n\nThe reason i know that all the (public LB) data is leaked is because I have a trained an XGB model which uses features from the future. The features from the future change both my CV by +0.012 and LB by +0.012. This indicates that every single user is leaking. (because if only some users were leaking then LB would not boost the same as CV. Note in this competition my CV and LB are always exactly the same).\n\n[1]: https://www.kaggle.com/code/cdeotte/lb-leak-probe",
      "votes": null
    },
    {
      "id": "2149025",
      "postDate": "02/17/2023 21:05:25",
      "content": "<p>UPDATE: My analysis above only applies to public LB. We do not know what is occurring on private LB.</p>",
      "rawMarkdown": "UPDATE: My analysis above only applies to public LB. We do not know what is occurring on private LB.",
      "votes": null
    },
    {
      "id": "2149043",
      "postDate": "02/17/2023 21:48:01",
      "content": "<p>A heads-up to all: we are confirming this issue with Chris' help. Once confirmed we will be fixing it ASAP; the hidden test set is not being leaked in any way, and there will be no long-term upside to doing this in your own notebooks. Please save your time for the competition's task!</p>\n<p>Thanks to Chris and others for pointing this out and helping us resolve it!</p>",
      "rawMarkdown": "A heads-up to all: we are confirming this issue with Chris' help. Once confirmed we will be fixing it ASAP; the hidden test set is not being leaked in any way, and there will be no long-term upside to doing this in your own notebooks. Please save your time for the competition's task!\n\nThanks to Chris and others for pointing this out and helping us resolve it!",
      "votes": null
    },
    {
      "id": "2149081",
      "postDate": "02/17/2023 23:25:39",
      "content": "<p>My experience supports the hypothesis posted by Chris: notebooks assuming the correct ordering (first, '0-4, then '5-12', then '13'-22) throw an error.</p>\n<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , this was driving me crazy. I used like 10 submissions just to debug my code, only to find out about this.</p>\n<p>Btw it is interesting to see that such a leak does not improve CV and LB by a lot (in absolute terms). I would expect it to be more.</p>",
      "rawMarkdown": "My experience supports the hypothesis posted by Chris: notebooks assuming the correct ordering (first, '0-4, then '5-12', then '13'-22) throw an error.\n\nThanks @cdeotte , this was driving me crazy. I used like 10 submissions just to debug my code, only to find out about this.\n\nBtw it is interesting to see that such a leak does not improve CV and LB by a lot (in absolute terms). I would expect it to be more.",
      "votes": null
    },
    {
      "id": "2149090",
      "postDate": "02/17/2023 23:43:47",
      "content": "<p>I was also frustrated trying to save data in the correct order. I had many failed submissions the first week and could not understand why. Eventually I concluded that the data was in random order. Then i trained 3 models, one forward, one backward, and one without forward or backward. During submission, I would detect which of the 3 scenarios occurs for each user (for each level_group) and then use the corresponding model for inference. My submission LB score matched the CV for the backward model which made me conclude that 100% of the data is backward.</p>",
      "rawMarkdown": "I was also frustrated trying to save data in the correct order. I had many failed submissions the first week and could not understand why. Eventually I concluded that the data was in random order. Then i trained 3 models, one forward, one backward, and one without forward or backward. During submission, I would detect which of the 3 scenarios occurs for each user (for each level_group) and then use the corresponding model for inference. My submission LB score matched the CV for the backward model which made me conclude that 100% of the data is backward.",
      "votes": null
    },
    {
      "id": "2149126",
      "postDate": "02/18/2023 00:53:27",
      "content": "<p>Great stuff &amp; impressive investigation. thank you for sharing this</p>",
      "rawMarkdown": "Great stuff & impressive investigation. thank you for sharing this",
      "votes": null
    },
    {
      "id": "2149138",
      "postDate": "02/18/2023 01:47:30",
      "content": "<p>Thanks for pointing it out, Chris.</p>",
      "rawMarkdown": "Thanks for pointing it out, Chris.",
      "votes": null
    },
    {
      "id": "2149163",
      "postDate": "02/18/2023 02:49:44",
      "content": "<p>Thanks for the awsome investigation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>\n<p>I have quite similar experience with you regarding leak info.</p>\n<p>In my case, by just using one future feature, can boost the score by 0.01</p>\n<p>It is the time gap between current level group and the subsequent one.</p>\n<p>I guess it is the question answering period truncated by the host, a longer period implies more trials and a poorer accuracy.</p>\n<p>Time gap between current  level group and the previous one does not use future info, but only boosts the score a little.</p>",
      "rawMarkdown": "Thanks for the awsome investigation @cdeotte.\n\nI have quite similar experience with you regarding leak info.\n\nIn my case, by just using one future feature, can boost the score by 0.01\n\nIt is the time gap between current level group and the subsequent one.\n\nI guess it is the question answering period truncated by the host, a longer period implies more trials and a poorer accuracy.\n\nTime gap between current  level group and the previous one does not use future info, but only boosts the score a little.",
      "votes": null
    },
    {
      "id": "2149166",
      "postDate": "02/18/2023 03:01:52",
      "content": "<p>Problem should be fixed. We will rerun all notebooks pre-fix at the start of next week, when we have engineering support on hand.</p>\n<p>Thanks very much for your help, Chris! We'll be reaching out privately.</p>\n<p>Thanks all! Please let us know if you all run across any additional issues, or anything that seems related.</p>",
      "rawMarkdown": "Problem should be fixed. We will rerun all notebooks pre-fix at the start of next week, when we have engineering support on hand.\n\nThanks very much for your help, Chris! We'll be reaching out privately.\n\nThanks all! Please let us know if you all run across any additional issues, or anything that seems related.",
      "votes": null
    },
    {
      "id": "2149387",
      "postDate": "02/18/2023 08:44:27",
      "content": "<p>This is a brilliant example of \"reverse engineering\"! Thanks for pointing it out!</p>",
      "rawMarkdown": "This is a brilliant example of \"reverse engineering\"! Thanks for pointing it out!",
      "votes": null
    },
    {
      "id": "2149590",
      "postDate": "02/18/2023 13:04:50",
      "content": "<p>Thanks a lot for fixing this Phil. And Chris for reporting!<br>\nExactly what is the fix? People should not be needing to guess what they need to predict.</p>\n<p>Can you share the exact specifications of how the data would be presented?<br>\nIdeally if using data from the past is allowed it could also be provided right away.</p>",
      "rawMarkdown": "Thanks a lot for fixing this Phil. And Chris for reporting!\nExactly what is the fix? People should not be needing to guess what they need to predict.\n\nCan you share the exact specifications of how the data would be presented?\nIdeally if using data from the past is allowed it could also be provided right away.",
      "votes": null
    },
    {
      "id": "2149954",
      "postDate": "02/18/2023 20:03:00",
      "content": "<p>Great, thanks a lot!</p>",
      "rawMarkdown": "Great, thanks a lot!",
      "votes": null
    },
    {
      "id": "2149959",
      "postDate": "02/18/2023 20:17:28",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> How does the Kaggle API handle users who repeat <code>level_group = '0-4'</code> or <code>5-12</code> after they finish later <code>level_groups</code>. See plot below. There are 2% train users (i.e. 229 out of 11779) who do this. Assuming 2% of test users do this also, does the Kaggle API still give us all of <code>level_group = '0-4'</code> information to predict questions 1 thru 3. If so, then the Kaggle API is giving us the future to help predict the past. We can save the future level 0-4 information and use it to help predict questions 4-18 from <code>level_groups '5-12' and '13-22'</code>. </p>\n<p>For example, one specific way to use this information is we can compute the time gap size between the two <code>level_group = '0-4'</code> sessions, then subtract the active time spent on <code>level_group = '5-12' and '13-22'</code> and deduce how much time they spent answering questions after <code>level_group = '5-12' and '13-22'</code>. Knowing how long users spend during the question answering portion can achieve at least <code>+0.010</code> boost to CV and LB (from prior time leak experiments).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/leak_user2.png\" alt=\"\"></p>",
      "rawMarkdown": "philculliton How does the Kaggle API handle users who repeat `level_group = '0-4'` or `5-12` after they finish later `level_groups`. See plot below. There are 2% train users (i.e. 229 out of 11779) who do this. Assuming 2% of test users do this also, does the Kaggle API still give us all of `level_group = '0-4'` information to predict questions 1 thru 3. If so, then the Kaggle API is giving us the future to help predict the past. We can save the future level 0-4 information and use it to help predict questions 4-18 from `level_groups '5-12' and '13-22'`. \n\nFor example, one specific way to use this information is we can compute the time gap size between the two `level_group = '0-4'` sessions, then subtract the active time spent on `level_group = '5-12' and '13-22'` and deduce how much time they spent answering questions after `level_group = '5-12' and '13-22'`. Knowing how long users spend during the question answering portion can achieve at least `+0.010` boost to CV and LB (from prior time leak experiments).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/leak_user2.png)",
      "votes": null
    },
    {
      "id": "2150575",
      "postDate": "02/19/2023 11:59:17",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!",
      "votes": null
    },
    {
      "id": "2153137",
      "postDate": "02/21/2023 08:13:21",
      "content": "<p>Hi Chris good luck</p>",
      "rawMarkdown": "Hi Chris good luck",
      "votes": null
    },
    {
      "id": "2153727",
      "postDate": "02/21/2023 15:46:11",
      "content": "<p>Hi, yes! The fix was ensuring that the ordering of the incoming data was correct - 0-4, 5-12, 13-22. This is the way the data is presented; you use the 0-4 data to predict question correctness at the end of that segment, and so on.</p>\n<p>You are allowed to use data from previous segments in a session. It's not possible with the API to present it all at once.</p>\n<p>We'll update the Data / Evaluation pages to make this clearer.</p>",
      "rawMarkdown": "Hi, yes! The fix was ensuring that the ordering of the incoming data was correct - 0-4, 5-12, 13-22. This is the way the data is presented; you use the 0-4 data to predict question correctness at the end of that segment, and so on.\n\nYou are allowed to use data from previous segments in a session. It's not possible with the API to present it all at once.\n\nWe'll update the Data / Evaluation pages to make this clearer.",
      "votes": null
    },
    {
      "id": "2153957",
      "postDate": "02/21/2023 18:27:07",
      "content": "<p>Phil, have all notebooks/scores been rerun yet with fixed lb?</p>",
      "rawMarkdown": "Phil, have all notebooks/scores been rerun yet with fixed lb?",
      "votes": null
    },
    {
      "id": "2155796",
      "postDate": "02/22/2023 22:54:54",
      "content": "<p>probably no since Chris still has 0.703</p>",
      "rawMarkdown": "probably no since Chris still has 0.703",
      "votes": null
    },
    {
      "id": "2157671",
      "postDate": "02/24/2023 08:08:51",
      "content": "<p>Hi, Chris. </p>\n<p>It seems we still don't get answer for this question, right?</p>",
      "rawMarkdown": "Hi, Chris. \n\nIt seems we still don't get answer for this question, right?",
      "votes": null
    },
    {
      "id": "2157688",
      "postDate": "02/24/2023 08:24:55",
      "content": "<p>Correct. We did not receive an answer. <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> How does the Kaggle API handle these 2% of data? Does the API give us a time leak, or are these 2% of users removed from test data, or are these 2% of users truncated and only provide the first visit to each level group?</p>",
      "rawMarkdown": "Correct. We did not receive an answer. @philculliton How does the Kaggle API handle these 2% of data? Does the API give us a time leak, or are these 2% of users removed from test data, or are these 2% of users truncated and only provide the first visit to each level group?",
      "votes": null
    },
    {
      "id": "2160628",
      "postDate": "02/26/2023 20:07:46",
      "content": "<p>Are there any updates?</p>",
      "rawMarkdown": "Are there any updates?",
      "votes": null
    },
    {
      "id": "2160813",
      "postDate": "02/27/2023 02:21:24",
      "content": "<p>So the bug is fixed or still have some new leaks?😂</p>",
      "rawMarkdown": "So the bug is fixed or still have some new leaks?😂",
      "votes": null
    },
    {
      "id": "2160857",
      "postDate": "02/27/2023 03:25:15",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> hello? :D</p>",
      "rawMarkdown": "philculliton hello? :D",
      "votes": null
    },
    {
      "id": "2165769",
      "postDate": "03/02/2023 12:49:05",
      "content": "<p>I have the same question. <br>\nIt is not clear for me if the fix has been done already or not.</p>",
      "rawMarkdown": "I have the same question. \nIt is not clear for me if the fix has been done already or not.",
      "votes": null
    },
    {
      "id": "2180189",
      "postDate": "03/13/2023 16:40:09",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> could you make it already clear about reruns pls? were they done?</p>",
      "rawMarkdown": "philculliton @sohier @maggiemd could you make it already clear about reruns pls? were they done?",
      "votes": null
    },
    {
      "id": "2183777",
      "postDate": "03/15/2023 22:37:13",
      "content": "<p>Am I a ghost already? :(</p>",
      "rawMarkdown": "Am I a ghost already? :(",
      "votes": null
    },
    {
      "id": "2183783",
      "postDate": "03/15/2023 22:51:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kvlmll\" target=\"_blank\">@kvlmll</a> my understanding is that there have been <strong>no reruns</strong> (i.e. updates to Kaggles' LB scores but the API has been fixed around Feb 15th). None of my previous leaky submissions nor leaky LB score results have been updated. Furthermore, i posted a leak notebook <a href=\"https://www.kaggle.com/code/cdeotte/lb-leak-probe\" target=\"_blank\">here</a> version 1 which scores LB 0.414 using leak. If this notebook is rescored, it should score LB 0.226</p>\n<p>Additionally it is my understanding that all submissions now (Feb 15th onward) and moving forward cannot access the reverse order leak (but can maybe access the 2% leak). Note that 1st place team \"French Touch\" submitted after the update. So I believe that \"French Touch\" LB score of 0.708 is legitimate and does <strong>not use</strong> the leak. </p>\n<p>My LB score of 0.703 was submitted before the update and does <strong>use</strong> the reverse order leak. Without leak, my LB score is only LB 0.692</p>",
      "rawMarkdown": "Hi @kvlmll my understanding is that there have been **no reruns** (i.e. updates to Kaggles' LB scores but the API has been fixed around Feb 15th). None of my previous leaky submissions nor leaky LB score results have been updated. Furthermore, i posted a leak notebook [here][1] version 1 which scores LB 0.414 using leak. If this notebook is rescored, it should score LB 0.226\n\nAdditionally it is my understanding that all submissions now (Feb 15th onward) and moving forward cannot access the reverse order leak (but can maybe access the 2% leak). Note that 1st place team \"French Touch\" submitted after the update. So I believe that \"French Touch\" LB score of 0.708 is legitimate and does **not use** the leak. \n  \nMy LB score of 0.703 was submitted before the update and does **use** the reverse order leak. Without leak, my LB score is only LB 0.692\n\n[1]: https://www.kaggle.com/code/cdeotte/lb-leak-probe",
      "votes": null
    },
    {
      "id": "2187223",
      "postDate": "03/18/2023 14:07:21",
      "content": "<p>0.708 is really a high score that I can't fetch.</p>",
      "rawMarkdown": "0.708 is really a high score that I can't fetch.",
      "votes": null
    },
    {
      "id": "2187285",
      "postDate": "03/18/2023 15:38:51",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for your great investigation. I have just 1 question and hope you can help. You said \"My LB score of 0.703 does use the leak.\". To my understanding your 0.703 score is AFTER Kaggle fixed the incorrect API order issue. So my question is that: do you mean the \"leak\" of 0.703 sub is just from 2% of the users with repeated session_group?</p>",
      "rawMarkdown": "cdeotte Thanks for your great investigation. I have just 1 question and hope you can help. You said \"My LB score of 0.703 does use the leak.\". To my understanding your 0.703 score is AFTER Kaggle fixed the incorrect API order issue. So my question is that: do you mean the \"leak\" of 0.703 sub is just from 2% of the users with repeated session_group?",
      "votes": null
    },
    {
      "id": "2187289",
      "postDate": "03/18/2023 15:43:05",
      "content": "<p>My LB 0.703 was made around Feb 8th, one week <strong>BEFORE</strong> Kaggle fixed the leak (and uses the full reverse order leak). I have not tried exploiting the 2% leak.</p>",
      "rawMarkdown": "My LB 0.703 was made around Feb 8th, one week **BEFORE** Kaggle fixed the leak (and uses the full reverse order leak). I have not tried exploiting the 2% leak.",
      "votes": null
    },
    {
      "id": "2187387",
      "postDate": "03/18/2023 17:27:22",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . That's strange that Kaggle didn't RERUN all our submissions! So the current public LB score is invalid (because your 0.703 can't be there anymore). But if they don't rerun it, the private score will STILL BE LEAKY AND GOOD after private LB is revealed? I think <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> may have to provide us an answer?</p>",
      "rawMarkdown": "thanks @cdeotte . That's strange that Kaggle didn't RERUN all our submissions! So the current public LB score is invalid (because your 0.703 can't be there anymore). But if they don't rerun it, the private score will STILL BE LEAKY AND GOOD after private LB is revealed? I think @philculliton may have to provide us an answer?",
      "votes": null
    },
    {
      "id": "2187395",
      "postDate": "03/18/2023 17:31:54",
      "content": "<p>I suppose they will do rerun anyway of course</p>",
      "rawMarkdown": "I suppose they will do rerun anyway of course",
      "votes": null
    },
    {
      "id": "2187410",
      "postDate": "03/18/2023 17:40:55",
      "content": "<p>My guess is that i am the only LB submission with a leak. Nobody knew about the leak when i made it public. And all other teams had LB 0.693 and below on Feb 15th after the update.</p>\n<p>After the update, teams starting increasing their LB. I believe all progress regarding submissions greater than LB 0.693 is by using information from previous levels. Before Kaggle fixed the leak, nobody could use information from previous levels.</p>",
      "rawMarkdown": "My guess is that i am the only LB submission with a leak. Nobody knew about the leak when i made it public. And all other teams had LB 0.693 and below on Feb 15th after the update.\n\nAfter the update, teams starting increasing their LB. I believe all progress regarding submissions greater than LB 0.693 is by using information from previous levels. Before Kaggle fixed the leak, nobody could use information from previous levels.",
      "votes": null
    },
    {
      "id": "2188248",
      "postDate": "03/19/2023 12:36:24",
      "content": "<p>Interesting. So do you guess top1 didn't use api leak? Or we don't know it for sure?</p>",
      "rawMarkdown": "Interesting. So do you guess top1 didn't use api leak? Or we don't know it for sure?",
      "votes": null
    },
    {
      "id": "2190457",
      "postDate": "03/21/2023 08:46:36",
      "content": "<p>To make it clear: 0.708 used data leak (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202</a>) and did not use API leak.<br>\nSorry for not letting you know sooner, but this has been estimated to be the best position to minimize disruption and ensure fairness.</p>",
      "rawMarkdown": "To make it clear: 0.708 used data leak (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202) and did not use API leak.\nSorry for not letting you know sooner, but this has been estimated to be the best position to minimize disruption and ensure fairness.",
      "votes": null
    },
    {
      "id": "2218145",
      "postDate": "04/11/2023 13:14:50",
      "content": "<p>Thanks Chris!</p>",
      "rawMarkdown": "Thanks Chris!",
      "votes": null
    },
    {
      "id": "2221569",
      "postDate": "04/14/2023 11:38:46",
      "content": "<p>I am so saved by your awareness. Thank you very much.</p>",
      "rawMarkdown": "I am so saved by your awareness. Thank you very much.",
      "votes": null
    },
    {
      "id": "2232425",
      "postDate": "04/24/2023 10:27:52",
      "content": "<p>Hi Chris, I am sorry if I missed it, but is there an answer for this somewhere in the forum? Currently, I am a bit doubtful in process these users (not sure I should discard them totally, truncate their sequences, or let them be)</p>",
      "rawMarkdown": "Hi Chris, I am sorry if I missed it, but is there an answer for this somewhere in the forum? Currently, I am a bit doubtful in process these users (not sure I should discard them totally, truncate their sequences, or let them be)",
      "votes": null
    },
    {
      "id": "2232442",
      "postDate": "04/24/2023 10:42:59",
      "content": "<p>No there is not an answer for this. My guess is that this leak exists. </p>",
      "rawMarkdown": "No there is not an answer for this. My guess is that this leak exists.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2148847,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "02/17/2023 17:58:15",
      "content": "<p>Wow this is new! How do you know that the Kaggle API gives data in such order? I have a lot of questions about the the Kaggle API's data fetching process (batch size, run-time, etc.) but don't know where to find the details. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2148855,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "02/17/2023 18:03:59",
      "content": "<p>Thanks Chris! We'll take a look.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149018,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "02/17/2023 20:51:53",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> We've checked the data and API inputs and we're having trouble seeing what could be causing this. Would you mind sharing an example? Must be something we're missing.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2149020,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/17/2023 21:00:41",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> The simplest example is <a href=\"https://www.kaggle.com/code/cdeotte/lb-leak-probe\" target=\"_blank\">here</a>. This submission notebook by default will submit all zeros and score LB 0.226. This submission remembers whenever it sees <code>level_group = '5-12'</code> and <code>level_group = '13-22'</code>. The first time it observes one leak (i.e. seeing a user <code>level_group = '0-4'</code> <strong>AFTER</strong> it sees either <code>5-12</code> or <code>13-22</code>), it starts predicting 1s thereafter. We see that this notebook scores LB 0.414 which proves that is <strong>observes at least 1 leak</strong> on public LB. (because LB 0.414 is the LB score of submitting all 1s. Furthermore we know the leak occurs very early since we achieved full LB 0.414)</p>\n<p>The reason i know that all the (public LB) data is leaked is because I have a trained an XGB model which uses features from the future. The features from the future change both my CV by +0.012 and LB by +0.012. This indicates that every single user is leaking. (because if only some users were leaking then LB would not boost the same as CV. Note in this competition my CV and LB are always exactly the same).</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2149025,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/17/2023 21:05:25",
              "content": "<p>UPDATE: My analysis above only applies to public LB. We do not know what is occurring on private LB.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2148988,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "02/17/2023 20:09:04",
      "content": "<p>Thanks for pointing it out Chris. I am curious, is your best score using the leak ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2148992,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/17/2023 20:12:32",
          "content": "<p>Yes. My current LB of 703 has CV 704 and uses information from the future.</p>\n<p>My best CV without leak is 695. My best CV with leak is 707.  I have not submitted these two models to the LB yet. (Still debugging some submit errors).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2149163,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "02/18/2023 02:49:44",
              "content": "<p>Thanks for the awsome investigation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>\n<p>I have quite similar experience with you regarding leak info.</p>\n<p>In my case, by just using one future feature, can boost the score by 0.01</p>\n<p>It is the time gap between current level group and the subsequent one.</p>\n<p>I guess it is the question answering period truncated by the host, a longer period implies more trials and a poorer accuracy.</p>\n<p>Time gap between current  level group and the previous one does not use future info, but only boosts the score a little.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2149043,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "02/17/2023 21:48:01",
      "content": "<p>A heads-up to all: we are confirming this issue with Chris' help. Once confirmed we will be fixing it ASAP; the hidden test set is not being leaked in any way, and there will be no long-term upside to doing this in your own notebooks. Please save your time for the competition's task!</p>\n<p>Thanks to Chris and others for pointing this out and helping us resolve it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2149081,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "02/17/2023 23:25:39",
      "content": "<p>My experience supports the hypothesis posted by Chris: notebooks assuming the correct ordering (first, '0-4, then '5-12', then '13'-22) throw an error.</p>\n<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , this was driving me crazy. I used like 10 submissions just to debug my code, only to find out about this.</p>\n<p>Btw it is interesting to see that such a leak does not improve CV and LB by a lot (in absolute terms). I would expect it to be more.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149090,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/17/2023 23:43:47",
          "content": "<p>I was also frustrated trying to save data in the correct order. I had many failed submissions the first week and could not understand why. Eventually I concluded that the data was in random order. Then i trained 3 models, one forward, one backward, and one without forward or backward. During submission, I would detect which of the 3 scenarios occurs for each user (for each level_group) and then use the corresponding model for inference. My submission LB score matched the CV for the backward model which made me conclude that 100% of the data is backward.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2149126,
              "author_name": "narsil",
              "author_url": "",
              "post_date": "02/18/2023 00:53:27",
              "content": "<p>Great stuff &amp; impressive investigation. thank you for sharing this</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2149387,
              "author_name": "andreaslup",
              "author_url": "",
              "post_date": "02/18/2023 08:44:27",
              "content": "<p>This is a brilliant example of \"reverse engineering\"! Thanks for pointing it out!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2149138,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "02/18/2023 01:47:30",
      "content": "<p>Thanks for pointing it out, Chris.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2149166,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "02/18/2023 03:01:52",
      "content": "<p>Problem should be fixed. We will rerun all notebooks pre-fix at the start of next week, when we have engineering support on hand.</p>\n<p>Thanks very much for your help, Chris! We'll be reaching out privately.</p>\n<p>Thanks all! Please let us know if you all run across any additional issues, or anything that seems related.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149590,
          "author_name": "carloshuertas",
          "author_url": "",
          "post_date": "02/18/2023 13:04:50",
          "content": "<p>Thanks a lot for fixing this Phil. And Chris for reporting!<br>\nExactly what is the fix? People should not be needing to guess what they need to predict.</p>\n<p>Can you share the exact specifications of how the data would be presented?<br>\nIdeally if using data from the past is allowed it could also be provided right away.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2153727,
              "author_name": "philculliton",
              "author_url": "",
              "post_date": "02/21/2023 15:46:11",
              "content": "<p>Hi, yes! The fix was ensuring that the ordering of the incoming data was correct - 0-4, 5-12, 13-22. This is the way the data is presented; you use the 0-4 data to predict question correctness at the end of that segment, and so on.</p>\n<p>You are allowed to use data from previous segments in a session. It's not possible with the API to present it all at once.</p>\n<p>We'll update the Data / Evaluation pages to make this clearer.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2153957,
                  "author_name": "joaopmpeinado",
                  "author_url": "",
                  "post_date": "02/21/2023 18:27:07",
                  "content": "<p>Phil, have all notebooks/scores been rerun yet with fixed lb?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2155796,
                      "author_name": "kvlmll",
                      "author_url": "",
                      "post_date": "02/22/2023 22:54:54",
                      "content": "<p>probably no since Chris still has 0.703</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2160628,
                          "author_name": "sggpls",
                          "author_url": "",
                          "post_date": "02/26/2023 20:07:46",
                          "content": "<p>Are there any updates?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2160857,
                              "author_name": "kvlmll",
                              "author_url": "",
                              "post_date": "02/27/2023 03:25:15",
                              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> hello? :D</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2180189,
          "author_name": "kvlmll",
          "author_url": "",
          "post_date": "03/13/2023 16:40:09",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> could you make it already clear about reruns pls? were they done?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2183777,
              "author_name": "kvlmll",
              "author_url": "",
              "post_date": "03/15/2023 22:37:13",
              "content": "<p>Am I a ghost already? :(</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2183783,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "03/15/2023 22:51:15",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kvlmll\" target=\"_blank\">@kvlmll</a> my understanding is that there have been <strong>no reruns</strong> (i.e. updates to Kaggles' LB scores but the API has been fixed around Feb 15th). None of my previous leaky submissions nor leaky LB score results have been updated. Furthermore, i posted a leak notebook <a href=\"https://www.kaggle.com/code/cdeotte/lb-leak-probe\" target=\"_blank\">here</a> version 1 which scores LB 0.414 using leak. If this notebook is rescored, it should score LB 0.226</p>\n<p>Additionally it is my understanding that all submissions now (Feb 15th onward) and moving forward cannot access the reverse order leak (but can maybe access the 2% leak). Note that 1st place team \"French Touch\" submitted after the update. So I believe that \"French Touch\" LB score of 0.708 is legitimate and does <strong>not use</strong> the leak. </p>\n<p>My LB score of 0.703 was submitted before the update and does <strong>use</strong> the reverse order leak. Without leak, my LB score is only LB 0.692</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2187223,
                  "author_name": "librauee",
                  "author_url": "",
                  "post_date": "03/18/2023 14:07:21",
                  "content": "<p>0.708 is really a high score that I can't fetch.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2187285,
                  "author_name": "khahuras",
                  "author_url": "",
                  "post_date": "03/18/2023 15:38:51",
                  "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for your great investigation. I have just 1 question and hope you can help. You said \"My LB score of 0.703 does use the leak.\". To my understanding your 0.703 score is AFTER Kaggle fixed the incorrect API order issue. So my question is that: do you mean the \"leak\" of 0.703 sub is just from 2% of the users with repeated session_group?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2187289,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "03/18/2023 15:43:05",
                      "content": "<p>My LB 0.703 was made around Feb 8th, one week <strong>BEFORE</strong> Kaggle fixed the leak (and uses the full reverse order leak). I have not tried exploiting the 2% leak.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2187387,
                          "author_name": "khahuras",
                          "author_url": "",
                          "post_date": "03/18/2023 17:27:22",
                          "content": "<p>thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . That's strange that Kaggle didn't RERUN all our submissions! So the current public LB score is invalid (because your 0.703 can't be there anymore). But if they don't rerun it, the private score will STILL BE LEAKY AND GOOD after private LB is revealed? I think <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> may have to provide us an answer?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2187395,
                              "author_name": "kvlmll",
                              "author_url": "",
                              "post_date": "03/18/2023 17:31:54",
                              "content": "<p>I suppose they will do rerun anyway of course</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2187410,
                                  "author_name": "cdeotte",
                                  "author_url": "",
                                  "post_date": "03/18/2023 17:40:55",
                                  "content": "<p>My guess is that i am the only LB submission with a leak. Nobody knew about the leak when i made it public. And all other teams had LB 0.693 and below on Feb 15th after the update.</p>\n<p>After the update, teams starting increasing their LB. I believe all progress regarding submissions greater than LB 0.693 is by using information from previous levels. Before Kaggle fixed the leak, nobody could use information from previous levels.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2188248,
                                      "author_name": "sggpls",
                                      "author_url": "",
                                      "post_date": "03/19/2023 12:36:24",
                                      "content": "<p>Interesting. So do you guess top1 didn't use api leak? Or we don't know it for sure?</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2149954,
      "author_name": "ritiksaxena90",
      "author_url": "",
      "post_date": "02/18/2023 20:03:00",
      "content": "<p>Great, thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2149959,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/18/2023 20:17:28",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> How does the Kaggle API handle users who repeat <code>level_group = '0-4'</code> or <code>5-12</code> after they finish later <code>level_groups</code>. See plot below. There are 2% train users (i.e. 229 out of 11779) who do this. Assuming 2% of test users do this also, does the Kaggle API still give us all of <code>level_group = '0-4'</code> information to predict questions 1 thru 3. If so, then the Kaggle API is giving us the future to help predict the past. We can save the future level 0-4 information and use it to help predict questions 4-18 from <code>level_groups '5-12' and '13-22'</code>. </p>\n<p>For example, one specific way to use this information is we can compute the time gap size between the two <code>level_group = '0-4'</code> sessions, then subtract the active time spent on <code>level_group = '5-12' and '13-22'</code> and deduce how much time they spent answering questions after <code>level_group = '5-12' and '13-22'</code>. Knowing how long users spend during the question answering portion can achieve at least <code>+0.010</code> boost to CV and LB (from prior time leak experiments).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/leak_user2.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2157671,
          "author_name": "hookman",
          "author_url": "",
          "post_date": "02/24/2023 08:08:51",
          "content": "<p>Hi, Chris. </p>\n<p>It seems we still don't get answer for this question, right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2157688,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/24/2023 08:24:55",
              "content": "<p>Correct. We did not receive an answer. <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> How does the Kaggle API handle these 2% of data? Does the API give us a time leak, or are these 2% of users removed from test data, or are these 2% of users truncated and only provide the first visit to each level group?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2232425,
                  "author_name": "shinomoriaoshi",
                  "author_url": "",
                  "post_date": "04/24/2023 10:27:52",
                  "content": "<p>Hi Chris, I am sorry if I missed it, but is there an answer for this somewhere in the forum? Currently, I am a bit doubtful in process these users (not sure I should discard them totally, truncate their sequences, or let them be)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2232442,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "04/24/2023 10:42:59",
                      "content": "<p>No there is not an answer for this. My guess is that this leak exists. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2150575,
      "author_name": "libingying",
      "author_url": "",
      "post_date": "02/19/2023 11:59:17",
      "content": "<p>thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2153137,
      "author_name": "asadbek1121",
      "author_url": "",
      "post_date": "02/21/2023 08:13:21",
      "content": "<p>Hi Chris good luck</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2160813,
      "author_name": "juzqyxs",
      "author_url": "",
      "post_date": "02/27/2023 02:21:24",
      "content": "<p>So the bug is fixed or still have some new leaks?😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 2165769,
          "author_name": "trasibulo",
          "author_url": "",
          "post_date": "03/02/2023 12:49:05",
          "content": "<p>I have the same question. <br>\nIt is not clear for me if the fix has been done already or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2190457,
      "author_name": "pdnartreb",
      "author_url": "",
      "post_date": "03/21/2023 08:46:36",
      "content": "<p>To make it clear: 0.708 used data leak (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202</a>) and did not use API leak.<br>\nSorry for not letting you know sooner, but this has been estimated to be the best position to minimize disruption and ensure fairness.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2218145,
      "author_name": "byungeunhwang",
      "author_url": "",
      "post_date": "04/11/2023 13:14:50",
      "content": "<p>Thanks Chris!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2221569,
      "author_name": "risakashiwabara",
      "author_url": "",
      "post_date": "04/14/2023 11:38:46",
      "content": "<p>I am so saved by your awareness. Thank you very much.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2148806": "In Kaggle's API example notebook, all of `level_group = '0-4'` comes **before** `level_group = '5-12'` which comes **before** `level_group = '13-22'`. This is the **correct** ordering in time.\n\nDuring submission, I have noticed that the Kaggle API gives us the levels **backward**. First we see all `level_group = '13-22'`, then we see all `level_group = '5-12'`, then we see all `level_group = '0-4'`. This is **incorrect** ordering in time. This means we can build features from the future before predicting the present.\n\nSohier, Maggie, Natalie, Alex, is this **test data leak** intentional or accidental?\n@nrambis @alexmlfranklin @sohier @maggiemd\n\n# UPDATE\nThis time leak has been fixed. Thank you Kaggle staff for correction",
    "2148847": "Wow this is new! How do you know that the Kaggle API gives data in such order? I have a lot of questions about the the Kaggle API's data fetching process (batch size, run-time, etc.) but don't know where to find the details.",
    "2148855": "Thanks Chris! We'll take a look.",
    "2148988": "Thanks for pointing it out Chris. I am curious, is your best score using the leak ?",
    "2148992": "Yes. My current LB of 703 has CV 704 and uses information from the future.\n\nMy best CV without leak is 695. My best CV with leak is 707.  I have not submitted these two models to the LB yet. (Still debugging some submit errors).",
    "2149018": "cdeotte We've checked the data and API inputs and we're having trouble seeing what could be causing this. Would you mind sharing an example? Must be something we're missing.",
    "2149020": "philculliton The simplest example is [here][1]. This submission notebook by default will submit all zeros and score LB 0.226. This submission remembers whenever it sees `level_group = '5-12'` and `level_group = '13-22'`. The first time it observes one leak (i.e. seeing a user `level_group = '0-4'` **AFTER** it sees either `5-12` or `13-22`), it starts predicting 1s thereafter. We see that this notebook scores LB 0.414 which proves that is **observes at least 1 leak** on public LB. (because LB 0.414 is the LB score of submitting all 1s. Furthermore we know the leak occurs very early since we achieved full LB 0.414)\n\nThe reason i know that all the (public LB) data is leaked is because I have a trained an XGB model which uses features from the future. The features from the future change both my CV by +0.012 and LB by +0.012. This indicates that every single user is leaking. (because if only some users were leaking then LB would not boost the same as CV. Note in this competition my CV and LB are always exactly the same).\n\n[1]: https://www.kaggle.com/code/cdeotte/lb-leak-probe",
    "2149025": "UPDATE: My analysis above only applies to public LB. We do not know what is occurring on private LB.",
    "2149043": "A heads-up to all: we are confirming this issue with Chris' help. Once confirmed we will be fixing it ASAP; the hidden test set is not being leaked in any way, and there will be no long-term upside to doing this in your own notebooks. Please save your time for the competition's task!\n\nThanks to Chris and others for pointing this out and helping us resolve it!",
    "2149081": "My experience supports the hypothesis posted by Chris: notebooks assuming the correct ordering (first, '0-4, then '5-12', then '13'-22) throw an error.\n\nThanks @cdeotte , this was driving me crazy. I used like 10 submissions just to debug my code, only to find out about this.\n\nBtw it is interesting to see that such a leak does not improve CV and LB by a lot (in absolute terms). I would expect it to be more.",
    "2149090": "I was also frustrated trying to save data in the correct order. I had many failed submissions the first week and could not understand why. Eventually I concluded that the data was in random order. Then i trained 3 models, one forward, one backward, and one without forward or backward. During submission, I would detect which of the 3 scenarios occurs for each user (for each level_group) and then use the corresponding model for inference. My submission LB score matched the CV for the backward model which made me conclude that 100% of the data is backward.",
    "2149126": "Great stuff & impressive investigation. thank you for sharing this",
    "2149138": "Thanks for pointing it out, Chris.",
    "2149163": "Thanks for the awsome investigation @cdeotte.\n\nI have quite similar experience with you regarding leak info.\n\nIn my case, by just using one future feature, can boost the score by 0.01\n\nIt is the time gap between current level group and the subsequent one.\n\nI guess it is the question answering period truncated by the host, a longer period implies more trials and a poorer accuracy.\n\nTime gap between current  level group and the previous one does not use future info, but only boosts the score a little.",
    "2149166": "Problem should be fixed. We will rerun all notebooks pre-fix at the start of next week, when we have engineering support on hand.\n\nThanks very much for your help, Chris! We'll be reaching out privately.\n\nThanks all! Please let us know if you all run across any additional issues, or anything that seems related.",
    "2149387": "This is a brilliant example of \"reverse engineering\"! Thanks for pointing it out!",
    "2149590": "Thanks a lot for fixing this Phil. And Chris for reporting!\nExactly what is the fix? People should not be needing to guess what they need to predict.\n\nCan you share the exact specifications of how the data would be presented?\nIdeally if using data from the past is allowed it could also be provided right away.",
    "2149954": "Great, thanks a lot!",
    "2149959": "philculliton How does the Kaggle API handle users who repeat `level_group = '0-4'` or `5-12` after they finish later `level_groups`. See plot below. There are 2% train users (i.e. 229 out of 11779) who do this. Assuming 2% of test users do this also, does the Kaggle API still give us all of `level_group = '0-4'` information to predict questions 1 thru 3. If so, then the Kaggle API is giving us the future to help predict the past. We can save the future level 0-4 information and use it to help predict questions 4-18 from `level_groups '5-12' and '13-22'`. \n\nFor example, one specific way to use this information is we can compute the time gap size between the two `level_group = '0-4'` sessions, then subtract the active time spent on `level_group = '5-12' and '13-22'` and deduce how much time they spent answering questions after `level_group = '5-12' and '13-22'`. Knowing how long users spend during the question answering portion can achieve at least `+0.010` boost to CV and LB (from prior time leak experiments).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/leak_user2.png)",
    "2150575": "thanks a lot!",
    "2153137": "Hi Chris good luck",
    "2153727": "Hi, yes! The fix was ensuring that the ordering of the incoming data was correct - 0-4, 5-12, 13-22. This is the way the data is presented; you use the 0-4 data to predict question correctness at the end of that segment, and so on.\n\nYou are allowed to use data from previous segments in a session. It's not possible with the API to present it all at once.\n\nWe'll update the Data / Evaluation pages to make this clearer.",
    "2153957": "Phil, have all notebooks/scores been rerun yet with fixed lb?",
    "2155796": "probably no since Chris still has 0.703",
    "2157671": "Hi, Chris. \n\nIt seems we still don't get answer for this question, right?",
    "2157688": "Correct. We did not receive an answer. @philculliton How does the Kaggle API handle these 2% of data? Does the API give us a time leak, or are these 2% of users removed from test data, or are these 2% of users truncated and only provide the first visit to each level group?",
    "2160628": "Are there any updates?",
    "2160813": "So the bug is fixed or still have some new leaks?😂",
    "2160857": "philculliton hello? :D",
    "2165769": "I have the same question. \nIt is not clear for me if the fix has been done already or not.",
    "2180189": "philculliton @sohier @maggiemd could you make it already clear about reruns pls? were they done?",
    "2183777": "Am I a ghost already? :(",
    "2183783": "Hi @kvlmll my understanding is that there have been **no reruns** (i.e. updates to Kaggles' LB scores but the API has been fixed around Feb 15th). None of my previous leaky submissions nor leaky LB score results have been updated. Furthermore, i posted a leak notebook [here][1] version 1 which scores LB 0.414 using leak. If this notebook is rescored, it should score LB 0.226\n\nAdditionally it is my understanding that all submissions now (Feb 15th onward) and moving forward cannot access the reverse order leak (but can maybe access the 2% leak). Note that 1st place team \"French Touch\" submitted after the update. So I believe that \"French Touch\" LB score of 0.708 is legitimate and does **not use** the leak. \n  \nMy LB score of 0.703 was submitted before the update and does **use** the reverse order leak. Without leak, my LB score is only LB 0.692\n\n[1]: https://www.kaggle.com/code/cdeotte/lb-leak-probe",
    "2187223": "0.708 is really a high score that I can't fetch.",
    "2187285": "cdeotte Thanks for your great investigation. I have just 1 question and hope you can help. You said \"My LB score of 0.703 does use the leak.\". To my understanding your 0.703 score is AFTER Kaggle fixed the incorrect API order issue. So my question is that: do you mean the \"leak\" of 0.703 sub is just from 2% of the users with repeated session_group?",
    "2187289": "My LB 0.703 was made around Feb 8th, one week **BEFORE** Kaggle fixed the leak (and uses the full reverse order leak). I have not tried exploiting the 2% leak.",
    "2187387": "thanks @cdeotte . That's strange that Kaggle didn't RERUN all our submissions! So the current public LB score is invalid (because your 0.703 can't be there anymore). But if they don't rerun it, the private score will STILL BE LEAKY AND GOOD after private LB is revealed? I think @philculliton may have to provide us an answer?",
    "2187395": "I suppose they will do rerun anyway of course",
    "2187410": "My guess is that i am the only LB submission with a leak. Nobody knew about the leak when i made it public. And all other teams had LB 0.693 and below on Feb 15th after the update.\n\nAfter the update, teams starting increasing their LB. I believe all progress regarding submissions greater than LB 0.693 is by using information from previous levels. Before Kaggle fixed the leak, nobody could use information from previous levels.",
    "2188248": "Interesting. So do you guess top1 didn't use api leak? Or we don't know it for sure?",
    "2190457": "To make it clear: 0.708 used data leak (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202) and did not use API leak.\nSorry for not letting you know sooner, but this has been estimated to be the best position to minimize disruption and ensure fairness.",
    "2218145": "Thanks Chris!",
    "2221569": "I am so saved by your awareness. Thank you very much.",
    "2232425": "Hi Chris, I am sorry if I missed it, but is there an answer for this somewhere in the forum? Currently, I am a bit doubtful in process these users (not sure I should discard them totally, truncate their sequences, or let them be)",
    "2232442": "No there is not an answer for this. My guess is that this leak exists."
  },
  "source": "meta"
}