{
  "id": 396202,
  "title": "Update on Leaked Competition Data",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/396202",
  "author_name": "Alex Franklin",
  "post_date": "2023-03-20T18:01:10.046000",
  "votes": 92,
  "comment_count": 117,
  "views": 0,
  "content": "<p>A Kaggler brought to our attention that the test data for this competition was unintentionally made available for a period of time. Although the leak was fairly limited, we can’t be certain that the test data is not compromised. Therefore, to be fair to all contestants <strong>we are releasing all of the leaked data</strong>, which contains the entire current test set.</p>\n<p>We will update the training data on the competition page to include the current test set, which will increase the size of the training data by 100%.</p>\n<p>The raw data from the game will also be made available and can be found at <a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">this site</a>, which you are welcome to use as supplemental data for this competition.</p>\n<p>Since we are doubling the size of the training set and linking to the raw data, <strong>we are also extending the competition by one month</strong> to give everyone a chance to incorporate the new data into their models.</p>\n<p>We will be creating a new public and private test set with data that was not previously part of the competition and was not part of the leak. The public leaderboard will be rerun to reflect scores on this new test set.</p>\n<p>The updates to the competition data should be complete by the end of the day.</p>\n<p>We would like to thank the Kaggler who brought this to our attention. We appreciate the integrity of the Kaggle community for consistently reporting issues which could have been exploited to gain an individual advantage in the competition. Thank you for bearing with us as we work to resolve this issue.</p>",
  "messages": [
    {
      "id": 2189733,
      "postDate": "2023-03-20T18:01:10.047Z",
      "content": "<p>A Kaggler brought to our attention that the test data for this competition was unintentionally made available for a period of time. Although the leak was fairly limited, we can’t be certain that the test data is not compromised. Therefore, to be fair to all contestants <strong>we are releasing all of the leaked data</strong>, which contains the entire current test set.</p>\n<p>We will update the training data on the competition page to include the current test set, which will increase the size of the training data by 100%.</p>\n<p>The raw data from the game will also be made available and can be found at <a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">this site</a>, which you are welcome to use as supplemental data for this competition.</p>\n<p>Since we are doubling the size of the training set and linking to the raw data, <strong>we are also extending the competition by one month</strong> to give everyone a chance to incorporate the new data into their models.</p>\n<p>We will be creating a new public and private test set with data that was not previously part of the competition and was not part of the leak. The public leaderboard will be rerun to reflect scores on this new test set.</p>\n<p>The updates to the competition data should be complete by the end of the day.</p>\n<p>We would like to thank the Kaggler who brought this to our attention. We appreciate the integrity of the Kaggle community for consistently reporting issues which could have been exploited to gain an individual advantage in the competition. Thank you for bearing with us as we work to resolve this issue.</p>",
      "rawMarkdown": "A Kaggler brought to our attention that the test data for this competition was unintentionally made available for a period of time. Although the leak was fairly limited, we can’t be certain that the test data is not compromised. Therefore, to be fair to all contestants **we are releasing all of the leaked data**, which contains the entire current test set.\n\nWe will update the training data on the competition page to include the current test set, which will increase the size of the training data by 100%.\n\nThe raw data from the game will also be made available and can be found at [this site](https://fielddaylab.wisc.edu/opengamedata/), which you are welcome to use as supplemental data for this competition.\n\nSince we are doubling the size of the training set and linking to the raw data, **we are also extending the competition by one month** to give everyone a chance to incorporate the new data into their models.\n\nWe will be creating a new public and private test set with data that was not previously part of the competition and was not part of the leak. The public leaderboard will be rerun to reflect scores on this new test set.\n\nThe updates to the competition data should be complete by the end of the day.\n\nWe would like to thank the Kaggler who brought this to our attention. We appreciate the integrity of the Kaggle community for consistently reporting issues which could have been exploited to gain an individual advantage in the competition. Thank you for bearing with us as we work to resolve this issue.",
      "votes": 91
    },
    {
      "id": 2190090,
      "postDate": "2023-03-21T03:03:46.213Z",
      "content": "<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>, just as <a href=\"https://www.kaggle.com/tanakaakinori\" target=\"_blank\">@tanakaakinori</a> and others have observed, there are differences in both the training data and the test data. I have also observed the changes in the test data that resulted in submission error, which according to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> that we may have a leak again.</p>\n<p>Please fix everything and let us know the details when you are done. Perhaps you should disable submission again so that we do not waste our times.</p>\n<p><strong>Update:</strong><br>\nBased on the replies to <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>'s post below, there was and still is a leak described as:<br>\n<code>2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition.</code></p>\n<p>We need an update on if or when that one will be fixed <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>. Thanks</p>",
      "rawMarkdown": "@alexmlfranklin, just as @tanakaakinori and others have observed, there are differences in both the training data and the test data. I have also observed the changes in the test data that resulted in submission error, which according to @cdeotte that we may have a leak again.\n\nPlease fix everything and let us know the details when you are done. Perhaps you should disable submission again so that we do not waste our times.\n\n**Update:**\nBased on the replies to @philculliton's post below, there was and still is a leak described as:\n`2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. `\n\nWe need an update on if or when that one will be fixed @philculliton. Thanks",
      "votes": 19,
      "replies": [
        {
          "id": 2191003,
          "postDate": "2023-03-21T16:29:15.293Z",
          "content": "<p>Hi! We do not have a (hidden) leak again. The issue people are describing exists only in the sample data, which we'll update.</p>\n<p>If you are seeing submission errors that are NOT related to the reordering of the tuple returned by the timeseries API, please feel free to reach out with a notebook that reproduces them.</p>",
          "rawMarkdown": "Hi! We do not have a (hidden) leak again. The issue people are describing exists only in the sample data, which we'll update.\n\nIf you are seeing submission errors that are NOT related to the reordering of the tuple returned by the timeseries API, please feel free to reach out with a notebook that reproduces them.",
          "votes": 6,
          "replies": [
            {
              "id": 2191077,
              "postDate": "2023-03-21T17:21:56.333Z",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, Is the sample test data updated yet? I just made a submission and got the \"Submission Scoring Error\" again. I used  <code>(test, sample_submission) in iter_test</code>:</p>",
              "rawMarkdown": "@philculliton, Is the sample test data updated yet? I just made a submission and got the \"Submission Scoring Error\" again. I used  `(test, sample_submission) in iter_test`:",
              "votes": 1
            },
            {
              "id": 2191347,
              "postDate": "2023-03-21T22:53:00.773Z",
              "content": "<p>Hi, I am still a bit confused. <br>\nBefore in the sample data, we have 1 session and 1 level_group per <code>test</code><br>\nNow we have 3 sessions and 3 level_groups per <code>test</code>.<br>\nDo you mean we have 3 sessions and 3 level_groups only in the sample data? Then how many do we have in the real test set?🤔</p>",
              "rawMarkdown": "Hi, I am still a bit confused. \nBefore in the sample data, we have 1 session and 1 level_group per `test`\nNow we have 3 sessions and 3 level_groups per `test`.\nDo you mean we have 3 sessions and 3 level_groups only in the sample data? Then how many do we have in the real test set?🤔"
            },
            {
              "id": 2191356,
              "postDate": "2023-03-21T23:05:45.543Z",
              "content": "<p>No. This is a mistake. Kaggle is fixing this. The new API will have <strong>one</strong> level group and <strong>one</strong> session_id per iteration loop. Kaggle is fixing this now. See other comments in this discussion</p>",
              "rawMarkdown": "No. This is a mistake. Kaggle is fixing this. The new API will have **one** level group and **one** session_id per iteration loop. Kaggle is fixing this now. See other comments in this discussion",
              "votes": 4
            },
            {
              "id": 2191365,
              "postDate": "2023-03-21T23:19:15.777Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>! Thanks again for all the fixes!<br>\nWill you make another announcement when things are settled?</p>",
              "rawMarkdown": "Hi @philculliton! Thanks again for all the fixes!\nWill you make another announcement when things are settled?"
            },
            {
              "id": 2191390,
              "postDate": "2023-03-21T23:39:57.670Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nIn my understanding. 2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. Is that right?</p>",
              "rawMarkdown": "@cdeotte \nIn my understanding. 2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. Is that right?",
              "votes": 1
            },
            {
              "id": 2191405,
              "postDate": "2023-03-21T23:50:00.533Z",
              "content": "<blockquote>\n  <p>In my understanding. 2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. Is that right?</p>\n</blockquote>\n<p>Yes, it is my understanding that this still exists. This was not fixed before the most recent data change (and not fixed after). And Kaggle staff has not commented about how they will address this yet.</p>",
              "rawMarkdown": ">In my understanding. 2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. Is that right?\n\nYes, it is my understanding that this still exists. This was not fixed before the most recent data change (and not fixed after). And Kaggle staff has not commented about how they will address this yet.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2194193,
      "postDate": "2023-03-23T19:44:27.610Z",
      "content": "<p>The problem seems unsolved…<br>\n<a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> </p>",
      "rawMarkdown": "The problem seems unsolved...\n@alexmlfranklin @nrambis @philculliton @maggiemd ",
      "votes": 10,
      "replies": [
        {
          "id": 2195487,
          "postDate": "2023-03-24T17:11:23.640Z",
          "content": "<p>0.753 is leak again? 👀</p>",
          "rawMarkdown": "0.753 is leak again? 👀",
          "votes": 1,
          "replies": [
            {
              "id": 2195517,
              "postDate": "2023-03-24T17:36:48.340Z",
              "content": "<p>I read Bertrand's post as the api problem is still not solved: the iter_test still returns all the test data. </p>\n<p>Re how we get 0.753 you will have to wait till after the competition end.</p>",
              "rawMarkdown": "I read Bertrand's post as the api problem is still not solved: the iter_test still returns all the test data. \n\nRe how we get 0.753 you will have to wait till after the competition end.",
              "votes": 4
            },
            {
              "id": 2198374,
              "postDate": "2023-03-27T00:39:10.863Z",
              "content": "<p>the main problem mentioned in this topic is Leaked Competition Data 🙃</p>",
              "rawMarkdown": "the main problem mentioned in this topic is Leaked Competition Data 🙃",
              "votes": 2
            },
            {
              "id": 2198603,
              "postDate": "2023-03-27T07:13:46.260Z",
              "content": "<p>Several of us made comments about the api change, including me here: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191719\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191719</a></p>\n<p>Other comments about this issue:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2193890\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2193890</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190039\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190039</a></p>\n<p>I honestly don't get  this:</p>\n<ul>\n<li>How fixing a bug had to result in swapping arguments?</li>\n<li>Why we can't get a fixed code when running notebooks before submission?</li>\n</ul>",
              "rawMarkdown": "Several of us made comments about the api change, including me here: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191719\n\nOther comments about this issue:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2193890\n\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190039\n\nI honestly don't get  this:\n- How fixing a bug had to result in swapping arguments?\n- Why we can't get a fixed code when running notebooks before submission?\n\n",
              "votes": 1
            },
            {
              "id": 2198837,
              "postDate": "2023-03-27T10:20:08.827Z",
              "content": "<p>good questions! in addition: why we had no reruns still after promising to do it when Chris found the first api bug? seems like this comp is not too important for kaggle </p>",
              "rawMarkdown": "good questions! in addition: why we had no reruns still after promising to do it when Chris found the first api bug? seems like this comp is not too important for kaggle ",
              "votes": 6
            },
            {
              "id": 2199032,
              "postDate": "2023-03-27T13:02:08.763Z",
              "content": "<p>Indeed lots of LB is from runs prior to the data reset. None of that is legit now.</p>",
              "rawMarkdown": "Indeed lots of LB is from runs prior to the data reset. None of that is legit now.",
              "votes": 1
            },
            {
              "id": 2199039,
              "postDate": "2023-03-27T13:06:50.727Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2199072,
              "postDate": "2023-03-27T13:27:50.427Z",
              "content": "<p>Hi. We did not update the LB because we had reports of other issues that we thought we might need to resolve; then the test data leak was revealed. It took some time to fix. We will be re-running the LB now that the public test set has been updated, and the sample data has been fixed.</p>",
              "rawMarkdown": "Hi. We did not update the LB because we had reports of other issues that we thought we might need to resolve; then the test data leak was revealed. It took some time to fix. We will be re-running the LB now that the public test set has been updated, and the sample data has been fixed.",
              "votes": 2
            },
            {
              "id": 2199075,
              "postDate": "2023-03-27T13:30:02.140Z",
              "content": "<p>Hi! The API is complex; the dataframes need to be served in a particular order to be fully supported. We needed to switch to avoid potential issues with the API.</p>\n<p>What's the fixed code you're referring to? The public sample data should no longer be returning all at once.</p>",
              "rawMarkdown": "Hi! The API is complex; the dataframes need to be served in a particular order to be fully supported. We needed to switch to avoid potential issues with the API.\n\nWhat's the fixed code you're referring to? The public sample data should no longer be returning all at once.",
              "votes": 1
            },
            {
              "id": 2199664,
              "postDate": "2023-03-27T23:21:08.607Z",
              "content": "<p>could you explain in a few words how will you rerun old submissions with the new api?</p>",
              "rawMarkdown": "could you explain in a few words how will you rerun old submissions with the new api?",
              "votes": 2
            },
            {
              "id": 2200342,
              "postDate": "2023-03-28T13:36:12.447Z",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> before the data reset, iteration in inference code looked like:</p>\n<pre><code> (sample_submission, test)  iter_test:\n</code></pre>\n<p>After it looks like:</p>\n<pre><code> (test, sample_submission)  iter_test:\n</code></pre>\n<p>It has been days we have been complaining about the argument order change. In particular it means that code submitted before the data reset will no longer run.</p>",
              "rawMarkdown": "@philculliton before the data reset, iteration in inference code looked like:\n\n```python\nfor (sample_submission, test) in iter_test:\n\n```\n\nAfter it looks like:\n\n```python\nfor (test, sample_submission) in iter_test:\n\n```\nIt has been days we have been complaining about the argument order change. In particular it means that code submitted before the data reset will no longer run.",
              "votes": 3
            },
            {
              "id": 2200492,
              "postDate": "2023-03-28T15:32:03.407Z",
              "content": "<p>Thanks. Yes, as mentioned in my original note, the order that the dataframes are returned in needed to be changed. We were seeing bugs that indicated this was causing problems in the API. Rather than let those problems continue, we made the update.</p>",
              "rawMarkdown": "Thanks. Yes, as mentioned in my original note, the order that the dataframes are returned in needed to be changed. We were seeing bugs that indicated this was causing problems in the API. Rather than let those problems continue, we made the update.",
              "votes": 1
            },
            {
              "id": 2200563,
              "postDate": "2023-03-28T16:06:38.540Z",
              "content": "<p>Then you get that the code written before the data reset will not longer run, right?</p>",
              "rawMarkdown": "Then you get that the code written before the data reset will not longer run, right?",
              "votes": 3
            },
            {
              "id": 2201564,
              "postDate": "2023-03-29T12:07:24.130Z",
              "content": "<blockquote>\n  <p>the order that the dataframes are returned in needed to be changed. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> we are discussing the order of the arguments to the iter_test object. We are not discussing the order in which you return the dataframe.</p>\n<p>It is weird that you don't acknowledge this order has changed and that it causes a major issue for rerunning past submissions.</p>",
              "rawMarkdown": "> the order that the dataframes are returned in needed to be changed. \n\n@philculliton we are discussing the order of the arguments to the iter_test object. We are not discussing the order in which you return the dataframe.\n\nIt is weird that you don't acknowledge this order has changed and that it causes a major issue for rerunning past submissions.\n\n\n\n",
              "votes": 2
            },
            {
              "id": 2201570,
              "postDate": "2023-03-29T12:13:48.463Z",
              "content": "<p>The order of the returned dataframes IS the order of the arguments in the iter_test object. That is precisely what I am saying. I noted this up front when the update was announced.</p>\n<p>I am aware that this causes issues for rerunning past submissions. It was, unfortunately, the better route to take due to reported bugs. These bugs were also noted up front when the update was announced, as the reason for the change.</p>",
              "rawMarkdown": "The order of the returned dataframes IS the order of the arguments in the iter_test object. That is precisely what I am saying. I noted this up front when the update was announced.\n\nI am aware that this causes issues for rerunning past submissions. It was, unfortunately, the better route to take due to reported bugs. These bugs were also noted up front when the update was announced, as the reason for the change.",
              "votes": 2
            },
            {
              "id": 2201578,
              "postDate": "2023-03-29T12:23:57.417Z",
              "content": "<p>ok, fair enough.</p>\n<p>I thought you were discussing the order in which level groups were returned.</p>",
              "rawMarkdown": "ok, fair enough.\n\nI thought you were discussing the order in which level groups were returned.\n\n"
            },
            {
              "id": 2201589,
              "postDate": "2023-03-29T12:36:40.807Z",
              "content": "<p>Just to be on the same page: Since it is not easily possible to rerun the notebooks, the public leaderboard is erroneous. Did I understand that correctly?</p>",
              "rawMarkdown": "Just to be on the same page: Since it is not easily possible to rerun the notebooks, the public leaderboard is erroneous. Did I understand that correctly?",
              "votes": 1
            },
            {
              "id": 2202623,
              "postDate": "2023-03-30T07:21:18.313Z",
              "content": "<p>So may I ask a question, for the lb 0.753, it is caused by api bug or leaked data? or it is just a normal score.</p>\n<p>I am quite confused here. <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> said the problem was unsolved. I am not sure what kind of problem it is and I thought it is leaked data as same as last time. But <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> mentioned that \"the iter_test still returns all the test data\". But as many people said there is <strong>only a bug in sample data</strong> that return all level_group data and <strong>in the test part it is fine</strong>. So it means that <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> found another bug to get all the test data in the iter_test?</p>",
              "rawMarkdown": "So may I ask a question, for the lb 0.753, it is caused by api bug or leaked data? or it is just a normal score.\n\nI am quite confused here. @pdnartreb said the problem was unsolved. I am not sure what kind of problem it is and I thought it is leaked data as same as last time. But @cpmpml mentioned that \"the iter_test still returns all the test data\". But as many people said there is **only a bug in sample data** that return all level_group data and **in the test part it is fine**. So it means that @pdnartreb @cpmpml found another bug to get all the test data in the iter_test?",
              "votes": 3
            },
            {
              "id": 2203402,
              "postDate": "2023-03-30T19:13:37.027Z",
              "content": "<blockquote>\n  <p>But as many people said there is only a bug in sample data that return all level_group data</p>\n</blockquote>\n<p>This is a bug in the code. The data is what it is. The code should iterate through session_id and level group.</p>\n<blockquote>\n  <p>for the lb 0.753, it is caused by api bug or leaked data? or it is just a normal score.</p>\n</blockquote>\n<p>This score is not legit. It was obtained using the leak that Bertrand discovered. The leak is now fixed and we have no way to reproduce that score anymore.  Same goes for the 0.72 submission.</p>\n<p>We only understood the issue very recently when we failed to reproduce these results.  We have asked Kaggle to remove these two submissions.</p>",
              "rawMarkdown": "> But as many people said there is only a bug in sample data that return all level_group data\n\nThis is a bug in the code. The data is what it is. The code should iterate through session_id and level group.\n\n> for the lb 0.753, it is caused by api bug or leaked data? or it is just a normal score.\n\nThis score is not legit. It was obtained using the leak that Bertrand discovered. The leak is now fixed and we have no way to reproduce that score anymore.  Same goes for the 0.72 submission.\n\nWe only understood the issue very recently when we failed to reproduce these results.  We have asked Kaggle to remove these two submissions.",
              "votes": 13
            },
            {
              "id": 2203670,
              "postDate": "2023-03-31T03:11:17.940Z",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Cool. Thanks for explaining everything.</p>",
              "rawMarkdown": "@cpmpml Cool. Thanks for explaining everything.",
              "votes": 2
            },
            {
              "id": 2209919,
              "postDate": "2023-04-05T03:14:53.937Z",
              "content": "<p>As I recall, the order has always been test_df, sample_sub for time-series code comps. I didn't see any explanation why it was opposite at first for this one. Not a helpful comment perhaps, except that the order should always be the same for such competitions.</p>",
              "rawMarkdown": "As I recall, the order has always been test_df, sample_sub for time-series code comps. I didn't see any explanation why it was opposite at first for this one. Not a helpful comment perhaps, except that the order should always be the same for such competitions."
            }
          ]
        }
      ]
    },
    {
      "id": 2190039,
      "postDate": "2023-03-21T01:36:00.407Z",
      "content": "<p>Hello! Thanks for the update.<br>\nI found some changes.</p>\n<ul>\n<li>The data size of 'train', 'train_label' has increased. <br>\nIt ran out of CPU notebook's 8GB RAM,  No problem for GPU's one.</li>\n<li><code>for (sample_submission, test) in iter_test:</code> was changed to for <code>(test,  sample_submission) in iter_test:</code></li>\n<li><code>for (test, sample_submission) in iter_test:</code>, old <code>test</code> used to have  single <code>session_id</code>, single <code>level_group</code>. <br>\n But now it contains multiple <code>session_id</code>, and all three <code>level_groups</code>.😲</li>\n</ul>",
      "rawMarkdown": "Hello! Thanks for the update.\nI found some changes.\n\n- The data size of 'train', 'train_label' has increased. \n    It ran out of CPU notebook's 8GB RAM,  No problem for GPU's one.\n- ` for (sample_submission, test) in iter_test: ` was changed to for `(test,  sample_submission) in iter_test:`\n- `for (test, sample_submission) in iter_test:`, old `test` used to have  single `session_id`, single `level_group`. \n     But now it contains multiple `session_id`, and all three `level_groups`.😲",
      "votes": 9,
      "replies": [
        {
          "id": 2190082,
          "postDate": "2023-03-21T02:49:14.147Z",
          "content": "<blockquote>\n  <p>But now it contains multiple session_id, and all three level_groups</p>\n</blockquote>\n<p>If this is the case then this is a time leak (i.e. this is the original time leak that Kaggle fixed on Feb 15th). We can use info from levels 5+ to predict questions from level 0-4. This gives at least <code>+0.010</code> boost on CV and LB!</p>",
          "rawMarkdown": ">But now it contains multiple session_id, and all three level_groups\n\nIf this is the case then this is a time leak (i.e. this is the original time leak that Kaggle fixed on Feb 15th). We can use info from levels 5+ to predict questions from level 0-4. This gives at least `+0.010` boost on CV and LB!",
          "votes": 10,
          "replies": [
            {
              "id": 2190110,
              "postDate": "2023-03-21T03:40:22.340Z",
              "content": "<p>OMG, I see, the same thing may have happened with the previous leak. <br>\nAs you say, it is not appropriate to use data to predict future behavior rather than solve problems.</p>\n<p>Looking at the CV in the new training notebook, from 0.696 improved to 0.709 (could be overfitting though).</p>\n<p>I am hesitant to submit this data as it is leaked data…</p>\n<p>I have learned a lot from your notes and discussions, thank you!</p>",
              "rawMarkdown": "OMG, I see, the same thing may have happened with the previous leak. \nAs you say, it is not appropriate to use data to predict future behavior rather than solve problems.\n\nLooking at the CV in the new training notebook, from 0.696 improved to 0.709 (could be overfitting though).\n\nI am hesitant to submit this data as it is leaked data...\n\nI have learned a lot from your notes and discussions, thank you!",
              "votes": 3
            },
            {
              "id": 2190529,
              "postDate": "2023-03-21T10:05:00.503Z",
              "content": "<p>doesn't that mean that we need to setup different feature creation strategy?<br>\nWe don't need to group by level group anymore as we have all the level groups at hand from the beginning (see <a href=\"https://www.kaggle.com/code/kmitsuhiro/api-update-iter-test-returns-all-test-data\" target=\"_blank\">here</a>)</p>",
              "rawMarkdown": "doesn't that mean that we need to setup different feature creation strategy?\nWe don't need to group by level group anymore as we have all the level groups at hand from the beginning (see [here](https://www.kaggle.com/code/kmitsuhiro/api-update-iter-test-returns-all-test-data))",
              "votes": 2
            },
            {
              "id": 2190900,
              "postDate": "2023-03-21T15:05:55.983Z",
              "content": "<p>yes, we'll have to try some new features. Also we don't need anymore to save data from previous groups to predict next groups.</p>",
              "rawMarkdown": "yes, we'll have to try some new features. Also we don't need anymore to save data from previous groups to predict next groups.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2190637,
          "postDate": "2023-03-21T11:43:59.293Z",
          "content": "<p>Thanks for the findings!</p>\n<blockquote>\n  <p>But now it contains multiple <code>session_id</code></p>\n</blockquote>\n<p>A related problem would then be, how many session IDs would be given in the actual test run (in the sample we see 3 only)? This determines how large the given dataframe size is per iteration, and could be important because we may run out of memory.</p>",
          "rawMarkdown": "Thanks for the findings!\n\n> But now it contains multiple `session_id`\n\nA related problem would then be, how many session IDs would be given in the actual test run (in the sample we see 3 only)? This determines how large the given dataframe size is per iteration, and could be important because we may run out of memory.",
          "votes": 1,
          "replies": [
            {
              "id": 2190989,
              "postDate": "2023-03-21T16:23:00.317Z",
              "content": "<p>Hi. This is an issue with the set of 3 provided public samples. The hidden test set does not have this problem. I will fix the sample set. Thanks for spotting it!</p>",
              "rawMarkdown": "Hi. This is an issue with the set of 3 provided public samples. The hidden test set does not have this problem. I will fix the sample set. Thanks for spotting it!",
              "votes": 7
            },
            {
              "id": 2191302,
              "postDate": "2023-03-21T21:20:12.470Z",
              "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> !</p>",
              "rawMarkdown": "Thanks a lot @philculliton !",
              "votes": 1
            },
            {
              "id": 2191393,
              "postDate": "2023-03-21T23:43:45.230Z",
              "content": "<p>Only with the sample set,<br>\nI understand clearly now.<br>\nThank you very much! <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a></p>",
              "rawMarkdown": "Only with the sample set,\nI understand clearly now.\nThank you very much! @philculliton",
              "votes": 1
            },
            {
              "id": 2191558,
              "postDate": "2023-03-22T04:03:21.097Z",
              "content": "<p>Thanks a lot :D</p>",
              "rawMarkdown": "Thanks a lot :D",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2199333,
      "postDate": "2023-03-27T16:33:36.787Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI have a question about the Kaggle API, <code>for (test, sample_submission,) in iter_test:</code><br>\nThanks for fixing the error that the sample test contains 3 level_groups. A few days ago I confirmed that both sample test and hidden test contain 1 level_group.</p>\n<p>But about the session_id…<br>\n<strong>The hidden test contains 1 session_id, but the sample test contains 3 session_ids.</strong><br>\nBefore the update on Mar 21, I remember that the old sample test contained 1 session_id.<br>\nThis complicates the Submission code. It looks like some baseline notebooks also needed to be changed.</p>\n<p>Is there a special reason for this? If not, I think it would be better for the sample test to contain 1 session_id.<br>\nI am a beginner Kaggler, so I'm sorry if my info is wrong.</p>",
      "rawMarkdown": "Hi @philculliton \nI have a question about the Kaggle API, `for (test, sample_submission,) in iter_test:`\nThanks for fixing the error that the sample test contains 3 level_groups. A few days ago I confirmed that both sample test and hidden test contain 1 level_group.\n\nBut about the session_id...\n**The hidden test contains 1 session_id, but the sample test contains 3 session_ids.**\nBefore the update on Mar 21, I remember that the old sample test contained 1 session_id.\nThis complicates the Submission code. It looks like some baseline notebooks also needed to be changed.\n\nIs there a special reason for this? If not, I think it would be better for the sample test to contain 1 session_id.\nI am a beginner Kaggler, so I'm sorry if my info is wrong.",
      "votes": 8,
      "replies": [
        {
          "id": 2204800,
          "postDate": "2023-04-01T00:57:40.917Z",
          "content": "<p>Could you find any answer to this issue? I'm also having troubles with this 😪</p>",
          "rawMarkdown": "Could you find any answer to this issue? I'm also having troubles with this 😪",
          "votes": 1,
          "replies": [
            {
              "id": 2205956,
              "postDate": "2023-04-02T06:02:29.587Z",
              "content": "<p>Thanks for your comments,@javihm77</p>\n<p>I was helped by the method shown in the discussion below.<br>\nI have used it to customize my submission code.<br>\nThanks for the helpful discussions.If not for these tips I would have had a very hard time.</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396468\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396468</a><br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396751\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396751</a><br>\n<a href=\"https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments\" target=\"_blank\">https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments</a><br>\n<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680/comments\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680/comments</a><br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732</a></p>\n<p>I hope they are helpful to you.</p>",
              "rawMarkdown": "Thanks for your comments,@javihm77\n\nI was helped by the method shown in the discussion below.\nI have used it to customize my submission code.\nThanks for the helpful discussions.If not for these tips I would have had a very hard time.\n\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396468\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396751\nhttps://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments\nhttps://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680/comments\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732\n\nI hope they are helpful to you.",
              "votes": 2
            },
            {
              "id": 2211092,
              "postDate": "2023-04-05T19:38:23.253Z",
              "content": "<p>I owe you one <a href=\"https://www.kaggle.com/tanakaakinori\" target=\"_blank\">@tanakaakinori</a> thanks a lot! you saved me a lot of time and stress, it works 🤝</p>",
              "rawMarkdown": "I owe you one @tanakaakinori thanks a lot! you saved me a lot of time and stress, it works 🤝",
              "votes": 1
            }
          ]
        },
        {
          "id": 2204935,
          "postDate": "2023-04-01T04:57:06.293Z",
          "content": "<p>according to kaggle staff reply (link below), the sample data has been fixed. However I am still facing the same submission issue as <a href=\"https://www.kaggle.com/tanakaakinori\" target=\"_blank\">@tanakaakinori</a> </p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2199072\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2199072</a></p>\n<p>It feels very frustrating that we spent so much of our very limited spare time on this competition, only for the submission sample data to change (through no fault of our own) and now, we also have to figure out how to properly update our submission code for the new sample data (where there are now 3 session_ids instead of 1).</p>\n<p>Plus there was no proper / clear announcement that submission sample data has been \"fixed\" (maybe a sticky comment at the top? updating the original post? putting out a code example that works for the new sample data?) - we have to trawl this forum thread to find out. This level of effort is going to dissuade all but the most committed kagglers from taking part in this competition. All the wasted time on admin issues could have been avoided, if there has been better communication with participants.</p>\n<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> </p>",
          "rawMarkdown": "according to kaggle staff reply (link below), the sample data has been fixed. However I am still facing the same submission issue as @tanakaakinori \n\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2199072\n\nIt feels very frustrating that we spent so much of our very limited spare time on this competition, only for the submission sample data to change (through no fault of our own) and now, we also have to figure out how to properly update our submission code for the new sample data (where there are now 3 session_ids instead of 1).\n\nPlus there was no proper / clear announcement that submission sample data has been \"fixed\" (maybe a sticky comment at the top? updating the original post? putting out a code example that works for the new sample data?) - we have to trawl this forum thread to find out. This level of effort is going to dissuade all but the most committed kagglers from taking part in this competition. All the wasted time on admin issues could have been avoided, if there has been better communication with participants.\n\n@philculliton @alexmlfranklin ",
          "votes": 6,
          "replies": [
            {
              "id": 2205962,
              "postDate": "2023-04-02T06:09:20.257Z",
              "content": "<p>Thanks for your comment, <a href=\"https://www.kaggle.com/ruhong\" target=\"_blank\">@ruhong</a></p>\n<p>I think it is disconcerting that the number of session_id behaves different ways between the sample test and hidden test.<br>\nIf improved, it might remove the barrier for beginners and those who are just starting to work on this competition and allow us to focus on their original task.</p>",
              "rawMarkdown": "Thanks for your comment, @ruhong\n\nI think it is disconcerting that the number of session_id behaves different ways between the sample test and hidden test.\nIf improved, it might remove the barrier for beginners and those who are just starting to work on this competition and allow us to focus on their original task."
            }
          ]
        }
      ]
    },
    {
      "id": 2189743,
      "postDate": "2023-03-20T18:09:58.660Z",
      "content": "<p>Hi all - thanks Alex for the update! Here are the steps you should expect to see today:</p>\n<ol>\n<li>Submissions will be disabled while I update the test data and evaluation metric. Please note that the order of the dataframes has changed - <code>iter_test</code> will now yield (test, sample_submission) rather than (sample_submission, test). This was in response to a bug report where the API was throwing exceptions for some people.</li>\n<li>I will reenable submissions once the test data and metric are updated.</li>\n<li>I will update the training data today, but it may take longer due to the increased volume of data.</li>\n<li>All previous notebooks will be re-run during the coming week.</li>\n</ol>",
      "rawMarkdown": "Hi all - thanks Alex for the update! Here are the steps you should expect to see today:\n1. Submissions will be disabled while I update the test data and evaluation metric. Please note that the order of the dataframes has changed - `iter_test` will now yield (test, sample_submission) rather than (sample_submission, test). This was in response to a bug report where the API was throwing exceptions for some people.\n2. I will reenable submissions once the test data and metric are updated.\n3. I will update the training data today, but it may take longer due to the increased volume of data.\n4. All previous notebooks will be re-run during the coming week.",
      "votes": 7,
      "replies": [
        {
          "id": 2189766,
          "postDate": "2023-03-20T18:34:17.233Z",
          "content": "<p>Phil, won't all notebook reruns fail because all notebooks use:</p>\n<pre><code>for (sample_submission, test) in iter_test:\n</code></pre>",
          "rawMarkdown": "Phil, won't all notebook reruns fail because all notebooks use:\n\n    for (sample_submission, test) in iter_test:",
          "votes": 11,
          "replies": [
            {
              "id": 2189998,
              "postDate": "2023-03-21T00:36:36.973Z",
              "content": "<p>I believe so. I just checked to see my 2 submission results and both failed whilst deducting the number of submissions available to me today. I was not aware of this announcement until after my submissions failed and I then searched for a reason. Copying <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> </p>",
              "rawMarkdown": "I believe so. I just checked to see my 2 submission results and both failed whilst deducting the number of submissions available to me today. I was not aware of this announcement until after my submissions failed and I then searched for a reason. Copying @philculliton ",
              "votes": 1
            },
            {
              "id": 2190623,
              "postDate": "2023-03-21T11:39:04.537Z",
              "content": "<p>Agree! 😧<br>\nAnd I don't get how fixing the API means the tuple order has to be changed…</p>",
              "rawMarkdown": "Agree! 😧\nAnd I don't get how fixing the API means the tuple order has to be changed...",
              "votes": 1
            },
            {
              "id": 2191000,
              "postDate": "2023-03-21T16:27:09.180Z",
              "content": "<p>Specifically, the order of the returned dataframes needed to change to eliminate the potential bug. This was an API issue that needed to be resolved this way.</p>",
              "rawMarkdown": "Specifically, the order of the returned dataframes needed to change to eliminate the potential bug. This was an API issue that needed to be resolved this way.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2191719,
      "postDate": "2023-03-22T06:52:28.423Z",
      "content": "<p>The iter_test still returns all the test data. When will this be fixed?</p>",
      "rawMarkdown": "The iter_test still returns all the test data. When will this be fixed?",
      "votes": 5,
      "replies": [
        {
          "id": 2191732,
          "postDate": "2023-03-22T07:08:41.120Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>,</p>\n<p>I submitted a notebook about 40min ago, which successfully pass now (get LB 0.692). I think iter_test returns all the test data only for 3 provided public samples, as discussed <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190989\" target=\"_blank\">here</a>. As for hidden test set, time series API provides <strong>one level_group of one session</strong> for each iteration. Some of code snippets of my submission is shown as follows:</p>\n<pre><code>for (test, sample_submission) in iter_test:\n    # Process the data and generate features\n\n    try:\n        # Run inference using uploaded models\n        # For one level_group of one session\n        y_pred = quick_infer((x, x_cat), models[cur_lv_gp])\n        y_pred = (y_pred &gt; best_thres).astype(np.int)\n\n        sample_submission.loc[:, \"correct\"] = y_pred\n    except:\n        # This part is used to pass public samples\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission)\n</code></pre>\n<p>Hope this helps, thanks!</p>",
          "rawMarkdown": "Hi @cpmpml,\n\nI submitted a notebook about 40min ago, which successfully pass now (get LB 0.692). I think iter_test returns all the test data only for 3 provided public samples, as discussed [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190989). As for hidden test set, time series API provides **one level_group of one session** for each iteration. Some of code snippets of my submission is shown as follows:\n\n```\nfor (test, sample_submission) in iter_test:\n    # Process the data and generate features\n    \n    try:\n        # Run inference using uploaded models\n        # For one level_group of one session\n        y_pred = quick_infer((x, x_cat), models[cur_lv_gp])\n        y_pred = (y_pred > best_thres).astype(np.int)\n\n        sample_submission.loc[:, \"correct\"] = y_pred\n    except:\n        # This part is used to pass public samples\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission)\n```\n\nHope this helps, thanks!",
          "votes": 8,
          "replies": [
            {
              "id": 2191806,
              "postDate": "2023-03-22T08:02:46.160Z",
              "content": "<p>Thanks. Are you sure you get only one session_d with one group at a time?</p>",
              "rawMarkdown": "Thanks. Are you sure you get only one session_d with one group at a time?",
              "votes": 2
            },
            {
              "id": 2192259,
              "postDate": "2023-03-22T14:48:20.860Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I think it is the case. Since trying his code snipet does not yield an error message.</p>",
              "rawMarkdown": "Hi @cpmpml I think it is the case. Since trying his code snipet does not yield an error message.",
              "votes": 4
            },
            {
              "id": 2196071,
              "postDate": "2023-03-25T06:01:10.517Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>,</p>\n<p>I think it's true for hidden test set. For further confirmation, I submit <a href=\"https://www.kaggle.com/code/abaojiang/check-new-hidden-test-leakage\" target=\"_blank\">this notebook</a> and get LB score equal to <strong>zero prediction baseline</strong>, which obtains 0.216 on new test set. Checking part is in version 2 and zero prediction baseline in version 3.</p>\n<p>If there's any mistake, please let me know. Thanks a lot!</p>",
              "rawMarkdown": "Hi @cpmpml,\n\nI think it's true for hidden test set. For further confirmation, I submit [this notebook](https://www.kaggle.com/code/abaojiang/check-new-hidden-test-leakage) and get LB score equal to **zero prediction baseline**, which obtains 0.216 on new test set. Checking part is in version 2 and zero prediction baseline in version 3.\n\nIf there's any mistake, please let me know. Thanks a lot!"
            }
          ]
        }
      ]
    },
    {
      "id": 2190027,
      "postDate": "2023-03-21T01:14:13.647Z",
      "content": "<p>It seems that the api will give all of the levels at once now. Has anyone else encountered this issue? </p>",
      "rawMarkdown": "It seems that the api will give all of the levels at once now. Has anyone else encountered this issue? ",
      "votes": 5,
      "replies": [
        {
          "id": 2190271,
          "postDate": "2023-03-21T06:41:02.293Z",
          "content": "<p>Yes I see the same, all 18 questions at once in the same iteration. So we need to rewrite the submission code.</p>",
          "rawMarkdown": "Yes I see the same, all 18 questions at once in the same iteration. So we need to rewrite the submission code.",
          "votes": 1,
          "replies": [
            {
              "id": 2190992,
              "postDate": "2023-03-21T16:23:33.277Z",
              "content": "<p>This is an issue with the sample data that does not exist in the hidden test set. I will update the sample data. Thanks!</p>",
              "rawMarkdown": "This is an issue with the sample data that does not exist in the hidden test set. I will update the sample data. Thanks!"
            },
            {
              "id": 2191301,
              "postDate": "2023-03-21T21:19:40.680Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2191975,
              "postDate": "2023-03-22T10:50:49.500Z",
              "content": "<p>Oh, I see. Then I will try to retrain and resubmit with no change to the code. I can see that the training is now using 23,562 unique sessions instead of the previous 11,779. It can only be completed with the GPU as the CPU runs out of memory. </p>",
              "rawMarkdown": "Oh, I see. Then I will try to retrain and resubmit with no change to the code. I can see that the training is now using 23,562 unique sessions instead of the previous 11,779. It can only be completed with the GPU as the CPU runs out of memory. "
            }
          ]
        }
      ]
    },
    {
      "id": 2262429,
      "postDate": "2023-05-16T23:25:07.303Z",
      "content": "<p>Hi!<br>\nSince the leaderboard update, some of my submissions are getting \"Submission Scoring Error\". Even when I have the exact same output code as other successful submissions, I still get the same error.<br>\nI have also confirmed that when I run it against test.csv, the output is correct and expected.<br>\nIs there something wrong with the Scoring process?<br>\nHas anyone else experienced a similar event?</p>",
      "rawMarkdown": "Hi!\nSince the leaderboard update, some of my submissions are getting \"Submission Scoring Error\". Even when I have the exact same output code as other successful submissions, I still get the same error.\nI have also confirmed that when I run it against test.csv, the output is correct and expected.\nIs there something wrong with the Scoring process?\nHas anyone else experienced a similar event?",
      "votes": 3
    },
    {
      "id": 2295066,
      "postDate": "2023-06-10T15:02:35.327Z",
      "content": "<p>I'm so happy to be joining my first official competition here! Hope I can learn a lot and share what I've learned</p>",
      "rawMarkdown": "I'm so happy to be joining my first official competition here! Hope I can learn a lot and share what I've learned",
      "votes": 2
    },
    {
      "id": 2206112,
      "postDate": "2023-04-02T09:50:54.863Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> ,<br>\ndoes the shared raw <a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">data</a> contain true labels ?</p>",
      "rawMarkdown": "Hi @alexmlfranklin ,\ndoes the shared raw [data](https://fielddaylab.wisc.edu/opengamedata/) contain true labels ?",
      "votes": 4,
      "replies": [
        {
          "id": 2207796,
          "postDate": "2023-04-03T16:26:40.417Z",
          "content": "<p>After a quick look, the event logs from this public data do contain additional events, and those events show how the player answered the questions we have to predict. There is, as far as I can see, NO direct labels, but they could be inferred from the text of the answer itself.</p>\n<p>For example, this event shows the correct answer to the question (of course, we still need to check that there was no wrong answer given before it):</p>\n<p><code>\"22020006390905390\"     \"JOWILDER\"      2022-03-27 10:44:07.747000      \"CUSTOM.11\" {\"cur_cmd_fqid\": \"tunic.capitol_0.hall.boss.chap1_finale_plaquefirst_1\", \"cur_cmd_type\": 1, \"event_custom\": 11, \"fqid\": \"chap1_finale\", \"http_user_agent\": \"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_6) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.1.2 Safari/605.1.15\", \"interacted_fqid\": \"tunic.entry_tunic\", \"level\": 4, \"name\": \"choice\", \"persistent_session_id\": \"22020006390905390\", \"remote_addr\": \"216.15.53.200\", \"room_coor\": [-569.5578218361513, -169.70583676184344], \"room_fqid\": \"tunic.capitol_0.hall\", \"screen_coor\": [92, 420], \"server_time\": \"2022-03-27T05:44:07\", \"subtype\": \"wildcard\", \"text\": \"Our shirt was around way before the women's basketball team!\", \"type\": \"click\"}  \"10\" \"0\"  None   \"\"   {}   {}  \"162\"\n</code></p>\n<p>So, with a bit of cleaning and processing this dataset may be used to increase the training data. </p>\n<p>Whether this open data actually contains our training or test data, is an open question. </p>",
          "rawMarkdown": "After a quick look, the event logs from this public data do contain additional events, and those events show how the player answered the questions we have to predict. There is, as far as I can see, NO direct labels, but they could be inferred from the text of the answer itself.\n\nFor example, this event shows the correct answer to the question (of course, we still need to check that there was no wrong answer given before it):\n\n`\"22020006390905390\"     \"JOWILDER\"      2022-03-27 10:44:07.747000      \"CUSTOM.11\" {\"cur_cmd_fqid\": \"tunic.capitol_0.hall.boss.chap1_finale_plaquefirst_1\", \"cur_cmd_type\": 1, \"event_custom\": 11, \"fqid\": \"chap1_finale\", \"http_user_agent\": \"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_6) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.1.2 Safari/605.1.15\", \"interacted_fqid\": \"tunic.entry_tunic\", \"level\": 4, \"name\": \"choice\", \"persistent_session_id\": \"22020006390905390\", \"remote_addr\": \"216.15.53.200\", \"room_coor\": [-569.5578218361513, -169.70583676184344], \"room_fqid\": \"tunic.capitol_0.hall\", \"screen_coor\": [92, 420], \"server_time\": \"2022-03-27T05:44:07\", \"subtype\": \"wildcard\", \"text\": \"Our shirt was around way before the women's basketball team!\", \"type\": \"click\"}  \"10\" \"0\"  None   \"\"   {}   {}  \"162\"\n`\n\nSo, with a bit of cleaning and processing this dataset may be used to increase the training data. \n\nWhether this open data actually contains our training or test data, is an open question. ",
          "votes": 8
        },
        {
          "id": 2212154,
          "postDate": "2023-04-06T14:58:28.497Z",
          "content": "<p>Hi Reacher, the raw data doesn't contain the labels directly but can be inferred from it (as Alex rightly pointed out). The test data for the competition is not on the open data site.</p>",
          "rawMarkdown": "Hi Reacher, the raw data doesn't contain the labels directly but can be inferred from it (as Alex rightly pointed out). The test data for the competition is not on the open data site.",
          "votes": 8,
          "replies": [
            {
              "id": 2213751,
              "postDate": "2023-04-07T20:29:23.960Z",
              "content": "<p>Thanks for the interesting discussion.<br>\nI checked the raw data and found 125883 session_ids. Of those, 3785 session_ids were also included in the train data. Also, as you pointed out, none were found in the hidden test.</p>",
              "rawMarkdown": "Thanks for the interesting discussion.\nI checked the raw data and found 125883 session_ids. Of those, 3785 session_ids were also included in the train data. Also, as you pointed out, none were found in the hidden test.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2195220,
      "postDate": "2023-03-24T13:58:13.847Z",
      "content": "<p>Today ,I have shared my first code on kaggle. </p>",
      "rawMarkdown": "Today ,I have shared my first code on kaggle. ",
      "votes": 2
    },
    {
      "id": 2204973,
      "postDate": "2023-04-01T05:43:30.667Z",
      "content": "<p>The iter_test still returns all the test data. When will this be fixed?❌</p>",
      "rawMarkdown": "The iter_test still returns all the test data. When will this be fixed?❌",
      "votes": 1
    },
    {
      "id": 2196187,
      "postDate": "2023-03-25T07:57:48.800Z",
      "content": "<p>the train.csv is so big, it is impossible to work on this .csv file with kaggle platform. Even if I was using colab and pycharm to work on it, it still took a very long time to read this .csv. Oh my god.</p>",
      "rawMarkdown": "the train.csv is so big, it is impossible to work on this .csv file with kaggle platform. Even if I was using colab and pycharm to work on it, it still took a very long time to read this .csv. Oh my god.",
      "votes": 1,
      "replies": [
        {
          "id": 2196433,
          "postDate": "2023-03-25T12:04:07.580Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kaianchan\" target=\"_blank\">@kaianchan</a>,</p>\n<p>It's still possible to play around with this new training set on Kaggle platform. Details are demonstrated by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396979\" target=\"_blank\">this thread</a>.</p>\n<p>Hope this helps, thanks!</p>",
          "rawMarkdown": "Hi @kaianchan,\n\nIt's still possible to play around with this new training set on Kaggle platform. Details are demonstrated by [@cdeotte](https://www.kaggle.com/cdeotte) in [this thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396979).\n\nHope this helps, thanks!",
          "votes": 1
        },
        {
          "id": 2197436,
          "postDate": "2023-03-26T06:45:45.823Z",
          "content": "<p>The train.csv can be downloaded in 953 MB - <a href=\"https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb\" target=\"_blank\">https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb</a></p>",
          "rawMarkdown": "The train.csv can be downloaded in 953 MB - https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb",
          "votes": 2
        },
        {
          "id": 2198564,
          "postDate": "2023-03-27T06:50:01.580Z",
          "content": "<p>I'd recommend trying <code>polars</code> instead of <code>pandas</code>. It's more efficient in memory usage and provides <a href=\"https://www.kaggle.com/code/demche/polars-memory-usage-optimization\" target=\"_blank\">more options</a> for memory usage optimization</p>",
          "rawMarkdown": "I'd recommend trying `polars` instead of `pandas`. It's more efficient in memory usage and provides [more options](https://www.kaggle.com/code/demche/polars-memory-usage-optimization) for memory usage optimization",
          "votes": 1,
          "replies": [
            {
              "id": 2199055,
              "postDate": "2023-03-27T13:16:15.740Z",
              "content": "<p>I second this.</p>",
              "rawMarkdown": "I second this."
            }
          ]
        },
        {
          "id": 2204970,
          "postDate": "2023-04-01T05:42:39.957Z",
          "content": "<p>Downloaded train.csv can be in 953 MB - <a href=\"https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb\" target=\"_blank\">https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb</a></p>",
          "rawMarkdown": "Downloaded train.csv can be in 953 MB - https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb",
          "votes": -1
        }
      ]
    },
    {
      "id": 2190286,
      "postDate": "2023-03-21T06:56:23.670Z",
      "content": "<p>Does the test size also double ?</p>",
      "rawMarkdown": "Does the test size also double ?",
      "votes": 1
    },
    {
      "id": 2190077,
      "postDate": "2023-03-21T02:40:58.217Z",
      "content": "<p>More data more work 😷</p>",
      "rawMarkdown": "More data more work 😷",
      "votes": 1,
      "replies": [
        {
          "id": 2260803,
          "postDate": "2023-05-15T22:16:08.457Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> </p>\n<p>How math \"more data\" improve score?</p>\n<p>Do you remember your score befor  and after ?</p>",
          "rawMarkdown": "Hi, @minhtu123 \n\nHow math \"more data\" improve score?\n\nDo you remember your score befor  and after ?"
        }
      ]
    },
    {
      "id": 2191962,
      "postDate": "2023-03-22T10:39:40.597Z",
      "content": "<p>Well, guys, as I understand all leaks (I mean 2% leak and API issue) were fixed. But leaderboard's notebooks were not reruned. Am I correct?</p>",
      "rawMarkdown": "Well, guys, as I understand all leaks (I mean 2% leak and API issue) were fixed. But leaderboard's notebooks were not reruned. Am I correct?",
      "votes": 2,
      "replies": [
        {
          "id": 2193433,
          "postDate": "2023-03-23T09:46:03.800Z",
          "content": "<p>I am curious too. As Chris said <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191405\" target=\"_blank\">here</a>, it seems to exist, Kaggle hasn't fixed the submission file as well, which makes the inference quite tricker.</p>",
          "rawMarkdown": "I am curious too. As Chris said [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191405), it seems to exist, Kaggle hasn't fixed the submission file as well, which makes the inference quite tricker.",
          "votes": 1,
          "replies": [
            {
              "id": 2197609,
              "postDate": "2023-03-26T08:51:00.220Z",
              "content": "<p>Yes, still don't understand status of competition… is it still leaked/bugged? It is meta-question.</p>",
              "rawMarkdown": "Yes, still don't understand status of competition... is it still leaked/bugged? It is meta-question.",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2190003,
      "postDate": "2023-03-21T00:38:41.273Z",
      "content": "<p>Thank you for the update regarding the data leak and fix.</p>\n<p>As I am writing the comment, it appears that the new training data set has been uploaded, its size comes to 4.72GB. Without any memory management, my notebook is not able to run due to the error:</p>\n<blockquote>\n  <p>Your notebook tried to allocate more memory than is available. It has restarted.</p>\n</blockquote>\n<p>Many kagglers have shared tips of downsizing the memory footprint, this error makes me believe now it is time to explore and implement those ideas.</p>\n<p>Additionally, basing on heuristics, many have observed data issues in both training and test sets (for example, <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395686\" target=\"_blank\">here</a>). Will this data update address these issues? How about the hidden test set?</p>",
      "rawMarkdown": "Thank you for the update regarding the data leak and fix.\n\nAs I am writing the comment, it appears that the new training data set has been uploaded, its size comes to 4.72GB. Without any memory management, my notebook is not able to run due to the error:\n>Your notebook tried to allocate more memory than is available. It has restarted.\n\nMany kagglers have shared tips of downsizing the memory footprint, this error makes me believe now it is time to explore and implement those ideas.\n\nAdditionally, basing on heuristics, many have observed data issues in both training and test sets (for example, [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250) and [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395686)). Will this data update address these issues? How about the hidden test set?",
      "votes": 2
    },
    {
      "id": 2194575,
      "postDate": "2023-03-24T03:45:40.953Z",
      "content": "<p>Thank you for the update and for addressing the issue promptly. We appreciate the transparency and effort to maintain the integrity of the competition. We look forward to incorporating the additional data and continuing to work on our models.</p>",
      "rawMarkdown": "Thank you for the update and for addressing the issue promptly. We appreciate the transparency and effort to maintain the integrity of the competition. We look forward to incorporating the additional data and continuing to work on our models."
    },
    {
      "id": 2190122,
      "postDate": "2023-03-21T03:54:26.453Z",
      "content": "<p>Thanks Dear <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>,<br>\nThanks for sharing such a valuable post. I appreciate the transparency and swift action taken by the organizers to address the leaked competition data. By releasing the entire current test set and extending the competition deadline, they have ensured fairness for all participants. It's great to see the Kaggle community's integrity in reporting such issues and the organizers' commitment to maintaining the integrity of the competition.👍</p>",
      "rawMarkdown": "Thanks Dear @alexmlfranklin,\nThanks for sharing such a valuable post. I appreciate the transparency and swift action taken by the organizers to address the leaked competition data. By releasing the entire current test set and extending the competition deadline, they have ensured fairness for all participants. It's great to see the Kaggle community's integrity in reporting such issues and the organizers' commitment to maintaining the integrity of the competition.👍",
      "votes": -1
    },
    {
      "id": 2236091,
      "postDate": "2023-04-26T14:24:28.090Z",
      "content": "<p>Hi<br>\nI am new here.</p>",
      "rawMarkdown": "Hi\nI am new here.",
      "votes": -4
    },
    {
      "id": 2197300,
      "postDate": "2023-03-26T04:07:04.570Z",
      "content": "<p>New platform for me</p>",
      "rawMarkdown": "New platform for me",
      "votes": -3
    },
    {
      "id": 2196497,
      "postDate": "2023-03-25T12:58:38.083Z",
      "content": "<p>Hi <br>\nI am new here.</p>",
      "rawMarkdown": "Hi \nI am new here.",
      "votes": -7
    },
    {
      "id": 2307273,
      "postDate": "2023-06-18T04:11:30.183Z",
      "content": "<p>Why is the time complexity of the baseline model very low, but my model has high time complexity and is not as good as its performance? I am very depressed. What perspective should I try to solve this problem from?</p>",
      "rawMarkdown": "Why is the time complexity of the baseline model very low, but my model has high time complexity and is not as good as its performance? I am very depressed. What perspective should I try to solve this problem from?"
    },
    {
      "id": 2280896,
      "postDate": "2023-05-30T12:53:29.930Z",
      "content": "<p>Hi! new here<br>\nI have one doubt related to the test dataset.<br>\nWe are getting the test dataset in level-group sessions. Let's suppose I got the test dataset for level-group 0-4 and we predicted the label for questions 1-3 using this dataset and now we will get the dataset for level-group 5-13. Do we also simultaneously get the <strong>true</strong> labels for level-group 0-4? (meaning how well they have truly performed in the previous level-group).</p>",
      "rawMarkdown": "Hi! new here\nI have one doubt related to the test dataset.\nWe are getting the test dataset in level-group sessions. Let's suppose I got the test dataset for level-group 0-4 and we predicted the label for questions 1-3 using this dataset and now we will get the dataset for level-group 5-13. Do we also simultaneously get the **true** labels for level-group 0-4? (meaning how well they have truly performed in the previous level-group)."
    },
    {
      "id": 2278217,
      "postDate": "2023-05-28T12:46:05.283Z",
      "content": "<p>Thank you for the advice you gave, it can be a great inspiration to me.</p>",
      "rawMarkdown": "Thank you for the advice you gave, it can be a great inspiration to me."
    },
    {
      "id": 2266347,
      "postDate": "2023-05-19T23:29:12.580Z",
      "content": "<p>at this site (<a href=\"https://fielddaylab.wisc.edu/opengamedata/):\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/):</a></p>\n<p>Research   ( <a href=\"https://arxiv.org/pdf/2210.09906.pdf\" target=\"_blank\">https://arxiv.org/pdf/2210.09906.pdf</a> )<br>\nThis resulted in four versions of the script:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2F3c5b4c4dc9b7414930dee18bd2f3dc7e%2F2023-05-20%20%2001.57.45.png?generation=1684537140471281&amp;alt=media\" alt=\"\"></p>\n<ol>\n<li>Snark + Humor (Original)</li>\n<li>Humor, no Snark </li>\n<li>Snark, no Humor</li>\n<li>Dry</li>\n</ol>\n<p>In train.csv:</p>\n<table>\n<thead>\n<tr>\n<th>session_id</th>\n<th>versions of the script</th>\n<th>text_x</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>22060015173801296</td>\n<td>1. Snark + Humor (Original)</td>\n<td>Did you do all of them?</td>\n</tr>\n<tr>\n<td>22060001580924636</td>\n<td>2. Humor, no Snark</td>\n<td>He's always trying to get you in trouble, and he doesn't like animals!</td>\n</tr>\n<tr>\n<td>22060014314686290</td>\n<td>2. Humor, no Snark</td>\n<td>He's always trying to get you in trouble, and he doesn't like animals!</td>\n</tr>\n<tr>\n<td>22060010020384030</td>\n<td>3. Snark, no Humor</td>\n<td>So? History is boring!</td>\n</tr>\n<tr>\n<td>22060013403542770</td>\n<td>3. Snark, no Humor</td>\n<td>So? History is boring!</td>\n</tr>\n<tr>\n<td>22060015102434092</td>\n<td>3. Snark, no Humor</td>\n<td>So? History is boring!</td>\n</tr>\n<tr>\n<td>22060008392175650</td>\n<td>4. Dry</td>\n<td>Yes! This old slip from 1916.</td>\n</tr>\n<tr>\n<td>22060008395913956</td>\n<td>4. Dry</td>\n<td>Yes! This old slip from 1916.</td>\n</tr>\n<tr>\n<td>22060011583687136</td>\n<td>4. Dry</td>\n<td>Yes! This old slip from 1916.</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://arxiv.org/pdf/2210.09906.pdf:\" target=\"_blank\">https://arxiv.org/pdf/2210.09906.pdf:</a></p>",
      "rawMarkdown": "at this site (https://fielddaylab.wisc.edu/opengamedata/):\n\nResearch   ( https://arxiv.org/pdf/2210.09906.pdf )\nThis resulted in four versions of the script:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2F3c5b4c4dc9b7414930dee18bd2f3dc7e%2F2023-05-20%20%2001.57.45.png?generation=1684537140471281&alt=media)\n1. Snark + Humor (Original)\n2. Humor, no Snark \n3. Snark, no Humor\n4. Dry\n\nIn train.csv:\n| session_id | versions of the script |text_x |\n| --- | --- | --- |\n| 22060015173801296| 1. Snark + Humor (Original)  |\tDid you do all of them? |\n| 22060001580924636 | 2. Humor, no Snark  |He's always trying to get you in trouble, and he doesn't like animals!|\n| 22060014314686290| 2. Humor, no Snark  |He's always trying to get you in trouble, and he doesn't like animals!|\n| 22060010020384030\t| 3. Snark, no Humor |\tSo? History is boring! |\n| 22060013403542770| 3. Snark, no Humor |\tSo? History is boring! |\n| 22060015102434092| 3. Snark, no Humor |\tSo? History is boring! |\n| 22060008392175650|  4. Dry|\t \tYes! This old slip from 1916.|\n| 22060008395913956| 4. Dry|\t \tYes! This old slip from 1916.|\n| 22060011583687136| 4. Dry|\t \tYes! This old slip from 1916.|\n\n\n\nhttps://arxiv.org/pdf/2210.09906.pdf:"
    },
    {
      "id": 2260800,
      "postDate": "2023-05-15T22:10:45.907Z",
      "content": "<p>Hello,  <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> :)</p>\n<p>Who can remembered:<br>\n does score increase,  after  increase train.csv  size from 10000 to 20000 people  ?</p>",
      "rawMarkdown": "Hello,  @alexmlfranklin :)\n\nWho can remembered:\n does score increase,  after  increase train.csv  size from 10000 to 20000 people  ?",
      "replies": [
        {
          "id": 2289906,
          "postDate": "2023-06-06T12:35:21.333Z",
          "content": "<p>Yes, both increased by about 0.001+</p>",
          "rawMarkdown": "Yes, both increased by about 0.001+"
        }
      ]
    },
    {
      "id": 2260796,
      "postDate": "2023-05-15T22:04:40.980Z",
      "content": "<p>Alex, hello:)</p>\n<p>Before multiple train size in 2 tims,<br>\nHow look lider bord?<br>\nIt's function save leaderboard to file.</p>\n<p>Do anybody have saved leaderboard, as it wos 2 month ago?</p>",
      "rawMarkdown": "Alex, hello:)\n\nBefore multiple train size in 2 tims,\nHow look lider bord?\nIt's function save leaderboard to file.\n\nDo anybody have saved leaderboard, as it wos 2 month ago?"
    },
    {
      "id": 2240538,
      "postDate": "2023-04-30T16:05:13.933Z",
      "content": "<p>lol, i was like what happened to my notebook, it could run properly before monthes.</p>",
      "rawMarkdown": "lol, i was like what happened to my notebook, it could run properly before monthes."
    },
    {
      "id": 2224165,
      "postDate": "2023-04-17T04:08:37.910Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ahmeddatascientist\" target=\"_blank\">@ahmeddatascientist</a> ,I am a beginner kaggler and I am wondering how the question_labels connect with the dataset \"train.csv\",so can you give me some instructions?</p>",
      "rawMarkdown": "Hi @ahmeddatascientist ,I am a beginner kaggler and I am wondering how the question_labels connect with the dataset \"train.csv\",so can you give me some instructions?"
    },
    {
      "id": 2210352,
      "postDate": "2023-04-05T10:29:22.593Z",
      "content": "<p>great job!</p>",
      "rawMarkdown": "great job!"
    },
    {
      "id": 2196192,
      "postDate": "2023-03-25T08:01:32.530Z",
      "content": "<p>After these updates test data includes only 3 session ids, same with sample_submission. iter_test also does only 1 iteration returning these 3 sessions at once. Maybe issue is specific to my account.</p>\n<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a></p>",
      "rawMarkdown": "After these updates test data includes only 3 session ids, same with sample_submission. iter_test also does only 1 iteration returning these 3 sessions at once. Maybe issue is specific to my account.\n\n@alexmlfranklin @nrambis @philculliton @maggiemd"
    },
    {
      "id": 2193890,
      "postDate": "2023-03-23T15:20:03.020Z",
      "content": "<p>does fixing the API means the tuple order has to be changed…</p>",
      "rawMarkdown": "does fixing the API means the tuple order has to be changed…",
      "replies": [
        {
          "id": 2194681,
          "postDate": "2023-03-24T05:07:18.740Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ankitkumarece20423\" target=\"_blank\">@ankitkumarece20423</a>,</p>\n<p>Maybe <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191000\" target=\"_blank\">this thread</a> gives the answer you want. Hope this helps.</p>",
          "rawMarkdown": "Hi @ankitkumarece20423,\n\nMaybe [this thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191000) gives the answer you want. Hope this helps."
        }
      ]
    },
    {
      "id": 2193039,
      "postDate": "2023-03-23T04:15:20.777Z",
      "content": "<p>which  competition the data is from ?</p>",
      "rawMarkdown": "which  competition the data is from ?",
      "replies": [
        {
          "id": 2193042,
          "postDate": "2023-03-23T04:16:30.003Z",
          "content": "<p>and what possible problems a leak like this could do ?</p>",
          "rawMarkdown": "and what possible problems a leak like this could do ?"
        }
      ]
    },
    {
      "id": 2190184,
      "postDate": "2023-03-21T05:04:50.453Z",
      "content": "<p>What does it look like when the submissions are closed? I had some submissions that ran successfully but failed during scoring (about an hour before time of posting this) and I'm just wondering if that's because submissions are closed right now and I just didn't know, or if I have a bug that needs to be fixed</p>",
      "rawMarkdown": "What does it look like when the submissions are closed? I had some submissions that ran successfully but failed during scoring (about an hour before time of posting this) and I'm just wondering if that's because submissions are closed right now and I just didn't know, or if I have a bug that needs to be fixed",
      "replies": [
        {
          "id": 2190243,
          "postDate": "2023-03-21T06:11:47.447Z",
          "content": "<p>Same question here. I just submit three of my notebooks which follow the rules of new API, but all of them threw exception!</p>",
          "rawMarkdown": "Same question here. I just submit three of my notebooks which follow the rules of new API, but all of them threw exception!",
          "votes": 1
        },
        {
          "id": 2190249,
          "postDate": "2023-03-21T06:18:41.833Z",
          "content": "<blockquote>\n  <p>What does it look like when the submissions are closed? I had some submissions that ran successfully but failed during scoring (about an hour before time of posting this) and I'm just wondering if that's because submissions are closed right now and I just didn't know, or if I have a bug that needs to be fixed</p>\n</blockquote>\n<p>Same issue here. I re-submitted me notebook which previously managed to run without problem but threw me “out of memory” exception this time. Is it due to that submission is disabled?</p>",
          "rawMarkdown": "> What does it look like when the submissions are closed? I had some submissions that ran successfully but failed during scoring (about an hour before time of posting this) and I'm just wondering if that's because submissions are closed right now and I just didn't know, or if I have a bug that needs to be fixed\n\nSame issue here. I re-submitted me notebook which previously managed to run without problem but threw me “out of memory” exception this time. Is it due to that submission is disabled?",
          "votes": 1,
          "replies": [
            {
              "id": 2191446,
              "postDate": "2023-03-22T01:30:45.860Z",
              "content": "<p>Same issue here, I cannot run the code on Kaggle, but a week ago it was executed successfully.</p>",
              "rawMarkdown": "Same issue here, I cannot run the code on Kaggle, but a week ago it was executed successfully."
            }
          ]
        }
      ]
    },
    {
      "id": 2190099,
      "postDate": "2023-03-21T03:20:09.510Z",
      "content": "<p>cool. 卷起来！</p>",
      "rawMarkdown": "cool. 卷起来！"
    },
    {
      "id": 2205204,
      "postDate": "2023-04-01T10:39:38.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2196380,
      "postDate": "2023-03-25T11:01:15.857Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 2190390,
      "postDate": "2023-03-21T08:01:39.343Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2251611,
      "postDate": "2023-05-09T14:18:09.510Z",
      "content": "<p>That's great! Thank you!</p>",
      "rawMarkdown": "That's great! Thank you!"
    },
    {
      "id": 2224598,
      "postDate": "2023-04-17T13:10:58.983Z",
      "content": "<p>Thanks for your notice.</p>",
      "rawMarkdown": "Thanks for your notice."
    },
    {
      "id": 2218136,
      "postDate": "2023-04-11T13:13:13.877Z",
      "content": "<p>Thanks this is helpful.</p>",
      "rawMarkdown": "Thanks this is helpful."
    },
    {
      "id": 2204224,
      "postDate": "2023-03-31T12:27:49.713Z",
      "content": "<p>Thanks for the update.</p>",
      "rawMarkdown": "Thanks for the update."
    },
    {
      "id": 2201058,
      "postDate": "2023-03-29T02:36:35.427Z",
      "content": "<p>Hello! Thanks for the update.</p>",
      "rawMarkdown": "Hello! Thanks for the update."
    },
    {
      "id": 2193332,
      "postDate": "2023-03-23T08:15:36.573Z",
      "content": "<p>Hello! Thanks for the update.</p>",
      "rawMarkdown": "Hello! Thanks for the update."
    },
    {
      "id": 2191403,
      "postDate": "2023-03-21T23:49:15.770Z",
      "content": "<p>Hello! Thanks for the update.</p>",
      "rawMarkdown": "Hello! Thanks for the update."
    }
  ],
  "comments": [
    {
      "id": 2190090,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2023-03-21T03:03:46.213000",
      "content": "<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>, just as <a href=\"https://www.kaggle.com/tanakaakinori\" target=\"_blank\">@tanakaakinori</a> and others have observed, there are differences in both the training data and the test data. I have also observed the changes in the test data that resulted in submission error, which according to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> that we may have a leak again.</p>\n<p>Please fix everything and let us know the details when you are done. Perhaps you should disable submission again so that we do not waste our times.</p>\n<p><strong>Update:</strong><br>\nBased on the replies to <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>'s post below, there was and still is a leak described as:<br>\n<code>2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition.</code></p>\n<p>We need an update on if or when that one will be fixed <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>. Thanks</p>",
      "votes": 19,
      "replies": [
        {
          "id": 2191003,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2023-03-21T16:29:15.293000",
          "content": "<p>Hi! We do not have a (hidden) leak again. The issue people are describing exists only in the sample data, which we'll update.</p>\n<p>If you are seeing submission errors that are NOT related to the reordering of the tuple returned by the timeseries API, please feel free to reach out with a notebook that reproduces them.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2191077,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2023-03-21T17:21:56.333000",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, Is the sample test data updated yet? I just made a submission and got the \"Submission Scoring Error\" again. I used  <code>(test, sample_submission) in iter_test</code>:</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2191347,
              "author_name": "__ChrisQ__",
              "author_url": "",
              "post_date": "2023-03-21T22:53:00.773000",
              "content": "<p>Hi, I am still a bit confused. <br>\nBefore in the sample data, we have 1 session and 1 level_group per <code>test</code><br>\nNow we have 3 sessions and 3 level_groups per <code>test</code>.<br>\nDo you mean we have 3 sessions and 3 level_groups only in the sample data? Then how many do we have in the real test set?🤔</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2191356,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-03-21T23:05:45.543000",
              "content": "<p>No. This is a mistake. Kaggle is fixing this. The new API will have <strong>one</strong> level group and <strong>one</strong> session_id per iteration loop. Kaggle is fixing this now. See other comments in this discussion</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2191365,
              "author_name": "Bosco Yung",
              "author_url": "",
              "post_date": "2023-03-21T23:19:15.777000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>! Thanks again for all the fixes!<br>\nWill you make another announcement when things are settled?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2191390,
              "author_name": "stallone",
              "author_url": "",
              "post_date": "2023-03-21T23:39:57.670000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nIn my understanding. 2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. Is that right?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2191405,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-03-21T23:50:00.533000",
              "content": "<blockquote>\n  <p>In my understanding. 2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. Is that right?</p>\n</blockquote>\n<p>Yes, it is my understanding that this still exists. This was not fixed before the most recent data change (and not fixed after). And Kaggle staff has not commented about how they will address this yet.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2194193,
      "author_name": "Bertrand P",
      "author_url": "",
      "post_date": "2023-03-23T19:44:27.610000",
      "content": "<p>The problem seems unsolved…<br>\n<a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> </p>",
      "votes": 10,
      "replies": [
        {
          "id": 2195487,
          "author_name": "empty",
          "author_url": "",
          "post_date": "2023-03-24T17:11:23.640000",
          "content": "<p>0.753 is leak again? 👀</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2195517,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-24T17:36:48.340000",
              "content": "<p>I read Bertrand's post as the api problem is still not solved: the iter_test still returns all the test data. </p>\n<p>Re how we get 0.753 you will have to wait till after the competition end.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2198374,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-03-27T00:39:10.863000",
              "content": "<p>the main problem mentioned in this topic is Leaked Competition Data 🙃</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2198603,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-27T07:13:46.260000",
              "content": "<p>Several of us made comments about the api change, including me here: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191719\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191719</a></p>\n<p>Other comments about this issue:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2193890\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2193890</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190039\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190039</a></p>\n<p>I honestly don't get  this:</p>\n<ul>\n<li>How fixing a bug had to result in swapping arguments?</li>\n<li>Why we can't get a fixed code when running notebooks before submission?</li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2198837,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-03-27T10:20:08.827000",
              "content": "<p>good questions! in addition: why we had no reruns still after promising to do it when Chris found the first api bug? seems like this comp is not too important for kaggle </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2199032,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-27T13:02:08.763000",
              "content": "<p>Indeed lots of LB is from runs prior to the data reset. None of that is legit now.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2199039,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-27T13:06:50.727000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2199072,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-27T13:27:50.427000",
              "content": "<p>Hi. We did not update the LB because we had reports of other issues that we thought we might need to resolve; then the test data leak was revealed. It took some time to fix. We will be re-running the LB now that the public test set has been updated, and the sample data has been fixed.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2199075,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-27T13:30:02.140000",
              "content": "<p>Hi! The API is complex; the dataframes need to be served in a particular order to be fully supported. We needed to switch to avoid potential issues with the API.</p>\n<p>What's the fixed code you're referring to? The public sample data should no longer be returning all at once.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2199664,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-03-27T23:21:08.607000",
              "content": "<p>could you explain in a few words how will you rerun old submissions with the new api?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2200342,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-28T13:36:12.447000",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> before the data reset, iteration in inference code looked like:</p>\n<pre><code> (sample_submission, test)  iter_test:\n</code></pre>\n<p>After it looks like:</p>\n<pre><code> (test, sample_submission)  iter_test:\n</code></pre>\n<p>It has been days we have been complaining about the argument order change. In particular it means that code submitted before the data reset will no longer run.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2200492,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-28T15:32:03.407000",
              "content": "<p>Thanks. Yes, as mentioned in my original note, the order that the dataframes are returned in needed to be changed. We were seeing bugs that indicated this was causing problems in the API. Rather than let those problems continue, we made the update.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2200563,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-28T16:06:38.540000",
              "content": "<p>Then you get that the code written before the data reset will not longer run, right?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2201564,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-29T12:07:24.130000",
              "content": "<blockquote>\n  <p>the order that the dataframes are returned in needed to be changed. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> we are discussing the order of the arguments to the iter_test object. We are not discussing the order in which you return the dataframe.</p>\n<p>It is weird that you don't acknowledge this order has changed and that it causes a major issue for rerunning past submissions.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2201570,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-29T12:13:48.463000",
              "content": "<p>The order of the returned dataframes IS the order of the arguments in the iter_test object. That is precisely what I am saying. I noted this up front when the update was announced.</p>\n<p>I am aware that this causes issues for rerunning past submissions. It was, unfortunately, the better route to take due to reported bugs. These bugs were also noted up front when the update was announced, as the reason for the change.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2201578,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-29T12:23:57.417000",
              "content": "<p>ok, fair enough.</p>\n<p>I thought you were discussing the order in which level groups were returned.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2201589,
              "author_name": "heiligerl",
              "author_url": "",
              "post_date": "2023-03-29T12:36:40.807000",
              "content": "<p>Just to be on the same page: Since it is not easily possible to rerun the notebooks, the public leaderboard is erroneous. Did I understand that correctly?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2202623,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2023-03-30T07:21:18.313000",
              "content": "<p>So may I ask a question, for the lb 0.753, it is caused by api bug or leaked data? or it is just a normal score.</p>\n<p>I am quite confused here. <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> said the problem was unsolved. I am not sure what kind of problem it is and I thought it is leaked data as same as last time. But <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> mentioned that \"the iter_test still returns all the test data\". But as many people said there is <strong>only a bug in sample data</strong> that return all level_group data and <strong>in the test part it is fine</strong>. So it means that <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> found another bug to get all the test data in the iter_test?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2203402,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-30T19:13:37.027000",
              "content": "<blockquote>\n  <p>But as many people said there is only a bug in sample data that return all level_group data</p>\n</blockquote>\n<p>This is a bug in the code. The data is what it is. The code should iterate through session_id and level group.</p>\n<blockquote>\n  <p>for the lb 0.753, it is caused by api bug or leaked data? or it is just a normal score.</p>\n</blockquote>\n<p>This score is not legit. It was obtained using the leak that Bertrand discovered. The leak is now fixed and we have no way to reproduce that score anymore.  Same goes for the 0.72 submission.</p>\n<p>We only understood the issue very recently when we failed to reproduce these results.  We have asked Kaggle to remove these two submissions.</p>",
              "votes": 13,
              "replies": []
            },
            {
              "id": 2203670,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2023-03-31T03:11:17.940000",
              "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Cool. Thanks for explaining everything.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2209919,
              "author_name": "JohnM",
              "author_url": "",
              "post_date": "2023-04-05T03:14:53.937000",
              "content": "<p>As I recall, the order has always been test_df, sample_sub for time-series code comps. I didn't see any explanation why it was opposite at first for this one. Not a helpful comment perhaps, except that the order should always be the same for such competitions.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2190039,
      "author_name": "AT",
      "author_url": "",
      "post_date": "2023-03-21T01:36:00.407000",
      "content": "<p>Hello! Thanks for the update.<br>\nI found some changes.</p>\n<ul>\n<li>The data size of 'train', 'train_label' has increased. <br>\nIt ran out of CPU notebook's 8GB RAM,  No problem for GPU's one.</li>\n<li><code>for (sample_submission, test) in iter_test:</code> was changed to for <code>(test,  sample_submission) in iter_test:</code></li>\n<li><code>for (test, sample_submission) in iter_test:</code>, old <code>test</code> used to have  single <code>session_id</code>, single <code>level_group</code>. <br>\n But now it contains multiple <code>session_id</code>, and all three <code>level_groups</code>.😲</li>\n</ul>",
      "votes": 9,
      "replies": [
        {
          "id": 2190082,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-03-21T02:49:14.147000",
          "content": "<blockquote>\n  <p>But now it contains multiple session_id, and all three level_groups</p>\n</blockquote>\n<p>If this is the case then this is a time leak (i.e. this is the original time leak that Kaggle fixed on Feb 15th). We can use info from levels 5+ to predict questions from level 0-4. This gives at least <code>+0.010</code> boost on CV and LB!</p>",
          "votes": 10,
          "replies": [
            {
              "id": 2190110,
              "author_name": "AT",
              "author_url": "",
              "post_date": "2023-03-21T03:40:22.340000",
              "content": "<p>OMG, I see, the same thing may have happened with the previous leak. <br>\nAs you say, it is not appropriate to use data to predict future behavior rather than solve problems.</p>\n<p>Looking at the CV in the new training notebook, from 0.696 improved to 0.709 (could be overfitting though).</p>\n<p>I am hesitant to submit this data as it is leaked data…</p>\n<p>I have learned a lot from your notes and discussions, thank you!</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2190529,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2023-03-21T10:05:00.503000",
              "content": "<p>doesn't that mean that we need to setup different feature creation strategy?<br>\nWe don't need to group by level group anymore as we have all the level groups at hand from the beginning (see <a href=\"https://www.kaggle.com/code/kmitsuhiro/api-update-iter-test-returns-all-test-data\" target=\"_blank\">here</a>)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2190900,
              "author_name": "Elias",
              "author_url": "",
              "post_date": "2023-03-21T15:05:55.983000",
              "content": "<p>yes, we'll have to try some new features. Also we don't need anymore to save data from previous groups to predict next groups.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2190637,
          "author_name": "Bosco Yung",
          "author_url": "",
          "post_date": "2023-03-21T11:43:59.293000",
          "content": "<p>Thanks for the findings!</p>\n<blockquote>\n  <p>But now it contains multiple <code>session_id</code></p>\n</blockquote>\n<p>A related problem would then be, how many session IDs would be given in the actual test run (in the sample we see 3 only)? This determines how large the given dataframe size is per iteration, and could be important because we may run out of memory.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2190989,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-21T16:23:00.317000",
              "content": "<p>Hi. This is an issue with the set of 3 provided public samples. The hidden test set does not have this problem. I will fix the sample set. Thanks for spotting it!</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2191302,
              "author_name": "Bosco Yung",
              "author_url": "",
              "post_date": "2023-03-21T21:20:12.470000",
              "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> !</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2191393,
              "author_name": "AT",
              "author_url": "",
              "post_date": "2023-03-21T23:43:45.230000",
              "content": "<p>Only with the sample set,<br>\nI understand clearly now.<br>\nThank you very much! <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2191558,
              "author_name": "Helmy",
              "author_url": "",
              "post_date": "2023-03-22T04:03:21.097000",
              "content": "<p>Thanks a lot :D</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2199333,
      "author_name": "AT",
      "author_url": "",
      "post_date": "2023-03-27T16:33:36.787000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI have a question about the Kaggle API, <code>for (test, sample_submission,) in iter_test:</code><br>\nThanks for fixing the error that the sample test contains 3 level_groups. A few days ago I confirmed that both sample test and hidden test contain 1 level_group.</p>\n<p>But about the session_id…<br>\n<strong>The hidden test contains 1 session_id, but the sample test contains 3 session_ids.</strong><br>\nBefore the update on Mar 21, I remember that the old sample test contained 1 session_id.<br>\nThis complicates the Submission code. It looks like some baseline notebooks also needed to be changed.</p>\n<p>Is there a special reason for this? If not, I think it would be better for the sample test to contain 1 session_id.<br>\nI am a beginner Kaggler, so I'm sorry if my info is wrong.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2204800,
          "author_name": "Javier Hernández",
          "author_url": "",
          "post_date": "2023-04-01T00:57:40.917000",
          "content": "<p>Could you find any answer to this issue? I'm also having troubles with this 😪</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2205956,
              "author_name": "AT",
              "author_url": "",
              "post_date": "2023-04-02T06:02:29.587000",
              "content": "<p>Thanks for your comments,@javihm77</p>\n<p>I was helped by the method shown in the discussion below.<br>\nI have used it to customize my submission code.<br>\nThanks for the helpful discussions.If not for these tips I would have had a very hard time.</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396468\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396468</a><br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396751\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396751</a><br>\n<a href=\"https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments\" target=\"_blank\">https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments</a><br>\n<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680/comments\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680/comments</a><br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732</a></p>\n<p>I hope they are helpful to you.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2211092,
              "author_name": "Javier Hernández",
              "author_url": "",
              "post_date": "2023-04-05T19:38:23.253000",
              "content": "<p>I owe you one <a href=\"https://www.kaggle.com/tanakaakinori\" target=\"_blank\">@tanakaakinori</a> thanks a lot! you saved me a lot of time and stress, it works 🤝</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2204935,
          "author_name": "ruhong",
          "author_url": "",
          "post_date": "2023-04-01T04:57:06.293000",
          "content": "<p>according to kaggle staff reply (link below), the sample data has been fixed. However I am still facing the same submission issue as <a href=\"https://www.kaggle.com/tanakaakinori\" target=\"_blank\">@tanakaakinori</a> </p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2199072\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2199072</a></p>\n<p>It feels very frustrating that we spent so much of our very limited spare time on this competition, only for the submission sample data to change (through no fault of our own) and now, we also have to figure out how to properly update our submission code for the new sample data (where there are now 3 session_ids instead of 1).</p>\n<p>Plus there was no proper / clear announcement that submission sample data has been \"fixed\" (maybe a sticky comment at the top? updating the original post? putting out a code example that works for the new sample data?) - we have to trawl this forum thread to find out. This level of effort is going to dissuade all but the most committed kagglers from taking part in this competition. All the wasted time on admin issues could have been avoided, if there has been better communication with participants.</p>\n<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> </p>",
          "votes": 6,
          "replies": [
            {
              "id": 2205962,
              "author_name": "AT",
              "author_url": "",
              "post_date": "2023-04-02T06:09:20.257000",
              "content": "<p>Thanks for your comment, <a href=\"https://www.kaggle.com/ruhong\" target=\"_blank\">@ruhong</a></p>\n<p>I think it is disconcerting that the number of session_id behaves different ways between the sample test and hidden test.<br>\nIf improved, it might remove the barrier for beginners and those who are just starting to work on this competition and allow us to focus on their original task.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2189743,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2023-03-20T18:09:58.660000",
      "content": "<p>Hi all - thanks Alex for the update! Here are the steps you should expect to see today:</p>\n<ol>\n<li>Submissions will be disabled while I update the test data and evaluation metric. Please note that the order of the dataframes has changed - <code>iter_test</code> will now yield (test, sample_submission) rather than (sample_submission, test). This was in response to a bug report where the API was throwing exceptions for some people.</li>\n<li>I will reenable submissions once the test data and metric are updated.</li>\n<li>I will update the training data today, but it may take longer due to the increased volume of data.</li>\n<li>All previous notebooks will be re-run during the coming week.</li>\n</ol>",
      "votes": 7,
      "replies": [
        {
          "id": 2189766,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-03-20T18:34:17.233000",
          "content": "<p>Phil, won't all notebook reruns fail because all notebooks use:</p>\n<pre><code>for (sample_submission, test) in iter_test:\n</code></pre>",
          "votes": 11,
          "replies": [
            {
              "id": 2189998,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2023-03-21T00:36:36.973000",
              "content": "<p>I believe so. I just checked to see my 2 submission results and both failed whilst deducting the number of submissions available to me today. I was not aware of this announcement until after my submissions failed and I then searched for a reason. Copying <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2190623,
              "author_name": "Bosco Yung",
              "author_url": "",
              "post_date": "2023-03-21T11:39:04.537000",
              "content": "<p>Agree! 😧<br>\nAnd I don't get how fixing the API means the tuple order has to be changed…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2191000,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-21T16:27:09.180000",
              "content": "<p>Specifically, the order of the returned dataframes needed to change to eliminate the potential bug. This was an API issue that needed to be resolved this way.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2191719,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2023-03-22T06:52:28.423000",
      "content": "<p>The iter_test still returns all the test data. When will this be fixed?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2191732,
          "author_name": "AbaoJiang",
          "author_url": "",
          "post_date": "2023-03-22T07:08:41.120000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>,</p>\n<p>I submitted a notebook about 40min ago, which successfully pass now (get LB 0.692). I think iter_test returns all the test data only for 3 provided public samples, as discussed <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2190989\" target=\"_blank\">here</a>. As for hidden test set, time series API provides <strong>one level_group of one session</strong> for each iteration. Some of code snippets of my submission is shown as follows:</p>\n<pre><code>for (test, sample_submission) in iter_test:\n    # Process the data and generate features\n\n    try:\n        # Run inference using uploaded models\n        # For one level_group of one session\n        y_pred = quick_infer((x, x_cat), models[cur_lv_gp])\n        y_pred = (y_pred &gt; best_thres).astype(np.int)\n\n        sample_submission.loc[:, \"correct\"] = y_pred\n    except:\n        # This part is used to pass public samples\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission)\n</code></pre>\n<p>Hope this helps, thanks!</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2191806,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-22T08:02:46.160000",
              "content": "<p>Thanks. Are you sure you get only one session_d with one group at a time?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2192259,
              "author_name": "Ulrich G.",
              "author_url": "",
              "post_date": "2023-03-22T14:48:20.860000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I think it is the case. Since trying his code snipet does not yield an error message.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2196071,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2023-03-25T06:01:10.517000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>,</p>\n<p>I think it's true for hidden test set. For further confirmation, I submit <a href=\"https://www.kaggle.com/code/abaojiang/check-new-hidden-test-leakage\" target=\"_blank\">this notebook</a> and get LB score equal to <strong>zero prediction baseline</strong>, which obtains 0.216 on new test set. Checking part is in version 2 and zero prediction baseline in version 3.</p>\n<p>If there's any mistake, please let me know. Thanks a lot!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2190027,
      "author_name": "Hao Wang",
      "author_url": "",
      "post_date": "2023-03-21T01:14:13.647000",
      "content": "<p>It seems that the api will give all of the levels at once now. Has anyone else encountered this issue? </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2190271,
          "author_name": "Elias",
          "author_url": "",
          "post_date": "2023-03-21T06:41:02.293000",
          "content": "<p>Yes I see the same, all 18 questions at once in the same iteration. So we need to rewrite the submission code.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2190992,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-03-21T16:23:33.277000",
              "content": "<p>This is an issue with the sample data that does not exist in the hidden test set. I will update the sample data. Thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2191301,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-21T21:19:40.680000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2191975,
              "author_name": "Elias",
              "author_url": "",
              "post_date": "2023-03-22T10:50:49.500000",
              "content": "<p>Oh, I see. Then I will try to retrain and resubmit with no change to the code. I can see that the training is now using 23,562 unique sessions instead of the previous 11,779. It can only be completed with the GPU as the CPU runs out of memory. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2262429,
      "author_name": "Takehiro Yano",
      "author_url": "",
      "post_date": "2023-05-16T23:25:07.303000",
      "content": "<p>Hi!<br>\nSince the leaderboard update, some of my submissions are getting \"Submission Scoring Error\". Even when I have the exact same output code as other successful submissions, I still get the same error.<br>\nI have also confirmed that when I run it against test.csv, the output is correct and expected.<br>\nIs there something wrong with the Scoring process?<br>\nHas anyone else experienced a similar event?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2295066,
      "author_name": "Hoping",
      "author_url": "",
      "post_date": "2023-06-10T15:02:35.327000",
      "content": "<p>I'm so happy to be joining my first official competition here! Hope I can learn a lot and share what I've learned</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2206112,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-04-02T09:50:54.863000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> ,<br>\ndoes the shared raw <a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">data</a> contain true labels ?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2207796,
          "author_name": "Alex Z",
          "author_url": "",
          "post_date": "2023-04-03T16:26:40.417000",
          "content": "<p>After a quick look, the event logs from this public data do contain additional events, and those events show how the player answered the questions we have to predict. There is, as far as I can see, NO direct labels, but they could be inferred from the text of the answer itself.</p>\n<p>For example, this event shows the correct answer to the question (of course, we still need to check that there was no wrong answer given before it):</p>\n<p><code>\"22020006390905390\"     \"JOWILDER\"      2022-03-27 10:44:07.747000      \"CUSTOM.11\" {\"cur_cmd_fqid\": \"tunic.capitol_0.hall.boss.chap1_finale_plaquefirst_1\", \"cur_cmd_type\": 1, \"event_custom\": 11, \"fqid\": \"chap1_finale\", \"http_user_agent\": \"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_6) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.1.2 Safari/605.1.15\", \"interacted_fqid\": \"tunic.entry_tunic\", \"level\": 4, \"name\": \"choice\", \"persistent_session_id\": \"22020006390905390\", \"remote_addr\": \"216.15.53.200\", \"room_coor\": [-569.5578218361513, -169.70583676184344], \"room_fqid\": \"tunic.capitol_0.hall\", \"screen_coor\": [92, 420], \"server_time\": \"2022-03-27T05:44:07\", \"subtype\": \"wildcard\", \"text\": \"Our shirt was around way before the women's basketball team!\", \"type\": \"click\"}  \"10\" \"0\"  None   \"\"   {}   {}  \"162\"\n</code></p>\n<p>So, with a bit of cleaning and processing this dataset may be used to increase the training data. </p>\n<p>Whether this open data actually contains our training or test data, is an open question. </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 2212154,
          "author_name": "Alex Franklin",
          "author_url": "",
          "post_date": "2023-04-06T14:58:28.497000",
          "content": "<p>Hi Reacher, the raw data doesn't contain the labels directly but can be inferred from it (as Alex rightly pointed out). The test data for the competition is not on the open data site.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2213751,
              "author_name": "AT",
              "author_url": "",
              "post_date": "2023-04-07T20:29:23.960000",
              "content": "<p>Thanks for the interesting discussion.<br>\nI checked the raw data and found 125883 session_ids. Of those, 3785 session_ids were also included in the train data. Also, as you pointed out, none were found in the hidden test.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2195220,
      "author_name": "Saurabh Mahadik",
      "author_url": "",
      "post_date": "2023-03-24T13:58:13.847000",
      "content": "<p>Today ,I have shared my first code on kaggle. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2204973,
      "author_name": "Rohit Jagdale",
      "author_url": "",
      "post_date": "2023-04-01T05:43:30.667000",
      "content": "<p>The iter_test still returns all the test data. When will this be fixed?❌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2196187,
      "author_name": "Ka Ian",
      "author_url": "",
      "post_date": "2023-03-25T07:57:48.800000",
      "content": "<p>the train.csv is so big, it is impossible to work on this .csv file with kaggle platform. Even if I was using colab and pycharm to work on it, it still took a very long time to read this .csv. Oh my god.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2196433,
          "author_name": "AbaoJiang",
          "author_url": "",
          "post_date": "2023-03-25T12:04:07.580000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kaianchan\" target=\"_blank\">@kaianchan</a>,</p>\n<p>It's still possible to play around with this new training set on Kaggle platform. Details are demonstrated by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396979\" target=\"_blank\">this thread</a>.</p>\n<p>Hope this helps, thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2197436,
          "author_name": "Vadim Kamaev",
          "author_url": "",
          "post_date": "2023-03-26T06:45:45.823000",
          "content": "<p>The train.csv can be downloaded in 953 MB - <a href=\"https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb\" target=\"_blank\">https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2198564,
          "author_name": "Demid Chernenko",
          "author_url": "",
          "post_date": "2023-03-27T06:50:01.580000",
          "content": "<p>I'd recommend trying <code>polars</code> instead of <code>pandas</code>. It's more efficient in memory usage and provides <a href=\"https://www.kaggle.com/code/demche/polars-memory-usage-optimization\" target=\"_blank\">more options</a> for memory usage optimization</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2199055,
              "author_name": "Mikkel Skovdal",
              "author_url": "",
              "post_date": "2023-03-27T13:16:15.740000",
              "content": "<p>I second this.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2204970,
          "author_name": "Rohit Jagdale",
          "author_url": "",
          "post_date": "2023-04-01T05:42:39.957000",
          "content": "<p>Downloaded train.csv can be in 953 MB - <a href=\"https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb\" target=\"_blank\">https://www.kaggle.com/code/vadimkamaev/reading-data-953-mb</a></p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2190286,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-03-21T06:56:23.670000",
      "content": "<p>Does the test size also double ?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2190077,
      "author_name": "minhtu.mt.mt",
      "author_url": "",
      "post_date": "2023-03-21T02:40:58.217000",
      "content": "<p>More data more work 😷</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2260803,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-05-15T22:16:08.457000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> </p>\n<p>How math \"more data\" improve score?</p>\n<p>Do you remember your score befor  and after ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2191962,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2023-03-22T10:39:40.597000",
      "content": "<p>Well, guys, as I understand all leaks (I mean 2% leak and API issue) were fixed. But leaderboard's notebooks were not reruned. Am I correct?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2193433,
          "author_name": "Minh Tri Phan",
          "author_url": "",
          "post_date": "2023-03-23T09:46:03.800000",
          "content": "<p>I am curious too. As Chris said <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191405\" target=\"_blank\">here</a>, it seems to exist, Kaggle hasn't fixed the submission file as well, which makes the inference quite tricker.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2197609,
              "author_name": "Sergey Bryansky",
              "author_url": "",
              "post_date": "2023-03-26T08:51:00.220000",
              "content": "<p>Yes, still don't understand status of competition… is it still leaked/bugged? It is meta-question.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2190003,
      "author_name": "fermion",
      "author_url": "",
      "post_date": "2023-03-21T00:38:41.273000",
      "content": "<p>Thank you for the update regarding the data leak and fix.</p>\n<p>As I am writing the comment, it appears that the new training data set has been uploaded, its size comes to 4.72GB. Without any memory management, my notebook is not able to run due to the error:</p>\n<blockquote>\n  <p>Your notebook tried to allocate more memory than is available. It has restarted.</p>\n</blockquote>\n<p>Many kagglers have shared tips of downsizing the memory footprint, this error makes me believe now it is time to explore and implement those ideas.</p>\n<p>Additionally, basing on heuristics, many have observed data issues in both training and test sets (for example, <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395686\" target=\"_blank\">here</a>). Will this data update address these issues? How about the hidden test set?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2194575,
      "author_name": "Siddharth Kumar",
      "author_url": "",
      "post_date": "2023-03-24T03:45:40.953000",
      "content": "<p>Thank you for the update and for addressing the issue promptly. We appreciate the transparency and effort to maintain the integrity of the competition. We look forward to incorporating the additional data and continuing to work on our models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2190122,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-03-21T03:54:26.453000",
      "content": "<p>Thanks Dear <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>,<br>\nThanks for sharing such a valuable post. I appreciate the transparency and swift action taken by the organizers to address the leaked competition data. By releasing the entire current test set and extending the competition deadline, they have ensured fairness for all participants. It's great to see the Kaggle community's integrity in reporting such issues and the organizers' commitment to maintaining the integrity of the competition.👍</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2236091,
      "author_name": "Yuansong Zhu",
      "author_url": "",
      "post_date": "2023-04-26T14:24:28.090000",
      "content": "<p>Hi<br>\nI am new here.</p>",
      "votes": -4,
      "replies": []
    },
    {
      "id": 2197300,
      "author_name": "DEVANAND",
      "author_url": "",
      "post_date": "2023-03-26T04:07:04.570000",
      "content": "<p>New platform for me</p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 2196497,
      "author_name": "Ashwin Varma",
      "author_url": "",
      "post_date": "2023-03-25T12:58:38.083000",
      "content": "<p>Hi <br>\nI am new here.</p>",
      "votes": -7,
      "replies": []
    },
    {
      "id": 2307273,
      "author_name": "ydxx",
      "author_url": "",
      "post_date": "2023-06-18T04:11:30.183000",
      "content": "<p>Why is the time complexity of the baseline model very low, but my model has high time complexity and is not as good as its performance? I am very depressed. What perspective should I try to solve this problem from?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2280896,
      "author_name": "Silent_Trick",
      "author_url": "",
      "post_date": "2023-05-30T12:53:29.930000",
      "content": "<p>Hi! new here<br>\nI have one doubt related to the test dataset.<br>\nWe are getting the test dataset in level-group sessions. Let's suppose I got the test dataset for level-group 0-4 and we predicted the label for questions 1-3 using this dataset and now we will get the dataset for level-group 5-13. Do we also simultaneously get the <strong>true</strong> labels for level-group 0-4? (meaning how well they have truly performed in the previous level-group).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2278217,
      "author_name": "imaginaire",
      "author_url": "",
      "post_date": "2023-05-28T12:46:05.283000",
      "content": "<p>Thank you for the advice you gave, it can be a great inspiration to me.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2266347,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-19T23:29:12.580000",
      "content": "<p>at this site (<a href=\"https://fielddaylab.wisc.edu/opengamedata/):\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/):</a></p>\n<p>Research   ( <a href=\"https://arxiv.org/pdf/2210.09906.pdf\" target=\"_blank\">https://arxiv.org/pdf/2210.09906.pdf</a> )<br>\nThis resulted in four versions of the script:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2F3c5b4c4dc9b7414930dee18bd2f3dc7e%2F2023-05-20%20%2001.57.45.png?generation=1684537140471281&amp;alt=media\" alt=\"\"></p>\n<ol>\n<li>Snark + Humor (Original)</li>\n<li>Humor, no Snark </li>\n<li>Snark, no Humor</li>\n<li>Dry</li>\n</ol>\n<p>In train.csv:</p>\n<table>\n<thead>\n<tr>\n<th>session_id</th>\n<th>versions of the script</th>\n<th>text_x</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>22060015173801296</td>\n<td>1. Snark + Humor (Original)</td>\n<td>Did you do all of them?</td>\n</tr>\n<tr>\n<td>22060001580924636</td>\n<td>2. Humor, no Snark</td>\n<td>He's always trying to get you in trouble, and he doesn't like animals!</td>\n</tr>\n<tr>\n<td>22060014314686290</td>\n<td>2. Humor, no Snark</td>\n<td>He's always trying to get you in trouble, and he doesn't like animals!</td>\n</tr>\n<tr>\n<td>22060010020384030</td>\n<td>3. Snark, no Humor</td>\n<td>So? History is boring!</td>\n</tr>\n<tr>\n<td>22060013403542770</td>\n<td>3. Snark, no Humor</td>\n<td>So? History is boring!</td>\n</tr>\n<tr>\n<td>22060015102434092</td>\n<td>3. Snark, no Humor</td>\n<td>So? History is boring!</td>\n</tr>\n<tr>\n<td>22060008392175650</td>\n<td>4. Dry</td>\n<td>Yes! This old slip from 1916.</td>\n</tr>\n<tr>\n<td>22060008395913956</td>\n<td>4. Dry</td>\n<td>Yes! This old slip from 1916.</td>\n</tr>\n<tr>\n<td>22060011583687136</td>\n<td>4. Dry</td>\n<td>Yes! This old slip from 1916.</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://arxiv.org/pdf/2210.09906.pdf:\" target=\"_blank\">https://arxiv.org/pdf/2210.09906.pdf:</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2260800,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-15T22:10:45.907000",
      "content": "<p>Hello,  <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> :)</p>\n<p>Who can remembered:<br>\n does score increase,  after  increase train.csv  size from 10000 to 20000 people  ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2289906,
          "author_name": "superslon",
          "author_url": "",
          "post_date": "2023-06-06T12:35:21.333000",
          "content": "<p>Yes, both increased by about 0.001+</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2260796,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-15T22:04:40.980000",
      "content": "<p>Alex, hello:)</p>\n<p>Before multiple train size in 2 tims,<br>\nHow look lider bord?<br>\nIt's function save leaderboard to file.</p>\n<p>Do anybody have saved leaderboard, as it wos 2 month ago?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2240538,
      "author_name": "Proknowor",
      "author_url": "",
      "post_date": "2023-04-30T16:05:13.933000",
      "content": "<p>lol, i was like what happened to my notebook, it could run properly before monthes.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2224165,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-17T04:08:37.910000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2210352,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-05T10:29:22.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2196192,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-25T08:01:32.530000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2193890,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-23T15:20:03.020000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2194681,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-24T05:07:18.740000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2193039,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-23T04:15:20.777000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2193042,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-23T04:16:30.003000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2190184,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-21T05:04:50.453000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2190243,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-21T06:11:47.447000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2190249,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-21T06:18:41.833000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2191446,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-22T01:30:45.860000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2190099,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-21T03:20:09.510000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2205204,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-01T10:39:38.150000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2196380,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-25T11:01:15.857000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2190390,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-21T08:01:39.343000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2251611,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-09T14:18:09.510000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2224598,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-17T13:10:58.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2218136,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-11T13:13:13.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2204224,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-31T12:27:49.713000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2201058,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-29T02:36:35.427000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2193332,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-23T08:15:36.573000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2191403,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-21T23:49:15.770000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2189733": "A Kaggler brought to our attention that the test data for this competition was unintentionally made available for a period of time. Although the leak was fairly limited, we can’t be certain that the test data is not compromised. Therefore, to be fair to all contestants **we are releasing all of the leaked data**, which contains the entire current test set.\n\nWe will update the training data on the competition page to include the current test set, which will increase the size of the training data by 100%.\n\nThe raw data from the game will also be made available and can be found at [this site](https://fielddaylab.wisc.edu/opengamedata/), which you are welcome to use as supplemental data for this competition.\n\nSince we are doubling the size of the training set and linking to the raw data, **we are also extending the competition by one month** to give everyone a chance to incorporate the new data into their models.\n\nWe will be creating a new public and private test set with data that was not previously part of the competition and was not part of the leak. The public leaderboard will be rerun to reflect scores on this new test set.\n\nThe updates to the competition data should be complete by the end of the day.\n\nWe would like to thank the Kaggler who brought this to our attention. We appreciate the integrity of the Kaggle community for consistently reporting issues which could have been exploited to gain an individual advantage in the competition. Thank you for bearing with us as we work to resolve this issue.",
    "2190090": "@alexmlfranklin, just as @tanakaakinori and others have observed, there are differences in both the training data and the test data. I have also observed the changes in the test data that resulted in submission error, which according to @cdeotte that we may have a leak again.\n\nPlease fix everything and let us know the details when you are done. Perhaps you should disable submission again so that we do not waste our times.\n\n**Update:**\nBased on the replies to @philculliton's post below, there was and still is a leak described as:\n`2% leak(users who repeat level_group = '0-4' or 5-12 after they finish later level_groups) still exists in this competition. `\n\nWe need an update on if or when that one will be fixed @philculliton. Thanks",
    "2194193": "The problem seems unsolved...\n@alexmlfranklin @nrambis @philculliton @maggiemd ",
    "2190039": "Hello! Thanks for the update.\nI found some changes.\n\n- The data size of 'train', 'train_label' has increased. \n    It ran out of CPU notebook's 8GB RAM,  No problem for GPU's one.\n- ` for (sample_submission, test) in iter_test: ` was changed to for `(test,  sample_submission) in iter_test:`\n- `for (test, sample_submission) in iter_test:`, old `test` used to have  single `session_id`, single `level_group`. \n     But now it contains multiple `session_id`, and all three `level_groups`.😲",
    "2199333": "Hi @philculliton \nI have a question about the Kaggle API, `for (test, sample_submission,) in iter_test:`\nThanks for fixing the error that the sample test contains 3 level_groups. A few days ago I confirmed that both sample test and hidden test contain 1 level_group.\n\nBut about the session_id...\n**The hidden test contains 1 session_id, but the sample test contains 3 session_ids.**\nBefore the update on Mar 21, I remember that the old sample test contained 1 session_id.\nThis complicates the Submission code. It looks like some baseline notebooks also needed to be changed.\n\nIs there a special reason for this? If not, I think it would be better for the sample test to contain 1 session_id.\nI am a beginner Kaggler, so I'm sorry if my info is wrong.",
    "2189743": "Hi all - thanks Alex for the update! Here are the steps you should expect to see today:\n1. Submissions will be disabled while I update the test data and evaluation metric. Please note that the order of the dataframes has changed - `iter_test` will now yield (test, sample_submission) rather than (sample_submission, test). This was in response to a bug report where the API was throwing exceptions for some people.\n2. I will reenable submissions once the test data and metric are updated.\n3. I will update the training data today, but it may take longer due to the increased volume of data.\n4. All previous notebooks will be re-run during the coming week.",
    "2191719": "The iter_test still returns all the test data. When will this be fixed?",
    "2190027": "It seems that the api will give all of the levels at once now. Has anyone else encountered this issue? ",
    "2262429": "Hi!\nSince the leaderboard update, some of my submissions are getting \"Submission Scoring Error\". Even when I have the exact same output code as other successful submissions, I still get the same error.\nI have also confirmed that when I run it against test.csv, the output is correct and expected.\nIs there something wrong with the Scoring process?\nHas anyone else experienced a similar event?",
    "2295066": "I'm so happy to be joining my first official competition here! Hope I can learn a lot and share what I've learned",
    "2206112": "Hi @alexmlfranklin ,\ndoes the shared raw [data](https://fielddaylab.wisc.edu/opengamedata/) contain true labels ?",
    "2195220": "Today ,I have shared my first code on kaggle. ",
    "2204973": "The iter_test still returns all the test data. When will this be fixed?❌",
    "2196187": "the train.csv is so big, it is impossible to work on this .csv file with kaggle platform. Even if I was using colab and pycharm to work on it, it still took a very long time to read this .csv. Oh my god.",
    "2190286": "Does the test size also double ?",
    "2190077": "More data more work 😷",
    "2191962": "Well, guys, as I understand all leaks (I mean 2% leak and API issue) were fixed. But leaderboard's notebooks were not reruned. Am I correct?",
    "2190003": "Thank you for the update regarding the data leak and fix.\n\nAs I am writing the comment, it appears that the new training data set has been uploaded, its size comes to 4.72GB. Without any memory management, my notebook is not able to run due to the error:\n>Your notebook tried to allocate more memory than is available. It has restarted.\n\nMany kagglers have shared tips of downsizing the memory footprint, this error makes me believe now it is time to explore and implement those ideas.\n\nAdditionally, basing on heuristics, many have observed data issues in both training and test sets (for example, [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250) and [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395686)). Will this data update address these issues? How about the hidden test set?",
    "2194575": "Thank you for the update and for addressing the issue promptly. We appreciate the transparency and effort to maintain the integrity of the competition. We look forward to incorporating the additional data and continuing to work on our models.",
    "2190122": "Thanks Dear @alexmlfranklin,\nThanks for sharing such a valuable post. I appreciate the transparency and swift action taken by the organizers to address the leaked competition data. By releasing the entire current test set and extending the competition deadline, they have ensured fairness for all participants. It's great to see the Kaggle community's integrity in reporting such issues and the organizers' commitment to maintaining the integrity of the competition.👍",
    "2236091": "Hi\nI am new here.",
    "2197300": "New platform for me",
    "2196497": "Hi \nI am new here.",
    "2307273": "Why is the time complexity of the baseline model very low, but my model has high time complexity and is not as good as its performance? I am very depressed. What perspective should I try to solve this problem from?",
    "2280896": "Hi! new here\nI have one doubt related to the test dataset.\nWe are getting the test dataset in level-group sessions. Let's suppose I got the test dataset for level-group 0-4 and we predicted the label for questions 1-3 using this dataset and now we will get the dataset for level-group 5-13. Do we also simultaneously get the **true** labels for level-group 0-4? (meaning how well they have truly performed in the previous level-group).",
    "2278217": "Thank you for the advice you gave, it can be a great inspiration to me.",
    "2266347": "at this site (https://fielddaylab.wisc.edu/opengamedata/):\n\nResearch   ( https://arxiv.org/pdf/2210.09906.pdf )\nThis resulted in four versions of the script:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2F3c5b4c4dc9b7414930dee18bd2f3dc7e%2F2023-05-20%20%2001.57.45.png?generation=1684537140471281&alt=media)\n1. Snark + Humor (Original)\n2. Humor, no Snark \n3. Snark, no Humor\n4. Dry\n\nIn train.csv:\n| session_id | versions of the script |text_x |\n| --- | --- | --- |\n| 22060015173801296| 1. Snark + Humor (Original)  |\tDid you do all of them? |\n| 22060001580924636 | 2. Humor, no Snark  |He's always trying to get you in trouble, and he doesn't like animals!|\n| 22060014314686290| 2. Humor, no Snark  |He's always trying to get you in trouble, and he doesn't like animals!|\n| 22060010020384030\t| 3. Snark, no Humor |\tSo? History is boring! |\n| 22060013403542770| 3. Snark, no Humor |\tSo? History is boring! |\n| 22060015102434092| 3. Snark, no Humor |\tSo? History is boring! |\n| 22060008392175650|  4. Dry|\t \tYes! This old slip from 1916.|\n| 22060008395913956| 4. Dry|\t \tYes! This old slip from 1916.|\n| 22060011583687136| 4. Dry|\t \tYes! This old slip from 1916.|\n\n\n\nhttps://arxiv.org/pdf/2210.09906.pdf:",
    "2260800": "Hello,  @alexmlfranklin :)\n\nWho can remembered:\n does score increase,  after  increase train.csv  size from 10000 to 20000 people  ?",
    "2260796": "Alex, hello:)\n\nBefore multiple train size in 2 tims,\nHow look lider bord?\nIt's function save leaderboard to file.\n\nDo anybody have saved leaderboard, as it wos 2 month ago?",
    "2240538": "lol, i was like what happened to my notebook, it could run properly before monthes.",
    "2224165": "Hi @ahmeddatascientist ,I am a beginner kaggler and I am wondering how the question_labels connect with the dataset \"train.csv\",so can you give me some instructions?",
    "2210352": "great job!",
    "2196192": "After these updates test data includes only 3 session ids, same with sample_submission. iter_test also does only 1 iteration returning these 3 sessions at once. Maybe issue is specific to my account.\n\n@alexmlfranklin @nrambis @philculliton @maggiemd",
    "2193890": "does fixing the API means the tuple order has to be changed…",
    "2193039": "which  competition the data is from ?",
    "2190184": "What does it look like when the submissions are closed? I had some submissions that ran successfully but failed during scoring (about an hour before time of posting this) and I'm just wondering if that's because submissions are closed right now and I just didn't know, or if I have a bug that needs to be fixed",
    "2190099": "cool. 卷起来！",
    "2205204": "",
    "2196380": "",
    "2190390": "",
    "2251611": "That's great! Thank you!",
    "2224598": "Thanks for your notice.",
    "2218136": "Thanks this is helpful.",
    "2204224": "Thanks for the update.",
    "2201058": "Hello! Thanks for the update.",
    "2193332": "Hello! Thanks for the update.",
    "2191403": "Hello! Thanks for the update."
  }
}