{
  "id": 402608,
  "title": "LB update incoming",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/402608",
  "author_name": "Phil Culliton",
  "post_date": "2023-04-19T01:39:46.800000",
  "votes": 6,
  "comment_count": 53,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>We are ready to begin running the leaderboard update discussed when the data was updated. It will occur over a period of a day or so. You will notice your old submissions being rerun, and the leaderboard changing.</p>\n<p>As we mentioned in our previous post, the API bug fixes that I implemented recently will likely cause many older submissions to fail. My apologies for this; it was the only path that ensured we were entirely solving the reported issues.</p>\n<p>We have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.</p>\n<p>Thanks to everyone who reported issues, and thanks to all for your patience.</p>\n<p>Note: The LB update is underway but will likely take several days as we will be running in batches. Old notebooks with code that does not conform to the API updates will likely error out. Thanks for your patience!</p>",
  "messages": [
    {
      "id": 2228354,
      "postDate": "2023-04-20T13:41:14.580Z",
      "content": "<p>One bug with one month to fix, we call this KAGGLE SPEED.</p>",
      "rawMarkdown": "One bug with one month to fix, we call this KAGGLE SPEED.",
      "votes": 8,
      "replies": [
        {
          "id": 2228494,
          "postDate": "2023-04-20T15:29:53.100Z",
          "content": "<p>Hi. I understand that's it taken some time. I apologize. I should note that this wasn't one bug, but a series of issues. A great deal of time has gone into finding a path forward, building the new competition data, working with the host to ensure that their goals were still being met, and attempting to reproduce, solve, and address all reported issues before we take the (quite large and time-consuming) step of updating the leaderboard.</p>\n<p>Updating a running competition with new data is a painstaking process that takes significant time and often creates secondary issues that need resolving. In this case we had no choice but to do an update. Our priority is, as always, making sure that the competition is viable.</p>",
          "rawMarkdown": "Hi. I understand that's it taken some time. I apologize. I should note that this wasn't one bug, but a series of issues. A great deal of time has gone into finding a path forward, building the new competition data, working with the host to ensure that their goals were still being met, and attempting to reproduce, solve, and address all reported issues before we take the (quite large and time-consuming) step of updating the leaderboard.\n\nUpdating a running competition with new data is a painstaking process that takes significant time and often creates secondary issues that need resolving. In this case we had no choice but to do an update. Our priority is, as always, making sure that the competition is viable.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2227281,
      "postDate": "2023-04-19T16:14:38.707Z",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Hi Phil, the mock public sample (i.e. commit API) is still <strong>not fixed</strong>. During each for-loop iteration, the mock public sample provides <strong>3 unique session_id</strong> in the <code>test dataframe</code> and <strong>1 unique session_id</strong> in the <code>sample_submission dataframe</code>.</p>\n<p>The correct implementation (to match submit API) is <strong>1 unique session_id</strong> in the <code>test dataframe</code> and <strong>1 unique session_id</strong> in the <code>sample_submission dataframe</code>.</p>",
      "rawMarkdown": "@philculliton Hi Phil, the mock public sample (i.e. commit API) is still **not fixed**. During each for-loop iteration, the mock public sample provides **3 unique session_id** in the `test dataframe` and **1 unique session_id** in the `sample_submission dataframe`.\n\nThe correct implementation (to match submit API) is **1 unique session_id** in the `test dataframe` and **1 unique session_id** in the `sample_submission dataframe`.",
      "votes": 6,
      "replies": [
        {
          "id": 2227373,
          "postDate": "2023-04-19T17:59:59.253Z",
          "content": "<p>Thanks Chris! Appreciate it - was an issue with the update. I've fixed it.</p>",
          "rawMarkdown": "Thanks Chris! Appreciate it - was an issue with the update. I've fixed it.",
          "votes": 2,
          "replies": [
            {
              "id": 2229324,
              "postDate": "2023-04-21T09:15:28.570Z",
              "content": "<p>I very much appreciate that you are trying to fix the submission, however, there seem to still be some problems. The problem now seems to be that the API requires the user to predict all 18 questions for one session_id in every iter_test(). I found this after running two similar notebooks, one that predicts the questions that should be predicted for the 'level_group' (fails) and one that predicts all 18 questions at each iteration (successes). This was hard to spot since the submission.csv looks good while testing the submission.</p>\n<p>If anyone else has succeeded in submitting a notebook predicting only the relevant questions for each level_group in each iter_test() iteration. Please show me how to do it.</p>",
              "rawMarkdown": "I very much appreciate that you are trying to fix the submission, however, there seem to still be some problems. The problem now seems to be that the API requires the user to predict all 18 questions for one session_id in every iter_test(). I found this after running two similar notebooks, one that predicts the questions that should be predicted for the 'level_group' (fails) and one that predicts all 18 questions at each iteration (successes). This was hard to spot since the submission.csv looks good while testing the submission.\n\nIf anyone else has succeeded in submitting a notebook predicting only the relevant questions for each level_group in each iter_test() iteration. Please show me how to do it.",
              "votes": 3
            },
            {
              "id": 2230524,
              "postDate": "2023-04-22T13:06:41.870Z",
              "content": "<p>Hi - I can't reproduce this issue. The public process tests out just fine for me. Would you mind sharing a notebook with me? Feel free to remove anything you don't want to share, I just need to see how the issue is being raised for you.</p>",
              "rawMarkdown": "Hi - I can't reproduce this issue. The public process tests out just fine for me. Would you mind sharing a notebook with me? Feel free to remove anything you don't want to share, I just need to see how the issue is being raised for you."
            },
            {
              "id": 2230683,
              "postDate": "2023-04-22T16:51:58.820Z",
              "content": "<p>Hi! I made two \"minimalistic\" notebooks to showcase this, I hope it is clear!<br>\nHere is my notebook, predicting 1's for all 18 questions, at each iteration (succeeds): <a href=\"https://www.kaggle.com/leonnorblad/predict-all-18-at-eah-iter\" target=\"_blank\">https://www.kaggle.com/leonnorblad/predict-all-18-at-eah-iter</a></p>\n<p>Here is my notebook, predicting 1's for questions that should be predicted for the current 'level_group' (fails): <a href=\"https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group\" target=\"_blank\">https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group</a></p>\n<p>I based the second notebook on how I believe the game is played and should be predicted (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796)</a>.   The player plays a couple of levels (we get data) and then we make a prediction, then it plays a few more levels (we get more data) and then we make predictions for the new questions. Have I misunderstood anything?</p>",
              "rawMarkdown": "Hi! I made two \"minimalistic\" notebooks to showcase this, I hope it is clear!\nHere is my notebook, predicting 1's for all 18 questions, at each iteration (succeeds): https://www.kaggle.com/leonnorblad/predict-all-18-at-eah-iter\n\nHere is my notebook, predicting 1's for questions that should be predicted for the current 'level_group' (fails): https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group\n\nI based the second notebook on how I believe the game is played and should be predicted (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796).   The player plays a couple of levels (we get data) and then we make a prediction, then it plays a few more levels (we get more data) and then we make predictions for the new questions. Have I misunderstood anything?",
              "votes": 2
            },
            {
              "id": 2230843,
              "postDate": "2023-04-22T19:43:16.523Z",
              "content": "<p>Thanks Leon! Would you mind sharing the notebooks with me?</p>\n<p>Your description of how the game is played is correct; when the notebooks are shared I'll look at your code to see where there might be an issue.</p>",
              "rawMarkdown": "Thanks Leon! Would you mind sharing the notebooks with me?\n\nYour description of how the game is played is correct; when the notebooks are shared I'll look at your code to see where there might be an issue."
            },
            {
              "id": 2230854,
              "postDate": "2023-04-22T20:02:53.863Z",
              "content": "<p>Sorry about that, they should be shared now :)</p>",
              "rawMarkdown": "Sorry about that, they should be shared now :)",
              "votes": 1
            },
            {
              "id": 2230898,
              "postDate": "2023-04-22T21:18:33.443Z",
              "content": "<p>Thanks very much Leon!</p>\n<p>Re: your first notebook - when you're running:</p>\n<pre><code>    for quest_nr in range(1,19):\n        sample_submission.loc[sample_submission.question == quest_nr, 'correct'] = 1\n</code></pre>\n<p>You're actually <em>not</em> setting anything for the questions that you don't currently have access to. <code>sample_submission.loc[sample_submission.question == quest_nr]</code> is only non-empty for the subset of questions you can currently access (you can check by looking at the <code>.shape</code> of the <code>sample_submission.loc[sample_submission.question == quest_nr]</code> for every <code>quest_nr</code> in each iteration, it shows the number of rows as <code>0</code> for any questions you can't currently access).</p>\n<p>So - this notebook appears to be setting <code>correct</code> only for the questions you have access to, not all 18 in each iteration. Which is the intent, so I think we're good on this notebook, unless there's something I'm missing! Please confirm the above and let me know.</p>\n<p>In your second notebook, everything appears to work as expected - you're getting access to, and setting answers for, the correct questions. Is the failure only happening when you submit to the private leaderboard?</p>",
              "rawMarkdown": "Thanks very much Leon!\n\nRe: your first notebook - when you're running:\n\n```\n    for quest_nr in range(1,19):\n        sample_submission.loc[sample_submission.question == quest_nr, 'correct'] = 1\n```\n\nYou're actually _not_ setting anything for the questions that you don't currently have access to. `sample_submission.loc[sample_submission.question == quest_nr]` is only non-empty for the subset of questions you can currently access (you can check by looking at the `.shape` of the `sample_submission.loc[sample_submission.question == quest_nr]` for every `quest_nr` in each iteration, it shows the number of rows as `0` for any questions you can't currently access).\n\nSo - this notebook appears to be setting `correct` only for the questions you have access to, not all 18 in each iteration. Which is the intent, so I think we're good on this notebook, unless there's something I'm missing! Please confirm the above and let me know.\n\nIn your second notebook, everything appears to work as expected - you're getting access to, and setting answers for, the correct questions. Is the failure only happening when you submit to the private leaderboard?"
            },
            {
              "id": 2231204,
              "postDate": "2023-04-23T06:31:13.410Z",
              "content": "<p>The first notebook was just a desperate attempt to make a successful submission. As you say it will only fill a value if a matching question exists, which is only the questions for the current 'level_group' when I test my submission by the three test sessions that we have access to.</p>\n<p>Is the failure only happening when you submit to the private leaderboard? -&gt; yes, everything looks fine when I run the notebook myself but I get the 'Submission Scoring Error' when trying to submit.</p>\n<p>The three test sessions that we have access to are fed (with iter_test) exactly how I expect them to be. Based on these two notebooks, I suspect that the test set in the private leaderboard is not fed the same way but instead in such a way that we needed to predict the 18 questions at each iteration (since the first notebook succeeds but the second notebook fails). This does not need to be the case, I just want to understand why the second notebook fails.</p>",
              "rawMarkdown": "The first notebook was just a desperate attempt to make a successful submission. As you say it will only fill a value if a matching question exists, which is only the questions for the current 'level_group' when I test my submission by the three test sessions that we have access to.\n\nIs the failure only happening when you submit to the private leaderboard? -> yes, everything looks fine when I run the notebook myself but I get the 'Submission Scoring Error' when trying to submit.\n\nThe three test sessions that we have access to are fed (with iter_test) exactly how I expect them to be. Based on these two notebooks, I suspect that the test set in the private leaderboard is not fed the same way but instead in such a way that we needed to predict the 18 questions at each iteration (since the first notebook succeeds but the second notebook fails). This does not need to be the case, I just want to understand why the second notebook fails.",
              "votes": 2
            },
            {
              "id": 2231382,
              "postDate": "2023-04-23T08:56:47.543Z",
              "content": "<p>I have the same issue. I get 'Submission Scoring Error' after 3~4hours from the submission, although everything looks fine when I run the notebook myself.</p>",
              "rawMarkdown": "I have the same issue. I get 'Submission Scoring Error' after 3~4hours from the submission, although everything looks fine when I run the notebook myself.",
              "votes": 1
            },
            {
              "id": 2231678,
              "postDate": "2023-04-23T14:38:50.467Z",
              "content": "<p>Thanks Leon! Very much appreciate your building and sharing the samples, and walking me through them.</p>\n<p>I dug deeper on your second notebook. <code>level_group_to_predict = test[\"level_group\"][0]</code> is throwing an exception (very) early with the private API (<code>0</code> is not an element in the index of the grouping you're trying to access, so it's a simple <code>KeyError</code>).</p>\n<p>You can entirely eliminate that code - the <code>sample_submission</code> for a given iteration will ONLY contain the questions you need to answer. If you simply fill in / predict <code>correct</code> for the questions in the <code>sample_submission</code> in each iteration of the loop, you should be fine.</p>\n<p>Alternatively, you could try using <code>level_group_to_predict = test[\"level_group\"].values[0]</code> if you just want the the <code>level_group</code> value for the first row in that sequence. This change appeared to work for me when testing your code.</p>\n<p>Please give this a shot and let me know if your issue persists.</p>",
              "rawMarkdown": "Thanks Leon! Very much appreciate your building and sharing the samples, and walking me through them.\n\nI dug deeper on your second notebook. `level_group_to_predict = test[\"level_group\"][0]` is throwing an exception (very) early with the private API (`0` is not an element in the index of the grouping you're trying to access, so it's a simple `KeyError`).\n\nYou can entirely eliminate that code - the `sample_submission` for a given iteration will ONLY contain the questions you need to answer. If you simply fill in / predict `correct` for the questions in the `sample_submission` in each iteration of the loop, you should be fine.\n\nAlternatively, you could try using `level_group_to_predict = test[\"level_group\"].values[0]` if you just want the the `level_group` value for the first row in that sequence. This change appeared to work for me when testing your code.\n\nPlease give this a shot and let me know if your issue persists.",
              "votes": 1
            },
            {
              "id": 2231686,
              "postDate": "2023-04-23T14:49:23.153Z",
              "content": "<p><a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a> - It looks like your submissions are erroring out before reaching the end of the loop. I can't tell what's happening in the code itself.</p>\n<p>It also appears as though you had a successful run of a notebook with the same name a few hours before the most recent one. Was there something you were doing differently there?</p>",
              "rawMarkdown": "@nynyny67 - It looks like your submissions are erroring out before reaching the end of the loop. I can't tell what's happening in the code itself.\n\nIt also appears as though you had a successful run of a notebook with the same name a few hours before the most recent one. Was there something you were doing differently there?",
              "votes": 1
            },
            {
              "id": 2231709,
              "postDate": "2023-04-23T15:11:24.050Z",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI commented out the line that change the 'sample_submission' like below.</p>\n<pre><code>#df_sample_submission.loc[df_sample_submission.session_id.str.contains(f'q{qid}'),'correct'] = int(prob&gt;threshold)\n</code></pre>\n<p>I also deleted some lines that prints some variables for the purpose of debugging. Maybe these lines contained some reference error, but I couldn't find suspecious parts.</p>\n<p>After I saw the submission process succeeded, I have activated this line and submission is now on progress. Before this study the submission fails in 3-4 hours but recent submission is running 5+ hours until now.<br>\nI don't know why the submissions are erroring before reaching the end of the loop, but the recent submissions continues longer at least.</p>",
              "rawMarkdown": "@philculliton \nI commented out the line that change the 'sample_submission' like below.\n```\n#df_sample_submission.loc[df_sample_submission.session_id.str.contains(f'q{qid}'),'correct'] = int(prob>threshold)\n```\nI also deleted some lines that prints some variables for the purpose of debugging. Maybe these lines contained some reference error, but I couldn't find suspecious parts.\n\nAfter I saw the submission process succeeded, I have activated this line and submission is now on progress. Before this study the submission fails in 3-4 hours but recent submission is running 5+ hours until now.\nI don't know why the submissions are erroring before reaching the end of the loop, but the recent submissions continues longer at least."
            },
            {
              "id": 2231738,
              "postDate": "2023-04-23T15:34:16.780Z",
              "content": "<p>Thank you very much! Yes, I can confirm that this works. I have updated my code for my second notebook and the submission succeeds!<br>\n(<a href=\"https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group\" target=\"_blank\">https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group</a>)</p>\n<p>Thank you for your time!</p>",
              "rawMarkdown": "Thank you very much! Yes, I can confirm that this works. I have updated my code for my second notebook and the submission succeeds!\n(https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group)\n\nThank you for your time!"
            },
            {
              "id": 2231765,
              "postDate": "2023-04-23T16:23:09.673Z",
              "content": "<p><a href=\"https://www.kaggle.com/leonnorblad\" target=\"_blank\">@leonnorblad</a> That's great! Glad to hear it worked, and I'm happy to help. Thank you for checking it out with me - let me know if you run into further issues!</p>",
              "rawMarkdown": "@leonnorblad That's great! Glad to hear it worked, and I'm happy to help. Thank you for checking it out with me - let me know if you run into further issues!"
            },
            {
              "id": 2231830,
              "postDate": "2023-04-23T18:03:35.557Z",
              "content": "<p>Great, <a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a>, glad to hear it's running longer. I don't see a bug in the code you posted, and I see that you had a successful submission in the interim. Please let me know if anything else fails!</p>",
              "rawMarkdown": "Great, @nynyny67, glad to hear it's running longer. I don't see a bug in the code you posted, and I see that you had a successful submission in the interim. Please let me know if anything else fails!",
              "votes": 1
            }
          ]
        },
        {
          "id": 2227606,
          "postDate": "2023-04-19T21:28:38.923Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I really admire your patience 😄</p>",
          "rawMarkdown": "@cdeotte I really admire your patience 😄",
          "votes": 6,
          "replies": [
            {
              "id": 2228258,
              "postDate": "2023-04-20T11:46:01.447Z",
              "rawMarkdown": "",
              "votes": 2,
              "isDeleted": true
            },
            {
              "id": 2228521,
              "postDate": "2023-04-20T15:46:33.460Z",
              "content": "<p>Yes, Chris, and several other community members, have been incredibly helpful. There are a tremendous number of moving parts when we do updates like this (including some manual fixes for the public API in this competition, which is unusual), and I apologize for the issues occurring. I very much appreciate when the community calls out the problems that they see, and their patience as we resolve them.</p>",
              "rawMarkdown": "Yes, Chris, and several other community members, have been incredibly helpful. There are a tremendous number of moving parts when we do updates like this (including some manual fixes for the public API in this competition, which is unusual), and I apologize for the issues occurring. I very much appreciate when the community calls out the problems that they see, and their patience as we resolve them.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2226500,
      "postDate": "2023-04-19T01:39:46.800Z",
      "content": "<p>Hi all!</p>\n<p>We are ready to begin running the leaderboard update discussed when the data was updated. It will occur over a period of a day or so. You will notice your old submissions being rerun, and the leaderboard changing.</p>\n<p>As we mentioned in our previous post, the API bug fixes that I implemented recently will likely cause many older submissions to fail. My apologies for this; it was the only path that ensured we were entirely solving the reported issues.</p>\n<p>We have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.</p>\n<p>Thanks to everyone who reported issues, and thanks to all for your patience.</p>\n<p>Note: The LB update is underway but will likely take several days as we will be running in batches. Old notebooks with code that does not conform to the API updates will likely error out. Thanks for your patience!</p>",
      "rawMarkdown": "Hi all!\n\nWe are ready to begin running the leaderboard update discussed when the data was updated. It will occur over a period of a day or so. You will notice your old submissions being rerun, and the leaderboard changing.\n\nAs we mentioned in our previous post, the API bug fixes that I implemented recently will likely cause many older submissions to fail. My apologies for this; it was the only path that ensured we were entirely solving the reported issues.\n\nWe have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.\n\nThanks to everyone who reported issues, and thanks to all for your patience.\n\nNote: The LB update is underway but will likely take several days as we will be running in batches. Old notebooks with code that does not conform to the API updates will likely error out. Thanks for your patience!",
      "votes": 6
    },
    {
      "id": 2244057,
      "postDate": "2023-05-03T11:58:36.217Z",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Hi Phil, is there any update? The rerun seems to be suspended for several days. I guess we need to know what's going on there.</p>",
      "rawMarkdown": "@philculliton Hi Phil, is there any update? The rerun seems to be suspended for several days. I guess we need to know what's going on there.",
      "votes": 1,
      "replies": [
        {
          "id": 2244119,
          "postDate": "2023-05-03T13:05:33.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> And it seems that the test data has been updated without any notification. I recently reran my notebook and my LB score dropped from 0.695 to 0.693 without any code changes. When I ran the previously open-sourced code with a LB score of 0.697, I also got 0.693. However, it appears that the LB still retains the ranking from before the test data update. For example, most people achieved a score of 0.697 with only 1-2 submissions, which seems to be using the previously open-sourced code. Is this unfair?</p>",
          "rawMarkdown": "@philculliton And it seems that the test data has been updated without any notification. I recently reran my notebook and my LB score dropped from 0.695 to 0.693 without any code changes. When I ran the previously open-sourced code with a LB score of 0.697, I also got 0.693. However, it appears that the LB still retains the ranking from before the test data update. For example, most people achieved a score of 0.697 with only 1-2 submissions, which seems to be using the previously open-sourced code. Is this unfair?",
          "votes": 2,
          "replies": [
            {
              "id": 2244208,
              "postDate": "2023-05-03T14:07:07.210Z",
              "content": "<p>Hi! <a href=\"https://www.kaggle.com/clement1999\" target=\"_blank\">@clement1999</a> We have not updated the hidden test set since the announcement of the last data update on March 20th. Was your 0.695 score from before that update?</p>\n<p><a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a> The rerun is taking quite a while (we have roughly 20,000 submissions to rerun and they are being run in batches and then checked for issues). I will post an update when it is complete.</p>",
              "rawMarkdown": "Hi! @clement1999 We have not updated the hidden test set since the announcement of the last data update on March 20th. Was your 0.695 score from before that update?\n\n@hookman The rerun is taking quite a while (we have roughly 20,000 submissions to rerun and they are being run in batches and then checked for issues). I will post an update when it is complete.",
              "votes": -3
            },
            {
              "id": 2244223,
              "postDate": "2023-05-03T14:19:09.713Z",
              "content": "<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403316#2243188\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403316#2243188</a><br>\n<a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> please check this</p>",
              "rawMarkdown": "https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403316#2243188\n@philculliton please check this"
            },
            {
              "id": 2244226,
              "postDate": "2023-05-03T14:20:17.183Z",
              "content": "<p>I have exact the same subs 18 days ago and yesterday. They have different lb scores :) </p>",
              "rawMarkdown": "I have exact the same subs 18 days ago and yesterday. They have different lb scores :) "
            },
            {
              "id": 2244229,
              "postDate": "2023-05-03T14:22:17.820Z",
              "content": "<p>Actually it seems that many people including me can't reproduce LB score with same code…</p>",
              "rawMarkdown": "Actually it seems that many people including me can't reproduce LB score with same code..."
            },
            {
              "id": 2244248,
              "postDate": "2023-05-03T14:34:53.773Z",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Thanks. By the way, I can also confirm that I cannot reproduce the lb score using exactly same models and codes. There must be something wrong about the hidden test set(maybe the dataset is the same, but the public lb percentage change? As far as I remembered, the public percentage is 50% before?)</p>",
              "rawMarkdown": "@philculliton Thanks. By the way, I can also confirm that I cannot reproduce the lb score using exactly same models and codes. There must be something wrong about the hidden test set(maybe the dataset is the same, but the public lb percentage change? As far as I remembered, the public percentage is 50% before?)"
            },
            {
              "id": 2244299,
              "postDate": "2023-05-03T15:01:16.093Z",
              "content": "<p>So we need another rerun 😀</p>",
              "rawMarkdown": "So we need another rerun 😀"
            },
            {
              "id": 2244576,
              "postDate": "2023-05-03T18:08:01.127Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kvlmll\" target=\"_blank\">@kvlmll</a>, I confirm that running exactly the same code the LB score changed and it happened 18 days ago. In my case, I get 0.001 lower now. I can even compare the training CV and the features importance and they are exactly the same. This means that the training data hasn't changed. Something changed in the hidden LB test data. Maybe the split public/private test data ?</p>",
              "rawMarkdown": "Hi @kvlmll, I confirm that running exactly the same code the LB score changed and it happened 18 days ago. In my case, I get 0.001 lower now. I can even compare the training CV and the features importance and they are exactly the same. This means that the training data hasn't changed. Something changed in the hidden LB test data. Maybe the split public/private test data ?"
            },
            {
              "id": 2244811,
              "postDate": "2023-05-03T22:41:54.690Z",
              "content": "<p>The ratio of public/private test data has returned to 50/50 from 36/64.</p>",
              "rawMarkdown": "The ratio of public/private test data has returned to 50/50 from 36/64."
            },
            {
              "id": 2244850,
              "postDate": "2023-05-03T23:30:49.080Z",
              "content": "<p>So do they change the test data again or it's only a description bug? Do I need to rerun my notebook recently submitted to get a different result? Will there be an official clarification for all these things?</p>",
              "rawMarkdown": "So do they change the test data again or it's only a description bug? Do I need to rerun my notebook recently submitted to get a different result? Will there be an official clarification for all these things?"
            },
            {
              "id": 2244865,
              "postDate": "2023-05-04T00:12:58.387Z",
              "content": "<p>I reran the notebook submitted when public / private split was 36/64 and LB does not change. So I think it is only<br>\na description bug.</p>",
              "rawMarkdown": "I reran the notebook submitted when public / private split was 36/64 and LB does not change. So I think it is only\na description bug.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2229448,
      "postDate": "2023-04-21T11:24:40.017Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, thank you for the hard work and letting us know.</p>\n<p>One question: Why does the new sample_submission.csv file maps question numbers to level_groups differently now?</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403129\" target=\"_blank\">see discussion</a></p>",
      "rawMarkdown": "Hi @philculliton, thank you for the hard work and letting us know.\n\nOne question: Why does the new sample_submission.csv file maps question numbers to level_groups differently now?\n\n[see discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403129)",
      "votes": 1,
      "replies": [
        {
          "id": 2231861,
          "postDate": "2023-04-23T18:38:22.023Z",
          "content": "<p>Hi! One of the fixes to the public API was grouping the data differently. As mentioned in the thread you linked, <code>session_level</code> is how we're grouping the data when it's being served - we had an issue in the public API with that column being set incorrectly.</p>",
          "rawMarkdown": "Hi! One of the fixes to the public API was grouping the data differently. As mentioned in the thread you linked, `session_level` is how we're grouping the data when it's being served - we had an issue in the public API with that column being set incorrectly."
        }
      ]
    },
    {
      "id": 2228448,
      "postDate": "2023-04-20T14:58:41.840Z",
      "content": "<p>Hello,</p>\n<p>It appears that the API was functioning normally 8 hours ago. However, in my latest code, I am experiencing an unexpected issue with the <code>sample_submission</code> from the <code>iter_test</code>. Specifically, there are some lines that are equivalent to <code>{'session_id': ['session_id'], 'correct': ['correct']}</code>, which generated afterwards. And this causes an error when I attempt to submit my work.</p>\n<p>I have double-checked my predictions to ensure that they are correct, but the issue persists.</p>",
      "rawMarkdown": "Hello,\n\nIt appears that the API was functioning normally 8 hours ago. However, in my latest code, I am experiencing an unexpected issue with the `sample_submission` from the `iter_test`. Specifically, there are some lines that are equivalent to `{'session_id': ['session_id'], 'correct': ['correct']}`, which generated afterwards. And this causes an error when I attempt to submit my work.\n\nI have double-checked my predictions to ensure that they are correct, but the issue persists.",
      "votes": 1
    },
    {
      "id": 2226546,
      "postDate": "2023-04-19T03:12:28.003Z",
      "content": "<p>Thanks for updating the mock public sample, <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> We have been looking forward to this.</p>",
      "rawMarkdown": "Thanks for updating the mock public sample, @philculliton We have been looking forward to this.",
      "votes": 1
    },
    {
      "id": 2246842,
      "postDate": "2023-05-05T13:59:25.310Z",
      "content": "<p>thanks , you are very good for us competitors</p>",
      "rawMarkdown": "thanks , you are very good for us competitors",
      "votes": 2
    },
    {
      "id": 2255846,
      "postDate": "2023-05-12T03:30:58.303Z",
      "content": "<p>Comment for contributer level</p>",
      "rawMarkdown": "Comment for contributer level",
      "replies": [
        {
          "id": 2269016,
          "postDate": "2023-05-22T07:03:53.413Z",
          "content": "<p>same to you hahah</p>",
          "rawMarkdown": "same to you hahah"
        }
      ]
    },
    {
      "id": 2239862,
      "postDate": "2023-04-30T01:02:18.543Z",
      "content": "<p>I got the same bug for all the day</p>",
      "rawMarkdown": "I got the same bug for all the day\n"
    },
    {
      "id": 2238079,
      "postDate": "2023-04-28T07:36:43.550Z",
      "content": "<p>I think it would be right to reset the leaderboard and invite everyone to restart the current notebooks</p>",
      "rawMarkdown": "I think it would be right to reset the leaderboard and invite everyone to restart the current notebooks"
    },
    {
      "id": 2230594,
      "postDate": "2023-04-22T14:36:36.607Z",
      "content": "<p>Hello everybody. I am trying to to create a  submission using the sample notebook provided, but the submission is being processed for over two hours already. Is it ok for this competition? </p>",
      "rawMarkdown": "Hello everybody. I am trying to to create a  submission using the sample notebook provided, but the submission is being processed for over two hours already. Is it ok for this competition? "
    },
    {
      "id": 2229659,
      "postDate": "2023-04-21T15:05:25.547Z",
      "content": "<p>i thought the leaderboard changing?</p>",
      "rawMarkdown": "i thought the leaderboard changing?",
      "replies": [
        {
          "id": 2230542,
          "postDate": "2023-04-22T13:22:26.873Z",
          "content": "<p>We have reports of new API issues and are pausing the LB updates to check them out.</p>",
          "rawMarkdown": "We have reports of new API issues and are pausing the LB updates to check them out.",
          "replies": [
            {
              "id": 2232494,
              "postDate": "2023-04-24T11:46:04.933Z",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Hi Phil, would you mind sharing the details and status of the new API issues, as well as when the LB updates are going to take place? TIA!</p>",
              "rawMarkdown": "@philculliton Hi Phil, would you mind sharing the details and status of the new API issues, as well as when the LB updates are going to take place? TIA!",
              "votes": 4
            },
            {
              "id": 2234216,
              "postDate": "2023-04-25T02:40:37.810Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> - the new reported issue appears to not be an API issue. The first batch of notebooks is rerunning now - the LB update is underway.</p>",
              "rawMarkdown": "Hi @hoangnguyen719 - the new reported issue appears to not be an API issue. The first batch of notebooks is rerunning now - the LB update is underway.",
              "votes": 3
            },
            {
              "id": 2236163,
              "postDate": "2023-04-26T15:35:50.413Z",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, can you please let us know when the LB updating is done? Thanks</p>",
              "rawMarkdown": "@philculliton, can you please let us know when the LB updating is done? Thanks",
              "votes": 4
            },
            {
              "id": 2237757,
              "postDate": "2023-04-27T23:49:34.520Z",
              "content": "<p>Yes i would also like to know when the LB is finished updating. Is the top LB score of 0.753 legitimate? It has been 2 days now, and the score hasn't changed. So I assume LB update finished and that LB 0.753 is legitimate.</p>",
              "rawMarkdown": "Yes i would also like to know when the LB is finished updating. Is the top LB score of 0.753 legitimate? It has been 2 days now, and the score hasn't changed. So I assume LB update finished and that LB 0.753 is legitimate.",
              "votes": 3
            },
            {
              "id": 2237806,
              "postDate": "2023-04-28T01:10:19.383Z",
              "content": "<p>It seems rerun has been stopped without a reason because some of my old submissions doesn't become error, and kaggle staff disappear again ?</p>",
              "rawMarkdown": "It seems rerun has been stopped without a reason because some of my old submissions doesn't become error, and kaggle staff disappear again ?",
              "votes": 7
            },
            {
              "id": 2238316,
              "postDate": "2023-04-28T12:07:18.830Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F911838%2Fc0afe65d71c1962375048c4ad0e5ebff%2Fw480.jpg?generation=1682683738711705&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F911838%2Fc0afe65d71c1962375048c4ad0e5ebff%2Fw480.jpg?generation=1682683738711705&alt=media)"
            }
          ]
        }
      ]
    },
    {
      "id": 2226751,
      "postDate": "2023-04-19T07:30:26.010Z",
      "content": "<p>Hello!</p>\n<blockquote>\n  <p>We have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.</p>\n</blockquote>\n<p>Now we have 3 session_ids in <code>test</code> and only 1 session_id in <code>sample_submission</code>.</p>\n<p>Example for first iteration:</p>\n<pre><code> jo_wilder\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n (test, sample_submission)  iter_test:\n    ()\n    ()\n    \n\n\n\n</code></pre>\n<p>How should we use <code>20090312143683264</code> and <code>20090312331414616</code> session_ids in <code>test</code>?</p>",
      "rawMarkdown": "Hello!\n\n> We have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.\n\nNow we have 3 session_ids in `test` and only 1 session_id in `sample_submission`.\n\nExample for first iteration:\n```python\nimport jo_wilder\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nfor (test, sample_submission) in iter_test:\n    print(f\"Test session_ids: {test['session_id'].unique()}\")\n    print(f\"Sample submission session_ids: {sample_submission['session_id'].unique()}\")\n    break\n\n# Test session_ids: [20090109393214576 20090312143683264 20090312331414616]\n# Sample submission session_ids: ['20090109393214576_q1' '20090109393214576_q2' '20090109393214576_q3']\n```\n\nHow should we use `20090312143683264` and `20090312331414616` session_ids in `test`?",
      "votes": 8,
      "isDeleted": true,
      "replies": [
        {
          "id": 2227374,
          "postDate": "2023-04-19T18:00:35.583Z",
          "content": "<p>Thanks! This should now be fixed.</p>",
          "rawMarkdown": "Thanks! This should now be fixed.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2228354,
      "author_name": "嘴爷",
      "author_url": "",
      "post_date": "2023-04-20T13:41:14.580000",
      "content": "<p>One bug with one month to fix, we call this KAGGLE SPEED.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2228494,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2023-04-20T15:29:53.100000",
          "content": "<p>Hi. I understand that's it taken some time. I apologize. I should note that this wasn't one bug, but a series of issues. A great deal of time has gone into finding a path forward, building the new competition data, working with the host to ensure that their goals were still being met, and attempting to reproduce, solve, and address all reported issues before we take the (quite large and time-consuming) step of updating the leaderboard.</p>\n<p>Updating a running competition with new data is a painstaking process that takes significant time and often creates secondary issues that need resolving. In this case we had no choice but to do an update. Our priority is, as always, making sure that the competition is viable.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2227281,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-04-19T16:14:38.707000",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Hi Phil, the mock public sample (i.e. commit API) is still <strong>not fixed</strong>. During each for-loop iteration, the mock public sample provides <strong>3 unique session_id</strong> in the <code>test dataframe</code> and <strong>1 unique session_id</strong> in the <code>sample_submission dataframe</code>.</p>\n<p>The correct implementation (to match submit API) is <strong>1 unique session_id</strong> in the <code>test dataframe</code> and <strong>1 unique session_id</strong> in the <code>sample_submission dataframe</code>.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2227373,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2023-04-19T17:59:59.253000",
          "content": "<p>Thanks Chris! Appreciate it - was an issue with the update. I've fixed it.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2229324,
              "author_name": "Leon Norblad",
              "author_url": "",
              "post_date": "2023-04-21T09:15:28.570000",
              "content": "<p>I very much appreciate that you are trying to fix the submission, however, there seem to still be some problems. The problem now seems to be that the API requires the user to predict all 18 questions for one session_id in every iter_test(). I found this after running two similar notebooks, one that predicts the questions that should be predicted for the 'level_group' (fails) and one that predicts all 18 questions at each iteration (successes). This was hard to spot since the submission.csv looks good while testing the submission.</p>\n<p>If anyone else has succeeded in submitting a notebook predicting only the relevant questions for each level_group in each iter_test() iteration. Please show me how to do it.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2230524,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-22T13:06:41.870000",
              "content": "<p>Hi - I can't reproduce this issue. The public process tests out just fine for me. Would you mind sharing a notebook with me? Feel free to remove anything you don't want to share, I just need to see how the issue is being raised for you.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2230683,
              "author_name": "Leon Norblad",
              "author_url": "",
              "post_date": "2023-04-22T16:51:58.820000",
              "content": "<p>Hi! I made two \"minimalistic\" notebooks to showcase this, I hope it is clear!<br>\nHere is my notebook, predicting 1's for all 18 questions, at each iteration (succeeds): <a href=\"https://www.kaggle.com/leonnorblad/predict-all-18-at-eah-iter\" target=\"_blank\">https://www.kaggle.com/leonnorblad/predict-all-18-at-eah-iter</a></p>\n<p>Here is my notebook, predicting 1's for questions that should be predicted for the current 'level_group' (fails): <a href=\"https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group\" target=\"_blank\">https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group</a></p>\n<p>I based the second notebook on how I believe the game is played and should be predicted (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796)</a>.   The player plays a couple of levels (we get data) and then we make a prediction, then it plays a few more levels (we get more data) and then we make predictions for the new questions. Have I misunderstood anything?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2230843,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-22T19:43:16.523000",
              "content": "<p>Thanks Leon! Would you mind sharing the notebooks with me?</p>\n<p>Your description of how the game is played is correct; when the notebooks are shared I'll look at your code to see where there might be an issue.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2230854,
              "author_name": "Leon Norblad",
              "author_url": "",
              "post_date": "2023-04-22T20:02:53.863000",
              "content": "<p>Sorry about that, they should be shared now :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2230898,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-22T21:18:33.443000",
              "content": "<p>Thanks very much Leon!</p>\n<p>Re: your first notebook - when you're running:</p>\n<pre><code>    for quest_nr in range(1,19):\n        sample_submission.loc[sample_submission.question == quest_nr, 'correct'] = 1\n</code></pre>\n<p>You're actually <em>not</em> setting anything for the questions that you don't currently have access to. <code>sample_submission.loc[sample_submission.question == quest_nr]</code> is only non-empty for the subset of questions you can currently access (you can check by looking at the <code>.shape</code> of the <code>sample_submission.loc[sample_submission.question == quest_nr]</code> for every <code>quest_nr</code> in each iteration, it shows the number of rows as <code>0</code> for any questions you can't currently access).</p>\n<p>So - this notebook appears to be setting <code>correct</code> only for the questions you have access to, not all 18 in each iteration. Which is the intent, so I think we're good on this notebook, unless there's something I'm missing! Please confirm the above and let me know.</p>\n<p>In your second notebook, everything appears to work as expected - you're getting access to, and setting answers for, the correct questions. Is the failure only happening when you submit to the private leaderboard?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2231204,
              "author_name": "Leon Norblad",
              "author_url": "",
              "post_date": "2023-04-23T06:31:13.410000",
              "content": "<p>The first notebook was just a desperate attempt to make a successful submission. As you say it will only fill a value if a matching question exists, which is only the questions for the current 'level_group' when I test my submission by the three test sessions that we have access to.</p>\n<p>Is the failure only happening when you submit to the private leaderboard? -&gt; yes, everything looks fine when I run the notebook myself but I get the 'Submission Scoring Error' when trying to submit.</p>\n<p>The three test sessions that we have access to are fed (with iter_test) exactly how I expect them to be. Based on these two notebooks, I suspect that the test set in the private leaderboard is not fed the same way but instead in such a way that we needed to predict the 18 questions at each iteration (since the first notebook succeeds but the second notebook fails). This does not need to be the case, I just want to understand why the second notebook fails.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2231382,
              "author_name": "nynyny67",
              "author_url": "",
              "post_date": "2023-04-23T08:56:47.543000",
              "content": "<p>I have the same issue. I get 'Submission Scoring Error' after 3~4hours from the submission, although everything looks fine when I run the notebook myself.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2231678,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-23T14:38:50.467000",
              "content": "<p>Thanks Leon! Very much appreciate your building and sharing the samples, and walking me through them.</p>\n<p>I dug deeper on your second notebook. <code>level_group_to_predict = test[\"level_group\"][0]</code> is throwing an exception (very) early with the private API (<code>0</code> is not an element in the index of the grouping you're trying to access, so it's a simple <code>KeyError</code>).</p>\n<p>You can entirely eliminate that code - the <code>sample_submission</code> for a given iteration will ONLY contain the questions you need to answer. If you simply fill in / predict <code>correct</code> for the questions in the <code>sample_submission</code> in each iteration of the loop, you should be fine.</p>\n<p>Alternatively, you could try using <code>level_group_to_predict = test[\"level_group\"].values[0]</code> if you just want the the <code>level_group</code> value for the first row in that sequence. This change appeared to work for me when testing your code.</p>\n<p>Please give this a shot and let me know if your issue persists.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2231686,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-23T14:49:23.153000",
              "content": "<p><a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a> - It looks like your submissions are erroring out before reaching the end of the loop. I can't tell what's happening in the code itself.</p>\n<p>It also appears as though you had a successful run of a notebook with the same name a few hours before the most recent one. Was there something you were doing differently there?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2231709,
              "author_name": "nynyny67",
              "author_url": "",
              "post_date": "2023-04-23T15:11:24.050000",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI commented out the line that change the 'sample_submission' like below.</p>\n<pre><code>#df_sample_submission.loc[df_sample_submission.session_id.str.contains(f'q{qid}'),'correct'] = int(prob&gt;threshold)\n</code></pre>\n<p>I also deleted some lines that prints some variables for the purpose of debugging. Maybe these lines contained some reference error, but I couldn't find suspecious parts.</p>\n<p>After I saw the submission process succeeded, I have activated this line and submission is now on progress. Before this study the submission fails in 3-4 hours but recent submission is running 5+ hours until now.<br>\nI don't know why the submissions are erroring before reaching the end of the loop, but the recent submissions continues longer at least.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2231738,
              "author_name": "Leon Norblad",
              "author_url": "",
              "post_date": "2023-04-23T15:34:16.780000",
              "content": "<p>Thank you very much! Yes, I can confirm that this works. I have updated my code for my second notebook and the submission succeeds!<br>\n(<a href=\"https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group\" target=\"_blank\">https://www.kaggle.com/code/leonnorblad/predict-only-questions-for-each-level-group</a>)</p>\n<p>Thank you for your time!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2231765,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-23T16:23:09.673000",
              "content": "<p><a href=\"https://www.kaggle.com/leonnorblad\" target=\"_blank\">@leonnorblad</a> That's great! Glad to hear it worked, and I'm happy to help. Thank you for checking it out with me - let me know if you run into further issues!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2231830,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-23T18:03:35.557000",
              "content": "<p>Great, <a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a>, glad to hear it's running longer. I don't see a bug in the code you posted, and I see that you had a successful submission in the interim. Please let me know if anything else fails!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2227606,
          "author_name": "empty",
          "author_url": "",
          "post_date": "2023-04-19T21:28:38.923000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I really admire your patience 😄</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2228258,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-20T11:46:01.447000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2228521,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-20T15:46:33.460000",
              "content": "<p>Yes, Chris, and several other community members, have been incredibly helpful. There are a tremendous number of moving parts when we do updates like this (including some manual fixes for the public API in this competition, which is unusual), and I apologize for the issues occurring. I very much appreciate when the community calls out the problems that they see, and their patience as we resolve them.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2244057,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2023-05-03T11:58:36.217000",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Hi Phil, is there any update? The rerun seems to be suspended for several days. I guess we need to know what's going on there.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244119,
          "author_name": "brannnn1995",
          "author_url": "",
          "post_date": "2023-05-03T13:05:33.677000",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> And it seems that the test data has been updated without any notification. I recently reran my notebook and my LB score dropped from 0.695 to 0.693 without any code changes. When I ran the previously open-sourced code with a LB score of 0.697, I also got 0.693. However, it appears that the LB still retains the ranking from before the test data update. For example, most people achieved a score of 0.697 with only 1-2 submissions, which seems to be using the previously open-sourced code. Is this unfair?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2244208,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-05-03T14:07:07.210000",
              "content": "<p>Hi! <a href=\"https://www.kaggle.com/clement1999\" target=\"_blank\">@clement1999</a> We have not updated the hidden test set since the announcement of the last data update on March 20th. Was your 0.695 score from before that update?</p>\n<p><a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a> The rerun is taking quite a while (we have roughly 20,000 submissions to rerun and they are being run in batches and then checked for issues). I will post an update when it is complete.</p>",
              "votes": -3,
              "replies": []
            },
            {
              "id": 2244223,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-05-03T14:19:09.713000",
              "content": "<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403316#2243188\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403316#2243188</a><br>\n<a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> please check this</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244226,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-05-03T14:20:17.183000",
              "content": "<p>I have exact the same subs 18 days ago and yesterday. They have different lb scores :) </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244229,
              "author_name": "嘴爷",
              "author_url": "",
              "post_date": "2023-05-03T14:22:17.820000",
              "content": "<p>Actually it seems that many people including me can't reproduce LB score with same code…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244248,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2023-05-03T14:34:53.773000",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Thanks. By the way, I can also confirm that I cannot reproduce the lb score using exactly same models and codes. There must be something wrong about the hidden test set(maybe the dataset is the same, but the public lb percentage change? As far as I remembered, the public percentage is 50% before?)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244299,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-05-03T15:01:16.093000",
              "content": "<p>So we need another rerun 😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244576,
              "author_name": "Elias",
              "author_url": "",
              "post_date": "2023-05-03T18:08:01.127000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kvlmll\" target=\"_blank\">@kvlmll</a>, I confirm that running exactly the same code the LB score changed and it happened 18 days ago. In my case, I get 0.001 lower now. I can even compare the training CV and the features importance and they are exactly the same. This means that the training data hasn't changed. Something changed in the hidden LB test data. Maybe the split public/private test data ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244811,
              "author_name": "theta",
              "author_url": "",
              "post_date": "2023-05-03T22:41:54.690000",
              "content": "<p>The ratio of public/private test data has returned to 50/50 from 36/64.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244850,
              "author_name": "Ya Xu",
              "author_url": "",
              "post_date": "2023-05-03T23:30:49.080000",
              "content": "<p>So do they change the test data again or it's only a description bug? Do I need to rerun my notebook recently submitted to get a different result? Will there be an official clarification for all these things?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2244865,
              "author_name": "theta",
              "author_url": "",
              "post_date": "2023-05-04T00:12:58.387000",
              "content": "<p>I reran the notebook submitted when public / private split was 36/64 and LB does not change. So I think it is only<br>\na description bug.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2229448,
      "author_name": "durvorezbariq",
      "author_url": "",
      "post_date": "2023-04-21T11:24:40.017000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, thank you for the hard work and letting us know.</p>\n<p>One question: Why does the new sample_submission.csv file maps question numbers to level_groups differently now?</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403129\" target=\"_blank\">see discussion</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2231861,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2023-04-23T18:38:22.023000",
          "content": "<p>Hi! One of the fixes to the public API was grouping the data differently. As mentioned in the thread you linked, <code>session_level</code> is how we're grouping the data when it's being served - we had an issue in the public API with that column being set incorrectly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2228448,
      "author_name": "Phu Gia Hoang",
      "author_url": "",
      "post_date": "2023-04-20T14:58:41.840000",
      "content": "<p>Hello,</p>\n<p>It appears that the API was functioning normally 8 hours ago. However, in my latest code, I am experiencing an unexpected issue with the <code>sample_submission</code> from the <code>iter_test</code>. Specifically, there are some lines that are equivalent to <code>{'session_id': ['session_id'], 'correct': ['correct']}</code>, which generated afterwards. And this causes an error when I attempt to submit my work.</p>\n<p>I have double-checked my predictions to ensure that they are correct, but the issue persists.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2226546,
      "author_name": "Thaweewat R",
      "author_url": "",
      "post_date": "2023-04-19T03:12:28.003000",
      "content": "<p>Thanks for updating the mock public sample, <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> We have been looking forward to this.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2246842,
      "author_name": "Haoyang Zhong",
      "author_url": "",
      "post_date": "2023-05-05T13:59:25.310000",
      "content": "<p>thanks , you are very good for us competitors</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2255846,
      "author_name": "Max Bodley",
      "author_url": "",
      "post_date": "2023-05-12T03:30:58.303000",
      "content": "<p>Comment for contributer level</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2269016,
          "author_name": "peixinxin",
          "author_url": "",
          "post_date": "2023-05-22T07:03:53.413000",
          "content": "<p>same to you hahah</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2239862,
      "author_name": "Elmehdi JABAR",
      "author_url": "",
      "post_date": "2023-04-30T01:02:18.543000",
      "content": "<p>I got the same bug for all the day</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2238079,
      "author_name": "Vadim Kamaev",
      "author_url": "",
      "post_date": "2023-04-28T07:36:43.550000",
      "content": "<p>I think it would be right to reset the leaderboard and invite everyone to restart the current notebooks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2230594,
      "author_name": "Mikhail Kivi",
      "author_url": "",
      "post_date": "2023-04-22T14:36:36.607000",
      "content": "<p>Hello everybody. I am trying to to create a  submission using the sample notebook provided, but the submission is being processed for over two hours already. Is it ok for this competition? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2229659,
      "author_name": "yu",
      "author_url": "",
      "post_date": "2023-04-21T15:05:25.547000",
      "content": "<p>i thought the leaderboard changing?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2230542,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2023-04-22T13:22:26.873000",
          "content": "<p>We have reports of new API issues and are pausing the LB updates to check them out.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2232494,
              "author_name": "Hoang Nguyen",
              "author_url": "",
              "post_date": "2023-04-24T11:46:04.933000",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Hi Phil, would you mind sharing the details and status of the new API issues, as well as when the LB updates are going to take place? TIA!</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2234216,
              "author_name": "Phil Culliton",
              "author_url": "",
              "post_date": "2023-04-25T02:40:37.810000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> - the new reported issue appears to not be an API issue. The first batch of notebooks is rerunning now - the LB update is underway.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2236163,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2023-04-26T15:35:50.413000",
              "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, can you please let us know when the LB updating is done? Thanks</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2237757,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-04-27T23:49:34.520000",
              "content": "<p>Yes i would also like to know when the LB is finished updating. Is the top LB score of 0.753 legitimate? It has been 2 days now, and the score hasn't changed. So I assume LB update finished and that LB 0.753 is legitimate.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2237806,
              "author_name": "嘴爷",
              "author_url": "",
              "post_date": "2023-04-28T01:10:19.383000",
              "content": "<p>It seems rerun has been stopped without a reason because some of my old submissions doesn't become error, and kaggle staff disappear again ?</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2238316,
              "author_name": "empty",
              "author_url": "",
              "post_date": "2023-04-28T12:07:18.830000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F911838%2Fc0afe65d71c1962375048c4ad0e5ebff%2Fw480.jpg?generation=1682683738711705&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2226751,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-19T07:30:26.010000",
      "content": "<p>Hello!</p>\n<blockquote>\n  <p>We have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.</p>\n</blockquote>\n<p>Now we have 3 session_ids in <code>test</code> and only 1 session_id in <code>sample_submission</code>.</p>\n<p>Example for first iteration:</p>\n<pre><code> jo_wilder\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n (test, sample_submission)  iter_test:\n    ()\n    ()\n    \n\n\n\n</code></pre>\n<p>How should we use <code>20090312143683264</code> and <code>20090312331414616</code> session_ids in <code>test</code>?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2227374,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2023-04-19T18:00:35.583000",
          "content": "<p>Thanks! This should now be fixed.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2228354": "One bug with one month to fix, we call this KAGGLE SPEED.",
    "2227281": "@philculliton Hi Phil, the mock public sample (i.e. commit API) is still **not fixed**. During each for-loop iteration, the mock public sample provides **3 unique session_id** in the `test dataframe` and **1 unique session_id** in the `sample_submission dataframe`.\n\nThe correct implementation (to match submit API) is **1 unique session_id** in the `test dataframe` and **1 unique session_id** in the `sample_submission dataframe`.",
    "2226500": "Hi all!\n\nWe are ready to begin running the leaderboard update discussed when the data was updated. It will occur over a period of a day or so. You will notice your old submissions being rerun, and the leaderboard changing.\n\nAs we mentioned in our previous post, the API bug fixes that I implemented recently will likely cause many older submissions to fail. My apologies for this; it was the only path that ensured we were entirely solving the reported issues.\n\nWe have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.\n\nThanks to everyone who reported issues, and thanks to all for your patience.\n\nNote: The LB update is underway but will likely take several days as we will be running in batches. Old notebooks with code that does not conform to the API updates will likely error out. Thanks for your patience!",
    "2244057": "@philculliton Hi Phil, is there any update? The rerun seems to be suspended for several days. I guess we need to know what's going on there.",
    "2229448": "Hi @philculliton, thank you for the hard work and letting us know.\n\nOne question: Why does the new sample_submission.csv file maps question numbers to level_groups differently now?\n\n[see discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/403129)",
    "2228448": "Hello,\n\nIt appears that the API was functioning normally 8 hours ago. However, in my latest code, I am experiencing an unexpected issue with the `sample_submission` from the `iter_test`. Specifically, there are some lines that are equivalent to `{'session_id': ['session_id'], 'correct': ['correct']}`, which generated afterwards. And this causes an error when I attempt to submit my work.\n\nI have double-checked my predictions to ensure that they are correct, but the issue persists.",
    "2226546": "Thanks for updating the mock public sample, @philculliton We have been looking forward to this.",
    "2246842": "thanks , you are very good for us competitors",
    "2255846": "Comment for contributer level",
    "2239862": "I got the same bug for all the day\n",
    "2238079": "I think it would be right to reset the leaderboard and invite everyone to restart the current notebooks",
    "2230594": "Hello everybody. I am trying to to create a  submission using the sample notebook provided, but the submission is being processed for over two hours already. Is it ok for this competition? ",
    "2229659": "i thought the leaderboard changing?",
    "2226751": "Hello!\n\n> We have also updated the mock public sample to be served correctly by the API. It was being served in larger chunks, as noted by several competitors.\n\nNow we have 3 session_ids in `test` and only 1 session_id in `sample_submission`.\n\nExample for first iteration:\n```python\nimport jo_wilder\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nfor (test, sample_submission) in iter_test:\n    print(f\"Test session_ids: {test['session_id'].unique()}\")\n    print(f\"Sample submission session_ids: {sample_submission['session_id'].unique()}\")\n    break\n\n# Test session_ids: [20090109393214576 20090312143683264 20090312331414616]\n# Sample submission session_ids: ['20090109393214576_q1' '20090109393214576_q2' '20090109393214576_q3']\n```\n\nHow should we use `20090312143683264` and `20090312331414616` session_ids in `test`?"
  }
}