{
  "id": 384801,
  "title": "Greetings from the Organizers!",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/384801",
  "author_name": "Natalie Rambis",
  "post_date": "2023-02-09T15:26:11.241000",
  "votes": 49,
  "comment_count": 38,
  "views": 0,
  "content": "<p>On behalf of Field Day Lab at University of Wisconsin - Madison, Vanderbilt University, and The Learning Agency Lab, welcome to Student Performance and Game Play: Predict student learning from Jo Wilder online educational game! We are thrilled to host this competition which tasks competitors with predicting assessment performance using data from Field Day’s online educational game Jo Wilder. </p>\n<p>Game-based learning allows for students to engage with educational content in a dynamic way that the typical classroom lesson doesn’t offer. Current research indicates that game-based learning could be effective in supporting learning outcomes, however few game-based learning  datasets makes it difficult to extend this research. Developing models that make real-time predictions about student performance on game objectives will support educational game developers to create more effective learning experiences for students. </p>\n<p>This competition provides access to data taken from <a href=\"https://pbswisconsineducation.org/jowilder/about/\" target=\"_blank\">Jo Wilder and the Capitol Case</a> (Jo Wilder) which is a free online point &amp; click game. Designed by Field Day Lab, Jo Wilder aligns with the Wisconsin Department of Public Instruction’s newly revised 3-5th grade Wisconsin Standards for Social Studies. The game utilizes historical analysis to encourage students to develop reading comprehension skills. Sections of the game assess the player’s comprehension of the story and the clues picked up along the way, which are used to create the measure of performance used in this competition. </p>\n<p>This competition will offer two tracks: a traditional accuracy track and a computationally efficient track. The computationally efficient track allows for environmentally friendly models that are easily adaptable in low-resource environments. The models, both traditional and efficient, will be fully open-sourced. </p>\n<p>Please don’t hesitate to reach out with questions regarding the data or competition.</p>\n<p>We hope everyone enjoys this competition, and we wish you all good luck!</p>\n<p>Hosts:<br>\nDavid Gagnon (davidgagnon), Field Day Lab at University of Wisconsin - Madison<br>\nDr. Scott Crossley (cookiecutters), Vanderbilt University<br>\nUlrich Boser (ulrichboser), The Learning Agency Lab<br>\nMeg Benner (megbenner), The Learning Agency Lab<br>\nPerpetual Baffour (pbaffour), The Learning Agency Lab<br>\nAlex Franklin (alexmlfranklin), The Learning Agency Lab<br>\nNatalie Rambis (nrambis), The Learning Agency Lab</p>",
  "messages": [
    {
      "id": 2136849,
      "postDate": "2023-02-09T15:26:11.240Z",
      "content": "<p>On behalf of Field Day Lab at University of Wisconsin - Madison, Vanderbilt University, and The Learning Agency Lab, welcome to Student Performance and Game Play: Predict student learning from Jo Wilder online educational game! We are thrilled to host this competition which tasks competitors with predicting assessment performance using data from Field Day’s online educational game Jo Wilder. </p>\n<p>Game-based learning allows for students to engage with educational content in a dynamic way that the typical classroom lesson doesn’t offer. Current research indicates that game-based learning could be effective in supporting learning outcomes, however few game-based learning  datasets makes it difficult to extend this research. Developing models that make real-time predictions about student performance on game objectives will support educational game developers to create more effective learning experiences for students. </p>\n<p>This competition provides access to data taken from <a href=\"https://pbswisconsineducation.org/jowilder/about/\" target=\"_blank\">Jo Wilder and the Capitol Case</a> (Jo Wilder) which is a free online point &amp; click game. Designed by Field Day Lab, Jo Wilder aligns with the Wisconsin Department of Public Instruction’s newly revised 3-5th grade Wisconsin Standards for Social Studies. The game utilizes historical analysis to encourage students to develop reading comprehension skills. Sections of the game assess the player’s comprehension of the story and the clues picked up along the way, which are used to create the measure of performance used in this competition. </p>\n<p>This competition will offer two tracks: a traditional accuracy track and a computationally efficient track. The computationally efficient track allows for environmentally friendly models that are easily adaptable in low-resource environments. The models, both traditional and efficient, will be fully open-sourced. </p>\n<p>Please don’t hesitate to reach out with questions regarding the data or competition.</p>\n<p>We hope everyone enjoys this competition, and we wish you all good luck!</p>\n<p>Hosts:<br>\nDavid Gagnon (davidgagnon), Field Day Lab at University of Wisconsin - Madison<br>\nDr. Scott Crossley (cookiecutters), Vanderbilt University<br>\nUlrich Boser (ulrichboser), The Learning Agency Lab<br>\nMeg Benner (megbenner), The Learning Agency Lab<br>\nPerpetual Baffour (pbaffour), The Learning Agency Lab<br>\nAlex Franklin (alexmlfranklin), The Learning Agency Lab<br>\nNatalie Rambis (nrambis), The Learning Agency Lab</p>",
      "rawMarkdown": "On behalf of Field Day Lab at University of Wisconsin - Madison, Vanderbilt University, and The Learning Agency Lab, welcome to Student Performance and Game Play: Predict student learning from Jo Wilder online educational game! We are thrilled to host this competition which tasks competitors with predicting assessment performance using data from Field Day’s online educational game Jo Wilder. \n\nGame-based learning allows for students to engage with educational content in a dynamic way that the typical classroom lesson doesn’t offer. Current research indicates that game-based learning could be effective in supporting learning outcomes, however few game-based learning  datasets makes it difficult to extend this research. Developing models that make real-time predictions about student performance on game objectives will support educational game developers to create more effective learning experiences for students. \n\nThis competition provides access to data taken from [Jo Wilder and the Capitol Case](https://pbswisconsineducation.org/jowilder/about/) (Jo Wilder) which is a free online point & click game. Designed by Field Day Lab, Jo Wilder aligns with the Wisconsin Department of Public Instruction’s newly revised 3-5th grade Wisconsin Standards for Social Studies. The game utilizes historical analysis to encourage students to develop reading comprehension skills. Sections of the game assess the player’s comprehension of the story and the clues picked up along the way, which are used to create the measure of performance used in this competition. \n\nThis competition will offer two tracks: a traditional accuracy track and a computationally efficient track. The computationally efficient track allows for environmentally friendly models that are easily adaptable in low-resource environments. The models, both traditional and efficient, will be fully open-sourced. \n\nPlease don’t hesitate to reach out with questions regarding the data or competition.\n\nWe hope everyone enjoys this competition, and we wish you all good luck!\n\nHosts:\nDavid Gagnon (davidgagnon), Field Day Lab at University of Wisconsin - Madison\nDr. Scott Crossley (cookiecutters), Vanderbilt University\nUlrich Boser (ulrichboser), The Learning Agency Lab\nMeg Benner (megbenner), The Learning Agency Lab\nPerpetual Baffour (pbaffour), The Learning Agency Lab\nAlex Franklin (alexmlfranklin), The Learning Agency Lab\nNatalie Rambis (nrambis), The Learning Agency Lab\n",
      "votes": 49
    },
    {
      "id": 2137042,
      "postDate": "2023-02-09T17:35:20.587Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> Thanks for this fun competition.</p>\n<p>Regarding the dataset and what we are predicting, what is the definition of \"correct answer\" and \"incorrrect answer\". During gameplay, it seems that we need to get every answer correct before continuing. So is \"correct answer\" mean \"get question correct on first try\"?</p>",
      "rawMarkdown": "Hi @nrambis Thanks for this fun competition.\n\nRegarding the dataset and what we are predicting, what is the definition of \"correct answer\" and \"incorrrect answer\". During gameplay, it seems that we need to get every answer correct before continuing. So is \"correct answer\" mean \"get question correct on first try\"?",
      "votes": 35,
      "replies": [
        {
          "id": 2137152,
          "postDate": "2023-02-09T18:58:31.833Z",
          "content": "<p>That's right, during gameplay the player must eventually select the correct answer to continue. A \"correct answer\" here indicates that <strong>the player got the answer correct on their first attempt</strong>. </p>",
          "rawMarkdown": "That's right, during gameplay the player must eventually select the correct answer to continue. A \"correct answer\" here indicates that **the player got the answer correct on their first attempt**. ",
          "votes": 48
        }
      ]
    },
    {
      "id": 2137727,
      "postDate": "2023-02-10T09:17:06.233Z",
      "content": "<p>Hi,  I have one question. </p>\n<p>In the Dataset Description, it is said:</p>\n<pre><code>At each checkpoint, you will have access to all previous test data for that section.\n</code></pre>\n<p>But in the time series api, when it comes to level_group 5-12, the test dataframe only contains the historical actions in level_group 5-12.  I guess we can get the level_group 0-4 at that time. So we may need to save the info of level_group 0-4 on my own during the iteration of time series api?</p>\n<p>Thanks in advanced.</p>",
      "rawMarkdown": "Hi,  I have one question. \n\nIn the Dataset Description, it is said:\n```\nAt each checkpoint, you will have access to all previous test data for that section.\n```\n\nBut in the time series api, when it comes to level_group 5-12, the test dataframe only contains the historical actions in level_group 5-12.  I guess we can get the level_group 0-4 at that time. So we may need to save the info of level_group 0-4 on my own during the iteration of time series api?\n\nThanks in advanced.",
      "votes": 7,
      "replies": [
        {
          "id": 2186858,
          "postDate": "2023-03-18T07:13:15.463Z",
          "content": "<p>Same question，have you got answer？</p>",
          "rawMarkdown": "Same question，have you got answer？"
        },
        {
          "id": 2190269,
          "postDate": "2023-03-21T06:38:06.830Z",
          "content": "<p>Same question</p>",
          "rawMarkdown": "Same question\n"
        },
        {
          "id": 2192992,
          "postDate": "2023-03-23T03:17:46Z",
          "content": "<p>Same question, can someone help?</p>",
          "rawMarkdown": "Same question, can someone help?"
        },
        {
          "id": 2196577,
          "postDate": "2023-03-25T14:03:39.410Z",
          "content": "<p>IMHO, we can save previous level_group information on my own and use them when the next level_group comes.</p>",
          "rawMarkdown": "IMHO, we can save previous level_group information on my own and use them when the next level_group comes."
        }
      ]
    },
    {
      "id": 2139879,
      "postDate": "2023-02-11T08:54:40.130Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> , thanks for hosting this competition. Expanding on <a href=\"https://www.kaggle.com/chrisqiu\" target=\"_blank\">@chrisqiu</a> question, there are also instances where the player restarted the game mid-way and re-do the questions from scratch. For example, session_id = 21100415142476300 has 8 checkpoints:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3123781%2Fbca062908f5f5597eb9f525145dce459%2FScreenshot%202023-02-11%20165320.png?generation=1676105618167188&amp;alt=media\" alt=\"\"></p>\n<p>In this case, are the answers recorded for the player's first or last attempt? </p>",
      "rawMarkdown": "Hi @nrambis , thanks for hosting this competition. Expanding on @chrisqiu question, there are also instances where the player restarted the game mid-way and re-do the questions from scratch. For example, session_id = 21100415142476300 has 8 checkpoints:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3123781%2Fbca062908f5f5597eb9f525145dce459%2FScreenshot%202023-02-11%20165320.png?generation=1676105618167188&alt=media)\n\nIn this case, are the answers recorded for the player's first or last attempt? ",
      "votes": 5,
      "replies": [
        {
          "id": 2147773,
          "postDate": "2023-02-16T21:42:28.997Z",
          "content": "<p>A player should not be able to navigate to previous parts of the game in the same session, but it appears that the game may sometimes not recognize a restart as a new session. These sessions may not provide as useful information as typical sessions with 3 checkpoints, but I believe this should only happen for a small portion of the data.</p>\n<p>To answer your question, in these cases the answers would be recorded for all of the attempts. In other words, the session would record a correct answer only if the player selected the correct option first every single time they answered the question.</p>",
          "rawMarkdown": "A player should not be able to navigate to previous parts of the game in the same session, but it appears that the game may sometimes not recognize a restart as a new session. These sessions may not provide as useful information as typical sessions with 3 checkpoints, but I believe this should only happen for a small portion of the data.\n\nTo answer your question, in these cases the answers would be recorded for all of the attempts. In other words, the session would record a correct answer only if the player selected the correct option first every single time they answered the question.",
          "votes": 3,
          "replies": [
            {
              "id": 2149342,
              "postDate": "2023-02-18T07:25:01.203Z",
              "content": "<blockquote>\n  <p>in these cases the answers would be recorded for all of the attempts. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> , sorry, I didn't fully understand it.  The session 20100110332615344 has 5 checkpoints for the level groups 0-4,5-12,13-22,0-4,5-12. Does the row 20100110332615344_q1 correspond to the first or the second attempt to answer the questions of level group 0-4?</p>",
              "rawMarkdown": ">  in these cases the answers would be recorded for all of the attempts. \n\n@alexmlfranklin , sorry, I didn't fully understand it.  The session 20100110332615344 has 5 checkpoints for the level groups 0-4,5-12,13-22,0-4,5-12. Does the row 20100110332615344_q1 correspond to the first or the second attempt to answer the questions of level group 0-4?",
              "votes": 1
            },
            {
              "id": 2152213,
              "postDate": "2023-02-20T16:42:50.687Z",
              "content": "<p>20100110332615344_q1 would have a value of 1 if the player answered question 1 correctly on their first attempt in <em>both</em> the first and second instance they encounter it. If <em>any</em> attempt was incorrect then 20100110332615344_q1 would have a value of 0. Does that help to clarify?</p>",
              "rawMarkdown": "20100110332615344_q1 would have a value of 1 if the player answered question 1 correctly on their first attempt in *both* the first and second instance they encounter it. If *any* attempt was incorrect then 20100110332615344_q1 would have a value of 0. Does that help to clarify?",
              "votes": 2
            },
            {
              "id": 2152951,
              "postDate": "2023-02-21T05:37:04.847Z",
              "content": "<p>Yes, it's clear now. Thank you.</p>",
              "rawMarkdown": "Yes, it's clear now. Thank you."
            },
            {
              "id": 2171818,
              "postDate": "2023-03-07T05:43:21.937Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> So if I am not wrong, each session_id should have 3 checkpoints in the ideal scenario and a total of 18 questions. But I see that there are 80 session ids where there are more than 3 checkpoints, basically seems like the user is able to go to the previous part and there are 30 session ids where there are exactly 6 checkpoints. My assumption is user is giving the test once just randomly in the first attempt, then going back and giving the correct answers. I say this because for session id with just 3 checkpoints 71% of the answers are correct, but where there are 6 checkpoints, just 51.85%, just like random.<br>\nI think it will be good idea to just drop such cases as they are less than 1%. Also, I would like to know if there is a way to identify how many questions are there in which level group?</p>",
              "rawMarkdown": "Hi @alexmlfranklin So if I am not wrong, each session_id should have 3 checkpoints in the ideal scenario and a total of 18 questions. But I see that there are 80 session ids where there are more than 3 checkpoints, basically seems like the user is able to go to the previous part and there are 30 session ids where there are exactly 6 checkpoints. My assumption is user is giving the test once just randomly in the first attempt, then going back and giving the correct answers. I say this because for session id with just 3 checkpoints 71% of the answers are correct, but where there are 6 checkpoints, just 51.85%, just like random.\nI think it will be good idea to just drop such cases as they are less than 1%. Also, I would like to know if there is a way to identify how many questions are there in which level group?"
            },
            {
              "id": 2174156,
              "postDate": "2023-03-08T22:22:51.167Z",
              "content": "<p>Yup that's right, a single playthrough of the game from beginning to end will include 3 checkpoints. To answer your question about level groups, there are 3, 10, and 5 questions in level groups 0-4, 5-12, and 13-22 respectively.</p>\n<p>I'll point to this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796\" target=\"_blank\">helpful post</a> by pjmathematician that lists the questions by group.</p>",
              "rawMarkdown": "Yup that's right, a single playthrough of the game from beginning to end will include 3 checkpoints. To answer your question about level groups, there are 3, 10, and 5 questions in level groups 0-4, 5-12, and 13-22 respectively.\n\nI'll point to this [helpful post](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796) by pjmathematician that lists the questions by group."
            },
            {
              "id": 2200037,
              "postDate": "2023-03-28T08:29:36.493Z",
              "content": "<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>, can I ask one more question concerning the above example with s=20100110332615344 and 5 groups 0-4,5-12,13-22,0-4*,5-12* (asterisk denotes the second occurrence of the group). I can't imagine how env.iter_test() generates such sessions. If the order is (s,0-4 + 0-4*), (s,5-12 + 5-12*),(s,13-22) then we have time leak. If the order is chronological  (s,0-4),(s,5-12),(s,13-22),(s,0-4*),(s,5-12*) then we make predictions for questions 1-13 twice, which is also strange since there are no two indexes 20100110332615344_q1 in the sample_submission I think.</p>",
              "rawMarkdown": "@alexmlfranklin, can I ask one more question concerning the above example with s=20100110332615344 and 5 groups 0-4,5-12,13-22,0-4\\*,5-12\\* (asterisk denotes the second occurrence of the group). I can't imagine how env.iter_test() generates such sessions. If the order is (s,0-4 + 0-4\\*), (s,5-12 + 5-12\\*),(s,13-22) then we have time leak. If the order is chronological  (s,0-4),(s,5-12),(s,13-22),(s,0-4\\*),(s,5-12\\*) then we make predictions for questions 1-13 twice, which is also strange since there are no two indexes 20100110332615344_q1 in the sample_submission I think.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2137318,
      "postDate": "2023-02-09T21:10:25.423Z",
      "content": "<p>I have a question about the order of the data output by the time series api.  </p>\n<p>Is it correct to assume that the data is provided in the correct order of \"0-4\", \"5-12\", and \"13-22\" in level_group for all sessions?<br>\nOr is it better to assume that omissions or shuffling of order may occur?</p>",
      "rawMarkdown": "I have a question about the order of the data output by the time series api.  \n\nIs it correct to assume that the data is provided in the correct order of \"0-4\", \"5-12\", and \"13-22\" in level_group for all sessions?\nOr is it better to assume that omissions or shuffling of order may occur?",
      "votes": 4
    },
    {
      "id": 2264174,
      "postDate": "2023-05-18T08:20:46.177Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> . Thanks for hosting this competition. </p>",
      "rawMarkdown": "Hi, @nrambis . Thanks for hosting this competition. ",
      "votes": 1
    },
    {
      "id": 2142031,
      "postDate": "2023-02-13T09:42:44.960Z",
      "content": "<p>Hello, is it possible to not count failed submissions towards submission attempts? I'm trying to debug why my submissions fail, while the initial commit passes without any problems and it is very frustrating due to each failed notebook accounts towards submission allowance while no detailed information on where the exception happened or what caused it. I've lost two days worth of submits and still have no idea of what's causing the exception.</p>",
      "rawMarkdown": "Hello, is it possible to not count failed submissions towards submission attempts? I'm trying to debug why my submissions fail, while the initial commit passes without any problems and it is very frustrating due to each failed notebook accounts towards submission allowance while no detailed information on where the exception happened or what caused it. I've lost two days worth of submits and still have no idea of what's causing the exception.",
      "votes": 2,
      "replies": [
        {
          "id": 2156765,
          "postDate": "2023-02-23T14:41:46.123Z",
          "content": "<p>Also my recent submissions fail. Could it be that the column \"session_level\" is not included in the test environments sample_submission, and that's why the output file is wrong format. The column is included in the competitions \"data\" tab, in \"sample_submission.csv\".</p>",
          "rawMarkdown": "Also my recent submissions fail. Could it be that the column \"session_level\" is not included in the test environments sample_submission, and that's why the output file is wrong format. The column is included in the competitions \"data\" tab, in \"sample_submission.csv\".",
          "votes": 1,
          "replies": [
            {
              "id": 2158636,
              "postDate": "2023-02-25T03:08:43.693Z",
              "content": "<p>Thanks for the help!</p>",
              "rawMarkdown": "Thanks for the help!"
            }
          ]
        }
      ]
    },
    {
      "id": 2139816,
      "postDate": "2023-02-11T07:30:37.290Z",
      "content": "<p>Hi, thanks for hosting the competition. I have two questions.</p>\n<ol>\n<li><p>I noticed that one session in the training set has only two checkpoints. session_id: <code>22090108192456930</code>. <br>\nFirst I think this user might quit the game before answering all the questions, but actually it is the first checkpoint that is missing (fqid: chap1_finale_c). And in labels the same user answers all the questions. <br>\nCan I assume there is a missing row for this user? <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2F5741cf69c7cb27d7009f33ca913ecacc%2F1.png?generation=1676100510691203&amp;alt=media\" alt=\"\"></p></li>\n<li><p>The second event of session '20090312431273200' has 1323 elapsed_time, but the third event has 831. This seems wrong because by inspecting the text I know these events are in order. So, is this because the user clicks too fast, or is it a wrong value?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2Fd9fdb7719ea5a35c16ddcad71f255d0c%2F2.png?generation=1676100576722265&amp;alt=media\" alt=\"\"></p></li>\n</ol>\n<p>Thank you very much.</p>",
      "rawMarkdown": "Hi, thanks for hosting the competition. I have two questions.\n\n1. I noticed that one session in the training set has only two checkpoints. session_id: `22090108192456930`. \nFirst I think this user might quit the game before answering all the questions, but actually it is the first checkpoint that is missing (fqid: chap1_finale_c). And in labels the same user answers all the questions. \nCan I assume there is a missing row for this user? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2F5741cf69c7cb27d7009f33ca913ecacc%2F1.png?generation=1676100510691203&alt=media)\n\n2. The second event of session '20090312431273200' has 1323 elapsed_time, but the third event has 831. This seems wrong because by inspecting the text I know these events are in order. So, is this because the user clicks too fast, or is it a wrong value?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2Fd9fdb7719ea5a35c16ddcad71f255d0c%2F2.png?generation=1676100576722265&alt=media)\n\n\nThank you very much.",
      "votes": 2,
      "replies": [
        {
          "id": 2147768,
          "postDate": "2023-02-16T21:34:55.863Z",
          "content": "<ol>\n<li><p>Yes, you can assume that there is a missing row for that user.</p></li>\n<li><p>The game's timestamps can sometimes be imprecise, which may be partly due to clicks happening quickly in succession. You can assume the <code>index</code> gives the correct order of events in cases where the <code>elapsed_time</code> is inconsistent.</p></li>\n</ol>",
          "rawMarkdown": "1. Yes, you can assume that there is a missing row for that user.\n\n2. The game's timestamps can sometimes be imprecise, which may be partly due to clicks happening quickly in succession. You can assume the `index` gives the correct order of events in cases where the `elapsed_time` is inconsistent.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2315242,
      "postDate": "2023-06-24T00:14:24.037Z",
      "content": "<p>I'm reaching out to draw your attention to a matter that has been causing me some concerns for the past five days. I initially reached out to you via email three days ago but have yet to receive a response.</p>\n<p>Understanding your busy schedule, I thought it prudent to follow up on this platform, given your invitation in this post encouraging participants to ask questions.</p>\n<p>Could you please assist by addressing the issue at your earliest convenience? Your guidance would be immensely appreciated.</p>",
      "rawMarkdown": "I'm reaching out to draw your attention to a matter that has been causing me some concerns for the past five days. I initially reached out to you via email three days ago but have yet to receive a response.\n\nUnderstanding your busy schedule, I thought it prudent to follow up on this platform, given your invitation in this post encouraging participants to ask questions.\n\nCould you please assist by addressing the issue at your earliest convenience? Your guidance would be immensely appreciated.\n"
    },
    {
      "id": 2308330,
      "postDate": "2023-06-18T20:04:43.570Z",
      "content": "<p>\". The models, both traditional and efficient, will be fully open-sourced\"; this is gold.</p>",
      "rawMarkdown": "\". The models, both traditional and efficient, will be fully open-sourced\"; this is gold."
    },
    {
      "id": 2301455,
      "postDate": "2023-06-14T00:31:34.760Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> and <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>, could you please clarify meaning of several parameters mentioned on the <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/efficiency-prize-evaluation\" target=\"_blank\">efficiency evaluation page</a>:</p>\n<ol>\n<li><p>\"<code>Benchmark</code>  is the score of the benchmark <code>sample_submission.csv</code>\". What is the <code>sample_submission.csv</code> and where to get its score?</p></li>\n<li><p>\"<code>RuntimeSeconds</code> is the number of seconds it takes for the submission to be evaluated\". To be evaluated means just the scoring time? Or that's the total notebook runtime, i.e. fitting+scoring?</p></li>\n</ol>\n<p>Thanks you in advance for the information.</p>",
      "rawMarkdown": "Dear @nrambis and @alexmlfranklin, could you please clarify meaning of several parameters mentioned on the [efficiency evaluation page](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/efficiency-prize-evaluation):\n\n1. \"`Benchmark`  is the score of the benchmark `sample_submission.csv`\". What is the `sample_submission.csv` and where to get its score?\n\n2. \"`RuntimeSeconds` is the number of seconds it takes for the submission to be evaluated\". To be evaluated means just the scoring time? Or that's the total notebook runtime, i.e. fitting+scoring?\n\nThanks you in advance for the information."
    },
    {
      "id": 2231897,
      "postDate": "2023-04-23T19:39:40.387Z",
      "content": "<p>Hi, my first competition! Excited to contribute.</p>",
      "rawMarkdown": "Hi, my first competition! Excited to contribute."
    },
    {
      "id": 2218139,
      "postDate": "2023-04-11T13:13:48.253Z",
      "content": "<p>this is helpful.</p>",
      "rawMarkdown": "this is helpful."
    },
    {
      "id": 2177038,
      "postDate": "2023-03-11T06:32:45.680Z",
      "content": "<p>thanks for hosting this competition, looks to be fun, will try the dataset.</p>",
      "rawMarkdown": "thanks for hosting this competition, looks to be fun, will try the dataset."
    },
    {
      "id": 2160818,
      "postDate": "2023-02-27T02:35:39.463Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> Thanks for this competition.<br>\nCould you tell us how much sessions will be use in the private LB? We only know 50% used in Public and 50% used in the Private. What's the Specific figures?</p>",
      "rawMarkdown": "Hello, @nrambis Thanks for this competition.\nCould you tell us how much sessions will be use in the private LB? We only know 50% used in Public and 50% used in the Private. What's the Specific figures?",
      "replies": [
        {
          "id": 2161259,
          "postDate": "2023-02-27T11:02:41.270Z",
          "content": "<p>Hi, the data section says:</p>\n<blockquote>\n  <p>Note that the hidden test set is roughly as large as the training set; you should expect it will take much longer to run on than the three test samples provided.</p>\n</blockquote>\n<p>Hence, I'd expect the number of sessions in the train and test set to be about the same.</p>",
          "rawMarkdown": "Hi, the data section says:\n\n>Note that the hidden test set is roughly as large as the training set; you should expect it will take much longer to run on than the three test samples provided.\n\nHence, I'd expect the number of sessions in the train and test set to be about the same.",
          "votes": 1,
          "replies": [
            {
              "id": 2162403,
              "postDate": "2023-02-28T07:22:31.283Z",
              "content": "<p>Thanks so much. Now I know how complex my model could be.</p>",
              "rawMarkdown": "Thanks so much. Now I know how complex my model could be.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2156769,
      "postDate": "2023-02-23T14:45:07.560Z",
      "content": "<p>Hi, </p>\n<p>Is the column \"session_level\" included in the test environment's sample_submission-object? My recent submission's seem to fail because of the wrong format in \"sample_submission.csv\". The first submissions got scored successfully.</p>",
      "rawMarkdown": "Hi, \n\nIs the column \"session_level\" included in the test environment's sample_submission-object? My recent submission's seem to fail because of the wrong format in \"sample_submission.csv\". The first submissions got scored successfully.",
      "replies": [
        {
          "id": 2158958,
          "postDate": "2023-02-25T10:07:29.007Z",
          "content": "<p>In my case, the problem was that the values in the \"correct\" column were \"numpy.int64\", when they need to be \"int\". Converting the types solved the problem and the submission went through.</p>",
          "rawMarkdown": "In my case, the problem was that the values in the \"correct\" column were \"numpy.int64\", when they need to be \"int\". Converting the types solved the problem and the submission went through."
        }
      ]
    },
    {
      "id": 2142702,
      "postDate": "2023-02-13T18:09:46.450Z",
      "content": "<p>Hi. The package isn't configured correctly to load in Reticulate<br>\nThe file seems to be compiled in Python 3.7 and seems to be compatible with r-miniconda python3.8<br>\nSee here: <a href=\"https://www.kaggle.com/nigelhenry/jowilder-in-r\" target=\"_blank\">https://www.kaggle.com/nigelhenry/jowilder-in-r</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F422095%2F85c09660a12f1ab00d5cf86161caa96d%2FScreenshot%202023-02-13%20130833.jpg?generation=1676311774076039&amp;alt=media\" alt=\"error message\"></p>",
      "rawMarkdown": "Hi. The package isn't configured correctly to load in Reticulate\nThe file seems to be compiled in Python 3.7 and seems to be compatible with r-miniconda python3.8\nSee here: https://www.kaggle.com/nigelhenry/jowilder-in-r\n\n![error message](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F422095%2F85c09660a12f1ab00d5cf86161caa96d%2FScreenshot%202023-02-13%20130833.jpg?generation=1676311774076039&alt=media)\n\n",
      "replies": [
        {
          "id": 2160422,
          "postDate": "2023-02-26T16:58:07.357Z",
          "content": "<p>it would be great to hear from the kaggle team. If you no longer plan to support access to the time series API using the Recticulate package , you can simply let us know </p>",
          "rawMarkdown": "it would be great to hear from the kaggle team. If you no longer plan to support access to the time series API using the Recticulate package , you can simply let us know "
        }
      ]
    },
    {
      "id": 2148102,
      "postDate": "2023-02-17T06:21:32.730Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2137533,
      "postDate": "2023-02-10T05:39:32.173Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2163767,
      "postDate": "2023-03-01T04:16:20.993Z",
      "content": "<p>Thanks for hosting competition.</p>",
      "rawMarkdown": "Thanks for hosting competition."
    }
  ],
  "comments": [
    {
      "id": 2137042,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-02-09T17:35:20.587000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> Thanks for this fun competition.</p>\n<p>Regarding the dataset and what we are predicting, what is the definition of \"correct answer\" and \"incorrrect answer\". During gameplay, it seems that we need to get every answer correct before continuing. So is \"correct answer\" mean \"get question correct on first try\"?</p>",
      "votes": 35,
      "replies": [
        {
          "id": 2137152,
          "author_name": "Alex Franklin",
          "author_url": "",
          "post_date": "2023-02-09T18:58:31.833000",
          "content": "<p>That's right, during gameplay the player must eventually select the correct answer to continue. A \"correct answer\" here indicates that <strong>the player got the answer correct on their first attempt</strong>. </p>",
          "votes": 48,
          "replies": []
        }
      ]
    },
    {
      "id": 2137727,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2023-02-10T09:17:06.233000",
      "content": "<p>Hi,  I have one question. </p>\n<p>In the Dataset Description, it is said:</p>\n<pre><code>At each checkpoint, you will have access to all previous test data for that section.\n</code></pre>\n<p>But in the time series api, when it comes to level_group 5-12, the test dataframe only contains the historical actions in level_group 5-12.  I guess we can get the level_group 0-4 at that time. So we may need to save the info of level_group 0-4 on my own during the iteration of time series api?</p>\n<p>Thanks in advanced.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2186858,
          "author_name": "LeLeCHAA",
          "author_url": "",
          "post_date": "2023-03-18T07:13:15.463000",
          "content": "<p>Same question，have you got answer？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2190269,
          "author_name": "Xứng BKA-Mi1",
          "author_url": "",
          "post_date": "2023-03-21T06:38:06.830000",
          "content": "<p>Same question</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2192992,
          "author_name": "Feiyi Dong",
          "author_url": "",
          "post_date": "2023-03-23T03:17:46",
          "content": "<p>Same question, can someone help?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2196577,
          "author_name": "ADAM.",
          "author_url": "",
          "post_date": "2023-03-25T14:03:39.410000",
          "content": "<p>IMHO, we can save previous level_group information on my own and use them when the next level_group comes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2139879,
      "author_name": "busybee",
      "author_url": "",
      "post_date": "2023-02-11T08:54:40.130000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> , thanks for hosting this competition. Expanding on <a href=\"https://www.kaggle.com/chrisqiu\" target=\"_blank\">@chrisqiu</a> question, there are also instances where the player restarted the game mid-way and re-do the questions from scratch. For example, session_id = 21100415142476300 has 8 checkpoints:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3123781%2Fbca062908f5f5597eb9f525145dce459%2FScreenshot%202023-02-11%20165320.png?generation=1676105618167188&amp;alt=media\" alt=\"\"></p>\n<p>In this case, are the answers recorded for the player's first or last attempt? </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2147773,
          "author_name": "Alex Franklin",
          "author_url": "",
          "post_date": "2023-02-16T21:42:28.997000",
          "content": "<p>A player should not be able to navigate to previous parts of the game in the same session, but it appears that the game may sometimes not recognize a restart as a new session. These sessions may not provide as useful information as typical sessions with 3 checkpoints, but I believe this should only happen for a small portion of the data.</p>\n<p>To answer your question, in these cases the answers would be recorded for all of the attempts. In other words, the session would record a correct answer only if the player selected the correct option first every single time they answered the question.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2149342,
              "author_name": "ln",
              "author_url": "",
              "post_date": "2023-02-18T07:25:01.203000",
              "content": "<blockquote>\n  <p>in these cases the answers would be recorded for all of the attempts. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> , sorry, I didn't fully understand it.  The session 20100110332615344 has 5 checkpoints for the level groups 0-4,5-12,13-22,0-4,5-12. Does the row 20100110332615344_q1 correspond to the first or the second attempt to answer the questions of level group 0-4?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2152213,
              "author_name": "Alex Franklin",
              "author_url": "",
              "post_date": "2023-02-20T16:42:50.687000",
              "content": "<p>20100110332615344_q1 would have a value of 1 if the player answered question 1 correctly on their first attempt in <em>both</em> the first and second instance they encounter it. If <em>any</em> attempt was incorrect then 20100110332615344_q1 would have a value of 0. Does that help to clarify?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2152951,
              "author_name": "ln",
              "author_url": "",
              "post_date": "2023-02-21T05:37:04.847000",
              "content": "<p>Yes, it's clear now. Thank you.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2171818,
              "author_name": "Prateek",
              "author_url": "",
              "post_date": "2023-03-07T05:43:21.937000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a> So if I am not wrong, each session_id should have 3 checkpoints in the ideal scenario and a total of 18 questions. But I see that there are 80 session ids where there are more than 3 checkpoints, basically seems like the user is able to go to the previous part and there are 30 session ids where there are exactly 6 checkpoints. My assumption is user is giving the test once just randomly in the first attempt, then going back and giving the correct answers. I say this because for session id with just 3 checkpoints 71% of the answers are correct, but where there are 6 checkpoints, just 51.85%, just like random.<br>\nI think it will be good idea to just drop such cases as they are less than 1%. Also, I would like to know if there is a way to identify how many questions are there in which level group?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2174156,
              "author_name": "Alex Franklin",
              "author_url": "",
              "post_date": "2023-03-08T22:22:51.167000",
              "content": "<p>Yup that's right, a single playthrough of the game from beginning to end will include 3 checkpoints. To answer your question about level groups, there are 3, 10, and 5 questions in level groups 0-4, 5-12, and 13-22 respectively.</p>\n<p>I'll point to this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796\" target=\"_blank\">helpful post</a> by pjmathematician that lists the questions by group.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2200037,
              "author_name": "ln",
              "author_url": "",
              "post_date": "2023-03-28T08:29:36.493000",
              "content": "<p><a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>, can I ask one more question concerning the above example with s=20100110332615344 and 5 groups 0-4,5-12,13-22,0-4*,5-12* (asterisk denotes the second occurrence of the group). I can't imagine how env.iter_test() generates such sessions. If the order is (s,0-4 + 0-4*), (s,5-12 + 5-12*),(s,13-22) then we have time leak. If the order is chronological  (s,0-4),(s,5-12),(s,13-22),(s,0-4*),(s,5-12*) then we make predictions for questions 1-13 twice, which is also strange since there are no two indexes 20100110332615344_q1 in the sample_submission I think.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2137318,
      "author_name": "T88",
      "author_url": "",
      "post_date": "2023-02-09T21:10:25.423000",
      "content": "<p>I have a question about the order of the data output by the time series api.  </p>\n<p>Is it correct to assume that the data is provided in the correct order of \"0-4\", \"5-12\", and \"13-22\" in level_group for all sessions?<br>\nOr is it better to assume that omissions or shuffling of order may occur?</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2264174,
      "author_name": "zhongjiajun",
      "author_url": "",
      "post_date": "2023-05-18T08:20:46.177000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> . Thanks for hosting this competition. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2142031,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-02-13T09:42:44.960000",
      "content": "<p>Hello, is it possible to not count failed submissions towards submission attempts? I'm trying to debug why my submissions fail, while the initial commit passes without any problems and it is very frustrating due to each failed notebook accounts towards submission allowance while no detailed information on where the exception happened or what caused it. I've lost two days worth of submits and still have no idea of what's causing the exception.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2156765,
          "author_name": "anaaak",
          "author_url": "",
          "post_date": "2023-02-23T14:41:46.123000",
          "content": "<p>Also my recent submissions fail. Could it be that the column \"session_level\" is not included in the test environments sample_submission, and that's why the output file is wrong format. The column is included in the competitions \"data\" tab, in \"sample_submission.csv\".</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2158636,
              "author_name": "Aaliyah Ali-Khan",
              "author_url": "",
              "post_date": "2023-02-25T03:08:43.693000",
              "content": "<p>Thanks for the help!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2139816,
      "author_name": "__ChrisQ__",
      "author_url": "",
      "post_date": "2023-02-11T07:30:37.290000",
      "content": "<p>Hi, thanks for hosting the competition. I have two questions.</p>\n<ol>\n<li><p>I noticed that one session in the training set has only two checkpoints. session_id: <code>22090108192456930</code>. <br>\nFirst I think this user might quit the game before answering all the questions, but actually it is the first checkpoint that is missing (fqid: chap1_finale_c). And in labels the same user answers all the questions. <br>\nCan I assume there is a missing row for this user? <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2F5741cf69c7cb27d7009f33ca913ecacc%2F1.png?generation=1676100510691203&amp;alt=media\" alt=\"\"></p></li>\n<li><p>The second event of session '20090312431273200' has 1323 elapsed_time, but the third event has 831. This seems wrong because by inspecting the text I know these events are in order. So, is this because the user clicks too fast, or is it a wrong value?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2Fd9fdb7719ea5a35c16ddcad71f255d0c%2F2.png?generation=1676100576722265&amp;alt=media\" alt=\"\"></p></li>\n</ol>\n<p>Thank you very much.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2147768,
          "author_name": "Alex Franklin",
          "author_url": "",
          "post_date": "2023-02-16T21:34:55.863000",
          "content": "<ol>\n<li><p>Yes, you can assume that there is a missing row for that user.</p></li>\n<li><p>The game's timestamps can sometimes be imprecise, which may be partly due to clicks happening quickly in succession. You can assume the <code>index</code> gives the correct order of events in cases where the <code>elapsed_time</code> is inconsistent.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2315242,
      "author_name": "Qurious",
      "author_url": "",
      "post_date": "2023-06-24T00:14:24.037000",
      "content": "<p>I'm reaching out to draw your attention to a matter that has been causing me some concerns for the past five days. I initially reached out to you via email three days ago but have yet to receive a response.</p>\n<p>Understanding your busy schedule, I thought it prudent to follow up on this platform, given your invitation in this post encouraging participants to ask questions.</p>\n<p>Could you please assist by addressing the issue at your earliest convenience? Your guidance would be immensely appreciated.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2308330,
      "author_name": "cli3",
      "author_url": "",
      "post_date": "2023-06-18T20:04:43.570000",
      "content": "<p>\". The models, both traditional and efficient, will be fully open-sourced\"; this is gold.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2301455,
      "author_name": "Oleksiy Kononenko",
      "author_url": "",
      "post_date": "2023-06-14T00:31:34.760000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> and <a href=\"https://www.kaggle.com/alexmlfranklin\" target=\"_blank\">@alexmlfranklin</a>, could you please clarify meaning of several parameters mentioned on the <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/efficiency-prize-evaluation\" target=\"_blank\">efficiency evaluation page</a>:</p>\n<ol>\n<li><p>\"<code>Benchmark</code>  is the score of the benchmark <code>sample_submission.csv</code>\". What is the <code>sample_submission.csv</code> and where to get its score?</p></li>\n<li><p>\"<code>RuntimeSeconds</code> is the number of seconds it takes for the submission to be evaluated\". To be evaluated means just the scoring time? Or that's the total notebook runtime, i.e. fitting+scoring?</p></li>\n</ol>\n<p>Thanks you in advance for the information.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2231897,
      "author_name": "Daniel  Evenschor",
      "author_url": "",
      "post_date": "2023-04-23T19:39:40.387000",
      "content": "<p>Hi, my first competition! Excited to contribute.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2218139,
      "author_name": "Byungeun Hwang",
      "author_url": "",
      "post_date": "2023-04-11T13:13:48.253000",
      "content": "<p>this is helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2177038,
      "author_name": "Sharan R H ",
      "author_url": "",
      "post_date": "2023-03-11T06:32:45.680000",
      "content": "<p>thanks for hosting this competition, looks to be fun, will try the dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2160818,
      "author_name": "biubiubiu~",
      "author_url": "",
      "post_date": "2023-02-27T02:35:39.463000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/nrambis\" target=\"_blank\">@nrambis</a> Thanks for this competition.<br>\nCould you tell us how much sessions will be use in the private LB? We only know 50% used in Public and 50% used in the Private. What's the Specific figures?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2161259,
          "author_name": "heiligerl",
          "author_url": "",
          "post_date": "2023-02-27T11:02:41.270000",
          "content": "<p>Hi, the data section says:</p>\n<blockquote>\n  <p>Note that the hidden test set is roughly as large as the training set; you should expect it will take much longer to run on than the three test samples provided.</p>\n</blockquote>\n<p>Hence, I'd expect the number of sessions in the train and test set to be about the same.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2162403,
              "author_name": "biubiubiu~",
              "author_url": "",
              "post_date": "2023-02-28T07:22:31.283000",
              "content": "<p>Thanks so much. Now I know how complex my model could be.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2156769,
      "author_name": "anaaak",
      "author_url": "",
      "post_date": "2023-02-23T14:45:07.560000",
      "content": "<p>Hi, </p>\n<p>Is the column \"session_level\" included in the test environment's sample_submission-object? My recent submission's seem to fail because of the wrong format in \"sample_submission.csv\". The first submissions got scored successfully.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2158958,
          "author_name": "anaaak",
          "author_url": "",
          "post_date": "2023-02-25T10:07:29.007000",
          "content": "<p>In my case, the problem was that the values in the \"correct\" column were \"numpy.int64\", when they need to be \"int\". Converting the types solved the problem and the submission went through.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2142702,
      "author_name": "Nigel A. R. Henry",
      "author_url": "",
      "post_date": "2023-02-13T18:09:46.450000",
      "content": "<p>Hi. The package isn't configured correctly to load in Reticulate<br>\nThe file seems to be compiled in Python 3.7 and seems to be compatible with r-miniconda python3.8<br>\nSee here: <a href=\"https://www.kaggle.com/nigelhenry/jowilder-in-r\" target=\"_blank\">https://www.kaggle.com/nigelhenry/jowilder-in-r</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F422095%2F85c09660a12f1ab00d5cf86161caa96d%2FScreenshot%202023-02-13%20130833.jpg?generation=1676311774076039&amp;alt=media\" alt=\"error message\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2160422,
          "author_name": "Nigel A. R. Henry",
          "author_url": "",
          "post_date": "2023-02-26T16:58:07.357000",
          "content": "<p>it would be great to hear from the kaggle team. If you no longer plan to support access to the time series API using the Recticulate package , you can simply let us know </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2148102,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-17T06:21:32.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2137533,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-10T05:39:32.173000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2163767,
      "author_name": "K Pradyumna",
      "author_url": "",
      "post_date": "2023-03-01T04:16:20.993000",
      "content": "<p>Thanks for hosting competition.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2136849": "On behalf of Field Day Lab at University of Wisconsin - Madison, Vanderbilt University, and The Learning Agency Lab, welcome to Student Performance and Game Play: Predict student learning from Jo Wilder online educational game! We are thrilled to host this competition which tasks competitors with predicting assessment performance using data from Field Day’s online educational game Jo Wilder. \n\nGame-based learning allows for students to engage with educational content in a dynamic way that the typical classroom lesson doesn’t offer. Current research indicates that game-based learning could be effective in supporting learning outcomes, however few game-based learning  datasets makes it difficult to extend this research. Developing models that make real-time predictions about student performance on game objectives will support educational game developers to create more effective learning experiences for students. \n\nThis competition provides access to data taken from [Jo Wilder and the Capitol Case](https://pbswisconsineducation.org/jowilder/about/) (Jo Wilder) which is a free online point & click game. Designed by Field Day Lab, Jo Wilder aligns with the Wisconsin Department of Public Instruction’s newly revised 3-5th grade Wisconsin Standards for Social Studies. The game utilizes historical analysis to encourage students to develop reading comprehension skills. Sections of the game assess the player’s comprehension of the story and the clues picked up along the way, which are used to create the measure of performance used in this competition. \n\nThis competition will offer two tracks: a traditional accuracy track and a computationally efficient track. The computationally efficient track allows for environmentally friendly models that are easily adaptable in low-resource environments. The models, both traditional and efficient, will be fully open-sourced. \n\nPlease don’t hesitate to reach out with questions regarding the data or competition.\n\nWe hope everyone enjoys this competition, and we wish you all good luck!\n\nHosts:\nDavid Gagnon (davidgagnon), Field Day Lab at University of Wisconsin - Madison\nDr. Scott Crossley (cookiecutters), Vanderbilt University\nUlrich Boser (ulrichboser), The Learning Agency Lab\nMeg Benner (megbenner), The Learning Agency Lab\nPerpetual Baffour (pbaffour), The Learning Agency Lab\nAlex Franklin (alexmlfranklin), The Learning Agency Lab\nNatalie Rambis (nrambis), The Learning Agency Lab\n",
    "2137042": "Hi @nrambis Thanks for this fun competition.\n\nRegarding the dataset and what we are predicting, what is the definition of \"correct answer\" and \"incorrrect answer\". During gameplay, it seems that we need to get every answer correct before continuing. So is \"correct answer\" mean \"get question correct on first try\"?",
    "2137727": "Hi,  I have one question. \n\nIn the Dataset Description, it is said:\n```\nAt each checkpoint, you will have access to all previous test data for that section.\n```\n\nBut in the time series api, when it comes to level_group 5-12, the test dataframe only contains the historical actions in level_group 5-12.  I guess we can get the level_group 0-4 at that time. So we may need to save the info of level_group 0-4 on my own during the iteration of time series api?\n\nThanks in advanced.",
    "2139879": "Hi @nrambis , thanks for hosting this competition. Expanding on @chrisqiu question, there are also instances where the player restarted the game mid-way and re-do the questions from scratch. For example, session_id = 21100415142476300 has 8 checkpoints:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3123781%2Fbca062908f5f5597eb9f525145dce459%2FScreenshot%202023-02-11%20165320.png?generation=1676105618167188&alt=media)\n\nIn this case, are the answers recorded for the player's first or last attempt? ",
    "2137318": "I have a question about the order of the data output by the time series api.  \n\nIs it correct to assume that the data is provided in the correct order of \"0-4\", \"5-12\", and \"13-22\" in level_group for all sessions?\nOr is it better to assume that omissions or shuffling of order may occur?",
    "2264174": "Hi, @nrambis . Thanks for hosting this competition. ",
    "2142031": "Hello, is it possible to not count failed submissions towards submission attempts? I'm trying to debug why my submissions fail, while the initial commit passes without any problems and it is very frustrating due to each failed notebook accounts towards submission allowance while no detailed information on where the exception happened or what caused it. I've lost two days worth of submits and still have no idea of what's causing the exception.",
    "2139816": "Hi, thanks for hosting the competition. I have two questions.\n\n1. I noticed that one session in the training set has only two checkpoints. session_id: `22090108192456930`. \nFirst I think this user might quit the game before answering all the questions, but actually it is the first checkpoint that is missing (fqid: chap1_finale_c). And in labels the same user answers all the questions. \nCan I assume there is a missing row for this user? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2F5741cf69c7cb27d7009f33ca913ecacc%2F1.png?generation=1676100510691203&alt=media)\n\n2. The second event of session '20090312431273200' has 1323 elapsed_time, but the third event has 831. This seems wrong because by inspecting the text I know these events are in order. So, is this because the user clicks too fast, or is it a wrong value?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4835451%2Fd9fdb7719ea5a35c16ddcad71f255d0c%2F2.png?generation=1676100576722265&alt=media)\n\n\nThank you very much.",
    "2315242": "I'm reaching out to draw your attention to a matter that has been causing me some concerns for the past five days. I initially reached out to you via email three days ago but have yet to receive a response.\n\nUnderstanding your busy schedule, I thought it prudent to follow up on this platform, given your invitation in this post encouraging participants to ask questions.\n\nCould you please assist by addressing the issue at your earliest convenience? Your guidance would be immensely appreciated.\n",
    "2308330": "\". The models, both traditional and efficient, will be fully open-sourced\"; this is gold.",
    "2301455": "Dear @nrambis and @alexmlfranklin, could you please clarify meaning of several parameters mentioned on the [efficiency evaluation page](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/efficiency-prize-evaluation):\n\n1. \"`Benchmark`  is the score of the benchmark `sample_submission.csv`\". What is the `sample_submission.csv` and where to get its score?\n\n2. \"`RuntimeSeconds` is the number of seconds it takes for the submission to be evaluated\". To be evaluated means just the scoring time? Or that's the total notebook runtime, i.e. fitting+scoring?\n\nThanks you in advance for the information.",
    "2231897": "Hi, my first competition! Excited to contribute.",
    "2218139": "this is helpful.",
    "2177038": "thanks for hosting this competition, looks to be fun, will try the dataset.",
    "2160818": "Hello, @nrambis Thanks for this competition.\nCould you tell us how much sessions will be use in the private LB? We only know 50% used in Public and 50% used in the Private. What's the Specific figures?",
    "2156769": "Hi, \n\nIs the column \"session_level\" included in the test environment's sample_submission-object? My recent submission's seem to fail because of the wrong format in \"sample_submission.csv\". The first submissions got scored successfully.",
    "2142702": "Hi. The package isn't configured correctly to load in Reticulate\nThe file seems to be compiled in Python 3.7 and seems to be compatible with r-miniconda python3.8\nSee here: https://www.kaggle.com/nigelhenry/jowilder-in-r\n\n![error message](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F422095%2F85c09660a12f1ab00d5cf86161caa96d%2FScreenshot%202023-02-13%20130833.jpg?generation=1676311774076039&alt=media)\n\n",
    "2148102": "",
    "2137533": "",
    "2163767": "Thanks for hosting competition."
  }
}