{
  "id": 388243,
  "title": "Can someone explain about why we are training one model for each question and how people came down to the following results:",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388243",
  "author_name": "",
  "post_date": "2023-02-16T15:22:22.713273500Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We train one model for each of 18 questions. Furthermore, we use data from level_groups = '0-4' to train model for questions 1-3, and level groups '5-12' to train questions 4 thru 13 and level groups '13-22' to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier level_groups and using that to predict future level_groups.</p>",
  "messages": [
    {
      "id": "2147377",
      "postDate": "02/16/2023 15:22:22",
      "content": "<p>We train one model for each of 18 questions. Furthermore, we use data from level_groups = '0-4' to train model for questions 1-3, and level groups '5-12' to train questions 4 thru 13 and level groups '13-22' to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier level_groups and using that to predict future level_groups.</p>",
      "rawMarkdown": "We train one model for each of 18 questions. Furthermore, we use data from level_groups = '0-4' to train model for questions 1-3, and level groups '5-12' to train questions 4 thru 13 and level groups '13-22' to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier level_groups and using that to predict future level_groups.",
      "votes": null
    },
    {
      "id": "2147917",
      "postDate": "02/17/2023 02:30:20",
      "content": "<p>I believe this comes from Chris's notebook here: <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook</a></p>\n<blockquote>\n  <p>Can someone explain about why we are training one model for each question</p>\n</blockquote>\n<p>We are supposed to predict whether a session answers correctly <strong>each</strong> of the 18 questions, meaning each session has 18 different target labels. Because in Chris's notebook he was using <code>XGBClassifier</code> which has only one output per model, he needs to build 18, one for each question.</p>\n<blockquote>\n  <p>how people came down to the following results</p>\n</blockquote>\n<p>As mentioned in the <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data\" target=\"_blank\">competition's description</a>, the test data will be fetched by <code>level_groups</code> (0-4, 5-12, 13-22) and contain all data from that <code>level_groups</code>. Each level group has a number of questions (0-4 has question 1 to 3, 5-12 has 4 to 13, and 13-22 has 14 through 18). Hence, Chris used the mentioned setup in training so it follows the same process in inference.</p>",
      "rawMarkdown": "I believe this comes from Chris's notebook here: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook\n\n>Can someone explain about why we are training one model for each question\n\nWe are supposed to predict whether a session answers correctly **each** of the 18 questions, meaning each session has 18 different target labels. Because in Chris's notebook he was using `XGBClassifier` which has only one output per model, he needs to build 18, one for each question.\n\n>how people came down to the following results\n\nAs mentioned in the [competition's description](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data), the test data will be fetched by `level_groups` (0-4, 5-12, 13-22) and contain all data from that `level_groups`. Each level group has a number of questions (0-4 has question 1 to 3, 5-12 has 4 to 13, and 13-22 has 14 through 18). Hence, Chris used the mentioned setup in training so it follows the same process in inference.",
      "votes": null
    },
    {
      "id": "2147981",
      "postDate": "02/17/2023 04:24:40",
      "content": "<p>Can you please share where it is written that level_group 0-4 has questions 1-3</p>",
      "rawMarkdown": "Can you please share where it is written that level_group 0-4 has questions 1-3",
      "votes": null
    },
    {
      "id": "2147984",
      "postDate": "02/17/2023 04:28:15",
      "content": "<p>doesn't level_group 0-4 mean that it includes level 0 to 4, how did you map it to the question number, I couldn't understand it through the description</p>",
      "rawMarkdown": "doesn't level_group 0-4 mean that it includes level 0 to 4, how did you map it to the question number, I couldn't understand it through the description",
      "votes": null
    },
    {
      "id": "2147991",
      "postDate": "02/17/2023 04:33:28",
      "content": "<p>level_group - which group of levels - and group of questions - this row belongs to (0-4, 5-12, 13-22)</p>\n<p>how do you infer the group of questions for each level group from here</p>",
      "rawMarkdown": "level_group - which group of levels - and group of questions - this row belongs to (0-4, 5-12, 13-22)\n\n\nhow do you infer the group of questions for each level group from here",
      "votes": null
    },
    {
      "id": "2148009",
      "postDate": "02/17/2023 04:53:13",
      "content": "<p>Now that I check again I couldn't find where in the competition doc it is stated explicitly.</p>\n<p>However you can still find it implied in the public test data. Specifically, when you iterate over the test api, it will spit out a sample_submission.csv and a test dataset <strong>for a group level</strong>. If you inspect the sample_submission.csv, you'll see that it contains sample answers for question 1 thru 3 if the group level is \"0-4\", 4 thru 13 for \"5-12\", and 14 to 18 for \"13-22\".</p>\n<p>If you're wondering why the level number (0-4) does not match the question index (1-3), \"level\" here simply means \"stage\" of the game, and after X stages you answer Y number of questions, with X and Y not having to be equal.</p>",
      "rawMarkdown": "Now that I check again I couldn't find where in the competition doc it is stated explicitly.\n\nHowever you can still find it implied in the public test data. Specifically, when you iterate over the test api, it will spit out a sample_submission.csv and a test dataset **for a group level**. If you inspect the sample_submission.csv, you'll see that it contains sample answers for question 1 thru 3 if the group level is \"0-4\", 4 thru 13 for \"5-12\", and 14 to 18 for \"13-22\".\n\nIf you're wondering why the level number (0-4) does not match the question index (1-3), \"level\" here simply means \"stage\" of the game, and after X stages you answer Y number of questions, with X and Y not having to be equal.",
      "votes": null
    },
    {
      "id": "2148108",
      "postDate": "02/17/2023 06:24:37",
      "content": "<p>Thank you so much, you actually cleared my doubts</p>",
      "rawMarkdown": "Thank you so much, you actually cleared my doubts",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2147917,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "02/17/2023 02:30:20",
      "content": "<p>I believe this comes from Chris's notebook here: <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook</a></p>\n<blockquote>\n  <p>Can someone explain about why we are training one model for each question</p>\n</blockquote>\n<p>We are supposed to predict whether a session answers correctly <strong>each</strong> of the 18 questions, meaning each session has 18 different target labels. Because in Chris's notebook he was using <code>XGBClassifier</code> which has only one output per model, he needs to build 18, one for each question.</p>\n<blockquote>\n  <p>how people came down to the following results</p>\n</blockquote>\n<p>As mentioned in the <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data\" target=\"_blank\">competition's description</a>, the test data will be fetched by <code>level_groups</code> (0-4, 5-12, 13-22) and contain all data from that <code>level_groups</code>. Each level group has a number of questions (0-4 has question 1 to 3, 5-12 has 4 to 13, and 13-22 has 14 through 18). Hence, Chris used the mentioned setup in training so it follows the same process in inference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2147981,
          "author_name": "ayushsengar999",
          "author_url": "",
          "post_date": "02/17/2023 04:24:40",
          "content": "<p>Can you please share where it is written that level_group 0-4 has questions 1-3</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2147984,
          "author_name": "ayushsengar999",
          "author_url": "",
          "post_date": "02/17/2023 04:28:15",
          "content": "<p>doesn't level_group 0-4 mean that it includes level 0 to 4, how did you map it to the question number, I couldn't understand it through the description</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2147991,
          "author_name": "ayushsengar999",
          "author_url": "",
          "post_date": "02/17/2023 04:33:28",
          "content": "<p>level_group - which group of levels - and group of questions - this row belongs to (0-4, 5-12, 13-22)</p>\n<p>how do you infer the group of questions for each level group from here</p>",
          "votes": null,
          "replies": [
            {
              "id": 2148009,
              "author_name": "hoangnguyen719",
              "author_url": "",
              "post_date": "02/17/2023 04:53:13",
              "content": "<p>Now that I check again I couldn't find where in the competition doc it is stated explicitly.</p>\n<p>However you can still find it implied in the public test data. Specifically, when you iterate over the test api, it will spit out a sample_submission.csv and a test dataset <strong>for a group level</strong>. If you inspect the sample_submission.csv, you'll see that it contains sample answers for question 1 thru 3 if the group level is \"0-4\", 4 thru 13 for \"5-12\", and 14 to 18 for \"13-22\".</p>\n<p>If you're wondering why the level number (0-4) does not match the question index (1-3), \"level\" here simply means \"stage\" of the game, and after X stages you answer Y number of questions, with X and Y not having to be equal.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2148108,
                  "author_name": "ayushsengar999",
                  "author_url": "",
                  "post_date": "02/17/2023 06:24:37",
                  "content": "<p>Thank you so much, you actually cleared my doubts</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2147377": "We train one model for each of 18 questions. Furthermore, we use data from level_groups = '0-4' to train model for questions 1-3, and level groups '5-12' to train questions 4 thru 13 and level groups '13-22' to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier level_groups and using that to predict future level_groups.",
    "2147917": "I believe this comes from Chris's notebook here: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook\n\n>Can someone explain about why we are training one model for each question\n\nWe are supposed to predict whether a session answers correctly **each** of the 18 questions, meaning each session has 18 different target labels. Because in Chris's notebook he was using `XGBClassifier` which has only one output per model, he needs to build 18, one for each question.\n\n>how people came down to the following results\n\nAs mentioned in the [competition's description](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data), the test data will be fetched by `level_groups` (0-4, 5-12, 13-22) and contain all data from that `level_groups`. Each level group has a number of questions (0-4 has question 1 to 3, 5-12 has 4 to 13, and 13-22 has 14 through 18). Hence, Chris used the mentioned setup in training so it follows the same process in inference.",
    "2147981": "Can you please share where it is written that level_group 0-4 has questions 1-3",
    "2147984": "doesn't level_group 0-4 mean that it includes level 0 to 4, how did you map it to the question number, I couldn't understand it through the description",
    "2147991": "level_group - which group of levels - and group of questions - this row belongs to (0-4, 5-12, 13-22)\n\n\nhow do you infer the group of questions for each level group from here",
    "2148009": "Now that I check again I couldn't find where in the competition doc it is stated explicitly.\n\nHowever you can still find it implied in the public test data. Specifically, when you iterate over the test api, it will spit out a sample_submission.csv and a test dataset **for a group level**. If you inspect the sample_submission.csv, you'll see that it contains sample answers for question 1 thru 3 if the group level is \"0-4\", 4 thru 13 for \"5-12\", and 14 to 18 for \"13-22\".\n\nIf you're wondering why the level number (0-4) does not match the question index (1-3), \"level\" here simply means \"stage\" of the game, and after X stages you answer Y number of questions, with X and Y not having to be equal.",
    "2148108": "Thank you so much, you actually cleared my doubts"
  },
  "source": "meta"
}