{
  "id": 424082,
  "title": "98th Place Solution",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/424082",
  "author_name": "ktokunaga",
  "post_date": "2023-07-12T13:27:20.306000",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Despite all the confusion, I would like to express my gratitude to the staff and the Kagglers who shared great information and knowledge. My result is not the top one, but I am posting this in the hope that it will be helpful to someone else and as a record for myself.</p>\n<h2>Overview</h2>\n<p>I created multiple XGboost models and blended their predictions as submission. Public Score was 0.703 - 0.704 and Private Score was 0.698 for all three.</p>\n<h2>Solution</h2>\n<p><a href=\"https://www.kaggle.com/code/pourchot/simple-xgb\" target=\"_blank\">Laurent's Notebook</a> was used as the baseline. The following features were added</p>\n<ul>\n<li>Identifiy conversations and actions that are essential to the progress of the game, and calculate the elapsed time between them and the number of Indexes.</li>\n<li>The number of Indexes and the sum of elapsed time diff using fqid, room, and level as keys, since the same fqid can occur in multiple rooms and levels.</li>\n<li>Number of times a very large elapsed time diff has occurred</li>\n<li>Whether the order of level_group is swapped or not, since there are cases in which the order of level_group is swapped as described in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250\" target=\"_blank\">AbaoJiang's Notebook</a>.</li>\n</ul>\n<p>Data from the previous level_group and predictions for the previous questions were also added as features</p>\n<ul>\n<li>e.g.) For the model with level_group = 5-12, the features created for level_group = 0-4 are added as they are.</li>\n<li>For the model predicting question t, I added the predictions for questions 1, 2, 3, …, t-1 as features.</li>\n</ul>\n<p>In this way, 1200 features were prepared for level_group = 0-4, 3000 for level_group 5-12, and 5500 for level_group = 13-22. From here, feature selection is performed by feature importance in terms of gains. I created multiple XGboost models (like 3 - 8 models), which vary which features to include and how much to reduce the number of features by feature selection. In many models, the number of features are around 800, 1500, 2500, respectively. The weighted average of these models was used as the final prediction. In the inference code, I sorted the test data as described in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/416963\" target=\"_blank\">Daniel's Notebook</a>. In many cases, this sorting improved scores over sorting by index even in GBDT models.</p>\n<p>The addition of the past level_group features and the addition of the past question predictions as features contributed in particular to the scores.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Tuning the hyperparameters for individual models did little to improve the scores</li>\n<li>Adjusting the thresholds for individual questions also did little to improve the scores</li>\n<li>I tried to use additional raw data available on the web, but could not improve scores.</li>\n<li>I tried to convert the coordinates of clicks in a room into a good feature, but could not improve scores.</li>\n<li>I also tried to incorporate the predictions of the previous question in two steps using stacking, but this did little to improve scores, so I adopted the simpler approach described above.</li>\n</ul>\n<h2>Reflection</h2>\n<ul>\n<li>Since I started feature selection, I calculated the CV score in a bad way, resulting in an inappropriate CV score. This caused an over-fitting to Public Score, because I did not know what to trust.</li>\n<li>I took the weighted average of multiple XGboosts as the final predicted value, but the models were so similar that it improved the Public Score but not the Private Score. Blending <a href=\"https://www.kaggle.com/code/vadimkamaev/catboost-new/notebook\" target=\"_blank\">VADIM’s publick notebook</a> and my XGboost model gives private score 0.700 but I could not select this one. I should have increased the diversity of the models instead of being drawn to the Public Score.</li>\n<li>In my submission list, I found a model with a Public Score of 0.695 but a Private Score of 0.704. I think it would have been impossible to choose this as my final submission because it was a fluke result!</li>\n</ul>\n<p>My code can be found in <a href=\"https://github.com/KazuakiTokunaga/kaggle-studentperformance\" target=\"_blank\">this github repository</a>.</p>",
  "messages": [
    {
      "id": 2342056,
      "postDate": "2023-07-12T13:27:20.307Z",
      "content": "<p>Despite all the confusion, I would like to express my gratitude to the staff and the Kagglers who shared great information and knowledge. My result is not the top one, but I am posting this in the hope that it will be helpful to someone else and as a record for myself.</p>\n<h2>Overview</h2>\n<p>I created multiple XGboost models and blended their predictions as submission. Public Score was 0.703 - 0.704 and Private Score was 0.698 for all three.</p>\n<h2>Solution</h2>\n<p><a href=\"https://www.kaggle.com/code/pourchot/simple-xgb\" target=\"_blank\">Laurent's Notebook</a> was used as the baseline. The following features were added</p>\n<ul>\n<li>Identifiy conversations and actions that are essential to the progress of the game, and calculate the elapsed time between them and the number of Indexes.</li>\n<li>The number of Indexes and the sum of elapsed time diff using fqid, room, and level as keys, since the same fqid can occur in multiple rooms and levels.</li>\n<li>Number of times a very large elapsed time diff has occurred</li>\n<li>Whether the order of level_group is swapped or not, since there are cases in which the order of level_group is swapped as described in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250\" target=\"_blank\">AbaoJiang's Notebook</a>.</li>\n</ul>\n<p>Data from the previous level_group and predictions for the previous questions were also added as features</p>\n<ul>\n<li>e.g.) For the model with level_group = 5-12, the features created for level_group = 0-4 are added as they are.</li>\n<li>For the model predicting question t, I added the predictions for questions 1, 2, 3, …, t-1 as features.</li>\n</ul>\n<p>In this way, 1200 features were prepared for level_group = 0-4, 3000 for level_group 5-12, and 5500 for level_group = 13-22. From here, feature selection is performed by feature importance in terms of gains. I created multiple XGboost models (like 3 - 8 models), which vary which features to include and how much to reduce the number of features by feature selection. In many models, the number of features are around 800, 1500, 2500, respectively. The weighted average of these models was used as the final prediction. In the inference code, I sorted the test data as described in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/416963\" target=\"_blank\">Daniel's Notebook</a>. In many cases, this sorting improved scores over sorting by index even in GBDT models.</p>\n<p>The addition of the past level_group features and the addition of the past question predictions as features contributed in particular to the scores.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Tuning the hyperparameters for individual models did little to improve the scores</li>\n<li>Adjusting the thresholds for individual questions also did little to improve the scores</li>\n<li>I tried to use additional raw data available on the web, but could not improve scores.</li>\n<li>I tried to convert the coordinates of clicks in a room into a good feature, but could not improve scores.</li>\n<li>I also tried to incorporate the predictions of the previous question in two steps using stacking, but this did little to improve scores, so I adopted the simpler approach described above.</li>\n</ul>\n<h2>Reflection</h2>\n<ul>\n<li>Since I started feature selection, I calculated the CV score in a bad way, resulting in an inappropriate CV score. This caused an over-fitting to Public Score, because I did not know what to trust.</li>\n<li>I took the weighted average of multiple XGboosts as the final predicted value, but the models were so similar that it improved the Public Score but not the Private Score. Blending <a href=\"https://www.kaggle.com/code/vadimkamaev/catboost-new/notebook\" target=\"_blank\">VADIM’s publick notebook</a> and my XGboost model gives private score 0.700 but I could not select this one. I should have increased the diversity of the models instead of being drawn to the Public Score.</li>\n<li>In my submission list, I found a model with a Public Score of 0.695 but a Private Score of 0.704. I think it would have been impossible to choose this as my final submission because it was a fluke result!</li>\n</ul>\n<p>My code can be found in <a href=\"https://github.com/KazuakiTokunaga/kaggle-studentperformance\" target=\"_blank\">this github repository</a>.</p>",
      "rawMarkdown": "Despite all the confusion, I would like to express my gratitude to the staff and the Kagglers who shared great information and knowledge. My result is not the top one, but I am posting this in the hope that it will be helpful to someone else and as a record for myself.\n\n## Overview\n\nI created multiple XGboost models and blended their predictions as submission. Public Score was 0.703 - 0.704 and Private Score was 0.698 for all three.\n\n## Solution\n\n[Laurent's Notebook](https://www.kaggle.com/code/pourchot/simple-xgb) was used as the baseline. The following features were added\n\n- Identifiy conversations and actions that are essential to the progress of the game, and calculate the elapsed time between them and the number of Indexes.\n- The number of Indexes and the sum of elapsed time diff using fqid, room, and level as keys, since the same fqid can occur in multiple rooms and levels.\n- Number of times a very large elapsed time diff has occurred\n- Whether the order of level_group is swapped or not, since there are cases in which the order of level_group is swapped as described in [AbaoJiang's Notebook](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250).\n\nData from the previous level_group and predictions for the previous questions were also added as features\n\n- e.g.) For the model with level_group = 5-12, the features created for level_group = 0-4 are added as they are.\n- For the model predicting question t, I added the predictions for questions 1, 2, 3, ..., t-1 as features.\n\nIn this way, 1200 features were prepared for level_group = 0-4, 3000 for level_group 5-12, and 5500 for level_group = 13-22. From here, feature selection is performed by feature importance in terms of gains. I created multiple XGboost models (like 3 - 8 models), which vary which features to include and how much to reduce the number of features by feature selection. In many models, the number of features are around 800, 1500, 2500, respectively. The weighted average of these models was used as the final prediction. In the inference code, I sorted the test data as described in [Daniel's Notebook](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/416963). In many cases, this sorting improved scores over sorting by index even in GBDT models.\n\nThe addition of the past level_group features and the addition of the past question predictions as features contributed in particular to the scores.\n\n## What did not work\n\n- Tuning the hyperparameters for individual models did little to improve the scores\n- Adjusting the thresholds for individual questions also did little to improve the scores\n- I tried to use additional raw data available on the web, but could not improve scores.\n- I tried to convert the coordinates of clicks in a room into a good feature, but could not improve scores.\n- I also tried to incorporate the predictions of the previous question in two steps using stacking, but this did little to improve scores, so I adopted the simpler approach described above.\n\n## Reflection\n\n- Since I started feature selection, I calculated the CV score in a bad way, resulting in an inappropriate CV score. This caused an over-fitting to Public Score, because I did not know what to trust.\n- I took the weighted average of multiple XGboosts as the final predicted value, but the models were so similar that it improved the Public Score but not the Private Score. Blending [VADIM’s publick notebook](https://www.kaggle.com/code/vadimkamaev/catboost-new/notebook) and my XGboost model gives private score 0.700 but I could not select this one. I should have increased the diversity of the models instead of being drawn to the Public Score.\n- In my submission list, I found a model with a Public Score of 0.695 but a Private Score of 0.704. I think it would have been impossible to choose this as my final submission because it was a fluke result!\n\nMy code can be found in [this github repository](https://github.com/KazuakiTokunaga/kaggle-studentperformance).",
      "votes": 8
    },
    {
      "id": 2342739,
      "postDate": "2023-07-13T05:41:22.383Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2342739,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-13T05:41:22.383000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2342056": "Despite all the confusion, I would like to express my gratitude to the staff and the Kagglers who shared great information and knowledge. My result is not the top one, but I am posting this in the hope that it will be helpful to someone else and as a record for myself.\n\n## Overview\n\nI created multiple XGboost models and blended their predictions as submission. Public Score was 0.703 - 0.704 and Private Score was 0.698 for all three.\n\n## Solution\n\n[Laurent's Notebook](https://www.kaggle.com/code/pourchot/simple-xgb) was used as the baseline. The following features were added\n\n- Identifiy conversations and actions that are essential to the progress of the game, and calculate the elapsed time between them and the number of Indexes.\n- The number of Indexes and the sum of elapsed time diff using fqid, room, and level as keys, since the same fqid can occur in multiple rooms and levels.\n- Number of times a very large elapsed time diff has occurred\n- Whether the order of level_group is swapped or not, since there are cases in which the order of level_group is swapped as described in [AbaoJiang's Notebook](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/395250).\n\nData from the previous level_group and predictions for the previous questions were also added as features\n\n- e.g.) For the model with level_group = 5-12, the features created for level_group = 0-4 are added as they are.\n- For the model predicting question t, I added the predictions for questions 1, 2, 3, ..., t-1 as features.\n\nIn this way, 1200 features were prepared for level_group = 0-4, 3000 for level_group 5-12, and 5500 for level_group = 13-22. From here, feature selection is performed by feature importance in terms of gains. I created multiple XGboost models (like 3 - 8 models), which vary which features to include and how much to reduce the number of features by feature selection. In many models, the number of features are around 800, 1500, 2500, respectively. The weighted average of these models was used as the final prediction. In the inference code, I sorted the test data as described in [Daniel's Notebook](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/416963). In many cases, this sorting improved scores over sorting by index even in GBDT models.\n\nThe addition of the past level_group features and the addition of the past question predictions as features contributed in particular to the scores.\n\n## What did not work\n\n- Tuning the hyperparameters for individual models did little to improve the scores\n- Adjusting the thresholds for individual questions also did little to improve the scores\n- I tried to use additional raw data available on the web, but could not improve scores.\n- I tried to convert the coordinates of clicks in a room into a good feature, but could not improve scores.\n- I also tried to incorporate the predictions of the previous question in two steps using stacking, but this did little to improve scores, so I adopted the simpler approach described above.\n\n## Reflection\n\n- Since I started feature selection, I calculated the CV score in a bad way, resulting in an inappropriate CV score. This caused an over-fitting to Public Score, because I did not know what to trust.\n- I took the weighted average of multiple XGboosts as the final predicted value, but the models were so similar that it improved the Public Score but not the Private Score. Blending [VADIM’s publick notebook](https://www.kaggle.com/code/vadimkamaev/catboost-new/notebook) and my XGboost model gives private score 0.700 but I could not select this one. I should have increased the diversity of the models instead of being drawn to the Public Score.\n- In my submission list, I found a model with a Public Score of 0.695 but a Private Score of 0.704. I think it would have been impossible to choose this as my final submission because it was a fluke result!\n\nMy code can be found in [this github repository](https://github.com/KazuakiTokunaga/kaggle-studentperformance).",
    "2342739": ""
  }
}