{
  "id": 420169,
  "title": "Solution Writeup and Lessons Learned",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420169",
  "author_name": "",
  "post_date": "2023-06-29T14:13:09.799342700Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>At first, thank you Kaggle and the hosts for hosting this interesting competition. Also, thanks the Kaggle community for all the sharings, discussions and kindness. I would like to share my solutions and the learning journey below (though the final ranking isn't that good 😓).</p>\n<h1>Overview</h1>\n<ul>\n<li>Add data from <a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">fielddaylab</a></li>\n<li>Use straified KFold based on #correct answers of each session</li>\n<li>Implement two-level modeling<ul>\n<li>Level1: The base model is mainly based on my first sharing of the DL model architecture <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565\" target=\"_blank\">here</a>.</li>\n<li>Level2: The stackers (<em>e.g.,</em> xgb, lgbm and catboost) are trained with the oof probas (meta features) of the base model and simple FE features.</li></ul></li>\n</ul>\n<h1>CV Scheme</h1>\n<p>Because the Macro-F1 score is derived on flattened prediction and groundtruth sequences, I think it would be better to somewhat retain the distribution of correctness when doing local CV. However, I find that it doesn't show superiority over the plain <code>GroupKFold</code> on <code>session_id</code>. Finally, I choose stratified KFold (k=5) based on #correct answers with 5 different seeds for stability of evaluation.</p>\n<h1>Modeling Process</h1>\n<p>As shown in the figure below, the modeling process on the highest level consists of two stages, <em>base model fitting</em> and <em>stacker training</em>. The subsequent subsections describe these phases respectively.<br>\n<a href=\"https://postimg.cc/1fK9Vsz5\" target=\"_blank\"><img src=\"https://i.postimg.cc/Kv64VYSL/3.png\" alt=\"3.png\"></a></p>\n<h2>Stage1 - Base Model Fitting</h2>\n<p>In stage1, the main model architecture is just the same as what I shared before. Nonetheless, more categorical features are taken into consideration, including <code>event_comb</code> (<em>i.e.,</em> <code>event_name</code> + <code>name</code>), <code>room_fqid</code>, <code>text_fqid</code>, <code>fqid</code>,  <code>level</code> and <code>page</code> (the corresponding embedding dimensions are shown in the figure). After the categorical embeddings are obtained, I use the difference of <code>elapsed_time</code> to tweak the scale of each event record (one column of the raw feature matrix), which can be thought of as the importance of that snapshot. Then a three-layer temporal convolution architecture is applied to extract the temporal pattern of the event sequence. After skip connection and readout, out-of-fold predictions (probas) are kept as the meta features for the stage2.</p>\n<h2>Stage2 - Stacker Training</h2>\n<p>I still remember on 20th, Jun I thought that it seemed impossible to break 0.697 on local CV with this DL model. Over the past three months, I worked hard to explore and try various of model architecture, but all of them fail to surpass the one as discussed above. Therefore, I decided to try some ensemble methods to see if performance could improve. After some trials, restacking based on the meta features from the level-1 model with some FE features made me nearly touch the score of 0.7 on local side. With further blending of three stackers, I got to pass 0.7 by a small margin. The number of features (meta + FE) used in each <code>level_group</code> is shown as follows:</p>\n<ul>\n<li><code>level_group</code> \"0-4\": 3 + 22</li>\n<li><code>level_group</code> \"5-12\": 13 + 26</li>\n<li><code>level_group</code> \"13-22\": 18 + 30</li>\n</ul>\n<h1>Performance Study</h1>\n<p>Following is the top performance report of my experiments. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Single DL Model</td>\n<td>0.6964</td>\n<td>0.697</td>\n<td>0.697</td>\n</tr>\n<tr>\n<td>Restack w/ Single XGB</td>\n<td>0.6988</td>\n<td>0.699</td>\n<td></td>\n</tr>\n<tr>\n<td>Restack w/ Single CatBoost</td>\n<td>0.6996</td>\n<td>0.698</td>\n<td>0.698</td>\n</tr>\n<tr>\n<td>Restack w/ CatBoost*1 + XGB*1 + LGBM*1</td>\n<td>0.6999</td>\n<td><strong>0.701</strong></td>\n<td>0.695</td>\n</tr>\n<tr>\n<td>Restack w/ CatBoost*1 + XGB*1 + LGBM*1</td>\n<td><strong>0.7000</strong></td>\n<td>0.699</td>\n<td>0.696</td>\n</tr>\n</tbody>\n</table>\n<p>Though the performance indeed improves when restacking plays in, I think that I over-tune my model during the last week of the competition, which leads to overfit and fail to generalize to private test set.</p>\n<h1>What Doesn't Work (on My Side)</h1>\n<ol>\n<li>Retrain base models with all sessions (still find it hard to tune and select the right checkpoints…).</li>\n<li>Use question embedding to do question-wise attention on either raw event sequence or the temporal pattern extracted by temporal convolution architecture.</li>\n<li>Add more numeric features at level-1 input data, such as <code>hover_duration</code> and the track of clicks (<em>i.e.,</em> distance and angle between consecutive clicks).</li>\n<li>Try more temporal pattern extractors (<em>e.g.,</em> Conv3D, <a href=\"https://arxiv.org/abs/2106.09305\" target=\"_blank\">SCINet</a>).</li>\n</ol>\n<h1>Conclusion</h1>\n<p>Although the final rank isn't good, I learn a lot from the Kaggle community and enjoy sharing some of my findings during this competitive journey. I'll keep progressing and try to clarify what I've lost (FE is crucial…). Thanks a lot for your patience!</p>",
  "messages": [
    {
      "id": "2322825",
      "postDate": "06/29/2023 14:13:09",
      "content": "<p>At first, thank you Kaggle and the hosts for hosting this interesting competition. Also, thanks the Kaggle community for all the sharings, discussions and kindness. I would like to share my solutions and the learning journey below (though the final ranking isn't that good 😓).</p>\n<h1>Overview</h1>\n<ul>\n<li>Add data from <a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">fielddaylab</a></li>\n<li>Use straified KFold based on #correct answers of each session</li>\n<li>Implement two-level modeling<ul>\n<li>Level1: The base model is mainly based on my first sharing of the DL model architecture <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565\" target=\"_blank\">here</a>.</li>\n<li>Level2: The stackers (<em>e.g.,</em> xgb, lgbm and catboost) are trained with the oof probas (meta features) of the base model and simple FE features.</li></ul></li>\n</ul>\n<h1>CV Scheme</h1>\n<p>Because the Macro-F1 score is derived on flattened prediction and groundtruth sequences, I think it would be better to somewhat retain the distribution of correctness when doing local CV. However, I find that it doesn't show superiority over the plain <code>GroupKFold</code> on <code>session_id</code>. Finally, I choose stratified KFold (k=5) based on #correct answers with 5 different seeds for stability of evaluation.</p>\n<h1>Modeling Process</h1>\n<p>As shown in the figure below, the modeling process on the highest level consists of two stages, <em>base model fitting</em> and <em>stacker training</em>. The subsequent subsections describe these phases respectively.<br>\n<a href=\"https://postimg.cc/1fK9Vsz5\" target=\"_blank\"><img src=\"https://i.postimg.cc/Kv64VYSL/3.png\" alt=\"3.png\"></a></p>\n<h2>Stage1 - Base Model Fitting</h2>\n<p>In stage1, the main model architecture is just the same as what I shared before. Nonetheless, more categorical features are taken into consideration, including <code>event_comb</code> (<em>i.e.,</em> <code>event_name</code> + <code>name</code>), <code>room_fqid</code>, <code>text_fqid</code>, <code>fqid</code>,  <code>level</code> and <code>page</code> (the corresponding embedding dimensions are shown in the figure). After the categorical embeddings are obtained, I use the difference of <code>elapsed_time</code> to tweak the scale of each event record (one column of the raw feature matrix), which can be thought of as the importance of that snapshot. Then a three-layer temporal convolution architecture is applied to extract the temporal pattern of the event sequence. After skip connection and readout, out-of-fold predictions (probas) are kept as the meta features for the stage2.</p>\n<h2>Stage2 - Stacker Training</h2>\n<p>I still remember on 20th, Jun I thought that it seemed impossible to break 0.697 on local CV with this DL model. Over the past three months, I worked hard to explore and try various of model architecture, but all of them fail to surpass the one as discussed above. Therefore, I decided to try some ensemble methods to see if performance could improve. After some trials, restacking based on the meta features from the level-1 model with some FE features made me nearly touch the score of 0.7 on local side. With further blending of three stackers, I got to pass 0.7 by a small margin. The number of features (meta + FE) used in each <code>level_group</code> is shown as follows:</p>\n<ul>\n<li><code>level_group</code> \"0-4\": 3 + 22</li>\n<li><code>level_group</code> \"5-12\": 13 + 26</li>\n<li><code>level_group</code> \"13-22\": 18 + 30</li>\n</ul>\n<h1>Performance Study</h1>\n<p>Following is the top performance report of my experiments. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Single DL Model</td>\n<td>0.6964</td>\n<td>0.697</td>\n<td>0.697</td>\n</tr>\n<tr>\n<td>Restack w/ Single XGB</td>\n<td>0.6988</td>\n<td>0.699</td>\n<td></td>\n</tr>\n<tr>\n<td>Restack w/ Single CatBoost</td>\n<td>0.6996</td>\n<td>0.698</td>\n<td>0.698</td>\n</tr>\n<tr>\n<td>Restack w/ CatBoost*1 + XGB*1 + LGBM*1</td>\n<td>0.6999</td>\n<td><strong>0.701</strong></td>\n<td>0.695</td>\n</tr>\n<tr>\n<td>Restack w/ CatBoost*1 + XGB*1 + LGBM*1</td>\n<td><strong>0.7000</strong></td>\n<td>0.699</td>\n<td>0.696</td>\n</tr>\n</tbody>\n</table>\n<p>Though the performance indeed improves when restacking plays in, I think that I over-tune my model during the last week of the competition, which leads to overfit and fail to generalize to private test set.</p>\n<h1>What Doesn't Work (on My Side)</h1>\n<ol>\n<li>Retrain base models with all sessions (still find it hard to tune and select the right checkpoints…).</li>\n<li>Use question embedding to do question-wise attention on either raw event sequence or the temporal pattern extracted by temporal convolution architecture.</li>\n<li>Add more numeric features at level-1 input data, such as <code>hover_duration</code> and the track of clicks (<em>i.e.,</em> distance and angle between consecutive clicks).</li>\n<li>Try more temporal pattern extractors (<em>e.g.,</em> Conv3D, <a href=\"https://arxiv.org/abs/2106.09305\" target=\"_blank\">SCINet</a>).</li>\n</ol>\n<h1>Conclusion</h1>\n<p>Although the final rank isn't good, I learn a lot from the Kaggle community and enjoy sharing some of my findings during this competitive journey. I'll keep progressing and try to clarify what I've lost (FE is crucial…). Thanks a lot for your patience!</p>",
      "rawMarkdown": "At first, thank you Kaggle and the hosts for hosting this interesting competition. Also, thanks the Kaggle community for all the sharings, discussions and kindness. I would like to share my solutions and the learning journey below (though the final ranking isn't that good 😓).\n\n# Overview\n* Add data from [fielddaylab](https://fielddaylab.wisc.edu/opengamedata/)\n* Use straified KFold based on #correct answers of each session\n* Implement two-level modeling\n    * Level1: The base model is mainly based on my first sharing of the DL model architecture [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565).\n   * Level2: The stackers (*e.g.,* xgb, lgbm and catboost) are trained with the oof probas (meta features) of the base model and simple FE features.\n\n# CV Scheme\nBecause the Macro-F1 score is derived on flattened prediction and groundtruth sequences, I think it would be better to somewhat retain the distribution of correctness when doing local CV. However, I find that it doesn't show superiority over the plain `GroupKFold` on `session_id`. Finally, I choose stratified KFold (k=5) based on #correct answers with 5 different seeds for stability of evaluation.\n\n# Modeling Process\nAs shown in the figure below, the modeling process on the highest level consists of two stages, *base model fitting* and *stacker training*. The subsequent subsections describe these phases respectively.\n[![3.png](https://i.postimg.cc/Kv64VYSL/3.png)](https://postimg.cc/1fK9Vsz5)\n## Stage1 - Base Model Fitting\nIn stage1, the main model architecture is just the same as what I shared before. Nonetheless, more categorical features are taken into consideration, including `event_comb` (*i.e.,* `event_name` + `name`), `room_fqid`, `text_fqid`, `fqid`,  `level` and `page` (the corresponding embedding dimensions are shown in the figure). After the categorical embeddings are obtained, I use the difference of `elapsed_time` to tweak the scale of each event record (one column of the raw feature matrix), which can be thought of as the importance of that snapshot. Then a three-layer temporal convolution architecture is applied to extract the temporal pattern of the event sequence. After skip connection and readout, out-of-fold predictions (probas) are kept as the meta features for the stage2.\n\n## Stage2 - Stacker Training\nI still remember on 20th, Jun I thought that it seemed impossible to break 0.697 on local CV with this DL model. Over the past three months, I worked hard to explore and try various of model architecture, but all of them fail to surpass the one as discussed above. Therefore, I decided to try some ensemble methods to see if performance could improve. After some trials, restacking based on the meta features from the level-1 model with some FE features made me nearly touch the score of 0.7 on local side. With further blending of three stackers, I got to pass 0.7 by a small margin. The number of features (meta + FE) used in each `level_group` is shown as follows:\n* `level_group` \"0-4\": 3 + 22\n* `level_group` \"5-12\": 13 + 26\n* `level_group` \"13-22\": 18 + 30\n\n# Performance Study\nFollowing is the top performance report of my experiments. \n\n| Model                      | CV     | LB    | PB    |\n| -------------------------- | ------ | ----- | ----- |\n| Single DL Model            | 0.6964 | 0.697 | 0.697 |\n| Restack w/ Single XGB      | 0.6988 | 0.699 | <span style=\"color:red\"><strong>0.698</strong></span> |\n| Restack w/ Single CatBoost | 0.6996 | 0.698 | 0.698 |\n| Restack w/ CatBoost\\*1 + XGB\\*1 + LGBM\\*1                           | 0.6999       | **0.701**      | 0.695      |\n| Restack w/ CatBoost\\*1 + XGB\\*1 + LGBM\\*1                           | **0.7000**       | 0.699      | 0.696      |\n\nThough the performance indeed improves when restacking plays in, I think that I over-tune my model during the last week of the competition, which leads to overfit and fail to generalize to private test set.\n\n# What Doesn't Work (on My Side)\n1. Retrain base models with all sessions (still find it hard to tune and select the right checkpoints...).\n2. Use question embedding to do question-wise attention on either raw event sequence or the temporal pattern extracted by temporal convolution architecture.\n3. Add more numeric features at level-1 input data, such as `hover_duration` and the track of clicks (*i.e.,* distance and angle between consecutive clicks).\n4. Try more temporal pattern extractors (*e.g.,* Conv3D, [SCINet](https://arxiv.org/abs/2106.09305)).\n\n# Conclusion\nAlthough the final rank isn't good, I learn a lot from the Kaggle community and enjoy sharing some of my findings during this competitive journey. I'll keep progressing and try to clarify what I've lost (FE is crucial...). Thanks a lot for your patience!",
      "votes": null
    },
    {
      "id": "2325363",
      "postDate": "07/01/2023 09:21:45",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>, the DL Notebook you posted is very well written and I learned a lot reading it and experimenting with it, but I couldn't get a competitive score out of it. Thanks for sharing.</p>",
      "rawMarkdown": "Hi @abaojiang, the DL Notebook you posted is very well written and I learned a lot reading it and experimenting with it, but I couldn't get a competitive score out of it. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2325551",
      "postDate": "07/01/2023 12:54:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a>,</p>\n<p>I also learned a lot from the notebook you published, especially the analysis related to abnormal event counts and session lengths.</p>\n<p>With insufficient experiments, I failed to improve the performance with this model architecture in the end. There's still so much to learn. Thanks a lot for your comments and all the sharings 😊  </p>",
      "rawMarkdown": "Hi @gehallak,\n\nI also learned a lot from the notebook you published, especially the analysis related to abnormal event counts and session lengths.\n\nWith insufficient experiments, I failed to improve the performance with this model architecture in the end. There's still so much to learn. Thanks a lot for your comments and all the sharings 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2325363,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "07/01/2023 09:21:45",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>, the DL Notebook you posted is very well written and I learned a lot reading it and experimenting with it, but I couldn't get a competitive score out of it. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2325551,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "07/01/2023 12:54:01",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a>,</p>\n<p>I also learned a lot from the notebook you published, especially the analysis related to abnormal event counts and session lengths.</p>\n<p>With insufficient experiments, I failed to improve the performance with this model architecture in the end. There's still so much to learn. Thanks a lot for your comments and all the sharings 😊  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2322825": "At first, thank you Kaggle and the hosts for hosting this interesting competition. Also, thanks the Kaggle community for all the sharings, discussions and kindness. I would like to share my solutions and the learning journey below (though the final ranking isn't that good 😓).\n\n# Overview\n* Add data from [fielddaylab](https://fielddaylab.wisc.edu/opengamedata/)\n* Use straified KFold based on #correct answers of each session\n* Implement two-level modeling\n    * Level1: The base model is mainly based on my first sharing of the DL model architecture [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565).\n   * Level2: The stackers (*e.g.,* xgb, lgbm and catboost) are trained with the oof probas (meta features) of the base model and simple FE features.\n\n# CV Scheme\nBecause the Macro-F1 score is derived on flattened prediction and groundtruth sequences, I think it would be better to somewhat retain the distribution of correctness when doing local CV. However, I find that it doesn't show superiority over the plain `GroupKFold` on `session_id`. Finally, I choose stratified KFold (k=5) based on #correct answers with 5 different seeds for stability of evaluation.\n\n# Modeling Process\nAs shown in the figure below, the modeling process on the highest level consists of two stages, *base model fitting* and *stacker training*. The subsequent subsections describe these phases respectively.\n[![3.png](https://i.postimg.cc/Kv64VYSL/3.png)](https://postimg.cc/1fK9Vsz5)\n## Stage1 - Base Model Fitting\nIn stage1, the main model architecture is just the same as what I shared before. Nonetheless, more categorical features are taken into consideration, including `event_comb` (*i.e.,* `event_name` + `name`), `room_fqid`, `text_fqid`, `fqid`,  `level` and `page` (the corresponding embedding dimensions are shown in the figure). After the categorical embeddings are obtained, I use the difference of `elapsed_time` to tweak the scale of each event record (one column of the raw feature matrix), which can be thought of as the importance of that snapshot. Then a three-layer temporal convolution architecture is applied to extract the temporal pattern of the event sequence. After skip connection and readout, out-of-fold predictions (probas) are kept as the meta features for the stage2.\n\n## Stage2 - Stacker Training\nI still remember on 20th, Jun I thought that it seemed impossible to break 0.697 on local CV with this DL model. Over the past three months, I worked hard to explore and try various of model architecture, but all of them fail to surpass the one as discussed above. Therefore, I decided to try some ensemble methods to see if performance could improve. After some trials, restacking based on the meta features from the level-1 model with some FE features made me nearly touch the score of 0.7 on local side. With further blending of three stackers, I got to pass 0.7 by a small margin. The number of features (meta + FE) used in each `level_group` is shown as follows:\n* `level_group` \"0-4\": 3 + 22\n* `level_group` \"5-12\": 13 + 26\n* `level_group` \"13-22\": 18 + 30\n\n# Performance Study\nFollowing is the top performance report of my experiments. \n\n| Model                      | CV     | LB    | PB    |\n| -------------------------- | ------ | ----- | ----- |\n| Single DL Model            | 0.6964 | 0.697 | 0.697 |\n| Restack w/ Single XGB      | 0.6988 | 0.699 | <span style=\"color:red\"><strong>0.698</strong></span> |\n| Restack w/ Single CatBoost | 0.6996 | 0.698 | 0.698 |\n| Restack w/ CatBoost\\*1 + XGB\\*1 + LGBM\\*1                           | 0.6999       | **0.701**      | 0.695      |\n| Restack w/ CatBoost\\*1 + XGB\\*1 + LGBM\\*1                           | **0.7000**       | 0.699      | 0.696      |\n\nThough the performance indeed improves when restacking plays in, I think that I over-tune my model during the last week of the competition, which leads to overfit and fail to generalize to private test set.\n\n# What Doesn't Work (on My Side)\n1. Retrain base models with all sessions (still find it hard to tune and select the right checkpoints...).\n2. Use question embedding to do question-wise attention on either raw event sequence or the temporal pattern extracted by temporal convolution architecture.\n3. Add more numeric features at level-1 input data, such as `hover_duration` and the track of clicks (*i.e.,* distance and angle between consecutive clicks).\n4. Try more temporal pattern extractors (*e.g.,* Conv3D, [SCINet](https://arxiv.org/abs/2106.09305)).\n\n# Conclusion\nAlthough the final rank isn't good, I learn a lot from the Kaggle community and enjoy sharing some of my findings during this competitive journey. I'll keep progressing and try to clarify what I've lost (FE is crucial...). Thanks a lot for your patience!",
    "2325363": "Hi @abaojiang, the DL Notebook you posted is very well written and I learned a lot reading it and experimenting with it, but I couldn't get a competitive score out of it. Thanks for sharing.",
    "2325551": "Hi @gehallak,\n\nI also learned a lot from the notebook you published, especially the analysis related to abnormal event counts and session lengths.\n\nWith insufficient experiments, I failed to improve the performance with this model architecture in the end. There's still so much to learn. Thanks a lot for your comments and all the sharings 😊"
  },
  "source": "meta"
}