{
  "id": 420528,
  "title": "8th Place Solution and Code",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/machine-not-learning-8th-place-solution-and-code",
  "author_name": "",
  "post_date": "2023-07-01T13:45:48.920Z",
  "votes": 30,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The competition was really exciting and it gave us a chance to practice feature engineering. I'm very thankful for the support and help from my team <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a>  <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> and <a href=\"https://www.kaggle.com/martasprg\" target=\"_blank\">@martasprg</a>. They were always there for me and together we made a big difference.</p>\n<p>I would like to thank the hosts, and special thanks to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> for identifying the issue of data leak, which made the competition right back on track.</p>\n<p>Special thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for his great starter notebooks and insights that helped me  in the early phase of the competition.</p>\n<h2>Overview</h2>\n<p>Here's an overview of what each of us worked on:<br>\n·         My main focus was on improving the XGBoost model and handling feature engineering.<br>\n·         Minh Tri Phan worked on a Transformer model with a CV (cross-validation) score of 0.699 and a public leaderboard (LB) score of 0.7.<br>\n·         Hoang processed the external data.<br>\n·         Martin worked on selecting the most relevant features.</p>\n<p>In our final submissions, we ensembled the XGBoost and Transformer models, which helped us achieve the gold position. Our ensemble submission had a public LB score of <strong>0.705</strong> and a private LB score of approximately <strong>0.7025</strong>. Additionally, we had two other submissions with single XGBoost models, where one had a public LB score of <strong>0.705</strong> and a private LB score of <strong>0.700</strong>.</p>\n<h2>My Part</h2>\n<p>Code: The code is a bit uncleaned, apologies for that. For any queries, contact me on <a href=\"https://www.linkedin.com/in/priyanshu-chaudhary-ba0b23199/\" target=\"_blank\">LinkedIn</a> <br>\n<strong>FE code:</strong> <a href=\"https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-11-external-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-11-external-data/notebook</a><br>\n<strong>Inference code:</strong> <a href=\"https://www.kaggle.com/code/chaudharypriyanshu/inference-xgb-25-11-17/notebook\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/inference-xgb-25-11-17/notebook</a><br>\n<strong>Training code:</strong> <a href=\"https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-9-training/notebook\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-9-training/notebook</a></p>\n<h3>Overview</h3>\n<p>I created a 5-fold XGBoost model for each question (a  total of 90 models). I used Kaggle kernels only to train XGBoost since it took only 45 mins on Kaggle’s P100 GPU to train all 90 models.<br>\nThe single XGBoost model achieved a Public leaderboard (LB) score of 0.705 and took 45-50 mins for inference, but it didn't perform as well on the private LB. When we included Hoang's external data, the model's score improved to <strong>0.704</strong> on the private LB. However, we decided not to use it because the public LB score was unusually low at <strong>0.702</strong>.</p>\n<h3>Feature engineering</h3>\n<ol>\n<li><p><strong>Session length:</strong> simply accounts for the total length of the session per level group.</p></li>\n<li><p><strong>Instance features:</strong>  I created Object click-based features (first object click, room coordinates of that click, I called them Instance features)that were most important and gave an improvement of 0.0007, when I added them with standard features. I created a total of 36 features since there were 12 instances where object clicks were present.</p></li>\n<li><p><strong>Magic bingo features:</strong> Inspired from the public notebooks. I created more such features for all 3 level groups and it improved the CV by <strong>0.0003</strong>.</p></li>\n<li><p><strong>Standard features:</strong></p>\n<p>a) <strong>Count features:</strong> I created count features based on <code>Fqid, text_Fqid, room_fqid, level, and event_comb</code>. These features capture the frequency of specific events or combinations. </p>\n<p>b) <strong>Binning of indexes:</strong> I performed binning on indexes with bin sizes of approximately 30 or 50 in sorted order. Raw indexes worked better on the private LB, while binned features yielded better results on the public LB.</p>\n<p>c) <strong>First and Sum features:</strong> I generated first and sum of elapsed_time_diff  for all categorical columns. I found that min, max, and std did not work well in my case. </p>\n<p>d) <strong>Aggregations based on hover duration.</strong></p></li>\n<li><p><strong>Top Level Group Features:</strong> Used top 15-25 features (according to feature importance), Duration and instance features across different level groups.</p></li>\n<li><p><strong>Meta features:</strong> Using past questions predictions to predict the current question. i.e. for question<code>t</code> I used all predictions for questions <code>(1 to t-1)</code>. Using them gave an improvement of around <strong>0.001</strong>.</p></li>\n</ol>\n<h3>Feature Selection (Martin's Part):</h3>\n<ol>\n<li>I eliminated features that had zero importance based on their Gain and Shapley feature importance scores.</li>\n<li>After performing feature selection, I made adjustments to the learning rate by reducing it from <strong>0.05</strong> to <strong>0.03</strong> and adding more features. </li>\n<li>Additionally, I removed duplicate features and features with more than <strong>95%</strong> values as null.</li>\n</ol>\n<h3>External data:</h3>\n<ol>\n<li>We used publicly available data. It had about 7500 sessions where all 18 questions were answered.</li>\n<li>Adding this external data improved our model's performance by 0.0005 in cross-validation and 0.002 on the leaderboard.</li>\n<li>Hoang also created processed external data that worked well on the private leaderboard (score of 0.704). If we had included it, our single XGBoost model could have reached a top 5. position. However, we decided not to use it because of its lower Public leaderboard score (a bad decision).</li>\n</ol>\n<h3>Inference:</h3>\n<ol>\n<li><p>We made improvements to retain the original order of the sequence during inference.</p></li>\n<li><p>We found that there are approx. 250 sessions with abnormal indexing(interestingly all of them are from the 5th and 6th December 2020)</p></li>\n<li><p>Created a function to preserve the original sequence for 99.5% of sequences, with only a small portion (0.5%) having misplaced events not more than 4-5 positions of the actual index.</p></li>\n<li><p>Reindexed these abnormal sessions which improved or scored on LB slightly.</p></li>\n</ol>\n<h3>Things not worked:</h3>\n<ol>\n<li>Ensemble with LGBM, Catboost didn’t work.</li>\n<li>Created a custom eval metric that uses benchmark true positives and negatives a model should have. It increased the CV by 0.001 but LB  decreased probably due to overfitting.</li>\n<li>Different thresholds for each question. (increased CV decreased LB).</li>\n</ol>\n<p>The below tables list our experiments with the best results.</p>\n<table>\n<thead>\n<tr>\n<th>External Data Used</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>final Sub</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No</td>\n<td>0.6996</td>\n<td>0.701</td>\n<td>0.698</td>\n<td>No</td>\n</tr>\n<tr>\n<td>No</td>\n<td>0.7001</td>\n<td>0.702</td>\n<td>0.700</td>\n<td>No</td>\n</tr>\n<tr>\n<td>No</td>\n<td>0.6996</td>\n<td>0.701</td>\n<td>0.698</td>\n<td>No</td>\n</tr>\n<tr>\n<td>Yes((Public ED)</td>\n<td>0.7015</td>\n<td>0.705</td>\n<td>0.700</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>Yes(Hoang's ED )</td>\n<td>0.7019</td>\n<td>0.702</td>\n<td>0.704</td>\n<td>No</td>\n</tr>\n<tr>\n<td>Yes (Hoang's ED)</td>\n<td>0.7022</td>\n<td>0.703</td>\n<td>0.703</td>\n<td>No</td>\n</tr>\n</tbody>\n</table>\n<h2>Tri's Part:</h2>\n<p>The model is shown in the following figure:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5005351%2F196a1b67b9b8bdf56ba47c91097e8d0d%2FPicture1.png?generation=1688196490904650&amp;alt=media\" alt=\"\"></p>\n<p>Particularly, it consists of 2 parts,<br>\n(i) Training a neural network, then extracting the embedding.<br>\n(ii) Concatenating the embedding from the neural network to a set of aggregated features, then training a gradient boosting model (XGBoost, CatBoost, LightGBM).</p>\n<h3>Neural network</h3>\n<p>I was inspired by the RIIID competition and <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>’s solution, in which he customized the multi-head attention mechanism to adopt the time gap between 2 actions. In my opinion, if we have to relate the problem to an NLP problem, RIIID competition is like a token classification task (e.g., NER), meanwhile, this competition is like a document classification task. Therefore, I decided to use a transformer and some other recurrent network types.</p>\n<p>I used the encoder-only structure as I didn’t see any motivation to have the decoder. <br>\nHowever, the transformer encoder alone didn’t work so well, so I decided to add some more (3) GRU layers in front of the encoder. The detailed architecture (Pytorch code) is given here (<a href=\"https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance/blob/main/Transformer/model.py\" target=\"_blank\">https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance/blob/main/Transformer/model.py</a>).</p>\n<h3>Some remarks about training:</h3>\n<ol>\n<li>I used 3 models for 3 level groups. At each level, I used the sequence of previous levels (e.g. The model for the 0-4 level uses the 0-4 sequence, the model for the 5-12 level uses the 0-4 and 5-12 sequences, and so on.)</li>\n<li>I used all the given features to train the model,</li>\n</ol>\n<pre><code>NUM_COLS = [, , , , , , ]\nTXT_COLS = [, , , , , , ]\n</code></pre>\n<ol>\n<li>I think the performance of a student, for example, in level 13-22 could carry some information to predict his/her performance in level 0-4. This is what I call the “global knowledge” of a student, and I want the network to capture that. Therefore, the neural network is trained in a multi-tasking manner, in which in the main output is the set of questions in the corresponding level (e.g., for level 0-4, the main output is 3-dimensional for questions 1, 2, and 3), the auxiliary head is used to predict all other questions. This trick helps to gain <strong>+0.002</strong> in CV.</li>\n</ol>\n<p>Overall, the NN gets <strong>0.695/0.700</strong> in CV and public LB (before the API crisis, after that I never check how the NN works in the public LB anymore as it was combined always with XGBoost)</p>\n<h3>Gradient Boosting</h3>\n<p>However, the NN in my case was not super satisfactory. I then decided to extract the embedding from the trained NN, concatenate them into a set of aggregated features, then use XGBoost to train the model. This helped me to get a huge boost in both CV and LB.</p>\n<p>Overall, the scores of this approach are shown below,</p>\n<table>\n<thead>\n<tr>\n<th>External Data Used</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No</td>\n<td>0.6993</td>\n<td>0.702</td>\n<td>0.697</td>\n</tr>\n<tr>\n<td>Yes</td>\n<td>0.6989</td>\n<td>0.701</td>\n<td>0.699</td>\n</tr>\n</tbody>\n</table>\n<p>Unfortunately, as I didn’t observe any gain in CV and public LB with external data, I decided not to choose that model to add to our model pool.</p>\n<p><strong>Links:</strong><br>\nTraining code: <a href=\"https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance----part-of--8th-solution.git\" target=\"_blank\">https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance----part-of--8th-solution.git</a></p>\n<p><strong>Inference code:</strong> <br>\n<strong>Without external data:</strong> <a href=\"https://www.kaggle.com/code/shinomoriaoshi/psp-v7b-infer\" target=\"_blank\">https://www.kaggle.com/code/shinomoriaoshi/psp-v7b-infer</a><br>\n<strong>With external data:</strong> <a href=\"https://www.kaggle.com/code/shinomoriaoshi/psp-v9a-infer\" target=\"_blank\">https://www.kaggle.com/code/shinomoriaoshi/psp-v9a-infer</a></p>\n<h2>Hoang's Part:</h2>\n<p>Hoang has described his work in a separate thread that describes the preprocessing of external data, experimental results and why to trust CV over LB.<br>\nLink to Hoang's Part: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420315\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420315</a></p>",
  "messages": [
    {
      "id": "2325249",
      "postDate": "07/01/2023 07:51:31",
      "content": "<p>The competition was really exciting and it gave us a chance to practice feature engineering. I'm very thankful for the support and help from my team <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a>  <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> and <a href=\"https://www.kaggle.com/martasprg\" target=\"_blank\">@martasprg</a>. They were always there for me and together we made a big difference.</p>\n<p>I would like to thank the hosts, and special thanks to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> for identifying the issue of data leak, which made the competition right back on track.</p>\n<p>Special thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for his great starter notebooks and insights that helped me  in the early phase of the competition.</p>\n<h2>Overview</h2>\n<p>Here's an overview of what each of us worked on:<br>\n·         My main focus was on improving the XGBoost model and handling feature engineering.<br>\n·         Minh Tri Phan worked on a Transformer model with a CV (cross-validation) score of 0.699 and a public leaderboard (LB) score of 0.7.<br>\n·         Hoang processed the external data.<br>\n·         Martin worked on selecting the most relevant features.</p>\n<p>In our final submissions, we ensembled the XGBoost and Transformer models, which helped us achieve the gold position. Our ensemble submission had a public LB score of <strong>0.705</strong> and a private LB score of approximately <strong>0.7025</strong>. Additionally, we had two other submissions with single XGBoost models, where one had a public LB score of <strong>0.705</strong> and a private LB score of <strong>0.700</strong>.</p>\n<h2>My Part</h2>\n<p>Code: The code is a bit uncleaned, apologies for that. For any queries, contact me on <a href=\"https://www.linkedin.com/in/priyanshu-chaudhary-ba0b23199/\" target=\"_blank\">LinkedIn</a> <br>\n<strong>FE code:</strong> <a href=\"https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-11-external-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-11-external-data/notebook</a><br>\n<strong>Inference code:</strong> <a href=\"https://www.kaggle.com/code/chaudharypriyanshu/inference-xgb-25-11-17/notebook\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/inference-xgb-25-11-17/notebook</a><br>\n<strong>Training code:</strong> <a href=\"https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-9-training/notebook\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-9-training/notebook</a></p>\n<h3>Overview</h3>\n<p>I created a 5-fold XGBoost model for each question (a  total of 90 models). I used Kaggle kernels only to train XGBoost since it took only 45 mins on Kaggle’s P100 GPU to train all 90 models.<br>\nThe single XGBoost model achieved a Public leaderboard (LB) score of 0.705 and took 45-50 mins for inference, but it didn't perform as well on the private LB. When we included Hoang's external data, the model's score improved to <strong>0.704</strong> on the private LB. However, we decided not to use it because the public LB score was unusually low at <strong>0.702</strong>.</p>\n<h3>Feature engineering</h3>\n<ol>\n<li><p><strong>Session length:</strong> simply accounts for the total length of the session per level group.</p></li>\n<li><p><strong>Instance features:</strong>  I created Object click-based features (first object click, room coordinates of that click, I called them Instance features)that were most important and gave an improvement of 0.0007, when I added them with standard features. I created a total of 36 features since there were 12 instances where object clicks were present.</p></li>\n<li><p><strong>Magic bingo features:</strong> Inspired from the public notebooks. I created more such features for all 3 level groups and it improved the CV by <strong>0.0003</strong>.</p></li>\n<li><p><strong>Standard features:</strong></p>\n<p>a) <strong>Count features:</strong> I created count features based on <code>Fqid, text_Fqid, room_fqid, level, and event_comb</code>. These features capture the frequency of specific events or combinations. </p>\n<p>b) <strong>Binning of indexes:</strong> I performed binning on indexes with bin sizes of approximately 30 or 50 in sorted order. Raw indexes worked better on the private LB, while binned features yielded better results on the public LB.</p>\n<p>c) <strong>First and Sum features:</strong> I generated first and sum of elapsed_time_diff  for all categorical columns. I found that min, max, and std did not work well in my case. </p>\n<p>d) <strong>Aggregations based on hover duration.</strong></p></li>\n<li><p><strong>Top Level Group Features:</strong> Used top 15-25 features (according to feature importance), Duration and instance features across different level groups.</p></li>\n<li><p><strong>Meta features:</strong> Using past questions predictions to predict the current question. i.e. for question<code>t</code> I used all predictions for questions <code>(1 to t-1)</code>. Using them gave an improvement of around <strong>0.001</strong>.</p></li>\n</ol>\n<h3>Feature Selection (Martin's Part):</h3>\n<ol>\n<li>I eliminated features that had zero importance based on their Gain and Shapley feature importance scores.</li>\n<li>After performing feature selection, I made adjustments to the learning rate by reducing it from <strong>0.05</strong> to <strong>0.03</strong> and adding more features. </li>\n<li>Additionally, I removed duplicate features and features with more than <strong>95%</strong> values as null.</li>\n</ol>\n<h3>External data:</h3>\n<ol>\n<li>We used publicly available data. It had about 7500 sessions where all 18 questions were answered.</li>\n<li>Adding this external data improved our model's performance by 0.0005 in cross-validation and 0.002 on the leaderboard.</li>\n<li>Hoang also created processed external data that worked well on the private leaderboard (score of 0.704). If we had included it, our single XGBoost model could have reached a top 5. position. However, we decided not to use it because of its lower Public leaderboard score (a bad decision).</li>\n</ol>\n<h3>Inference:</h3>\n<ol>\n<li><p>We made improvements to retain the original order of the sequence during inference.</p></li>\n<li><p>We found that there are approx. 250 sessions with abnormal indexing(interestingly all of them are from the 5th and 6th December 2020)</p></li>\n<li><p>Created a function to preserve the original sequence for 99.5% of sequences, with only a small portion (0.5%) having misplaced events not more than 4-5 positions of the actual index.</p></li>\n<li><p>Reindexed these abnormal sessions which improved or scored on LB slightly.</p></li>\n</ol>\n<h3>Things not worked:</h3>\n<ol>\n<li>Ensemble with LGBM, Catboost didn’t work.</li>\n<li>Created a custom eval metric that uses benchmark true positives and negatives a model should have. It increased the CV by 0.001 but LB  decreased probably due to overfitting.</li>\n<li>Different thresholds for each question. (increased CV decreased LB).</li>\n</ol>\n<p>The below tables list our experiments with the best results.</p>\n<table>\n<thead>\n<tr>\n<th>External Data Used</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>final Sub</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No</td>\n<td>0.6996</td>\n<td>0.701</td>\n<td>0.698</td>\n<td>No</td>\n</tr>\n<tr>\n<td>No</td>\n<td>0.7001</td>\n<td>0.702</td>\n<td>0.700</td>\n<td>No</td>\n</tr>\n<tr>\n<td>No</td>\n<td>0.6996</td>\n<td>0.701</td>\n<td>0.698</td>\n<td>No</td>\n</tr>\n<tr>\n<td>Yes((Public ED)</td>\n<td>0.7015</td>\n<td>0.705</td>\n<td>0.700</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>Yes(Hoang's ED )</td>\n<td>0.7019</td>\n<td>0.702</td>\n<td>0.704</td>\n<td>No</td>\n</tr>\n<tr>\n<td>Yes (Hoang's ED)</td>\n<td>0.7022</td>\n<td>0.703</td>\n<td>0.703</td>\n<td>No</td>\n</tr>\n</tbody>\n</table>\n<h2>Tri's Part:</h2>\n<p>The model is shown in the following figure:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5005351%2F196a1b67b9b8bdf56ba47c91097e8d0d%2FPicture1.png?generation=1688196490904650&amp;alt=media\" alt=\"\"></p>\n<p>Particularly, it consists of 2 parts,<br>\n(i) Training a neural network, then extracting the embedding.<br>\n(ii) Concatenating the embedding from the neural network to a set of aggregated features, then training a gradient boosting model (XGBoost, CatBoost, LightGBM).</p>\n<h3>Neural network</h3>\n<p>I was inspired by the RIIID competition and <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>’s solution, in which he customized the multi-head attention mechanism to adopt the time gap between 2 actions. In my opinion, if we have to relate the problem to an NLP problem, RIIID competition is like a token classification task (e.g., NER), meanwhile, this competition is like a document classification task. Therefore, I decided to use a transformer and some other recurrent network types.</p>\n<p>I used the encoder-only structure as I didn’t see any motivation to have the decoder. <br>\nHowever, the transformer encoder alone didn’t work so well, so I decided to add some more (3) GRU layers in front of the encoder. The detailed architecture (Pytorch code) is given here (<a href=\"https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance/blob/main/Transformer/model.py\" target=\"_blank\">https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance/blob/main/Transformer/model.py</a>).</p>\n<h3>Some remarks about training:</h3>\n<ol>\n<li>I used 3 models for 3 level groups. At each level, I used the sequence of previous levels (e.g. The model for the 0-4 level uses the 0-4 sequence, the model for the 5-12 level uses the 0-4 and 5-12 sequences, and so on.)</li>\n<li>I used all the given features to train the model,</li>\n</ol>\n<pre><code>NUM_COLS = [, , , , , , ]\nTXT_COLS = [, , , , , , ]\n</code></pre>\n<ol>\n<li>I think the performance of a student, for example, in level 13-22 could carry some information to predict his/her performance in level 0-4. This is what I call the “global knowledge” of a student, and I want the network to capture that. Therefore, the neural network is trained in a multi-tasking manner, in which in the main output is the set of questions in the corresponding level (e.g., for level 0-4, the main output is 3-dimensional for questions 1, 2, and 3), the auxiliary head is used to predict all other questions. This trick helps to gain <strong>+0.002</strong> in CV.</li>\n</ol>\n<p>Overall, the NN gets <strong>0.695/0.700</strong> in CV and public LB (before the API crisis, after that I never check how the NN works in the public LB anymore as it was combined always with XGBoost)</p>\n<h3>Gradient Boosting</h3>\n<p>However, the NN in my case was not super satisfactory. I then decided to extract the embedding from the trained NN, concatenate them into a set of aggregated features, then use XGBoost to train the model. This helped me to get a huge boost in both CV and LB.</p>\n<p>Overall, the scores of this approach are shown below,</p>\n<table>\n<thead>\n<tr>\n<th>External Data Used</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No</td>\n<td>0.6993</td>\n<td>0.702</td>\n<td>0.697</td>\n</tr>\n<tr>\n<td>Yes</td>\n<td>0.6989</td>\n<td>0.701</td>\n<td>0.699</td>\n</tr>\n</tbody>\n</table>\n<p>Unfortunately, as I didn’t observe any gain in CV and public LB with external data, I decided not to choose that model to add to our model pool.</p>\n<p><strong>Links:</strong><br>\nTraining code: <a href=\"https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance----part-of--8th-solution.git\" target=\"_blank\">https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance----part-of--8th-solution.git</a></p>\n<p><strong>Inference code:</strong> <br>\n<strong>Without external data:</strong> <a href=\"https://www.kaggle.com/code/shinomoriaoshi/psp-v7b-infer\" target=\"_blank\">https://www.kaggle.com/code/shinomoriaoshi/psp-v7b-infer</a><br>\n<strong>With external data:</strong> <a href=\"https://www.kaggle.com/code/shinomoriaoshi/psp-v9a-infer\" target=\"_blank\">https://www.kaggle.com/code/shinomoriaoshi/psp-v9a-infer</a></p>\n<h2>Hoang's Part:</h2>\n<p>Hoang has described his work in a separate thread that describes the preprocessing of external data, experimental results and why to trust CV over LB.<br>\nLink to Hoang's Part: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420315\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420315</a></p>",
      "rawMarkdown": "The competition was really exciting and it gave us a chance to practice feature engineering. I'm very thankful for the support and help from my team @shinomoriaoshi  @hoangnguyen719 and @martasprg. They were always there for me and together we made a big difference.\n\nI would like to thank the hosts, and special thanks to @cpmpml and @pdnartreb for identifying the issue of data leak, which made the competition right back on track.\n\nSpecial thanks to @cdeotte for his great starter notebooks and insights that helped me  in the early phase of the competition.\n\n##Overview\nHere's an overview of what each of us worked on:\n·         My main focus was on improving the XGBoost model and handling feature engineering.\n·         Minh Tri Phan worked on a Transformer model with a CV (cross-validation) score of 0.699 and a public leaderboard (LB) score of 0.7.\n·         Hoang processed the external data.\n·         Martin worked on selecting the most relevant features.\n\nIn our final submissions, we ensembled the XGBoost and Transformer models, which helped us achieve the gold position. Our ensemble submission had a public LB score of **0.705** and a private LB score of approximately **0.7025**. Additionally, we had two other submissions with single XGBoost models, where one had a public LB score of **0.705** and a private LB score of **0.700**.\n\n##My Part\n\nCode: The code is a bit uncleaned, apologies for that. For any queries, contact me on [LinkedIn](<https://www.linkedin.com/in/priyanshu-chaudhary-ba0b23199/) \n**FE code:** <https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-11-external-data/notebook>\n**Inference code:** <https://www.kaggle.com/code/chaudharypriyanshu/inference-xgb-25-11-17/notebook>\n**Training code:** <https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-9-training/notebook>\n\n### Overview\nI created a 5-fold XGBoost model for each question (a  total of 90 models). I used Kaggle kernels only to train XGBoost since it took only 45 mins on Kaggle’s P100 GPU to train all 90 models.\nThe single XGBoost model achieved a Public leaderboard (LB) score of 0.705 and took 45-50 mins for inference, but it didn't perform as well on the private LB. When we included Hoang's external data, the model's score improved to **0.704** on the private LB. However, we decided not to use it because the public LB score was unusually low at **0.702**.\n\n###Feature engineering \n\n1. **Session length:** simply accounts for the total length of the session per level group.\n\n2. **Instance features:**  I created Object click-based features (first object click, room coordinates of that click, I called them Instance features)that were most important and gave an improvement of 0.0007, when I added them with standard features. I created a total of 36 features since there were 12 instances where object clicks were present.\n\n3. **Magic bingo features:** Inspired from the public notebooks. I created more such features for all 3 level groups and it improved the CV by **0.0003**.\n\n4. **Standard features:**\n\n      a) **Count features:** I created count features based on `Fqid, text_Fqid, room_fqid, level, and event_comb`. These features capture the frequency of specific events or combinations. \n\n      b) **Binning of indexes:** I performed binning on indexes with bin sizes of approximately 30 or 50 in sorted order. Raw indexes worked better on the private LB, while binned features yielded better results on the public LB.\n\n      c) **First and Sum features:** I generated first and sum of elapsed_time_diff  for all categorical columns. I found that min, max, and std did not work well in my case. \n\n      d) **Aggregations based on hover duration.**\n\n5. **Top Level Group Features:** Used top 15-25 features (according to feature importance), Duration and instance features across different level groups.\n\n6. **Meta features:** Using past questions predictions to predict the current question. i.e. for question` t` I used all predictions for questions `(1 to t-1)`. Using them gave an improvement of around **0.001**.\n\n###Feature Selection (Martin's Part):\n1. I eliminated features that had zero importance based on their Gain and Shapley feature importance scores.\n2.  After performing feature selection, I made adjustments to the learning rate by reducing it from **0.05** to **0.03** and adding more features. \n3. Additionally, I removed duplicate features and features with more than **95%** values as null.\n\n\n###External data:\n1. We used publicly available data. It had about 7500 sessions where all 18 questions were answered.\n2. Adding this external data improved our model's performance by 0.0005 in cross-validation and 0.002 on the leaderboard.\n3. Hoang also created processed external data that worked well on the private leaderboard (score of 0.704). If we had included it, our single XGBoost model could have reached a top 5. position. However, we decided not to use it because of its lower Public leaderboard score (a bad decision).\n\n###Inference:\n\n1. We made improvements to retain the original order of the sequence during inference.\n\n2. We found that there are approx. 250 sessions with abnormal indexing(interestingly all of them are from the 5th and 6th December 2020)\n\n3. Created a function to preserve the original sequence for 99.5% of sequences, with only a small portion (0.5%) having misplaced events not more than 4-5 positions of the actual index.\n\n4. Reindexed these abnormal sessions which improved or scored on LB slightly.\n\n###Things not worked:\n1. Ensemble with LGBM, Catboost didn’t work.\n2. Created a custom eval metric that uses benchmark true positives and negatives a model should have. It increased the CV by 0.001 but LB  decreased probably due to overfitting.\n3. Different thresholds for each question. (increased CV decreased LB).\n\n\nThe below tables list our experiments with the best results.\n\n|External Data Used| CV |Public LB  | Private LB | final Sub|\n| --- | --- | --- | --- | --- |\n| No |0.6996  |0.701| 0.698 | No |\n| No|0.7001 |0.702| 0.700 | No |\n| No |0.6996  |0.701| 0.698 | No |\n| Yes((Public ED) |0.7015 |0.705| 0.700 | Yes |\n| Yes(Hoang's ED ) |0.7019 |0.702| 0.704 | No |\n| Yes (Hoang's ED) |0.7022 |0.703| 0.703 | No |\n\n##Tri's Part:\n\nThe model is shown in the following figure:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5005351%2F196a1b67b9b8bdf56ba47c91097e8d0d%2FPicture1.png?generation=1688196490904650&alt=media)\n\nParticularly, it consists of 2 parts,\n(i) Training a neural network, then extracting the embedding.\n(ii) Concatenating the embedding from the neural network to a set of aggregated features, then training a gradient boosting model (XGBoost, CatBoost, LightGBM).\n\n###Neural network\n\nI was inspired by the RIIID competition and @letranduckinh’s solution, in which he customized the multi-head attention mechanism to adopt the time gap between 2 actions. In my opinion, if we have to relate the problem to an NLP problem, RIIID competition is like a token classification task (e.g., NER), meanwhile, this competition is like a document classification task. Therefore, I decided to use a transformer and some other recurrent network types.\n\nI used the encoder-only structure as I didn’t see any motivation to have the decoder. \nHowever, the transformer encoder alone didn’t work so well, so I decided to add some more (3) GRU layers in front of the encoder. The detailed architecture (Pytorch code) is given here (<https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance/blob/main/Transformer/model.py>).\n\n###Some remarks about training:\n1. I used 3 models for 3 level groups. At each level, I used the sequence of previous levels (e.g. The model for the 0-4 level uses the 0-4 sequence, the model for the 5-12 level uses the 0-4 and 5-12 sequences, and so on.)\n2. I used all the given features to train the model,\n\n```python\nNUM_COLS = ['index', 'time_diff', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration']\nTXT_COLS = ['level', 'event_name', 'name', 'text', 'fqid', 'room_fqid', 'text_fqid']\n```\n\n3. I think the performance of a student, for example, in level 13-22 could carry some information to predict his/her performance in level 0-4. This is what I call the “global knowledge” of a student, and I want the network to capture that. Therefore, the neural network is trained in a multi-tasking manner, in which in the main output is the set of questions in the corresponding level (e.g., for level 0-4, the main output is 3-dimensional for questions 1, 2, and 3), the auxiliary head is used to predict all other questions. This trick helps to gain **+0.002** in CV.\n\nOverall, the NN gets **0.695/0.700** in CV and public LB (before the API crisis, after that I never check how the NN works in the public LB anymore as it was combined always with XGBoost)\n\n###Gradient Boosting\n\nHowever, the NN in my case was not super satisfactory. I then decided to extract the embedding from the trained NN, concatenate them into a set of aggregated features, then use XGBoost to train the model. This helped me to get a huge boost in both CV and LB.\n\nOverall, the scores of this approach are shown below,\n\n|External Data Used| CV |Public LB  | Private LB|\n| --- | --- | --- | --- |\n| No |0.6993  |0.702| 0.697 |\n| Yes |0.6989  |0.701| 0.699 |\n\n\nUnfortunately, as I didn’t observe any gain in CV and public LB with external data, I decided not to choose that model to add to our model pool.\n\n**Links:**\nTraining code: <https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance----part-of--8th-solution.git>\n\n**Inference code:** \n**Without external data:** <https://www.kaggle.com/code/shinomoriaoshi/psp-v7b-infer>\n**With external data:** <https://www.kaggle.com/code/shinomoriaoshi/psp-v9a-infer>\n\n## Hoang's Part: \nHoang has described his work in a separate thread that describes the preprocessing of external data, experimental results and why to trust CV over LB.\nLink to Hoang's Part: <https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420315>",
      "votes": null
    },
    {
      "id": "2325433",
      "postDate": "07/01/2023 10:40:55",
      "content": "<p>Thanks a lot for your work on this guys! So grateful I had a chance to collaborate and learn from you all.</p>",
      "rawMarkdown": "Thanks a lot for your work on this guys! So grateful I had a chance to collaborate and learn from you all.",
      "votes": null
    },
    {
      "id": "2327796",
      "postDate": "07/03/2023 07:20:22",
      "content": "<p>Congratulations!<br>\nThank you for sharing your solution so quickly.</p>\n<p>I am currently reviewing and looking back at the Discussions of the gold medalists in this competition.<br>\nI am not sure I understand the <strong>Magic bingo features</strong> that improved the CV by 0.0003. <br>\nI know you are busy and don't have much time, but I would appreciate it if you could tell me about the <strong>Magic bingo features</strong>.</p>",
      "rawMarkdown": "Congratulations!\nThank you for sharing your solution so quickly.\n\nI am currently reviewing and looking back at the Discussions of the gold medalists in this competition.\nI am not sure I understand the **Magic bingo features** that improved the CV by 0.0003. \nI know you are busy and don't have much time, but I would appreciate it if you could tell me about the **Magic bingo features**.",
      "votes": null
    },
    {
      "id": "2327857",
      "postDate": "07/03/2023 07:57:07",
      "content": "<p>Hey Ino,<br>\nIn the dataset, there are fqids, that have the word bingo in them. They signify that a user was asked to find something on the screen and press it, when the user presses the correct item bingo fqid pops up. Therefore these features indicate the time and number of clicks it took for the player to get the correct item so that bingo-based fqid appeared.<br>\nI'd advise going through one session and observing the fqid for level group 5-12, you will see some fqids that have word bingo in them.</p>",
      "rawMarkdown": "Hey Ino,\nIn the dataset, there are fqids, that have the word bingo in them. They signify that a user was asked to find something on the screen and press it, when the user presses the correct item bingo fqid pops up. Therefore these features indicate the time and number of clicks it took for the player to get the correct item so that bingo-based fqid appeared.\nI'd advise going through one session and observing the fqid for level group 5-12, you will see some fqids that have word bingo in them.",
      "votes": null
    },
    {
      "id": "2327877",
      "postDate": "07/03/2023 08:16:50",
      "content": "<p>Thank you for your quick and detailed response.<br>\nI understood that the word \"bingo\" was in the fqid, but did not know it would pop up when I pressed the correct item.<br>\nAs you said, I understand that features based on \"bingo\" words can help improve scores.<br>\nThank you so much for your help.</p>",
      "rawMarkdown": "Thank you for your quick and detailed response.\nI understood that the word \"bingo\" was in the fqid, but did not know it would pop up when I pressed the correct item.\nAs you said, I understand that features based on \"bingo\" words can help improve scores.\nThank you so much for your help.",
      "votes": null
    },
    {
      "id": "2330518",
      "postDate": "07/05/2023 04:09:20",
      "content": "<p>Congrats; nice correlation, Meta features on priors are the common theme , think that was the killer other than other smart features. And off course not being tempted by the lb 😀</p>\n<p>But nice nn approach as well by your team</p>",
      "rawMarkdown": "Congrats; nice correlation, Meta features on priors are the common theme , think that was the killer other than other smart features. And off course not being tempted by the lb 😀\n\nBut nice nn approach as well by your team",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2325433,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "07/01/2023 10:40:55",
      "content": "<p>Thanks a lot for your work on this guys! So grateful I had a chance to collaborate and learn from you all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2327796,
      "author_name": "inoway",
      "author_url": "",
      "post_date": "07/03/2023 07:20:22",
      "content": "<p>Congratulations!<br>\nThank you for sharing your solution so quickly.</p>\n<p>I am currently reviewing and looking back at the Discussions of the gold medalists in this competition.<br>\nI am not sure I understand the <strong>Magic bingo features</strong> that improved the CV by 0.0003. <br>\nI know you are busy and don't have much time, but I would appreciate it if you could tell me about the <strong>Magic bingo features</strong>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2327857,
          "author_name": "chaudharypriyanshu",
          "author_url": "",
          "post_date": "07/03/2023 07:57:07",
          "content": "<p>Hey Ino,<br>\nIn the dataset, there are fqids, that have the word bingo in them. They signify that a user was asked to find something on the screen and press it, when the user presses the correct item bingo fqid pops up. Therefore these features indicate the time and number of clicks it took for the player to get the correct item so that bingo-based fqid appeared.<br>\nI'd advise going through one session and observing the fqid for level group 5-12, you will see some fqids that have word bingo in them.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2327877,
              "author_name": "inoway",
              "author_url": "",
              "post_date": "07/03/2023 08:16:50",
              "content": "<p>Thank you for your quick and detailed response.<br>\nI understood that the word \"bingo\" was in the fqid, but did not know it would pop up when I pressed the correct item.<br>\nAs you said, I understand that features based on \"bingo\" words can help improve scores.<br>\nThank you so much for your help.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2330518,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/05/2023 04:09:20",
      "content": "<p>Congrats; nice correlation, Meta features on priors are the common theme , think that was the killer other than other smart features. And off course not being tempted by the lb 😀</p>\n<p>But nice nn approach as well by your team</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2325249": "The competition was really exciting and it gave us a chance to practice feature engineering. I'm very thankful for the support and help from my team @shinomoriaoshi  @hoangnguyen719 and @martasprg. They were always there for me and together we made a big difference.\n\nI would like to thank the hosts, and special thanks to @cpmpml and @pdnartreb for identifying the issue of data leak, which made the competition right back on track.\n\nSpecial thanks to @cdeotte for his great starter notebooks and insights that helped me  in the early phase of the competition.\n\n##Overview\nHere's an overview of what each of us worked on:\n·         My main focus was on improving the XGBoost model and handling feature engineering.\n·         Minh Tri Phan worked on a Transformer model with a CV (cross-validation) score of 0.699 and a public leaderboard (LB) score of 0.7.\n·         Hoang processed the external data.\n·         Martin worked on selecting the most relevant features.\n\nIn our final submissions, we ensembled the XGBoost and Transformer models, which helped us achieve the gold position. Our ensemble submission had a public LB score of **0.705** and a private LB score of approximately **0.7025**. Additionally, we had two other submissions with single XGBoost models, where one had a public LB score of **0.705** and a private LB score of **0.700**.\n\n##My Part\n\nCode: The code is a bit uncleaned, apologies for that. For any queries, contact me on [LinkedIn](<https://www.linkedin.com/in/priyanshu-chaudhary-ba0b23199/) \n**FE code:** <https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-11-external-data/notebook>\n**Inference code:** <https://www.kaggle.com/code/chaudharypriyanshu/inference-xgb-25-11-17/notebook>\n**Training code:** <https://www.kaggle.com/code/chaudharypriyanshu/mb-fb5-train-xgb-25-9-training/notebook>\n\n### Overview\nI created a 5-fold XGBoost model for each question (a  total of 90 models). I used Kaggle kernels only to train XGBoost since it took only 45 mins on Kaggle’s P100 GPU to train all 90 models.\nThe single XGBoost model achieved a Public leaderboard (LB) score of 0.705 and took 45-50 mins for inference, but it didn't perform as well on the private LB. When we included Hoang's external data, the model's score improved to **0.704** on the private LB. However, we decided not to use it because the public LB score was unusually low at **0.702**.\n\n###Feature engineering \n\n1. **Session length:** simply accounts for the total length of the session per level group.\n\n2. **Instance features:**  I created Object click-based features (first object click, room coordinates of that click, I called them Instance features)that were most important and gave an improvement of 0.0007, when I added them with standard features. I created a total of 36 features since there were 12 instances where object clicks were present.\n\n3. **Magic bingo features:** Inspired from the public notebooks. I created more such features for all 3 level groups and it improved the CV by **0.0003**.\n\n4. **Standard features:**\n\n      a) **Count features:** I created count features based on `Fqid, text_Fqid, room_fqid, level, and event_comb`. These features capture the frequency of specific events or combinations. \n\n      b) **Binning of indexes:** I performed binning on indexes with bin sizes of approximately 30 or 50 in sorted order. Raw indexes worked better on the private LB, while binned features yielded better results on the public LB.\n\n      c) **First and Sum features:** I generated first and sum of elapsed_time_diff  for all categorical columns. I found that min, max, and std did not work well in my case. \n\n      d) **Aggregations based on hover duration.**\n\n5. **Top Level Group Features:** Used top 15-25 features (according to feature importance), Duration and instance features across different level groups.\n\n6. **Meta features:** Using past questions predictions to predict the current question. i.e. for question` t` I used all predictions for questions `(1 to t-1)`. Using them gave an improvement of around **0.001**.\n\n###Feature Selection (Martin's Part):\n1. I eliminated features that had zero importance based on their Gain and Shapley feature importance scores.\n2.  After performing feature selection, I made adjustments to the learning rate by reducing it from **0.05** to **0.03** and adding more features. \n3. Additionally, I removed duplicate features and features with more than **95%** values as null.\n\n\n###External data:\n1. We used publicly available data. It had about 7500 sessions where all 18 questions were answered.\n2. Adding this external data improved our model's performance by 0.0005 in cross-validation and 0.002 on the leaderboard.\n3. Hoang also created processed external data that worked well on the private leaderboard (score of 0.704). If we had included it, our single XGBoost model could have reached a top 5. position. However, we decided not to use it because of its lower Public leaderboard score (a bad decision).\n\n###Inference:\n\n1. We made improvements to retain the original order of the sequence during inference.\n\n2. We found that there are approx. 250 sessions with abnormal indexing(interestingly all of them are from the 5th and 6th December 2020)\n\n3. Created a function to preserve the original sequence for 99.5% of sequences, with only a small portion (0.5%) having misplaced events not more than 4-5 positions of the actual index.\n\n4. Reindexed these abnormal sessions which improved or scored on LB slightly.\n\n###Things not worked:\n1. Ensemble with LGBM, Catboost didn’t work.\n2. Created a custom eval metric that uses benchmark true positives and negatives a model should have. It increased the CV by 0.001 but LB  decreased probably due to overfitting.\n3. Different thresholds for each question. (increased CV decreased LB).\n\n\nThe below tables list our experiments with the best results.\n\n|External Data Used| CV |Public LB  | Private LB | final Sub|\n| --- | --- | --- | --- | --- |\n| No |0.6996  |0.701| 0.698 | No |\n| No|0.7001 |0.702| 0.700 | No |\n| No |0.6996  |0.701| 0.698 | No |\n| Yes((Public ED) |0.7015 |0.705| 0.700 | Yes |\n| Yes(Hoang's ED ) |0.7019 |0.702| 0.704 | No |\n| Yes (Hoang's ED) |0.7022 |0.703| 0.703 | No |\n\n##Tri's Part:\n\nThe model is shown in the following figure:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5005351%2F196a1b67b9b8bdf56ba47c91097e8d0d%2FPicture1.png?generation=1688196490904650&alt=media)\n\nParticularly, it consists of 2 parts,\n(i) Training a neural network, then extracting the embedding.\n(ii) Concatenating the embedding from the neural network to a set of aggregated features, then training a gradient boosting model (XGBoost, CatBoost, LightGBM).\n\n###Neural network\n\nI was inspired by the RIIID competition and @letranduckinh’s solution, in which he customized the multi-head attention mechanism to adopt the time gap between 2 actions. In my opinion, if we have to relate the problem to an NLP problem, RIIID competition is like a token classification task (e.g., NER), meanwhile, this competition is like a document classification task. Therefore, I decided to use a transformer and some other recurrent network types.\n\nI used the encoder-only structure as I didn’t see any motivation to have the decoder. \nHowever, the transformer encoder alone didn’t work so well, so I decided to add some more (3) GRU layers in front of the encoder. The detailed architecture (Pytorch code) is given here (<https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance/blob/main/Transformer/model.py>).\n\n###Some remarks about training:\n1. I used 3 models for 3 level groups. At each level, I used the sequence of previous levels (e.g. The model for the 0-4 level uses the 0-4 sequence, the model for the 5-12 level uses the 0-4 and 5-12 sequences, and so on.)\n2. I used all the given features to train the model,\n\n```python\nNUM_COLS = ['index', 'time_diff', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration']\nTXT_COLS = ['level', 'event_name', 'name', 'text', 'fqid', 'room_fqid', 'text_fqid']\n```\n\n3. I think the performance of a student, for example, in level 13-22 could carry some information to predict his/her performance in level 0-4. This is what I call the “global knowledge” of a student, and I want the network to capture that. Therefore, the neural network is trained in a multi-tasking manner, in which in the main output is the set of questions in the corresponding level (e.g., for level 0-4, the main output is 3-dimensional for questions 1, 2, and 3), the auxiliary head is used to predict all other questions. This trick helps to gain **+0.002** in CV.\n\nOverall, the NN gets **0.695/0.700** in CV and public LB (before the API crisis, after that I never check how the NN works in the public LB anymore as it was combined always with XGBoost)\n\n###Gradient Boosting\n\nHowever, the NN in my case was not super satisfactory. I then decided to extract the embedding from the trained NN, concatenate them into a set of aggregated features, then use XGBoost to train the model. This helped me to get a huge boost in both CV and LB.\n\nOverall, the scores of this approach are shown below,\n\n|External Data Used| CV |Public LB  | Private LB|\n| --- | --- | --- | --- |\n| No |0.6993  |0.702| 0.697 |\n| Yes |0.6989  |0.701| 0.699 |\n\n\nUnfortunately, as I didn’t observe any gain in CV and public LB with external data, I decided not to choose that model to add to our model pool.\n\n**Links:**\nTraining code: <https://github.com/minhtriphan/Kaggle-competition---Predicting-Student-Performance----part-of--8th-solution.git>\n\n**Inference code:** \n**Without external data:** <https://www.kaggle.com/code/shinomoriaoshi/psp-v7b-infer>\n**With external data:** <https://www.kaggle.com/code/shinomoriaoshi/psp-v9a-infer>\n\n## Hoang's Part: \nHoang has described his work in a separate thread that describes the preprocessing of external data, experimental results and why to trust CV over LB.\nLink to Hoang's Part: <https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420315>",
    "2325433": "Thanks a lot for your work on this guys! So grateful I had a chance to collaborate and learn from you all.",
    "2327796": "Congratulations!\nThank you for sharing your solution so quickly.\n\nI am currently reviewing and looking back at the Discussions of the gold medalists in this competition.\nI am not sure I understand the **Magic bingo features** that improved the CV by 0.0003. \nI know you are busy and don't have much time, but I would appreciate it if you could tell me about the **Magic bingo features**.",
    "2327857": "Hey Ino,\nIn the dataset, there are fqids, that have the word bingo in them. They signify that a user was asked to find something on the screen and press it, when the user presses the correct item bingo fqid pops up. Therefore these features indicate the time and number of clicks it took for the player to get the correct item so that bingo-based fqid appeared.\nI'd advise going through one session and observing the fqid for level group 5-12, you will see some fqids that have word bingo in them.",
    "2327877": "Thank you for your quick and detailed response.\nI understood that the word \"bingo\" was in the fqid, but did not know it would pop up when I pressed the correct item.\nAs you said, I understand that features based on \"bingo\" words can help improve scores.\nThank you so much for your help.",
    "2330518": "Congrats; nice correlation, Meta features on priors are the common theme , think that was the killer other than other smart features. And off course not being tempted by the lb 😀\n\nBut nice nn approach as well by your team"
  },
  "source": "meta"
}