{
  "id": 420235,
  "title": "3rd Place Solution",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/stablegbt-nn-3rd-place-solution",
  "author_name": "",
  "post_date": "2023-07-06T17:35:44.813Z",
  "votes": 44,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks a lot to the hosts of the competition and my teammates ( <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a>, <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a>, and <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a>). I am thrilled that <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> and I will become GM, <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a> is one step closer to becoming a Master, and <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> will get his second gold medal after this competition!</p>\n<p>Here We will explain our overall solution, <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> also provided additional explanation of the parts he worked on: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274</a></p>\n<h1>Classification Task Formulation</h1>\n<p>In this competition, we are asked to predict 18 values for each session. Each session contains 3 level groups. There are multiple ways to model this.</p>\n<ol>\n<li>18 binary classifiers</li>\n<li>3 Level group classifiers, each one can be<br>\na. A multi-label classifier that predicts all values within a level group<br>\nb. A binary classifier that takes “question index” as a feature within a level group</li>\n<li>1 classifier that is<br>\na. A multi-label classifier that predicts 18 values within a session<br>\nb. A binary classifier that takes “question index” as a feature within a session</li>\n</ol>\n<p>For Gradient boosted tree models, method 2b &gt; method 3b &gt; method 1. Method 2a and 3a are ignored because training the multi-label task is a lot slower with Gradient boosted tree models.</p>\n<p>For NN models, we focus on the method 2a and 3a, because</p>\n<ul>\n<li>These 2 methods are not well handled by tree models</li>\n<li>Multi-label learning makes more sense, because of the F1 score setting of this competition. (some posts discuss we should not optimize for 1 question).</li>\n<li>Multi-label NN models are faster to train and infer.</li>\n</ul>\n<h1>Additional dataset generated from the raw data</h1>\n<p>We create an additional dataset from the raw data, it contains 11343 complete sessions.<br>\nThis dataset boosts the CV scores for GBT models  by about +0.001~2, but there is not much effect on the public and private scores, and it has both positive and negative outcomes.<br>\nHowever, it works very well for NN models, we see +0.002 improvement in both CV and public scores.</p>\n<h1>Validation</h1>\n<p>We are using 5-fold GroupKFold on session_id so that there won’t be any seen sessions in the validation set. Also we didn’t include additional data in our validation set.</p>\n<h1>Gradient Boosted Tree</h1>\n<p>Per question classifier is handled by <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a>, Per level, and All-in-1 classifier is handled by  <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a> <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a>. That’s why there are some inconsistencies in the data preprocessing steps, such as sort by index vs sort by time.</p>\n<h2>Per Question Classifiers</h2>\n<p>We create features for each level group and sorted by index. The features and the sorting methods differ from other models.</p>\n<pre><code>\ndf1 = df.(pl.col() == )\ndf2 = df.(pl.col() == )\ndf3 = df.(pl.col() == )\n\ndf1 = df1.sort(pl.col(), pl.col())\ndf2 = df2.sort(pl.col(), pl.col())\ndf3 = df3.sort(pl.col(), pl.col())\n</code></pre>\n<h5>The number of features:</h5>\n<ul>\n<li>Level group 0-4: 1,000 features</li>\n<li>Level group 5-12: 2,000 features</li>\n<li>Level group 13-22: 2,400 features</li>\n</ul>\n<h5>Feature Selection</h5>\n<p>We try feature selection with out-of-folds but the public scores tend to decrease, so we don’t select features about this model in the final submission.</p>\n<h5>The typical features</h5>\n<ul>\n<li>Elapsed time between the previous level group and the current level group.</li>\n<li>Elapsed time and index count between flag events.</li>\n<li>Prediction probabilities for previous questions.</li>\n<li>Sum of the most recent M (M=1,2,…) prediction probabilities.</li>\n</ul>\n<p>Flag events are events that must be passed during game progression. We extract them with reference to jo_wilder's source code, game playing, and the log data of users who have got perfect scores. </p>\n<h5>Single Best Model(5folds XGBoost)</h5>\n<p>CV: 0.702, Public LB: 0.700, Private LB: 0.701</p>\n<h2>Per Level Group Classifiers</h2>\n<p>In order to allow the level group models to utilize information from previous level groups, we first split the training data by:</p>\n<pre><code>\ndf1 = df.(pl.col() == )\ndf2 = df.((pl.col() == ) | (pl.col() == ))\ndf3 = df\n\ndf1 = df1.sort(pl.col(), pl.col())\ndf2 = df2.sort(pl.col(), pl.col())\ndf3 = df3.sort(pl.col(), pl.col())\n</code></pre>\n<p>Feature selection is then applied after feature engineering.</p>\n<h5>Features Engineering</h5>\n<ul>\n<li>Room distance and screen distance</li>\n</ul>\n<pre><code>    (pl.col() - pl.col().shift()).over([]).().alias(),\n    (pl.col() - pl.col().shift()).over([]).().alias(),    \n    (pl.col() - pl.col().shift()).over([]).().alias(),\n    (pl.col() - pl.col().shift()).over([]).().alias(),    \n</code></pre>\n<ul>\n<li>Final scene, checkpoint and answer time</li>\n</ul>\n<p>By playing the game manually, we know that students are only taking the quiz at the end of each level. The shorter time they used to finish the session of answering questions, the higher probability that they answered those questions correctly. Captured by features like:</p>\n<pre><code>                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                (pl.col().(pl.col() == ).() - pl.col().(pl.col() == ).()).alias()\n</code></pre>\n<ul>\n<li>Unnecessary moves</li>\n</ul>\n<p>Also from the experience of playing the game, we believe that there are many people who have played the game for more than one time. Would be great if we are have some feature to identify these players</p>\n<pre><code>unnecessary_data_values = {}\n q  ():\n    unnecessary_data_values[q] = {}\n     feature_type  [, , ]:\n        unnecessary_data_values[q][feature_type] = []\n        unique_values = (df.((pl.col() == q))[feature_type].unique())\n\n         val  unique_values:\n             df.((pl.col() == q) &amp; (pl.col(feature_type) == val))[].n_unique() &lt; :\n                unused_data_values[q][feature_type].append(val)\n</code></pre>\n<p>If they are playing for the first time, they likely have many unnecessary moves. Then we calculate the time / actions they have spent of these moves</p>\n<pre><code>     col  []:\n\n        aggs.extend([\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).count().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).count().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).count().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).().alias()  level  level_feature],\n        ])\n</code></pre>\n<ul>\n<li>Time / actions spent on tasks</li>\n</ul>\n<p>Another class of feature to filter out experienced  players is to measure how fast they finish the tasks before the quiz in every level group. For example the first task of the game is to find the notebook, our hypothesis is that an experienced player would spend less time and actions to finish it. And they have a higher chance to answer the quiz questions correctly.</p>\n<p>Two examples for chapter 1</p>\n<pre><code>                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias()\n</code></pre>\n<h5>Feature Selection</h5>\n<p>The selection is based on Catboost feature importance over the Catboost feature importance with shuffled labels. (Which is the idea of Null Importances <a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances</a>)</p>\n<ol>\n<li>Compute Catboost feature importance with the entire training data.</li>\n<li>Shuffle the training data labels and obtain the importance again for N times.</li>\n<li>Compute the final importance by the base importance divided by mean random importance.</li>\n<li>We then use <code>gp_minimize</code> to search for the best feature size based on 5-fold cross-validation.<br>\nIn the end, we have 233, 647, 693 features respectively for each level group.</li>\n</ol>\n<p>With Catboost 5-fold CV out of fold F1: 0.7019<br>\nWith Xgboost 5-fold CV out of fold F1: 0.7021</p>\n<p>Then feature engineering is applied to each of the data frames above. And the transformed data frames are used to train our level group models.</p>\n<h5>18-in-1 Classifiers</h5>\n<p>To train the 18-questions-in-1 classifier, we further concat the above 3 data frames together to form a large data frame.</p>\n<pre><code>\n\nall_df = pd.concat([\n    df1[FEATURES1 + []],\n    df2[FEATURES2 + []],\n    df3[FEATURES3 + []],\n], axis=)\n</code></pre>\n<p>This mega concatenation creates many null values because some features only exist in a particular level group. That’s why when building the features for this 18-in-1 classifier:<br>\nFirst, reuse the feature selection results from the Per Level Group case.<br>\nRerun feature selection again after the mega concatenation</p>\n<p>With Catboost 5-fold CV out of fold F1: 0.7002<br>\nWith Xgboost 5-fold CV out of fold F1: 0.7007</p>\n<h1>Neural Network</h1>\n<p>Model: Transformer + LSTM</p>\n<p>The pipeline of our NN is based on this public notebook: <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features</a></p>\n<h5>Numerical input:</h5>\n<ul>\n<li>np.log1p( elapsed_time_diff )</li>\n</ul>\n<h5>Categorical inputs:</h5>\n<ul>\n<li>event_comb, room_fqid, page, text_fqid, level</li>\n</ul>\n<h5>Transformer part (3 variants):</h5>\n<ul>\n<li>Type A: Conformer like transformer, with last query attention (<a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/218318\" target=\"_blank\">https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/218318</a>)</li>\n<li>Type B: Conformer like transformer, with last query attention</li>\n<li>Type C: Standard transformer</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fb48403960b8dc8cceff7d92e3d18a1bd%2FPSP%20NN.png?generation=1688066794674184&amp;alt=media\" alt=\"\"></p>\n<h5>Post Transformer LSTM:</h5>\n<ul>\n<li>1 Bidirectional LSTM + 1 LSTM layer</li>\n</ul>\n<h5>Pooling method:</h5>\n<ul>\n<li>Concat of sum, std, max, last</li>\n</ul>\n<h5>Training method:</h5>\n<ol>\n<li>As mentioned in the previous section, we train the model with multi-label, and there are two variants:<br>\na. One model per level group<br>\nb. Same model for ALL level groups</li>\n<li>We find that combining models trained with different settings can improve both the CV and public LB.</li>\n<li>Additional data was used for training, it improves both CV and public LB for NN</li>\n</ol>\n<h3>Best NN only ensemble (5 NN with different settings):</h3>\n<ul>\n<li>CV: 0.7028, Public LB: 0.701, Private LB: 0.704</li>\n<li>It turns out that NN doesn’t perform very well in Public LB, but does well in Private LB.</li>\n</ul>\n<h1>Submission Selection</h1>\n<p>We selected a submission with the highest LB, a submission with the highest CV, and a submission with a target on a reasonably high CV and a high variety of methods/models. </p>\n<p>Our best-selected sub is an ensemble of </p>\n<ul>\n<li>One level group Catboost, one 18-in1 Catboost, two 18-in1 Xgboost, and three NN.</li>\n<li>The NNs we selected are Type A per level group, Type B per level group, and Type C ALL level groups. This combination gives good diversity to the final ensemble.</li>\n<li>We ensemble GBT models and NN models on oof data separately with 2 standalone Logistic regression models, then combined them with GBT:NN = 6:4 ratio. </li>\n<li>The manual weighting in combining GBT and NN results is due to NN not performing well in public LB, so we didn't have enough confidence to give too much weight to our NN models as discussed below.    </li>\n</ul>\n<p>Best selected ensemble:</p>\n<ul>\n<li>CV: 0.7046, Public LB: 0.706, Private LB: 0.704</li>\n</ul>\n<h1>0.705 subs that we haven’t picked</h1>\n<p>We have three 705 private score submissions that are not selected. Our best-selected subs ranked 13th in all of our subs in terms of private score.</p>\n<p>Among these 705 private subs:</p>\n<ul>\n<li>Per-question GBT model + Group Level GBT model gives us the 705 private score, but not a high ensemble CV score. </li>\n<li>Per-level-group GBT + NN models with Logistic regression ensemble gives us the 705 private score, but not a high public score.</li>\n</ul>\n<h1>Observations:</h1>\n<ol>\n<li>NN models perform well in CV and private but very poorly in public, while GBT models fit the public so well, It is very strange…</li>\n<li>Single-question GBT models makes a lower CV ensemble but perform quite ok in both public and private</li>\n</ol>",
  "messages": [
    {
      "id": "2323273",
      "postDate": "06/29/2023 19:48:04",
      "content": "<p>Thanks a lot to the hosts of the competition and my teammates ( <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a>, <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a>, and <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a>). I am thrilled that <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> and I will become GM, <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a> is one step closer to becoming a Master, and <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> will get his second gold medal after this competition!</p>\n<p>Here We will explain our overall solution, <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> also provided additional explanation of the parts he worked on: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274</a></p>\n<h1>Classification Task Formulation</h1>\n<p>In this competition, we are asked to predict 18 values for each session. Each session contains 3 level groups. There are multiple ways to model this.</p>\n<ol>\n<li>18 binary classifiers</li>\n<li>3 Level group classifiers, each one can be<br>\na. A multi-label classifier that predicts all values within a level group<br>\nb. A binary classifier that takes “question index” as a feature within a level group</li>\n<li>1 classifier that is<br>\na. A multi-label classifier that predicts 18 values within a session<br>\nb. A binary classifier that takes “question index” as a feature within a session</li>\n</ol>\n<p>For Gradient boosted tree models, method 2b &gt; method 3b &gt; method 1. Method 2a and 3a are ignored because training the multi-label task is a lot slower with Gradient boosted tree models.</p>\n<p>For NN models, we focus on the method 2a and 3a, because</p>\n<ul>\n<li>These 2 methods are not well handled by tree models</li>\n<li>Multi-label learning makes more sense, because of the F1 score setting of this competition. (some posts discuss we should not optimize for 1 question).</li>\n<li>Multi-label NN models are faster to train and infer.</li>\n</ul>\n<h1>Additional dataset generated from the raw data</h1>\n<p>We create an additional dataset from the raw data, it contains 11343 complete sessions.<br>\nThis dataset boosts the CV scores for GBT models  by about +0.001~2, but there is not much effect on the public and private scores, and it has both positive and negative outcomes.<br>\nHowever, it works very well for NN models, we see +0.002 improvement in both CV and public scores.</p>\n<h1>Validation</h1>\n<p>We are using 5-fold GroupKFold on session_id so that there won’t be any seen sessions in the validation set. Also we didn’t include additional data in our validation set.</p>\n<h1>Gradient Boosted Tree</h1>\n<p>Per question classifier is handled by <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a>, Per level, and All-in-1 classifier is handled by  <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a> <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a>. That’s why there are some inconsistencies in the data preprocessing steps, such as sort by index vs sort by time.</p>\n<h2>Per Question Classifiers</h2>\n<p>We create features for each level group and sorted by index. The features and the sorting methods differ from other models.</p>\n<pre><code>\ndf1 = df.(pl.col() == )\ndf2 = df.(pl.col() == )\ndf3 = df.(pl.col() == )\n\ndf1 = df1.sort(pl.col(), pl.col())\ndf2 = df2.sort(pl.col(), pl.col())\ndf3 = df3.sort(pl.col(), pl.col())\n</code></pre>\n<h5>The number of features:</h5>\n<ul>\n<li>Level group 0-4: 1,000 features</li>\n<li>Level group 5-12: 2,000 features</li>\n<li>Level group 13-22: 2,400 features</li>\n</ul>\n<h5>Feature Selection</h5>\n<p>We try feature selection with out-of-folds but the public scores tend to decrease, so we don’t select features about this model in the final submission.</p>\n<h5>The typical features</h5>\n<ul>\n<li>Elapsed time between the previous level group and the current level group.</li>\n<li>Elapsed time and index count between flag events.</li>\n<li>Prediction probabilities for previous questions.</li>\n<li>Sum of the most recent M (M=1,2,…) prediction probabilities.</li>\n</ul>\n<p>Flag events are events that must be passed during game progression. We extract them with reference to jo_wilder's source code, game playing, and the log data of users who have got perfect scores. </p>\n<h5>Single Best Model(5folds XGBoost)</h5>\n<p>CV: 0.702, Public LB: 0.700, Private LB: 0.701</p>\n<h2>Per Level Group Classifiers</h2>\n<p>In order to allow the level group models to utilize information from previous level groups, we first split the training data by:</p>\n<pre><code>\ndf1 = df.(pl.col() == )\ndf2 = df.((pl.col() == ) | (pl.col() == ))\ndf3 = df\n\ndf1 = df1.sort(pl.col(), pl.col())\ndf2 = df2.sort(pl.col(), pl.col())\ndf3 = df3.sort(pl.col(), pl.col())\n</code></pre>\n<p>Feature selection is then applied after feature engineering.</p>\n<h5>Features Engineering</h5>\n<ul>\n<li>Room distance and screen distance</li>\n</ul>\n<pre><code>    (pl.col() - pl.col().shift()).over([]).().alias(),\n    (pl.col() - pl.col().shift()).over([]).().alias(),    \n    (pl.col() - pl.col().shift()).over([]).().alias(),\n    (pl.col() - pl.col().shift()).over([]).().alias(),    \n</code></pre>\n<ul>\n<li>Final scene, checkpoint and answer time</li>\n</ul>\n<p>By playing the game manually, we know that students are only taking the quiz at the end of each level. The shorter time they used to finish the session of answering questions, the higher probability that they answered those questions correctly. Captured by features like:</p>\n<pre><code>                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                (pl.col().(pl.col() == ).() - pl.col().(pl.col() == ).()).alias()\n</code></pre>\n<ul>\n<li>Unnecessary moves</li>\n</ul>\n<p>Also from the experience of playing the game, we believe that there are many people who have played the game for more than one time. Would be great if we are have some feature to identify these players</p>\n<pre><code>unnecessary_data_values = {}\n q  ():\n    unnecessary_data_values[q] = {}\n     feature_type  [, , ]:\n        unnecessary_data_values[q][feature_type] = []\n        unique_values = (df.((pl.col() == q))[feature_type].unique())\n\n         val  unique_values:\n             df.((pl.col() == q) &amp; (pl.col(feature_type) == val))[].n_unique() &lt; :\n                unused_data_values[q][feature_type].append(val)\n</code></pre>\n<p>If they are playing for the first time, they likely have many unnecessary moves. Then we calculate the time / actions they have spent of these moves</p>\n<pre><code>     col  []:\n\n        aggs.extend([\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).count().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).count().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).count().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).().alias()  level  level_feature],\n             *[pl.col(col).((pl.col() == level) &amp; (pl.col().is_in(unused_data_values[level][]))).().alias()  level  level_feature],\n        ])\n</code></pre>\n<ul>\n<li>Time / actions spent on tasks</li>\n</ul>\n<p>Another class of feature to filter out experienced  players is to measure how fast they finish the tasks before the quiz in every level group. For example the first task of the game is to find the notebook, our hypothesis is that an experienced player would spend less time and actions to finish it. And they have a higher chance to answer the quiz questions correctly.</p>\n<p>Two examples for chapter 1</p>\n<pre><code>                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias(),\n                pl.col().((pl.col() == ) | (pl.col() == )).apply( s: s.() - s.()).alias()\n</code></pre>\n<h5>Feature Selection</h5>\n<p>The selection is based on Catboost feature importance over the Catboost feature importance with shuffled labels. (Which is the idea of Null Importances <a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances</a>)</p>\n<ol>\n<li>Compute Catboost feature importance with the entire training data.</li>\n<li>Shuffle the training data labels and obtain the importance again for N times.</li>\n<li>Compute the final importance by the base importance divided by mean random importance.</li>\n<li>We then use <code>gp_minimize</code> to search for the best feature size based on 5-fold cross-validation.<br>\nIn the end, we have 233, 647, 693 features respectively for each level group.</li>\n</ol>\n<p>With Catboost 5-fold CV out of fold F1: 0.7019<br>\nWith Xgboost 5-fold CV out of fold F1: 0.7021</p>\n<p>Then feature engineering is applied to each of the data frames above. And the transformed data frames are used to train our level group models.</p>\n<h5>18-in-1 Classifiers</h5>\n<p>To train the 18-questions-in-1 classifier, we further concat the above 3 data frames together to form a large data frame.</p>\n<pre><code>\n\nall_df = pd.concat([\n    df1[FEATURES1 + []],\n    df2[FEATURES2 + []],\n    df3[FEATURES3 + []],\n], axis=)\n</code></pre>\n<p>This mega concatenation creates many null values because some features only exist in a particular level group. That’s why when building the features for this 18-in-1 classifier:<br>\nFirst, reuse the feature selection results from the Per Level Group case.<br>\nRerun feature selection again after the mega concatenation</p>\n<p>With Catboost 5-fold CV out of fold F1: 0.7002<br>\nWith Xgboost 5-fold CV out of fold F1: 0.7007</p>\n<h1>Neural Network</h1>\n<p>Model: Transformer + LSTM</p>\n<p>The pipeline of our NN is based on this public notebook: <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features</a></p>\n<h5>Numerical input:</h5>\n<ul>\n<li>np.log1p( elapsed_time_diff )</li>\n</ul>\n<h5>Categorical inputs:</h5>\n<ul>\n<li>event_comb, room_fqid, page, text_fqid, level</li>\n</ul>\n<h5>Transformer part (3 variants):</h5>\n<ul>\n<li>Type A: Conformer like transformer, with last query attention (<a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/218318\" target=\"_blank\">https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/218318</a>)</li>\n<li>Type B: Conformer like transformer, with last query attention</li>\n<li>Type C: Standard transformer</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fb48403960b8dc8cceff7d92e3d18a1bd%2FPSP%20NN.png?generation=1688066794674184&amp;alt=media\" alt=\"\"></p>\n<h5>Post Transformer LSTM:</h5>\n<ul>\n<li>1 Bidirectional LSTM + 1 LSTM layer</li>\n</ul>\n<h5>Pooling method:</h5>\n<ul>\n<li>Concat of sum, std, max, last</li>\n</ul>\n<h5>Training method:</h5>\n<ol>\n<li>As mentioned in the previous section, we train the model with multi-label, and there are two variants:<br>\na. One model per level group<br>\nb. Same model for ALL level groups</li>\n<li>We find that combining models trained with different settings can improve both the CV and public LB.</li>\n<li>Additional data was used for training, it improves both CV and public LB for NN</li>\n</ol>\n<h3>Best NN only ensemble (5 NN with different settings):</h3>\n<ul>\n<li>CV: 0.7028, Public LB: 0.701, Private LB: 0.704</li>\n<li>It turns out that NN doesn’t perform very well in Public LB, but does well in Private LB.</li>\n</ul>\n<h1>Submission Selection</h1>\n<p>We selected a submission with the highest LB, a submission with the highest CV, and a submission with a target on a reasonably high CV and a high variety of methods/models. </p>\n<p>Our best-selected sub is an ensemble of </p>\n<ul>\n<li>One level group Catboost, one 18-in1 Catboost, two 18-in1 Xgboost, and three NN.</li>\n<li>The NNs we selected are Type A per level group, Type B per level group, and Type C ALL level groups. This combination gives good diversity to the final ensemble.</li>\n<li>We ensemble GBT models and NN models on oof data separately with 2 standalone Logistic regression models, then combined them with GBT:NN = 6:4 ratio. </li>\n<li>The manual weighting in combining GBT and NN results is due to NN not performing well in public LB, so we didn't have enough confidence to give too much weight to our NN models as discussed below.    </li>\n</ul>\n<p>Best selected ensemble:</p>\n<ul>\n<li>CV: 0.7046, Public LB: 0.706, Private LB: 0.704</li>\n</ul>\n<h1>0.705 subs that we haven’t picked</h1>\n<p>We have three 705 private score submissions that are not selected. Our best-selected subs ranked 13th in all of our subs in terms of private score.</p>\n<p>Among these 705 private subs:</p>\n<ul>\n<li>Per-question GBT model + Group Level GBT model gives us the 705 private score, but not a high ensemble CV score. </li>\n<li>Per-level-group GBT + NN models with Logistic regression ensemble gives us the 705 private score, but not a high public score.</li>\n</ul>\n<h1>Observations:</h1>\n<ol>\n<li>NN models perform well in CV and private but very poorly in public, while GBT models fit the public so well, It is very strange…</li>\n<li>Single-question GBT models makes a lower CV ensemble but perform quite ok in both public and private</li>\n</ol>",
      "rawMarkdown": "Thanks a lot to the hosts of the competition and my teammates ( @kingychiu, @tangtunyu, and @yyykrk). I am thrilled that @kingychiu and I will become GM, @tangtunyu is one step closer to becoming a Master, and @yyykrk will get his second gold medal after this competition!\n\nHere We will explain our overall solution, @yyykrk also provided additional explanation of the parts he worked on: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274\n\n# Classification Task Formulation\nIn this competition, we are asked to predict 18 values for each session. Each session contains 3 level groups. There are multiple ways to model this.\n1. 18 binary classifiers\n2. 3 Level group classifiers, each one can be\na. A multi-label classifier that predicts all values within a level group\nb. A binary classifier that takes “question index” as a feature within a level group\n3. 1 classifier that is\na. A multi-label classifier that predicts 18 values within a session\nb. A binary classifier that takes “question index” as a feature within a session\n\nFor Gradient boosted tree models, method 2b > method 3b > method 1. Method 2a and 3a are ignored because training the multi-label task is a lot slower with Gradient boosted tree models.\n\nFor NN models, we focus on the method 2a and 3a, because\n- These 2 methods are not well handled by tree models\n- Multi-label learning makes more sense, because of the F1 score setting of this competition. (some posts discuss we should not optimize for 1 question).\n- Multi-label NN models are faster to train and infer.\n\n\n# Additional dataset generated from the raw data\nWe create an additional dataset from the raw data, it contains 11343 complete sessions.\nThis dataset boosts the CV scores for GBT models  by about +0.001~2, but there is not much effect on the public and private scores, and it has both positive and negative outcomes.\nHowever, it works very well for NN models, we see +0.002 improvement in both CV and public scores.\n\n\n# Validation\nWe are using 5-fold GroupKFold on session_id so that there won’t be any seen sessions in the validation set. Also we didn’t include additional data in our validation set.\n\n\n# Gradient Boosted Tree\nPer question classifier is handled by @yyykrk, Per level, and All-in-1 classifier is handled by  @tangtunyu @kingychiu. That’s why there are some inconsistencies in the data preprocessing steps, such as sort by index vs sort by time.\n\n## Per Question Classifiers\nWe create features for each level group and sorted by index. The features and the sorting methods differ from other models.\n\n```python\n# Code in polars\ndf1 = df.filter(pl.col(\"level_group\") == \"0-4\")\ndf2 = df.filter(pl.col(\"level_group\") == \"5-12\")\ndf3 = df.filter(pl.col(\"level_group\") == \"13-22\")\n\ndf1 = df1.sort(pl.col(\"session_id\"), pl.col(\"index\"))\ndf2 = df2.sort(pl.col(\"session_id\"), pl.col(\"index\"))\ndf3 = df3.sort(pl.col(\"session_id\"), pl.col(\"index\"))\n```\n\n##### The number of features:\n- Level group 0-4: 1,000 features\n- Level group 5-12: 2,000 features\n- Level group 13-22: 2,400 features\n\n##### Feature Selection\nWe try feature selection with out-of-folds but the public scores tend to decrease, so we don’t select features about this model in the final submission.\n\n##### The typical features\n- Elapsed time between the previous level group and the current level group.\n- Elapsed time and index count between flag events.\n- Prediction probabilities for previous questions.\n- Sum of the most recent M (M=1,2,...) prediction probabilities.\n\nFlag events are events that must be passed during game progression. We extract them with reference to jo_wilder's source code, game playing, and the log data of users who have got perfect scores. \n\n##### Single Best Model(5folds XGBoost)\nCV: 0.702, Public LB: 0.700, Private LB: 0.701\n\n\n## Per Level Group Classifiers\nIn order to allow the level group models to utilize information from previous level groups, we first split the training data by:\n```python\n# Code in polars\ndf1 = df.filter(pl.col(\"level_group\") == \"0-4\")\ndf2 = df.filter((pl.col(\"level_group\") == \"0-4\") | (pl.col(\"level_group\") == \"5-12\"))\ndf3 = df\n\ndf1 = df1.sort(pl.col(\"session_id\"), pl.col(\"elapsed_time\"))\ndf2 = df2.sort(pl.col(\"session_id\"), pl.col(\"elapsed_time\"))\ndf3 = df3.sort(pl.col(\"session_id\"), pl.col(\"elapsed_time\"))\n```\n\nFeature selection is then applied after feature engineering.\n\n##### Features Engineering\n\n- Room distance and screen distance\n\n```python\n    (pl.col(\"room_coor_x\") - pl.col(\"room_coor_x\").shift(1)).over([\"session_id\"]).pow(2).alias(\"room_coor_x_dis\"),\n    (pl.col(\"room_coor_y\") - pl.col(\"room_coor_y\").shift(1)).over([\"session_id\"]).pow(2).alias(\"room_coor_y_dis\"),    \n    (pl.col(\"screen_coor_x\") - pl.col(\"screen_coor_x\").shift(1)).over([\"session_id\"]).pow(2).alias(\"screen_coor_x_dis\"),\n    (pl.col(\"screen_coor_y\") - pl.col(\"screen_coor_y\").shift(1)).over([\"session_id\"]).pow(2).alias(\"screen_coor_y_dis\"),    \n\n```\n\n- Final scene, checkpoint and answer time\n\nBy playing the game manually, we know that students are only taking the quiz at the end of each level. The shorter time they used to finish the session of answering questions, the higher probability that they answered those questions correctly. Captured by features like:\n\n```python\n                pl.col(\"index\").filter((pl.col(\"fqid\") == \"chap2_finale_c\") | (pl.col(\"event_name\") == \"checkpoint\")).apply(lambda s: s.max() - s.min()).alias(\"chap2_answer_indexCount\"),\n                (pl.col(\"elapsed_time\").filter(pl.col(\"level_group\") == \"5-12\").min() - pl.col(\"elapsed_time\").filter(pl.col(\"level_group\") == \"0-4\").max()).alias(\"chap1_answer_time\")\n\n```\n\n- Unnecessary moves\n\nAlso from the experience of playing the game, we believe that there are many people who have played the game for more than one time. Would be great if we are have some feature to identify these players\n\n```python\nunnecessary_data_values = {}\nfor q in range(23):\n    unnecessary_data_values[q] = {}\n    for feature_type in ['text', 'fqid', 'text_fqid']:\n        unnecessary_data_values[q][feature_type] = []\n        unique_values = list(df.filter((pl.col(\"level\") == q))[feature_type].unique())\n        \n        for val in unique_values:\n            if df.filter((pl.col(\"level\") == q) & (pl.col(feature_type) == val))['session_id'].n_unique() < 23000:\n                unused_data_values[q][feature_type].append(val)\n```\n\nIf they are playing for the first time, they likely have many unnecessary moves. Then we calculate the time / actions they have spent of these moves\n\n```python\n    for col in ['elapsed_time_diff']:\n\n        aggs.extend([\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text\").is_in(unused_data_values[level][\"text\"]))).count().alias(f\"level_{level}_unused_text_{col}_counts\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"fqid\").is_in(unused_data_values[level][\"fqid\"]))).count().alias(f\"level_{level}_unused_fqid_{col}_counts\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text_fqid\").is_in(unused_data_values[level][\"text_fqid\"]))).count().alias(f\"level_{level}_unused_text_fqid_{col}_counts\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text\").is_in(unused_data_values[level][\"text\"]))).sum().alias(f\"level_{level}_unused_text_{col}_sum\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"fqid\").is_in(unused_data_values[level][\"fqid\"]))).sum().alias(f\"level_{level}_unused_fqid_{col}_sum\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text_fqid\").is_in(unused_data_values[level][\"text_fqid\"]))).sum().alias(f\"level_{level}_unused_text_fqid_{col}_sum\") for level in level_feature],\n        ])\n    \n```\n\n- Time / actions spent on tasks\n\nAnother class of feature to filter out experienced  players is to measure how fast they finish the tasks before the quiz in every level group. For example the first task of the game is to find the notebook, our hypothesis is that an experienced player would spend less time and actions to finish it. And they have a higher chance to answer the quiz questions correctly.\n\nTwo examples for chapter 1\n```python\n                pl.col(\"elapsed_time\").filter((pl.col(\"text\") == \"Now where did I put my notebook?\") | (pl.col(\"text\") == \"Found it!\")).apply(lambda s: s.max() - s.min()).alias(\"find_notebook_duration\"),\n                pl.col(\"index\").filter((pl.col(\"text\") == \"Now where did I put my notebook?\") | (pl.col(\"text\") == \"Found it!\")).apply(lambda s: s.max() - s.min()).alias(\"find_notebook_indexCount\"),\n                pl.col(\"elapsed_time\").filter((pl.col(\"text\") == \"Found it!\") | (pl.col(\"text\") == \"Let's get started. The Wisconsin Wonders exhibit opens tomorrow!\")).apply(lambda s: s.max() - s.min()).alias(\"go_upstairs_duration\"),\n                pl.col(\"index\").filter((pl.col(\"text\") == \"Found it!\") | (pl.col(\"text\") == \"Let's get started. The Wisconsin Wonders exhibit opens tomorrow!\")).apply(lambda s: s.max() - s.min()).alias(\"go_upstairs_events\")\n```\n\n##### Feature Selection\nThe selection is based on Catboost feature importance over the Catboost feature importance with shuffled labels. (Which is the idea of Null Importances https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances)\n1. Compute Catboost feature importance with the entire training data.\n2. Shuffle the training data labels and obtain the importance again for N times.\n3. Compute the final importance by the base importance divided by mean random importance.\n4. We then use `gp_minimize` to search for the best feature size based on 5-fold cross-validation.\nIn the end, we have 233, 647, 693 features respectively for each level group.\n\nWith Catboost 5-fold CV out of fold F1: 0.7019\nWith Xgboost 5-fold CV out of fold F1: 0.7021\n\nThen feature engineering is applied to each of the data frames above. And the transformed data frames are used to train our level group models.\n\n##### 18-in-1 Classifiers\nTo train the 18-questions-in-1 classifier, we further concat the above 3 data frames together to form a large data frame.\n```python\n# Code in pandas\n\nall_df = pd.concat([\n    df1[FEATURES1 + [\"q\"]],\n    df2[FEATURES2 + [\"q\"]],\n    df3[FEATURES3 + [\"q\"]],\n], axis=0)\n```\nThis mega concatenation creates many null values because some features only exist in a particular level group. That’s why when building the features for this 18-in-1 classifier:\nFirst, reuse the feature selection results from the Per Level Group case.\nRerun feature selection again after the mega concatenation\n\nWith Catboost 5-fold CV out of fold F1: 0.7002\nWith Xgboost 5-fold CV out of fold F1: 0.7007\n\n\n# Neural Network\n\nModel: Transformer + LSTM\n\nThe pipeline of our NN is based on this public notebook: https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\n\n##### Numerical input:\n- np.log1p( elapsed_time_diff )\n##### Categorical inputs:\n- event_comb, room_fqid, page, text_fqid, level\n\n\n##### Transformer part (3 variants):\n- Type A: Conformer like transformer, with last query attention (https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/218318)\n- Type B: Conformer like transformer, with last query attention\n- Type C: Standard transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fb48403960b8dc8cceff7d92e3d18a1bd%2FPSP%20NN.png?generation=1688066794674184&alt=media)\n\n##### Post Transformer LSTM:\n- 1 Bidirectional LSTM + 1 LSTM layer\n\n\n##### Pooling method:\n- Concat of sum, std, max, last\n\n##### Training method:\n1. As mentioned in the previous section, we train the model with multi-label, and there are two variants:\na. One model per level group\nb. Same model for ALL level groups\n2. We find that combining models trained with different settings can improve both the CV and public LB.\n3. Additional data was used for training, it improves both CV and public LB for NN\n\n\n### Best NN only ensemble (5 NN with different settings):\n- CV: 0.7028, Public LB: 0.701, Private LB: 0.704\n- It turns out that NN doesn’t perform very well in Public LB, but does well in Private LB.\n\n\n# Submission Selection\nWe selected a submission with the highest LB, a submission with the highest CV, and a submission with a target on a reasonably high CV and a high variety of methods/models. \n\nOur best-selected sub is an ensemble of \n- One level group Catboost, one 18-in1 Catboost, two 18-in1 Xgboost, and three NN.\n- The NNs we selected are Type A per level group, Type B per level group, and Type C ALL level groups. This combination gives good diversity to the final ensemble.\n- We ensemble GBT models and NN models on oof data separately with 2 standalone Logistic regression models, then combined them with GBT:NN = 6:4 ratio. \n- The manual weighting in combining GBT and NN results is due to NN not performing well in public LB, so we didn't have enough confidence to give too much weight to our NN models as discussed below.\t\n\nBest selected ensemble:\n- CV: 0.7046, Public LB: 0.706, Private LB: 0.704\n\n\n# 0.705 subs that we haven’t picked\nWe have three 705 private score submissions that are not selected. Our best-selected subs ranked 13th in all of our subs in terms of private score.\n\nAmong these 705 private subs:\n- Per-question GBT model + Group Level GBT model gives us the 705 private score, but not a high ensemble CV score. \n- Per-level-group GBT + NN models with Logistic regression ensemble gives us the 705 private score, but not a high public score.\n\n\n# Observations:\n1. NN models perform well in CV and private but very poorly in public, while GBT models fit the public so well, It is very strange…\n2. Single-question GBT models makes a lower CV ensemble but perform quite ok in both public and private",
      "votes": null
    },
    {
      "id": "2323357",
      "postDate": "06/29/2023 22:08:52",
      "content": "<p>Great solution, thanks for sharing. Congrats on the result.</p>\n<p>There are many conclusions you made that we independently reached too. This is reasuusring in a way.</p>\n<p>I wonder how we can compare our submissions with your unselected 0.705. I hope we'll get more digits in score display soon.</p>",
      "rawMarkdown": "Great solution, thanks for sharing. Congrats on the result.\n\nThere are many conclusions you made that we independently reached too. This is reasuusring in a way.\n\nI wonder how we can compare our submissions with your unselected 0.705. I hope we'll get more digits in score display soon.",
      "votes": null
    },
    {
      "id": "2323624",
      "postDate": "06/30/2023 04:55:18",
      "content": "<p><a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> cool solution! Congrats with 3 place and gold medal!</p>",
      "rawMarkdown": "wimwim cool solution! Congrats with 3 place and gold medal!",
      "votes": null
    },
    {
      "id": "2323992",
      "postDate": "06/30/2023 09:53:21",
      "content": "<p>Congrats my friend <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> and <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> on becoming GM. Very strong solution!</p>",
      "rawMarkdown": "Congrats my friend @wimwim and @kingychiu on becoming GM. Very strong solution!",
      "votes": null
    },
    {
      "id": "2325039",
      "postDate": "07/01/2023 04:54:39",
      "content": "<p>Thank your for your posting!</p>\n<p>\"We try feature selection with out-of-folds but the public scores tend to decrease, \" means that you performed feature selection for each fold?</p>",
      "rawMarkdown": "Thank your for your posting!\n\n\"We try feature selection with out-of-folds but the public scores tend to decrease, \" means that you performed feature selection for each fold?",
      "votes": null
    },
    {
      "id": "2325048",
      "postDate": "07/01/2023 05:23:58",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> on both the strong finish and becoming GM. </p>\n<p>Interesting choice of NN architectures, more specifically last query attention. Also, for categorical inputs do you embed them? Any chance you would be able to make code public for these NN :)</p>",
      "rawMarkdown": "Congratulations @wimwim on both the strong finish and becoming GM. \n\nInteresting choice of NN architectures, more specifically last query attention. Also, for categorical inputs do you embed them? Any chance you would be able to make code public for these NN :)",
      "votes": null
    },
    {
      "id": "2325218",
      "postDate": "07/01/2023 07:21:14",
      "content": "<p>congratulation for both new GM, its such a great achievement..<br>\nthe solution is well written with detailed information</p>",
      "rawMarkdown": "congratulation for both new GM, its such a great achievement..\nthe solution is well written with detailed information",
      "votes": null
    },
    {
      "id": "2325369",
      "postDate": "07/01/2023 09:26:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> , thank you for the question.<br>\nYou are right, I embed categorical inputs, here are the nodebooks to train these NN.</p>\n<p>Type A per level group : <a href=\"https://www.kaggle.com/wimwim/psp-transformer-type-a-grp-level\" target=\"_blank\">https://www.kaggle.com/wimwim/psp-transformer-type-a-grp-level</a><br>\nType B per level group: <a href=\"https://www.kaggle.com/wimwim/psp-transformer-type-b-grp-level\" target=\"_blank\">https://www.kaggle.com/wimwim/psp-transformer-type-b-grp-level</a><br>\nType C ALL level groups: <a href=\"https://www.kaggle.com/wimwim/psp-transformer-type-c-18in1\" target=\"_blank\">https://www.kaggle.com/wimwim/psp-transformer-type-c-18in1</a></p>",
      "rawMarkdown": "Hi @tezdhar , thank you for the question.\nYou are right, I embed categorical inputs, here are the nodebooks to train these NN.\n\nType A per level group : https://www.kaggle.com/wimwim/psp-transformer-type-a-grp-level\nType B per level group: https://www.kaggle.com/wimwim/psp-transformer-type-b-grp-level\nType C ALL level groups: https://www.kaggle.com/wimwim/psp-transformer-type-c-18in1",
      "votes": null
    },
    {
      "id": "2325377",
      "postDate": "07/01/2023 09:33:20",
      "content": "<p>Yes, and you can find more information about this section in <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> 's solution write-up:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274</a></p>",
      "rawMarkdown": "Yes, and you can find more information about this section in @yyykrk 's solution write-up:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274",
      "votes": null
    },
    {
      "id": "2325380",
      "postDate": "07/01/2023 09:38:47",
      "content": "<p>Thank you my friend!</p>",
      "rawMarkdown": "Thank you my friend!",
      "votes": null
    },
    {
      "id": "2325486",
      "postDate": "07/01/2023 12:04:30",
      "content": "<p>The 3rd place solution used gradient boosted tree (GBT) models and neural network (NN) models to predict student performance. They improved the models through feature engineering and achieved a high F1 score of 0.702 in cross-validation.</p>",
      "rawMarkdown": "The 3rd place solution used gradient boosted tree (GBT) models and neural network (NN) models to predict student performance. They improved the models through feature engineering and achieved a high F1 score of 0.702 in cross-validation.",
      "votes": null
    },
    {
      "id": "2325503",
      "postDate": "07/01/2023 12:18:19",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "2325654",
      "postDate": "07/01/2023 13:59:19",
      "content": "<p><a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> Congratulations on your impressive performance! The detailed explanation of your solution and the different methods used for classification and feature engineering is quite insightful. Well done!</p>",
      "rawMarkdown": "wimwim Congratulations on your impressive performance! The detailed explanation of your solution and the different methods used for classification and feature engineering is quite insightful. Well done!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2323357,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/29/2023 22:08:52",
      "content": "<p>Great solution, thanks for sharing. Congrats on the result.</p>\n<p>There are many conclusions you made that we independently reached too. This is reasuusring in a way.</p>\n<p>I wonder how we can compare our submissions with your unselected 0.705. I hope we'll get more digits in score display soon.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323624,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "06/30/2023 04:55:18",
      "content": "<p><a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> cool solution! Congrats with 3 place and gold medal!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323992,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "06/30/2023 09:53:21",
      "content": "<p>Congrats my friend <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> and <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> on becoming GM. Very strong solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2325380,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "07/01/2023 09:38:47",
          "content": "<p>Thank you my friend!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2325039,
      "author_name": "aesoptacit",
      "author_url": "",
      "post_date": "07/01/2023 04:54:39",
      "content": "<p>Thank your for your posting!</p>\n<p>\"We try feature selection with out-of-folds but the public scores tend to decrease, \" means that you performed feature selection for each fold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2325377,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "07/01/2023 09:33:20",
          "content": "<p>Yes, and you can find more information about this section in <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> 's solution write-up:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2325503,
              "author_name": "aesoptacit",
              "author_url": "",
              "post_date": "07/01/2023 12:18:19",
              "content": "<p>Thank you so much!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2325048,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "07/01/2023 05:23:58",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> on both the strong finish and becoming GM. </p>\n<p>Interesting choice of NN architectures, more specifically last query attention. Also, for categorical inputs do you embed them? Any chance you would be able to make code public for these NN :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2325369,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "07/01/2023 09:26:29",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> , thank you for the question.<br>\nYou are right, I embed categorical inputs, here are the nodebooks to train these NN.</p>\n<p>Type A per level group : <a href=\"https://www.kaggle.com/wimwim/psp-transformer-type-a-grp-level\" target=\"_blank\">https://www.kaggle.com/wimwim/psp-transformer-type-a-grp-level</a><br>\nType B per level group: <a href=\"https://www.kaggle.com/wimwim/psp-transformer-type-b-grp-level\" target=\"_blank\">https://www.kaggle.com/wimwim/psp-transformer-type-b-grp-level</a><br>\nType C ALL level groups: <a href=\"https://www.kaggle.com/wimwim/psp-transformer-type-c-18in1\" target=\"_blank\">https://www.kaggle.com/wimwim/psp-transformer-type-c-18in1</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2325218,
      "author_name": "luthfiholic",
      "author_url": "",
      "post_date": "07/01/2023 07:21:14",
      "content": "<p>congratulation for both new GM, its such a great achievement..<br>\nthe solution is well written with detailed information</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2325486,
      "author_name": "poojach7611",
      "author_url": "",
      "post_date": "07/01/2023 12:04:30",
      "content": "<p>The 3rd place solution used gradient boosted tree (GBT) models and neural network (NN) models to predict student performance. They improved the models through feature engineering and achieved a high F1 score of 0.702 in cross-validation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2325654,
      "author_name": "akshayvyas02",
      "author_url": "",
      "post_date": "07/01/2023 13:59:19",
      "content": "<p><a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> Congratulations on your impressive performance! The detailed explanation of your solution and the different methods used for classification and feature engineering is quite insightful. Well done!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2323273": "Thanks a lot to the hosts of the competition and my teammates ( @kingychiu, @tangtunyu, and @yyykrk). I am thrilled that @kingychiu and I will become GM, @tangtunyu is one step closer to becoming a Master, and @yyykrk will get his second gold medal after this competition!\n\nHere We will explain our overall solution, @yyykrk also provided additional explanation of the parts he worked on: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274\n\n# Classification Task Formulation\nIn this competition, we are asked to predict 18 values for each session. Each session contains 3 level groups. There are multiple ways to model this.\n1. 18 binary classifiers\n2. 3 Level group classifiers, each one can be\na. A multi-label classifier that predicts all values within a level group\nb. A binary classifier that takes “question index” as a feature within a level group\n3. 1 classifier that is\na. A multi-label classifier that predicts 18 values within a session\nb. A binary classifier that takes “question index” as a feature within a session\n\nFor Gradient boosted tree models, method 2b > method 3b > method 1. Method 2a and 3a are ignored because training the multi-label task is a lot slower with Gradient boosted tree models.\n\nFor NN models, we focus on the method 2a and 3a, because\n- These 2 methods are not well handled by tree models\n- Multi-label learning makes more sense, because of the F1 score setting of this competition. (some posts discuss we should not optimize for 1 question).\n- Multi-label NN models are faster to train and infer.\n\n\n# Additional dataset generated from the raw data\nWe create an additional dataset from the raw data, it contains 11343 complete sessions.\nThis dataset boosts the CV scores for GBT models  by about +0.001~2, but there is not much effect on the public and private scores, and it has both positive and negative outcomes.\nHowever, it works very well for NN models, we see +0.002 improvement in both CV and public scores.\n\n\n# Validation\nWe are using 5-fold GroupKFold on session_id so that there won’t be any seen sessions in the validation set. Also we didn’t include additional data in our validation set.\n\n\n# Gradient Boosted Tree\nPer question classifier is handled by @yyykrk, Per level, and All-in-1 classifier is handled by  @tangtunyu @kingychiu. That’s why there are some inconsistencies in the data preprocessing steps, such as sort by index vs sort by time.\n\n## Per Question Classifiers\nWe create features for each level group and sorted by index. The features and the sorting methods differ from other models.\n\n```python\n# Code in polars\ndf1 = df.filter(pl.col(\"level_group\") == \"0-4\")\ndf2 = df.filter(pl.col(\"level_group\") == \"5-12\")\ndf3 = df.filter(pl.col(\"level_group\") == \"13-22\")\n\ndf1 = df1.sort(pl.col(\"session_id\"), pl.col(\"index\"))\ndf2 = df2.sort(pl.col(\"session_id\"), pl.col(\"index\"))\ndf3 = df3.sort(pl.col(\"session_id\"), pl.col(\"index\"))\n```\n\n##### The number of features:\n- Level group 0-4: 1,000 features\n- Level group 5-12: 2,000 features\n- Level group 13-22: 2,400 features\n\n##### Feature Selection\nWe try feature selection with out-of-folds but the public scores tend to decrease, so we don’t select features about this model in the final submission.\n\n##### The typical features\n- Elapsed time between the previous level group and the current level group.\n- Elapsed time and index count between flag events.\n- Prediction probabilities for previous questions.\n- Sum of the most recent M (M=1,2,...) prediction probabilities.\n\nFlag events are events that must be passed during game progression. We extract them with reference to jo_wilder's source code, game playing, and the log data of users who have got perfect scores. \n\n##### Single Best Model(5folds XGBoost)\nCV: 0.702, Public LB: 0.700, Private LB: 0.701\n\n\n## Per Level Group Classifiers\nIn order to allow the level group models to utilize information from previous level groups, we first split the training data by:\n```python\n# Code in polars\ndf1 = df.filter(pl.col(\"level_group\") == \"0-4\")\ndf2 = df.filter((pl.col(\"level_group\") == \"0-4\") | (pl.col(\"level_group\") == \"5-12\"))\ndf3 = df\n\ndf1 = df1.sort(pl.col(\"session_id\"), pl.col(\"elapsed_time\"))\ndf2 = df2.sort(pl.col(\"session_id\"), pl.col(\"elapsed_time\"))\ndf3 = df3.sort(pl.col(\"session_id\"), pl.col(\"elapsed_time\"))\n```\n\nFeature selection is then applied after feature engineering.\n\n##### Features Engineering\n\n- Room distance and screen distance\n\n```python\n    (pl.col(\"room_coor_x\") - pl.col(\"room_coor_x\").shift(1)).over([\"session_id\"]).pow(2).alias(\"room_coor_x_dis\"),\n    (pl.col(\"room_coor_y\") - pl.col(\"room_coor_y\").shift(1)).over([\"session_id\"]).pow(2).alias(\"room_coor_y_dis\"),    \n    (pl.col(\"screen_coor_x\") - pl.col(\"screen_coor_x\").shift(1)).over([\"session_id\"]).pow(2).alias(\"screen_coor_x_dis\"),\n    (pl.col(\"screen_coor_y\") - pl.col(\"screen_coor_y\").shift(1)).over([\"session_id\"]).pow(2).alias(\"screen_coor_y_dis\"),    \n\n```\n\n- Final scene, checkpoint and answer time\n\nBy playing the game manually, we know that students are only taking the quiz at the end of each level. The shorter time they used to finish the session of answering questions, the higher probability that they answered those questions correctly. Captured by features like:\n\n```python\n                pl.col(\"index\").filter((pl.col(\"fqid\") == \"chap2_finale_c\") | (pl.col(\"event_name\") == \"checkpoint\")).apply(lambda s: s.max() - s.min()).alias(\"chap2_answer_indexCount\"),\n                (pl.col(\"elapsed_time\").filter(pl.col(\"level_group\") == \"5-12\").min() - pl.col(\"elapsed_time\").filter(pl.col(\"level_group\") == \"0-4\").max()).alias(\"chap1_answer_time\")\n\n```\n\n- Unnecessary moves\n\nAlso from the experience of playing the game, we believe that there are many people who have played the game for more than one time. Would be great if we are have some feature to identify these players\n\n```python\nunnecessary_data_values = {}\nfor q in range(23):\n    unnecessary_data_values[q] = {}\n    for feature_type in ['text', 'fqid', 'text_fqid']:\n        unnecessary_data_values[q][feature_type] = []\n        unique_values = list(df.filter((pl.col(\"level\") == q))[feature_type].unique())\n        \n        for val in unique_values:\n            if df.filter((pl.col(\"level\") == q) & (pl.col(feature_type) == val))['session_id'].n_unique() < 23000:\n                unused_data_values[q][feature_type].append(val)\n```\n\nIf they are playing for the first time, they likely have many unnecessary moves. Then we calculate the time / actions they have spent of these moves\n\n```python\n    for col in ['elapsed_time_diff']:\n\n        aggs.extend([\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text\").is_in(unused_data_values[level][\"text\"]))).count().alias(f\"level_{level}_unused_text_{col}_counts\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"fqid\").is_in(unused_data_values[level][\"fqid\"]))).count().alias(f\"level_{level}_unused_fqid_{col}_counts\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text_fqid\").is_in(unused_data_values[level][\"text_fqid\"]))).count().alias(f\"level_{level}_unused_text_fqid_{col}_counts\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text\").is_in(unused_data_values[level][\"text\"]))).sum().alias(f\"level_{level}_unused_text_{col}_sum\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"fqid\").is_in(unused_data_values[level][\"fqid\"]))).sum().alias(f\"level_{level}_unused_fqid_{col}_sum\") for level in level_feature],\n             *[pl.col(col).filter((pl.col(\"level\") == level) & (pl.col(\"text_fqid\").is_in(unused_data_values[level][\"text_fqid\"]))).sum().alias(f\"level_{level}_unused_text_fqid_{col}_sum\") for level in level_feature],\n        ])\n    \n```\n\n- Time / actions spent on tasks\n\nAnother class of feature to filter out experienced  players is to measure how fast they finish the tasks before the quiz in every level group. For example the first task of the game is to find the notebook, our hypothesis is that an experienced player would spend less time and actions to finish it. And they have a higher chance to answer the quiz questions correctly.\n\nTwo examples for chapter 1\n```python\n                pl.col(\"elapsed_time\").filter((pl.col(\"text\") == \"Now where did I put my notebook?\") | (pl.col(\"text\") == \"Found it!\")).apply(lambda s: s.max() - s.min()).alias(\"find_notebook_duration\"),\n                pl.col(\"index\").filter((pl.col(\"text\") == \"Now where did I put my notebook?\") | (pl.col(\"text\") == \"Found it!\")).apply(lambda s: s.max() - s.min()).alias(\"find_notebook_indexCount\"),\n                pl.col(\"elapsed_time\").filter((pl.col(\"text\") == \"Found it!\") | (pl.col(\"text\") == \"Let's get started. The Wisconsin Wonders exhibit opens tomorrow!\")).apply(lambda s: s.max() - s.min()).alias(\"go_upstairs_duration\"),\n                pl.col(\"index\").filter((pl.col(\"text\") == \"Found it!\") | (pl.col(\"text\") == \"Let's get started. The Wisconsin Wonders exhibit opens tomorrow!\")).apply(lambda s: s.max() - s.min()).alias(\"go_upstairs_events\")\n```\n\n##### Feature Selection\nThe selection is based on Catboost feature importance over the Catboost feature importance with shuffled labels. (Which is the idea of Null Importances https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances)\n1. Compute Catboost feature importance with the entire training data.\n2. Shuffle the training data labels and obtain the importance again for N times.\n3. Compute the final importance by the base importance divided by mean random importance.\n4. We then use `gp_minimize` to search for the best feature size based on 5-fold cross-validation.\nIn the end, we have 233, 647, 693 features respectively for each level group.\n\nWith Catboost 5-fold CV out of fold F1: 0.7019\nWith Xgboost 5-fold CV out of fold F1: 0.7021\n\nThen feature engineering is applied to each of the data frames above. And the transformed data frames are used to train our level group models.\n\n##### 18-in-1 Classifiers\nTo train the 18-questions-in-1 classifier, we further concat the above 3 data frames together to form a large data frame.\n```python\n# Code in pandas\n\nall_df = pd.concat([\n    df1[FEATURES1 + [\"q\"]],\n    df2[FEATURES2 + [\"q\"]],\n    df3[FEATURES3 + [\"q\"]],\n], axis=0)\n```\nThis mega concatenation creates many null values because some features only exist in a particular level group. That’s why when building the features for this 18-in-1 classifier:\nFirst, reuse the feature selection results from the Per Level Group case.\nRerun feature selection again after the mega concatenation\n\nWith Catboost 5-fold CV out of fold F1: 0.7002\nWith Xgboost 5-fold CV out of fold F1: 0.7007\n\n\n# Neural Network\n\nModel: Transformer + LSTM\n\nThe pipeline of our NN is based on this public notebook: https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\n\n##### Numerical input:\n- np.log1p( elapsed_time_diff )\n##### Categorical inputs:\n- event_comb, room_fqid, page, text_fqid, level\n\n\n##### Transformer part (3 variants):\n- Type A: Conformer like transformer, with last query attention (https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/218318)\n- Type B: Conformer like transformer, with last query attention\n- Type C: Standard transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fb48403960b8dc8cceff7d92e3d18a1bd%2FPSP%20NN.png?generation=1688066794674184&alt=media)\n\n##### Post Transformer LSTM:\n- 1 Bidirectional LSTM + 1 LSTM layer\n\n\n##### Pooling method:\n- Concat of sum, std, max, last\n\n##### Training method:\n1. As mentioned in the previous section, we train the model with multi-label, and there are two variants:\na. One model per level group\nb. Same model for ALL level groups\n2. We find that combining models trained with different settings can improve both the CV and public LB.\n3. Additional data was used for training, it improves both CV and public LB for NN\n\n\n### Best NN only ensemble (5 NN with different settings):\n- CV: 0.7028, Public LB: 0.701, Private LB: 0.704\n- It turns out that NN doesn’t perform very well in Public LB, but does well in Private LB.\n\n\n# Submission Selection\nWe selected a submission with the highest LB, a submission with the highest CV, and a submission with a target on a reasonably high CV and a high variety of methods/models. \n\nOur best-selected sub is an ensemble of \n- One level group Catboost, one 18-in1 Catboost, two 18-in1 Xgboost, and three NN.\n- The NNs we selected are Type A per level group, Type B per level group, and Type C ALL level groups. This combination gives good diversity to the final ensemble.\n- We ensemble GBT models and NN models on oof data separately with 2 standalone Logistic regression models, then combined them with GBT:NN = 6:4 ratio. \n- The manual weighting in combining GBT and NN results is due to NN not performing well in public LB, so we didn't have enough confidence to give too much weight to our NN models as discussed below.\t\n\nBest selected ensemble:\n- CV: 0.7046, Public LB: 0.706, Private LB: 0.704\n\n\n# 0.705 subs that we haven’t picked\nWe have three 705 private score submissions that are not selected. Our best-selected subs ranked 13th in all of our subs in terms of private score.\n\nAmong these 705 private subs:\n- Per-question GBT model + Group Level GBT model gives us the 705 private score, but not a high ensemble CV score. \n- Per-level-group GBT + NN models with Logistic regression ensemble gives us the 705 private score, but not a high public score.\n\n\n# Observations:\n1. NN models perform well in CV and private but very poorly in public, while GBT models fit the public so well, It is very strange…\n2. Single-question GBT models makes a lower CV ensemble but perform quite ok in both public and private",
    "2323357": "Great solution, thanks for sharing. Congrats on the result.\n\nThere are many conclusions you made that we independently reached too. This is reasuusring in a way.\n\nI wonder how we can compare our submissions with your unselected 0.705. I hope we'll get more digits in score display soon.",
    "2323624": "wimwim cool solution! Congrats with 3 place and gold medal!",
    "2323992": "Congrats my friend @wimwim and @kingychiu on becoming GM. Very strong solution!",
    "2325039": "Thank your for your posting!\n\n\"We try feature selection with out-of-folds but the public scores tend to decrease, \" means that you performed feature selection for each fold?",
    "2325048": "Congratulations @wimwim on both the strong finish and becoming GM. \n\nInteresting choice of NN architectures, more specifically last query attention. Also, for categorical inputs do you embed them? Any chance you would be able to make code public for these NN :)",
    "2325218": "congratulation for both new GM, its such a great achievement..\nthe solution is well written with detailed information",
    "2325369": "Hi @tezdhar , thank you for the question.\nYou are right, I embed categorical inputs, here are the nodebooks to train these NN.\n\nType A per level group : https://www.kaggle.com/wimwim/psp-transformer-type-a-grp-level\nType B per level group: https://www.kaggle.com/wimwim/psp-transformer-type-b-grp-level\nType C ALL level groups: https://www.kaggle.com/wimwim/psp-transformer-type-c-18in1",
    "2325377": "Yes, and you can find more information about this section in @yyykrk 's solution write-up:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420274",
    "2325380": "Thank you my friend!",
    "2325486": "The 3rd place solution used gradient boosted tree (GBT) models and neural network (NN) models to predict student performance. They improved the models through feature engineering and achieved a high F1 score of 0.702 in cross-validation.",
    "2325503": "Thank you so much!",
    "2325654": "wimwim Congratulations on your impressive performance! The detailed explanation of your solution and the different methods used for classification and feature engineering is quite insightful. Well done!"
  },
  "source": "meta"
}