{
  "id": 420217,
  "title": "1st Place Solution",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420217",
  "author_name": "Bertrand P",
  "post_date": "2023-06-29T17:06:37.315000",
  "votes": 168,
  "comment_count": 73,
  "views": 0,
  "content": "<p>Unbelievable to write this!</p>\n<h1>Thanks!</h1>\n<p>As it is the usage, we first <strong>thank the host and Kaggle</strong>. These are special thanks because you and us have had a special link in this competition as we gave you more work by reporting data leaks. No doubt you tried to do your best. You are right to animate this community and to trust in it. You are part of it. Please take care of this community that is able to build so much together by sharing. As all of us you have made mistakes and we hope you will learn from them.</p>\n<p>We also want to <strong>thank all of you</strong>, Kagglers. We love and are grateful to be part of our group/community. Thanks for sharing and for the collective learning experience.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview</a>,</li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data</a>.</li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>Our solution is essentially a blend of a XGBoost and a NN models. Both heavily rely on duration that appeared to be a powerful leverage. Time was aggregated in different ways and combined with counts for the GBDT while it is transformed via a custom TimeEmbedding block based on 1D convolutions that produce a representation combined with user event representations for the NN.<br>\nRobustness and efficiency founded our work. XGBoost models were validated on 10 bags of 5 folds and features incorporated only if the mean of the CV of these 10 bags was greater than the level of noise we quantified while we opted for a majority/consensus strategy to build the NN, i.e. validate choices only if 4 of 5 folds were improved. The 3rd place of the efficiency LB was achieved with a lightweight NN accelerated via TF Lite.</p>\n<h1>Details of the submission</h1>\n<h2>Code</h2>\n<p>After publishing this write-up we decided to open our code: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332</a>.<br>\nIt is composed by several parts: <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training\" target=\"_blank\">how to train the XGBoost models</a>, how to <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining\" target=\"_blank\">pretrain</a> and <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">train</a> the NN models and the <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference\" target=\"_blank\">inference notebook</a> used to win this competition.</p>\n<h2>Data</h2>\n<p>Looking at the 1st data released showed that there aren't a lot of sessions so not a lot of sequences. Moreover these are long sequences. This is not ideal for a deep learning approach.  <br>\nExploring the Field Day Lab research instructed that the Jo Wilder application was built to help learning to read and that way more than 11,500 learners had played this game.  <br>\nThese 2 ideas led to search for a bigger dataset. In 1 Google search and 3 clicks we came up to the open data portal (<a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/</a>) which contains a lot of sessions. 1 hour and 3 bash commands latter we knew that the train set was in part in the open data. So we took a week to <strong>build a pipeline that extracts 98 % of the sessions of the train set perfectly and with minor errors for the last 2 %</strong>. Our data are even better than the comp data because we knew before the host confirmation that for the sessions with 2 games the target was skewed (0 if wrong in 1 of the 2 games when we aim at predicting the responses for the 1st game). It seems that fixing these targets can bring a significant boost up to +0.002.</p>\n<p>We took 1 more week to build a GBDT/XGBoost baseline that would have scored top 10 given the CV score, with the use of the supplemental data (~20,000 sessions) that gave +0.003/0.004 at that time. As we simulated the API locally (see after), we used some training sessions to infer and noticed that it scored 0.718. We were hoping that the LB sessions were not part of the open data portal but our 1st submission, LB 0.708, immediately showed to us that we had rebuilt about a half of the data and especially the targets in the public LB, because 0.708 = (0.698 + 0.718) / 2. The host and Kaggle have been immediately informed. You know what happened next (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/415820)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/415820)</a>.  <br>\nAfter the release of the LB data we measured that we perfectly rebuilt ~7000 sessions over the ~11,500 of the LB data.</p>\n<p>We spent the first month exploring the data until we understood/knew it pretty well. For example we even reconstituted sessions for what might be schools (several games on 1 IP session), extracted every single session with at least 1 answer, …</p>\n<p>After the update we made a first submission that scored 0.72. This was shocking because this meant that some leaked data were remaining. A few days later we noticed that the open data was not totally similar with the state we found it 1 month before. A file was missing. So we returned to the host and Kaggle to give them more work.</p>\n<p><strong>This process/work led us to perfectly understand the data model</strong> (that changed since the 1st release of the game). This also allowed us to deeply understand the data itself.</p>\n<p>Note that we only used the sessions for which we had responses to all questions of the 2 1st level groups. 1) This is more consistant with the sessions we want to predict (game from the beginning to the end) and 2) this approach preserves performance (vs all data) while reducing training time.</p>\n<p>Our dataset is constituted by <strong>37323 complete sessions (23562 comp + logs) in a total of 66376 sessions</strong>.</p>\n<p>The supplemental data (that we fully added 1 month ago) gave us consistently <strong>CV +0.002</strong>.</p>\n<h2>Model</h2>\n<p>Our solution is mainly an ensemble of GBDT + NN models.</p>\n<h3>Trust your validation</h3>\n<p>We think that <strong>the main reason of the robustness of our solution is that we only relied on CV</strong> for decision making. No choice had been made on LB.</p>\n<p>Probing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API. This is a small set. In our experiments 5,000 sessions is the minimum to guarantee a stable CV/LB alignment. A set less than 2,000 is very noisy so <strong>robustness was the way to go</strong>.</p>\n<p>We <strong>only added features that improved the CV for sure</strong>. This is not easy to delete features that you believe in but this is needed as science is not a matter of belief. There are several ways to do so: for example monitor all folds in a CV (and accept only on majority or consensus), monitor several bags (composition of CV to not overfit validation), …</p>\n<p>For the GBDT approach, we mainly validated on the mean of 10 bags (we defined a bag as a composition of the folds). As we estimated the noise to be ~0.0003, only improvements greater than the noise have been considered. For the NN as we needed to iterate quicker we only used a single bag and only incorporated &gt; 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.</p>\n<h3>Metric</h3>\n<p>We experimented a lot on finding a threshold by question but found that this approach is less robust than a single threshold. We mainly used 0.625 as global threshold despite our highest LB scores that were obtained with a threshold per question.</p>\n<h3>GBDT</h3>\n<p>We prototyped a baseline with <strong>XGBoost because of the structured/tabular nature of the data</strong>. The feature engineering process is interesting to understand what is predictive and to understand the causation, i.e. how the features or decision criteria that enable to predict correctly.</p>\n<p>Generally speaking we followed 3 ways to build features: <strong>business knowledge</strong>, our <strong>intuition</strong> playing the game and a meticulous <strong>exploration of the data</strong>.<br>\nBusiness knowledge refers to using expert knowledge. Reading the papers of the researchers that built this game allow to understand the game beyond usage. For example, Jo Wilder has been built to improve the players reading skills. So this means that the text duration should be important. These are like killer features.</p>\n<p>We exclusively made use of Polars because of the CPU constraints and to simply learn it.<br>\nOur features (663, 1993, 3734 for each level_group) are mainly <strong>durations and counts for different aggregations</strong>: how much time in a level, in a room, reading a text, interacting in some way (event type), how many events in a level_group, how many events of each type, how many events of each type in a room or a level, …  <br>\nWe also built a few notebook dedicated features: how many type of events on the notebook in a level, …  <br>\nDespite our efforts we weren't able to extract useful information from the coordinates, the only few features of this type had been mean and std for some events in the activities (journal interactions for example).</p>\n<p>We considered that injecting targets predicted in the previous level groups was a compression of the signal, meaning a loss of information, so we used, for each session, <strong>all interactions from the beginning of the game/session</strong>. This led to a +0.002 at the time of this choice.</p>\n<p>After the API needed to order the data, we noticed that <strong>models trained both on original order and on index order</strong> but validated on index order (inference order) improved our scores. This leads to more variety that was needed to <strong>improve stability and robustness</strong>. The same goes for the composition of the validation sets: usage of several bags (composition of validation sets) based on the comp data but also on the extracted data improved our scores. We detected late that increasing the number of folds from 5 to 10 could also be leveraged.</p>\n<p>The code for GBDT allows to switch from XGBoost to LightGBM and CatBoost with a simple variable parameter but despite the good scores (~0.001 less than XGBoost), this did not bring to ensemble so we sticked to only XGBoost.</p>\n<p>We experimented a lot around feature selection but were unable to build a stable strategy. So instead of a top-down approach consisting in deleting useless features, we adopted a bottom-up approach choosing carefully each group of features.</p>\n<p>Our <strong>XGBoost models score CV ~0.7025 +/-0.0003</strong> and blending 5 of them (the only XGBoost we still have with correct score) scores <strong>LB 0.704</strong>.</p>\n<h3>NN</h3>\n<p>After achieving a good score with gradient boosting and having understood well the data we focused on deep learning.</p>\n<p>The <strong>first attempt was with Transformers</strong>. The 1st results were disappointed: CV 0.685 with 2 hours / fold (as far as we can remember). Transformers are very computationally intensive. Resources: <a href=\"https://arxiv.org/pdf/1912.09363.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.09363.pdf</a>, <a href=\"https://arxiv.org/pdf/2001.08317.pdf\" target=\"_blank\">https://arxiv.org/pdf/2001.08317.pdf</a>, <a href=\"https://arxiv.org/pdf/1711.03905.pdf\" target=\"_blank\">https://arxiv.org/pdf/1711.03905.pdf</a>, <a href=\"https://arxiv.org/pdf/1907.00235.pdf\" target=\"_blank\">https://arxiv.org/pdf/1907.00235.pdf</a>, …</p>\n<p>We then gave a try to <strong>Conv1D</strong>. In one day we had a very simple model that scored as Transformers but <strong>10x faster</strong> allowing to iterate quicker. So we pushed this approach and could seamlessly scaled it beyond our expectations.</p>\n<p>Difficult to share the <strong>tens or hundreds of experimentations</strong> needed to achieve the final solution which is both based on a simple architecture and a slightly complex training pipeline.</p>\n<h4>Architecture roots</h4>\n<p>We browsed the literature based on the question: how to model time in deep learning?  <br>\nThis research made us come to the idea of <strong>time-aware events</strong> (i.e. <a href=\"https://proceedings.mlr.press/v126/zhang20c/zhang20c.pdf\" target=\"_blank\">https://proceedings.mlr.press/v126/zhang20c/zhang20c.pdf</a>) and back to <strong>WaveNet</strong> (<a href=\"https://arxiv.org/pdf/1609.03499.pdf\" target=\"_blank\">https://arxiv.org/pdf/1609.03499.pdf</a>) because it uses <strong>Conv1D to model long sequences with considerations on causation</strong>.  <br>\nOther papers also inspired us: <a href=\"https://arxiv.org/pdf/1703.04691.pdf\" target=\"_blank\">https://arxiv.org/pdf/1703.04691.pdf</a> build on top of WaveNet paper for time series, <a href=\"https://idus.us.es/bitstream/handle/11441/114701/Short-Term%20Load%20Forecasting%20Using%20Encoder-Decoder%20WaveNet.pdf?sequence=1&amp;isAllowed=y\" target=\"_blank\">https://idus.us.es/bitstream/handle/11441/114701/Short-Term%20Load%20Forecasting%20Using%20Encoder-Decoder%20WaveNet.pdf?sequence=1&amp;isAllowed=y</a> also build on top of WaveNet.  <br>\nWe also have to mention the excellent work that <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> shared (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565</a> and <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part)\" target=\"_blank\">https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part)</a>. It inspired our research and maybe successfully biased it.</p>\n<p>Let's focus on the model of our efficiency submission that is also one of our final ensemble and which performance is nearly the same as models with a few more features.</p>\n<h4>Feature representations</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Ff692423c63172494ea4df97b01d65d33%2Fdata.png?generation=1702113769960411&amp;alt=media\" alt=\"\"></p>\n<p><strong>5 features as inputs: duration, text_fqid, room_fqid, fqid, event_name + name</strong> (this is the event type from the original data model as far as we remember). Each of these information is encoded/embedded into a vector representation (d_model = 24) to be the merged. The 4 <strong>categorical features feed a classical Embedding layer and the duration a TimeEmbedding</strong> which is a custom block.</p>\n<p>Developing the GBDT solution showed that the <strong>duration</strong> was crucial, so we put a crucial amount of time trying to model it greatly. The TimeEmbedding layer is a composition of 4x ConvBlock which is inspired by the Transformer main block: Conv1D -&gt; skip connection -&gt; layer norm -&gt; dropout.</p>\n<pre><code> (tf.keras.layers.):\n     ():\n        (, ).__init__()\n        .conv_blocks = [(d_model, dropout_rate=dropout_rate)  _  range(n_blocks)]\n\n     ():\n        x = tf.expand_dims(inputs, axis=-)\n         conv_block  .\n            x = conv_block(x)\n         x\n</code></pre>\n<pre><code> (tf.keras.layers.):\n     ():\n        (, ).__init__()\n        .conv1d = tf.keras.layers.D(d_model, kernel_size=, padding=, activation=)\n        .layer_norm = tf.keras.layers.()\n        .dropout = tf.keras.layers.(rate=dropout_rate)\n\n     ():\n        x = .conv1d(inputs)\n        x = x + inputs\n        x = .layer_norm(x)\n        outputs = .dropout(x)\n         outputs\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F9407390ffd7f226b7be0f57c37051f77%2Ftime_embedding.png?generation=1702113795144186&amp;alt=media\" alt=\"\"></p>\n<h4>Time-aware events</h4>\n<p>As said, the goal of building these representations was to model time-aware events. We considered the <strong>categorical features as events</strong> because they represent the user interactions with business entities of the game. We then tried to incorporate duration to make them time-awared. Our main intuition showed to be the best. It is a <strong>simple solution based on operation priority to represent that the duration should be associated to each event before associated them together</strong>: duration * event_1 + duration * event_2 + … which had been factorized to duration * (event_1 + event_2 + …).</p>\n<pre><code> (tf.keras.Model):\n     ():\n        (ConvNet, ).__init__(name=name)\n        .input_dims = input_dims\n        .n_outputs = n_outputs\n        .d_model = d_model\n        .n_blocks = n_blocks\n        .event_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .room_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .text_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .fqid_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .duration_embedding = TimeEmbedding(n_blocks=n_blocks, d_model=d_model, dropout_rate=)\n        .gap = tf.keras.layers.GlobalAveragePooling1D()\n\n     ():\n        event = .event_embedding(inputs[])\n        room = .room_embedding(inputs[])\n        text = .text_embedding(inputs[])\n        fqid = .fqid_embedding(inputs[])\n        duration = .duration_embedding(inputs[])\n        x = duration * (event + room + text + fqid)\n        outputs = .gap(x)\n         outputs\n\n     ():\n        config = ().get_config().copy()\n        config.update({\n            : .input_dims,\n            : .n_outputs,\n            : .d_model,\n            : .n_blocks,\n            : ._name,\n        })\n         config\n\n\n     ():\n         cls(**config)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Fca563daa4f9ef7bf8d8b20e0c85c7951%2Ftime_aware_events_1.png?generation=1702113822015438&amp;alt=media\" alt=\"\"><br>\nThe 2 representations are equivalent: either you can think time-aware events as a combination of time and sub-events or as a combination of sub-events and time.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F1f65d399dfd93f128ca6c887d0f0be89%2Ftime_aware_events_2.png?generation=1702113853644045&amp;alt=media\" alt=\"\"></p>\n<h4>Training pipeline</h4>\n<p>The training pipeline is not totally straight forward.</p>\n<p><a href=\"https://www.kaggle.com/dongyk\" target=\"_blank\">@dongyk</a> published great schematics that can be useful to illustrate what is explained bellow: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166</a>.</p>\n<h5>1st step (pre-training?)</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F970baa66fc37b301bb8960cc92cf8ab4%2Fpre_training.png?generation=1702113879947155&amp;alt=media\" alt=\"\"></p>\n<p>The best approach for us consists in a kind of <strong>backbone that represents the events of a level_group</strong>.</p>\n<p>This backbone is trained on all the data available for this level_group (i.e. on complete + incomplete sessions). It is associated with a temporary SimpleHead optimizing BCE loss.</p>\n<pre><code> (tf.keras.):\n     ():\n        (, ).__init__(name=name)\n        .ffs = [tf.keras.layers.(units, activation=)  units  n_units]\n        .out = tf.keras.layers.(n_outputs, activation=)\n\n     ():\n        x = inputs\n         ff  .\n            x = ff(x)\n        outputs = .out(x)\n         outputs\n</code></pre>\n<p>This approach allows to score <strong>CV 0.70025 +/- 0.0005</strong>.</p>\n<h5>2nd step (training?)</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F98e32800c54ffab63d98d90d220fbbcd%2Fend_2_end.png?generation=1702113900174294&amp;alt=media\" alt=\"\"></p>\n<p><strong>The weights of each of the 3 backbones (1 by level_group) are freezed</strong> for the 2nd level of training to speedup training but also because it is more stable and efficient. These backbones can be thought as \"embedders\".</p>\n<p>During this 2nd step, <strong>all the submodels that composed the solution were trained on all complete sessions in an end-to-end setup</strong>. The input data are 3 sequences of the 5 features, 1 for each of the 3 level groups. Each \"embedders\" outputs a 24 dim-vector representation. These outputs are the inputs of a head in which enters the representation of level_group '0-4' to predict the 3 first questions and the concatenation of the previous and the current representations for level_groups '5-12' and '13-22' to make use of all information.</p>\n<p>Proceeding like this allows to optimize the overall performance and to monitor it based on the F1 score that is the score of the competition. This means we optimized BCE with F1 score as a metric.</p>\n<p>Our winning submission uses a simple <strong>MLP head</strong> but also a <strong>skip head</strong> (512 -&gt; 512 -&gt; 512 allow it for example). <strong>MMoE</strong> did not improve the simplest approaches.</p>\n<p>This approach allows to score <strong>CV 0.70175 +/- 0.0003</strong> which is <strong>comparable to the GBDT solution</strong>.</p>\n<h3>Inference</h3>\n<h4>Build a simulator</h4>\n<p>Early in the competition we built a simulator of the API. Doing so we never experimented any submission error. Maybe trying to keep ideas and code as simple as possible was also key to debug easily.</p>\n<h4>Efficiency</h4>\n<p>We invested the efficiency part of the challenge for GBDT as well as NNs.  <br>\nUsing <strong>Treelite</strong> for XGBoost allow us to divide by 2 the execution time.  <br>\nOur deep learning models were lights: <strong>400,000 weights</strong> for the end-to-end model which combines every parts/sub-models. Having already used <strong>TF Lite</strong> we knew it could be a game changer. Converting our models led to a significant boost in inference time without any performance loss (we do not remember exactly but we think it is at least <strong>6x faster</strong> on our local inference simulator).  <br>\nBeginning to explore pruning as well as hard quantization showed that the performance loss would be significant (which is OK in production but not in a competition) so we sticked to a simple TF Lite conversion.</p>\n<p>We have not leveraged what seems to be a problem in the efficiency metric. As we identified the private test sessions to be the 1450/1500 first served by the API we tried to just predict the others to check which time was used (public for public and private for private). Doing so we gain a place but choose to not use this.</p>\n<p>Our <strong>efficiency submission is a NN that scores public LB 0.702 and private LB 0.699 in less than 5 minutes</strong>.</p>\n<h4>Ensemble</h4>\n<p>We experimented a lot of ensembling alternatives. In the end we sticked to a simple average 50/50 GBDT/NN with:</p>\n<ul>\n<li>2 kinds of GBDT: trained on original order + trained on index order (validated on index order that is the inference case),</li>\n<li>3 kinds of NNs: trained on original order + trained on index order with 5 or all features.</li>\n</ul>\n<p>As our models are lightweight we were able to build a hugh ensemble: <strong>2 x 4 x 10 folds XGBoost + 3 x 4 x 5 folds NNs</strong>. The bottleneck for us is the 8 Go RAM constraint.</p>\n<p>The winning submission scores <strong>CV 0.705, public LB 0.705 and private LB 0.705</strong>.</p>\n<h2>Conclusion</h2>\n<p>The main achievement of our work is that it is a good solution for the researchers, learners and children that can benefit of it and we hope it will contributes to progresses for a better learning experience. Up to you guys!</p>\n<p>Thanks if you read until here!<br>\nIf you have any question do not hesitate to ask. We will do our best to respond.</p>\n<h2>Presentation to the host</h2>\n<p>A video presentation to the host has been recorded and can be available on demand. Feel free to ask via PM.</p>\n<h1>Sources</h1>\n<p>Below are the main sources that we used. More sources can be found in section <em>Details of the submission</em> above.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332</a>,</li>\n<li><a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/</a>,</li>\n<li><a href=\"https://arxiv.org/pdf/1609.03499.pdf\" target=\"_blank\">https://arxiv.org/pdf/1609.03499.pdf</a>,</li>\n<li><a href=\"https://www.tensorflow.org/lite/guide\" target=\"_blank\">https://www.tensorflow.org/lite/guide</a></li>\n</ul>",
  "messages": [
    {
      "id": 2323109,
      "postDate": "2023-06-29T17:06:37.317Z",
      "content": "<p>Unbelievable to write this!</p>\n<h1>Thanks!</h1>\n<p>As it is the usage, we first <strong>thank the host and Kaggle</strong>. These are special thanks because you and us have had a special link in this competition as we gave you more work by reporting data leaks. No doubt you tried to do your best. You are right to animate this community and to trust in it. You are part of it. Please take care of this community that is able to build so much together by sharing. As all of us you have made mistakes and we hope you will learn from them.</p>\n<p>We also want to <strong>thank all of you</strong>, Kagglers. We love and are grateful to be part of our group/community. Thanks for sharing and for the collective learning experience.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview</a>,</li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data</a>.</li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>Our solution is essentially a blend of a XGBoost and a NN models. Both heavily rely on duration that appeared to be a powerful leverage. Time was aggregated in different ways and combined with counts for the GBDT while it is transformed via a custom TimeEmbedding block based on 1D convolutions that produce a representation combined with user event representations for the NN.<br>\nRobustness and efficiency founded our work. XGBoost models were validated on 10 bags of 5 folds and features incorporated only if the mean of the CV of these 10 bags was greater than the level of noise we quantified while we opted for a majority/consensus strategy to build the NN, i.e. validate choices only if 4 of 5 folds were improved. The 3rd place of the efficiency LB was achieved with a lightweight NN accelerated via TF Lite.</p>\n<h1>Details of the submission</h1>\n<h2>Code</h2>\n<p>After publishing this write-up we decided to open our code: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332</a>.<br>\nIt is composed by several parts: <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training\" target=\"_blank\">how to train the XGBoost models</a>, how to <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining\" target=\"_blank\">pretrain</a> and <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">train</a> the NN models and the <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference\" target=\"_blank\">inference notebook</a> used to win this competition.</p>\n<h2>Data</h2>\n<p>Looking at the 1st data released showed that there aren't a lot of sessions so not a lot of sequences. Moreover these are long sequences. This is not ideal for a deep learning approach.  <br>\nExploring the Field Day Lab research instructed that the Jo Wilder application was built to help learning to read and that way more than 11,500 learners had played this game.  <br>\nThese 2 ideas led to search for a bigger dataset. In 1 Google search and 3 clicks we came up to the open data portal (<a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/</a>) which contains a lot of sessions. 1 hour and 3 bash commands latter we knew that the train set was in part in the open data. So we took a week to <strong>build a pipeline that extracts 98 % of the sessions of the train set perfectly and with minor errors for the last 2 %</strong>. Our data are even better than the comp data because we knew before the host confirmation that for the sessions with 2 games the target was skewed (0 if wrong in 1 of the 2 games when we aim at predicting the responses for the 1st game). It seems that fixing these targets can bring a significant boost up to +0.002.</p>\n<p>We took 1 more week to build a GBDT/XGBoost baseline that would have scored top 10 given the CV score, with the use of the supplemental data (~20,000 sessions) that gave +0.003/0.004 at that time. As we simulated the API locally (see after), we used some training sessions to infer and noticed that it scored 0.718. We were hoping that the LB sessions were not part of the open data portal but our 1st submission, LB 0.708, immediately showed to us that we had rebuilt about a half of the data and especially the targets in the public LB, because 0.708 = (0.698 + 0.718) / 2. The host and Kaggle have been immediately informed. You know what happened next (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/415820)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/415820)</a>.  <br>\nAfter the release of the LB data we measured that we perfectly rebuilt ~7000 sessions over the ~11,500 of the LB data.</p>\n<p>We spent the first month exploring the data until we understood/knew it pretty well. For example we even reconstituted sessions for what might be schools (several games on 1 IP session), extracted every single session with at least 1 answer, …</p>\n<p>After the update we made a first submission that scored 0.72. This was shocking because this meant that some leaked data were remaining. A few days later we noticed that the open data was not totally similar with the state we found it 1 month before. A file was missing. So we returned to the host and Kaggle to give them more work.</p>\n<p><strong>This process/work led us to perfectly understand the data model</strong> (that changed since the 1st release of the game). This also allowed us to deeply understand the data itself.</p>\n<p>Note that we only used the sessions for which we had responses to all questions of the 2 1st level groups. 1) This is more consistant with the sessions we want to predict (game from the beginning to the end) and 2) this approach preserves performance (vs all data) while reducing training time.</p>\n<p>Our dataset is constituted by <strong>37323 complete sessions (23562 comp + logs) in a total of 66376 sessions</strong>.</p>\n<p>The supplemental data (that we fully added 1 month ago) gave us consistently <strong>CV +0.002</strong>.</p>\n<h2>Model</h2>\n<p>Our solution is mainly an ensemble of GBDT + NN models.</p>\n<h3>Trust your validation</h3>\n<p>We think that <strong>the main reason of the robustness of our solution is that we only relied on CV</strong> for decision making. No choice had been made on LB.</p>\n<p>Probing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API. This is a small set. In our experiments 5,000 sessions is the minimum to guarantee a stable CV/LB alignment. A set less than 2,000 is very noisy so <strong>robustness was the way to go</strong>.</p>\n<p>We <strong>only added features that improved the CV for sure</strong>. This is not easy to delete features that you believe in but this is needed as science is not a matter of belief. There are several ways to do so: for example monitor all folds in a CV (and accept only on majority or consensus), monitor several bags (composition of CV to not overfit validation), …</p>\n<p>For the GBDT approach, we mainly validated on the mean of 10 bags (we defined a bag as a composition of the folds). As we estimated the noise to be ~0.0003, only improvements greater than the noise have been considered. For the NN as we needed to iterate quicker we only used a single bag and only incorporated &gt; 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.</p>\n<h3>Metric</h3>\n<p>We experimented a lot on finding a threshold by question but found that this approach is less robust than a single threshold. We mainly used 0.625 as global threshold despite our highest LB scores that were obtained with a threshold per question.</p>\n<h3>GBDT</h3>\n<p>We prototyped a baseline with <strong>XGBoost because of the structured/tabular nature of the data</strong>. The feature engineering process is interesting to understand what is predictive and to understand the causation, i.e. how the features or decision criteria that enable to predict correctly.</p>\n<p>Generally speaking we followed 3 ways to build features: <strong>business knowledge</strong>, our <strong>intuition</strong> playing the game and a meticulous <strong>exploration of the data</strong>.<br>\nBusiness knowledge refers to using expert knowledge. Reading the papers of the researchers that built this game allow to understand the game beyond usage. For example, Jo Wilder has been built to improve the players reading skills. So this means that the text duration should be important. These are like killer features.</p>\n<p>We exclusively made use of Polars because of the CPU constraints and to simply learn it.<br>\nOur features (663, 1993, 3734 for each level_group) are mainly <strong>durations and counts for different aggregations</strong>: how much time in a level, in a room, reading a text, interacting in some way (event type), how many events in a level_group, how many events of each type, how many events of each type in a room or a level, …  <br>\nWe also built a few notebook dedicated features: how many type of events on the notebook in a level, …  <br>\nDespite our efforts we weren't able to extract useful information from the coordinates, the only few features of this type had been mean and std for some events in the activities (journal interactions for example).</p>\n<p>We considered that injecting targets predicted in the previous level groups was a compression of the signal, meaning a loss of information, so we used, for each session, <strong>all interactions from the beginning of the game/session</strong>. This led to a +0.002 at the time of this choice.</p>\n<p>After the API needed to order the data, we noticed that <strong>models trained both on original order and on index order</strong> but validated on index order (inference order) improved our scores. This leads to more variety that was needed to <strong>improve stability and robustness</strong>. The same goes for the composition of the validation sets: usage of several bags (composition of validation sets) based on the comp data but also on the extracted data improved our scores. We detected late that increasing the number of folds from 5 to 10 could also be leveraged.</p>\n<p>The code for GBDT allows to switch from XGBoost to LightGBM and CatBoost with a simple variable parameter but despite the good scores (~0.001 less than XGBoost), this did not bring to ensemble so we sticked to only XGBoost.</p>\n<p>We experimented a lot around feature selection but were unable to build a stable strategy. So instead of a top-down approach consisting in deleting useless features, we adopted a bottom-up approach choosing carefully each group of features.</p>\n<p>Our <strong>XGBoost models score CV ~0.7025 +/-0.0003</strong> and blending 5 of them (the only XGBoost we still have with correct score) scores <strong>LB 0.704</strong>.</p>\n<h3>NN</h3>\n<p>After achieving a good score with gradient boosting and having understood well the data we focused on deep learning.</p>\n<p>The <strong>first attempt was with Transformers</strong>. The 1st results were disappointed: CV 0.685 with 2 hours / fold (as far as we can remember). Transformers are very computationally intensive. Resources: <a href=\"https://arxiv.org/pdf/1912.09363.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.09363.pdf</a>, <a href=\"https://arxiv.org/pdf/2001.08317.pdf\" target=\"_blank\">https://arxiv.org/pdf/2001.08317.pdf</a>, <a href=\"https://arxiv.org/pdf/1711.03905.pdf\" target=\"_blank\">https://arxiv.org/pdf/1711.03905.pdf</a>, <a href=\"https://arxiv.org/pdf/1907.00235.pdf\" target=\"_blank\">https://arxiv.org/pdf/1907.00235.pdf</a>, …</p>\n<p>We then gave a try to <strong>Conv1D</strong>. In one day we had a very simple model that scored as Transformers but <strong>10x faster</strong> allowing to iterate quicker. So we pushed this approach and could seamlessly scaled it beyond our expectations.</p>\n<p>Difficult to share the <strong>tens or hundreds of experimentations</strong> needed to achieve the final solution which is both based on a simple architecture and a slightly complex training pipeline.</p>\n<h4>Architecture roots</h4>\n<p>We browsed the literature based on the question: how to model time in deep learning?  <br>\nThis research made us come to the idea of <strong>time-aware events</strong> (i.e. <a href=\"https://proceedings.mlr.press/v126/zhang20c/zhang20c.pdf\" target=\"_blank\">https://proceedings.mlr.press/v126/zhang20c/zhang20c.pdf</a>) and back to <strong>WaveNet</strong> (<a href=\"https://arxiv.org/pdf/1609.03499.pdf\" target=\"_blank\">https://arxiv.org/pdf/1609.03499.pdf</a>) because it uses <strong>Conv1D to model long sequences with considerations on causation</strong>.  <br>\nOther papers also inspired us: <a href=\"https://arxiv.org/pdf/1703.04691.pdf\" target=\"_blank\">https://arxiv.org/pdf/1703.04691.pdf</a> build on top of WaveNet paper for time series, <a href=\"https://idus.us.es/bitstream/handle/11441/114701/Short-Term%20Load%20Forecasting%20Using%20Encoder-Decoder%20WaveNet.pdf?sequence=1&amp;isAllowed=y\" target=\"_blank\">https://idus.us.es/bitstream/handle/11441/114701/Short-Term%20Load%20Forecasting%20Using%20Encoder-Decoder%20WaveNet.pdf?sequence=1&amp;isAllowed=y</a> also build on top of WaveNet.  <br>\nWe also have to mention the excellent work that <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> shared (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565</a> and <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part)\" target=\"_blank\">https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part)</a>. It inspired our research and maybe successfully biased it.</p>\n<p>Let's focus on the model of our efficiency submission that is also one of our final ensemble and which performance is nearly the same as models with a few more features.</p>\n<h4>Feature representations</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Ff692423c63172494ea4df97b01d65d33%2Fdata.png?generation=1702113769960411&amp;alt=media\" alt=\"\"></p>\n<p><strong>5 features as inputs: duration, text_fqid, room_fqid, fqid, event_name + name</strong> (this is the event type from the original data model as far as we remember). Each of these information is encoded/embedded into a vector representation (d_model = 24) to be the merged. The 4 <strong>categorical features feed a classical Embedding layer and the duration a TimeEmbedding</strong> which is a custom block.</p>\n<p>Developing the GBDT solution showed that the <strong>duration</strong> was crucial, so we put a crucial amount of time trying to model it greatly. The TimeEmbedding layer is a composition of 4x ConvBlock which is inspired by the Transformer main block: Conv1D -&gt; skip connection -&gt; layer norm -&gt; dropout.</p>\n<pre><code> (tf.keras.layers.):\n     ():\n        (, ).__init__()\n        .conv_blocks = [(d_model, dropout_rate=dropout_rate)  _  range(n_blocks)]\n\n     ():\n        x = tf.expand_dims(inputs, axis=-)\n         conv_block  .\n            x = conv_block(x)\n         x\n</code></pre>\n<pre><code> (tf.keras.layers.):\n     ():\n        (, ).__init__()\n        .conv1d = tf.keras.layers.D(d_model, kernel_size=, padding=, activation=)\n        .layer_norm = tf.keras.layers.()\n        .dropout = tf.keras.layers.(rate=dropout_rate)\n\n     ():\n        x = .conv1d(inputs)\n        x = x + inputs\n        x = .layer_norm(x)\n        outputs = .dropout(x)\n         outputs\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F9407390ffd7f226b7be0f57c37051f77%2Ftime_embedding.png?generation=1702113795144186&amp;alt=media\" alt=\"\"></p>\n<h4>Time-aware events</h4>\n<p>As said, the goal of building these representations was to model time-aware events. We considered the <strong>categorical features as events</strong> because they represent the user interactions with business entities of the game. We then tried to incorporate duration to make them time-awared. Our main intuition showed to be the best. It is a <strong>simple solution based on operation priority to represent that the duration should be associated to each event before associated them together</strong>: duration * event_1 + duration * event_2 + … which had been factorized to duration * (event_1 + event_2 + …).</p>\n<pre><code> (tf.keras.Model):\n     ():\n        (ConvNet, ).__init__(name=name)\n        .input_dims = input_dims\n        .n_outputs = n_outputs\n        .d_model = d_model\n        .n_blocks = n_blocks\n        .event_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .room_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .text_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .fqid_embedding = tf.keras.layers.Embedding(input_dims[], d_model, mask_zero=)\n        .duration_embedding = TimeEmbedding(n_blocks=n_blocks, d_model=d_model, dropout_rate=)\n        .gap = tf.keras.layers.GlobalAveragePooling1D()\n\n     ():\n        event = .event_embedding(inputs[])\n        room = .room_embedding(inputs[])\n        text = .text_embedding(inputs[])\n        fqid = .fqid_embedding(inputs[])\n        duration = .duration_embedding(inputs[])\n        x = duration * (event + room + text + fqid)\n        outputs = .gap(x)\n         outputs\n\n     ():\n        config = ().get_config().copy()\n        config.update({\n            : .input_dims,\n            : .n_outputs,\n            : .d_model,\n            : .n_blocks,\n            : ._name,\n        })\n         config\n\n\n     ():\n         cls(**config)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Fca563daa4f9ef7bf8d8b20e0c85c7951%2Ftime_aware_events_1.png?generation=1702113822015438&amp;alt=media\" alt=\"\"><br>\nThe 2 representations are equivalent: either you can think time-aware events as a combination of time and sub-events or as a combination of sub-events and time.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F1f65d399dfd93f128ca6c887d0f0be89%2Ftime_aware_events_2.png?generation=1702113853644045&amp;alt=media\" alt=\"\"></p>\n<h4>Training pipeline</h4>\n<p>The training pipeline is not totally straight forward.</p>\n<p><a href=\"https://www.kaggle.com/dongyk\" target=\"_blank\">@dongyk</a> published great schematics that can be useful to illustrate what is explained bellow: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166</a>.</p>\n<h5>1st step (pre-training?)</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F970baa66fc37b301bb8960cc92cf8ab4%2Fpre_training.png?generation=1702113879947155&amp;alt=media\" alt=\"\"></p>\n<p>The best approach for us consists in a kind of <strong>backbone that represents the events of a level_group</strong>.</p>\n<p>This backbone is trained on all the data available for this level_group (i.e. on complete + incomplete sessions). It is associated with a temporary SimpleHead optimizing BCE loss.</p>\n<pre><code> (tf.keras.):\n     ():\n        (, ).__init__(name=name)\n        .ffs = [tf.keras.layers.(units, activation=)  units  n_units]\n        .out = tf.keras.layers.(n_outputs, activation=)\n\n     ():\n        x = inputs\n         ff  .\n            x = ff(x)\n        outputs = .out(x)\n         outputs\n</code></pre>\n<p>This approach allows to score <strong>CV 0.70025 +/- 0.0005</strong>.</p>\n<h5>2nd step (training?)</h5>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F98e32800c54ffab63d98d90d220fbbcd%2Fend_2_end.png?generation=1702113900174294&amp;alt=media\" alt=\"\"></p>\n<p><strong>The weights of each of the 3 backbones (1 by level_group) are freezed</strong> for the 2nd level of training to speedup training but also because it is more stable and efficient. These backbones can be thought as \"embedders\".</p>\n<p>During this 2nd step, <strong>all the submodels that composed the solution were trained on all complete sessions in an end-to-end setup</strong>. The input data are 3 sequences of the 5 features, 1 for each of the 3 level groups. Each \"embedders\" outputs a 24 dim-vector representation. These outputs are the inputs of a head in which enters the representation of level_group '0-4' to predict the 3 first questions and the concatenation of the previous and the current representations for level_groups '5-12' and '13-22' to make use of all information.</p>\n<p>Proceeding like this allows to optimize the overall performance and to monitor it based on the F1 score that is the score of the competition. This means we optimized BCE with F1 score as a metric.</p>\n<p>Our winning submission uses a simple <strong>MLP head</strong> but also a <strong>skip head</strong> (512 -&gt; 512 -&gt; 512 allow it for example). <strong>MMoE</strong> did not improve the simplest approaches.</p>\n<p>This approach allows to score <strong>CV 0.70175 +/- 0.0003</strong> which is <strong>comparable to the GBDT solution</strong>.</p>\n<h3>Inference</h3>\n<h4>Build a simulator</h4>\n<p>Early in the competition we built a simulator of the API. Doing so we never experimented any submission error. Maybe trying to keep ideas and code as simple as possible was also key to debug easily.</p>\n<h4>Efficiency</h4>\n<p>We invested the efficiency part of the challenge for GBDT as well as NNs.  <br>\nUsing <strong>Treelite</strong> for XGBoost allow us to divide by 2 the execution time.  <br>\nOur deep learning models were lights: <strong>400,000 weights</strong> for the end-to-end model which combines every parts/sub-models. Having already used <strong>TF Lite</strong> we knew it could be a game changer. Converting our models led to a significant boost in inference time without any performance loss (we do not remember exactly but we think it is at least <strong>6x faster</strong> on our local inference simulator).  <br>\nBeginning to explore pruning as well as hard quantization showed that the performance loss would be significant (which is OK in production but not in a competition) so we sticked to a simple TF Lite conversion.</p>\n<p>We have not leveraged what seems to be a problem in the efficiency metric. As we identified the private test sessions to be the 1450/1500 first served by the API we tried to just predict the others to check which time was used (public for public and private for private). Doing so we gain a place but choose to not use this.</p>\n<p>Our <strong>efficiency submission is a NN that scores public LB 0.702 and private LB 0.699 in less than 5 minutes</strong>.</p>\n<h4>Ensemble</h4>\n<p>We experimented a lot of ensembling alternatives. In the end we sticked to a simple average 50/50 GBDT/NN with:</p>\n<ul>\n<li>2 kinds of GBDT: trained on original order + trained on index order (validated on index order that is the inference case),</li>\n<li>3 kinds of NNs: trained on original order + trained on index order with 5 or all features.</li>\n</ul>\n<p>As our models are lightweight we were able to build a hugh ensemble: <strong>2 x 4 x 10 folds XGBoost + 3 x 4 x 5 folds NNs</strong>. The bottleneck for us is the 8 Go RAM constraint.</p>\n<p>The winning submission scores <strong>CV 0.705, public LB 0.705 and private LB 0.705</strong>.</p>\n<h2>Conclusion</h2>\n<p>The main achievement of our work is that it is a good solution for the researchers, learners and children that can benefit of it and we hope it will contributes to progresses for a better learning experience. Up to you guys!</p>\n<p>Thanks if you read until here!<br>\nIf you have any question do not hesitate to ask. We will do our best to respond.</p>\n<h2>Presentation to the host</h2>\n<p>A video presentation to the host has been recorded and can be available on demand. Feel free to ask via PM.</p>\n<h1>Sources</h1>\n<p>Below are the main sources that we used. More sources can be found in section <em>Details of the submission</em> above.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332</a>,</li>\n<li><a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/</a>,</li>\n<li><a href=\"https://arxiv.org/pdf/1609.03499.pdf\" target=\"_blank\">https://arxiv.org/pdf/1609.03499.pdf</a>,</li>\n<li><a href=\"https://www.tensorflow.org/lite/guide\" target=\"_blank\">https://www.tensorflow.org/lite/guide</a></li>\n</ul>",
      "rawMarkdown": "Unbelievable to write this!\n\n# Thanks!\n\nAs it is the usage, we first **thank the host and Kaggle**. These are special thanks because you and us have had a special link in this competition as we gave you more work by reporting data leaks. No doubt you tried to do your best. You are right to animate this community and to trust in it. You are part of it. Please take care of this community that is able to build so much together by sharing. As all of us you have made mistakes and we hope you will learn from them.\n\nWe also want to **thank all of you**, Kagglers. We love and are grateful to be part of our group/community. Thanks for sharing and for the collective learning experience.\n\n# Context\n\n- Business context: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview,\n- Data context: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data.\n\n# Overview of the Approach\n\nOur solution is essentially a blend of a XGBoost and a NN models. Both heavily rely on duration that appeared to be a powerful leverage. Time was aggregated in different ways and combined with counts for the GBDT while it is transformed via a custom TimeEmbedding block based on 1D convolutions that produce a representation combined with user event representations for the NN.\nRobustness and efficiency founded our work. XGBoost models were validated on 10 bags of 5 folds and features incorporated only if the mean of the CV of these 10 bags was greater than the level of noise we quantified while we opted for a majority/consensus strategy to build the NN, i.e. validate choices only if 4 of 5 folds were improved. The 3rd place of the efficiency LB was achieved with a lightweight NN accelerated via TF Lite.\n\n# Details of the submission\n\n## Code\n\nAfter publishing this write-up we decided to open our code: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332.\nIt is composed by several parts: [how to train the XGBoost models](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training), how to [pretrain](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining) and [train](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training) the NN models and the [inference notebook](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference) used to win this competition.\n\n## Data\n\nLooking at the 1st data released showed that there aren't a lot of sessions so not a lot of sequences. Moreover these are long sequences. This is not ideal for a deep learning approach.  \nExploring the Field Day Lab research instructed that the Jo Wilder application was built to help learning to read and that way more than 11,500 learners had played this game.  \nThese 2 ideas led to search for a bigger dataset. In 1 Google search and 3 clicks we came up to the open data portal (https://fielddaylab.wisc.edu/opengamedata/) which contains a lot of sessions. 1 hour and 3 bash commands latter we knew that the train set was in part in the open data. So we took a week to **build a pipeline that extracts 98 % of the sessions of the train set perfectly and with minor errors for the last 2 %**. Our data are even better than the comp data because we knew before the host confirmation that for the sessions with 2 games the target was skewed (0 if wrong in 1 of the 2 games when we aim at predicting the responses for the 1st game). It seems that fixing these targets can bring a significant boost up to +0.002.\n\nWe took 1 more week to build a GBDT/XGBoost baseline that would have scored top 10 given the CV score, with the use of the supplemental data (~20,000 sessions) that gave +0.003/0.004 at that time. As we simulated the API locally (see after), we used some training sessions to infer and noticed that it scored 0.718. We were hoping that the LB sessions were not part of the open data portal but our 1st submission, LB 0.708, immediately showed to us that we had rebuilt about a half of the data and especially the targets in the public LB, because 0.708 = (0.698 + 0.718) / 2. The host and Kaggle have been immediately informed. You know what happened next (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/415820).  \nAfter the release of the LB data we measured that we perfectly rebuilt ~7000 sessions over the ~11,500 of the LB data.\n\nWe spent the first month exploring the data until we understood/knew it pretty well. For example we even reconstituted sessions for what might be schools (several games on 1 IP session), extracted every single session with at least 1 answer, ...\n\nAfter the update we made a first submission that scored 0.72. This was shocking because this meant that some leaked data were remaining. A few days later we noticed that the open data was not totally similar with the state we found it 1 month before. A file was missing. So we returned to the host and Kaggle to give them more work.\n\n**This process/work led us to perfectly understand the data model** (that changed since the 1st release of the game). This also allowed us to deeply understand the data itself.\n\nNote that we only used the sessions for which we had responses to all questions of the 2 1st level groups. 1) This is more consistant with the sessions we want to predict (game from the beginning to the end) and 2) this approach preserves performance (vs all data) while reducing training time.\n\nOur dataset is constituted by **37323 complete sessions (23562 comp + logs) in a total of 66376 sessions**.\n\nThe supplemental data (that we fully added 1 month ago) gave us consistently **CV +0.002**.\n\n## Model\n\nOur solution is mainly an ensemble of GBDT + NN models.\n\n### Trust your validation\n\nWe think that **the main reason of the robustness of our solution is that we only relied on CV** for decision making. No choice had been made on LB.\n\nProbing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API. This is a small set. In our experiments 5,000 sessions is the minimum to guarantee a stable CV/LB alignment. A set less than 2,000 is very noisy so **robustness was the way to go**.\n\nWe **only added features that improved the CV for sure**. This is not easy to delete features that you believe in but this is needed as science is not a matter of belief. There are several ways to do so: for example monitor all folds in a CV (and accept only on majority or consensus), monitor several bags (composition of CV to not overfit validation), ...\n\nFor the GBDT approach, we mainly validated on the mean of 10 bags (we defined a bag as a composition of the folds). As we estimated the noise to be ~0.0003, only improvements greater than the noise have been considered. For the NN as we needed to iterate quicker we only used a single bag and only incorporated > 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.\n\n### Metric\n\nWe experimented a lot on finding a threshold by question but found that this approach is less robust than a single threshold. We mainly used 0.625 as global threshold despite our highest LB scores that were obtained with a threshold per question.\n\n### GBDT\n\nWe prototyped a baseline with **XGBoost because of the structured/tabular nature of the data**. The feature engineering process is interesting to understand what is predictive and to understand the causation, i.e. how the features or decision criteria that enable to predict correctly.\n\nGenerally speaking we followed 3 ways to build features: **business knowledge**, our **intuition** playing the game and a meticulous **exploration of the data**.\nBusiness knowledge refers to using expert knowledge. Reading the papers of the researchers that built this game allow to understand the game beyond usage. For example, Jo Wilder has been built to improve the players reading skills. So this means that the text duration should be important. These are like killer features.\n\nWe exclusively made use of Polars because of the CPU constraints and to simply learn it.\nOur features (663, 1993, 3734 for each level\\_group) are mainly **durations and counts for different aggregations**: how much time in a level, in a room, reading a text, interacting in some way (event type), how many events in a level_group, how many events of each type, how many events of each type in a room or a level, ...  \nWe also built a few notebook dedicated features: how many type of events on the notebook in a level, ...  \nDespite our efforts we weren't able to extract useful information from the coordinates, the only few features of this type had been mean and std for some events in the activities (journal interactions for example).\n\nWe considered that injecting targets predicted in the previous level groups was a compression of the signal, meaning a loss of information, so we used, for each session, **all interactions from the beginning of the game/session**. This led to a +0.002 at the time of this choice.\n\nAfter the API needed to order the data, we noticed that **models trained both on original order and on index order** but validated on index order (inference order) improved our scores. This leads to more variety that was needed to **improve stability and robustness**. The same goes for the composition of the validation sets: usage of several bags (composition of validation sets) based on the comp data but also on the extracted data improved our scores. We detected late that increasing the number of folds from 5 to 10 could also be leveraged.\n\nThe code for GBDT allows to switch from XGBoost to LightGBM and CatBoost with a simple variable parameter but despite the good scores (~0.001 less than XGBoost), this did not bring to ensemble so we sticked to only XGBoost.\n\nWe experimented a lot around feature selection but were unable to build a stable strategy. So instead of a top-down approach consisting in deleting useless features, we adopted a bottom-up approach choosing carefully each group of features.\n\nOur **XGBoost models score CV ~0.7025 +/-0.0003** and blending 5 of them (the only XGBoost we still have with correct score) scores **LB 0.704**.\n\n### NN\n\nAfter achieving a good score with gradient boosting and having understood well the data we focused on deep learning.\n\nThe **first attempt was with Transformers**. The 1st results were disappointed: CV 0.685 with 2 hours / fold (as far as we can remember). Transformers are very computationally intensive. Resources: https://arxiv.org/pdf/1912.09363.pdf, https://arxiv.org/pdf/2001.08317.pdf, https://arxiv.org/pdf/1711.03905.pdf, https://arxiv.org/pdf/1907.00235.pdf, ...\n\nWe then gave a try to **Conv1D**. In one day we had a very simple model that scored as Transformers but **10x faster** allowing to iterate quicker. So we pushed this approach and could seamlessly scaled it beyond our expectations.\n\nDifficult to share the **tens or hundreds of experimentations** needed to achieve the final solution which is both based on a simple architecture and a slightly complex training pipeline.\n\n#### Architecture roots\n\nWe browsed the literature based on the question: how to model time in deep learning?  \nThis research made us come to the idea of **time-aware events** (i.e. https://proceedings.mlr.press/v126/zhang20c/zhang20c.pdf) and back to **WaveNet** (https://arxiv.org/pdf/1609.03499.pdf) because it uses **Conv1D to model long sequences with considerations on causation**.  \nOther papers also inspired us: https://arxiv.org/pdf/1703.04691.pdf build on top of WaveNet paper for time series, https://idus.us.es/bitstream/handle/11441/114701/Short-Term%20Load%20Forecasting%20Using%20Encoder-Decoder%20WaveNet.pdf?sequence=1&isAllowed=y also build on top of WaveNet.  \nWe also have to mention the excellent work that @abaojiang shared (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565 and https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part). It inspired our research and maybe successfully biased it.\n\nLet's focus on the model of our efficiency submission that is also one of our final ensemble and which performance is nearly the same as models with a few more features.\n\n#### Feature representations\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Ff692423c63172494ea4df97b01d65d33%2Fdata.png?generation=1702113769960411&alt=media)\n\n**5 features as inputs: duration, text\\_fqid, room\\_fqid, fqid, event\\_name + name** (this is the event type from the original data model as far as we remember). Each of these information is encoded/embedded into a vector representation (d_model = 24) to be the merged. The 4 **categorical features feed a classical Embedding layer and the duration a TimeEmbedding** which is a custom block.\n\nDeveloping the GBDT solution showed that the **duration** was crucial, so we put a crucial amount of time trying to model it greatly. The TimeEmbedding layer is a composition of 4x ConvBlock which is inspired by the Transformer main block: Conv1D -> skip connection -> layer norm -> dropout.\n\n```\nclass TimeEmbedding(tf.keras.layers.Layer):\n    def __init__(self, n_blocks, d_model, dropout_rate):\n        super(TimeEmbedding, self).__init__()\n        self.conv_blocks = [ConvBlock(d_model, dropout_rate=dropout_rate) for _ in range(n_blocks)]\n        \n    def call(self, inputs):\n        x = tf.expand_dims(inputs, axis=-1)\n        for conv_block in self.conv_blocks:\n            x = conv_block(x)\n        return x\n```\n\n```\nclass ConvBlock(tf.keras.layers.Layer):\n    def __init__(self, d_model, dropout_rate):\n        super(ConvBlock, self).__init__()\n        self.conv1d = tf.keras.layers.Conv1D(d_model, kernel_size=5, padding='same', activation='gelu')\n        self.layer_norm = tf.keras.layers.LayerNormalization()\n        self.dropout = tf.keras.layers.Dropout(rate=dropout_rate)\n        \n    def call(self, inputs):\n        x = self.conv1d(inputs)\n        x = x + inputs\n        x = self.layer_norm(x)\n        outputs = self.dropout(x)\n        return outputs\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F9407390ffd7f226b7be0f57c37051f77%2Ftime_embedding.png?generation=1702113795144186&alt=media)\n\n#### Time-aware events\n\nAs said, the goal of building these representations was to model time-aware events. We considered the **categorical features as events** because they represent the user interactions with business entities of the game. We then tried to incorporate duration to make them time-awared. Our main intuition showed to be the best. It is a **simple solution based on operation priority to represent that the duration should be associated to each event before associated them together**: duration * event_1 + duration * event_2 + ... which had been factorized to duration * (event_1 + event_2 + ...).\n\n```\nclass ConvNet(tf.keras.Model):\n    def __init__(self, input_dims, n_outputs, d_model, n_blocks=4, name=None):\n        super(ConvNet, self).__init__(name=name)\n        self.input_dims = input_dims\n        self.n_outputs = n_outputs\n        self.d_model = d_model\n        self.n_blocks = n_blocks\n        self.event_embedding = tf.keras.layers.Embedding(input_dims['event_name_name'], d_model, mask_zero=True)\n        self.room_embedding = tf.keras.layers.Embedding(input_dims['room_fqid'], d_model, mask_zero=True)\n        self.text_embedding = tf.keras.layers.Embedding(input_dims['text'], d_model, mask_zero=True)\n        self.fqid_embedding = tf.keras.layers.Embedding(input_dims['fqid'], d_model, mask_zero=True)\n        self.duration_embedding = TimeEmbedding(n_blocks=n_blocks, d_model=d_model, dropout_rate=0.2)\n        self.gap = tf.keras.layers.GlobalAveragePooling1D()\n        \n    def call(self, inputs):\n        event = self.event_embedding(inputs['event_name_name'])\n        room = self.room_embedding(inputs['room_fqid'])\n        text = self.text_embedding(inputs['text'])\n        fqid = self.fqid_embedding(inputs['fqid'])\n        duration = self.duration_embedding(inputs['duration'])\n        x = duration * (event + room + text + fqid)\n        outputs = self.gap(x)\n        return outputs\n\n    def get_config(self):\n        config = super().get_config().copy()\n        config.update({\n            'input_dims': self.input_dims,\n            'n_outputs': self.n_outputs,\n            'd_model': self.d_model,\n            'n_blocks': self.n_blocks,\n            'name': self._name,\n        })\n        return config\n\n    @classmethod\n    def from_config(cls, config):\n        return cls(**config)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Fca563daa4f9ef7bf8d8b20e0c85c7951%2Ftime_aware_events_1.png?generation=1702113822015438&alt=media)\nThe 2 representations are equivalent: either you can think time-aware events as a combination of time and sub-events or as a combination of sub-events and time.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F1f65d399dfd93f128ca6c887d0f0be89%2Ftime_aware_events_2.png?generation=1702113853644045&alt=media)\n \n#### Training pipeline\n\nThe training pipeline is not totally straight forward.\n\n@dongyk published great schematics that can be useful to illustrate what is explained bellow: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166.\n\n##### 1st step (pre-training?)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F970baa66fc37b301bb8960cc92cf8ab4%2Fpre_training.png?generation=1702113879947155&alt=media)\n\nThe best approach for us consists in a kind of **backbone that represents the events of a level_group**.\n\nThis backbone is trained on all the data available for this level_group (i.e. on complete + incomplete sessions). It is associated with a temporary SimpleHead optimizing BCE loss.\n\n```\nclass SimpleHead(tf.keras.Model):\n    def __init__(self, n_units, n_outputs, name=None):\n        super(SimpleHead, self).__init__(name=name)\n        self.ffs = [tf.keras.layers.Dense(units, activation='gelu') for units in n_units]\n        self.out = tf.keras.layers.Dense(n_outputs, activation='sigmoid')\n        \n    def call(self, inputs):\n        x = inputs\n        for ff in self.ffs:\n            x = ff(x)\n        outputs = self.out(x)\n        return outputs\n```\n\nThis approach allows to score **CV 0.70025 +/- 0.0005**.\n\n##### 2nd step (training?)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F98e32800c54ffab63d98d90d220fbbcd%2Fend_2_end.png?generation=1702113900174294&alt=media)\n\n**The weights of each of the 3 backbones (1 by level_group) are freezed** for the 2nd level of training to speedup training but also because it is more stable and efficient. These backbones can be thought as \"embedders\".\n\nDuring this 2nd step, **all the submodels that composed the solution were trained on all complete sessions in an end-to-end setup**. The input data are 3 sequences of the 5 features, 1 for each of the 3 level groups. Each \"embedders\" outputs a 24 dim-vector representation. These outputs are the inputs of a head in which enters the representation of level\\_group '0-4' to predict the 3 first questions and the concatenation of the previous and the current representations for level_groups '5-12' and '13-22' to make use of all information.\n\nProceeding like this allows to optimize the overall performance and to monitor it based on the F1 score that is the score of the competition. This means we optimized BCE with F1 score as a metric.\n\nOur winning submission uses a simple **MLP head** but also a **skip head** (512 -> 512 -> 512 allow it for example). **MMoE** did not improve the simplest approaches.\n\nThis approach allows to score **CV 0.70175 +/- 0.0003** which is **comparable to the GBDT solution**.\n\n### Inference\n\n#### Build a simulator\n\nEarly in the competition we built a simulator of the API. Doing so we never experimented any submission error. Maybe trying to keep ideas and code as simple as possible was also key to debug easily.\n\n#### Efficiency\n\nWe invested the efficiency part of the challenge for GBDT as well as NNs.  \nUsing **Treelite** for XGBoost allow us to divide by 2 the execution time.  \nOur deep learning models were lights: **400,000 weights** for the end-to-end model which combines every parts/sub-models. Having already used **TF Lite** we knew it could be a game changer. Converting our models led to a significant boost in inference time without any performance loss (we do not remember exactly but we think it is at least **6x faster** on our local inference simulator).  \nBeginning to explore pruning as well as hard quantization showed that the performance loss would be significant (which is OK in production but not in a competition) so we sticked to a simple TF Lite conversion.\n\nWe have not leveraged what seems to be a problem in the efficiency metric. As we identified the private test sessions to be the 1450/1500 first served by the API we tried to just predict the others to check which time was used (public for public and private for private). Doing so we gain a place but choose to not use this.\n\nOur **efficiency submission is a NN that scores public LB 0.702 and private LB 0.699 in less than 5 minutes**.\n\n#### Ensemble\n\nWe experimented a lot of ensembling alternatives. In the end we sticked to a simple average 50/50 GBDT/NN with:\n\n  * 2 kinds of GBDT: trained on original order + trained on index order (validated on index order that is the inference case),\n  * 3 kinds of NNs: trained on original order + trained on index order with 5 or all features.\n  \nAs our models are lightweight we were able to build a hugh ensemble: **2 x 4 x 10 folds XGBoost + 3 x 4 x 5 folds NNs**. The bottleneck for us is the 8 Go RAM constraint.\n\nThe winning submission scores **CV 0.705, public LB 0.705 and private LB 0.705**.\n\n## Conclusion\n\nThe main achievement of our work is that it is a good solution for the researchers, learners and children that can benefit of it and we hope it will contributes to progresses for a better learning experience. Up to you guys!\n\nThanks if you read until here!\nIf you have any question do not hesitate to ask. We will do our best to respond.\n\n## Presentation to the host\n\nA video presentation to the host has been recorded and can be available on demand. Feel free to ask via PM.\n\n# Sources\n\nBelow are the main sources that we used. More sources can be found in section *Details of the submission* above.\n\n- https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332,\n- https://fielddaylab.wisc.edu/opengamedata/,\n- https://arxiv.org/pdf/1609.03499.pdf,\n- https://www.tensorflow.org/lite/guide",
      "votes": 168
    },
    {
      "id": 2323315,
      "postDate": "2023-06-29T20:59:06.750Z",
      "content": "<p>Huge congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a>! Thanks for sharing such a great solution. Double congrats for <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for finally break the 2nd place curse and get your first winning!</p>",
      "rawMarkdown": "Huge congrats @cpmpml and @pdnartreb! Thanks for sharing such a great solution. Double congrats for @cpmpml for finally break the 2nd place curse and get your first winning!",
      "votes": 7
    },
    {
      "id": 2323974,
      "postDate": "2023-06-30T09:41:15.223Z",
      "content": "<p>Congrats, impressive solution! Thank you for your detailed explanation. </p>\n<p>Will you share the kernel? I can't wait to learn.</p>",
      "rawMarkdown": "Congrats, impressive solution! Thank you for your detailed explanation. \n\nWill you share the kernel? I can't wait to learn.",
      "votes": 3,
      "replies": [
        {
          "id": 2324001,
          "postDate": "2023-06-30T10:01:06.463Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/zrongchu\" target=\"_blank\">@zrongchu</a>!<br>\nWe have decided to share our code: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332</a>.</p>",
          "rawMarkdown": "Thanks @zrongchu!\nWe have decided to share our code: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332.",
          "votes": 3,
          "replies": [
            {
              "id": 2324042,
              "postDate": "2023-06-30T10:49:01.050Z",
              "content": "<p>Thanks! <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> , <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
              "rawMarkdown": "Thanks! @pdnartreb , @cpmpml ",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2339570,
      "postDate": "2023-07-10T22:35:08.770Z",
      "content": "<p>Greetings, <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> !</p>\n<p>I wanted to express my sincere appreciation and congratulations for your outstanding work in this recent competition. Your solution is truly impressive and it's evident that you put a tremendous amount of effort and expertise into developing it.</p>\n<p>Your focus on understanding the data model and extracting valuable insights is impressive. The combination of GBDT and NN models in your ensemble and search of external feature engineering demonstrates your deep understanding of the problem and your ability to leverage different techniques effectively.</p>\n<p>I have a question regarding your approach: In the training pipeline, you mentioned that the backbone, representing the events of a level group, is trained on all the available data for that group. Could you elaborate on how you handle the incomplete sessions during this training phase? How do you ensure that the backbone captures the most relevant information from both complete and incomplete sessions?</p>\n<p>Once again, congratulations on your remarkable achievement, and thank you for your contributions to the Kaggle community. Your solution is an inspiration to fellow data scientists and serves as a testament to your expertise.</p>\n<p>Best regards,keep up with the great work🦾🚀</p>",
      "rawMarkdown": "Greetings, @pdnartreb and @cpmpml !\n\nI wanted to express my sincere appreciation and congratulations for your outstanding work in this recent competition. Your solution is truly impressive and it's evident that you put a tremendous amount of effort and expertise into developing it.\n\nYour focus on understanding the data model and extracting valuable insights is impressive. The combination of GBDT and NN models in your ensemble and search of external feature engineering demonstrates your deep understanding of the problem and your ability to leverage different techniques effectively.\n\nI have a question regarding your approach: In the training pipeline, you mentioned that the backbone, representing the events of a level group, is trained on all the available data for that group. Could you elaborate on how you handle the incomplete sessions during this training phase? How do you ensure that the backbone captures the most relevant information from both complete and incomplete sessions?\n\nOnce again, congratulations on your remarkable achievement, and thank you for your contributions to the Kaggle community. Your solution is an inspiration to fellow data scientists and serves as a testament to your expertise.\n\nBest regards,keep up with the great work🦾🚀",
      "votes": 1,
      "replies": [
        {
          "id": 2340147,
          "postDate": "2023-07-11T09:58:40.450Z",
          "content": "<p>I am not sure I understand what you don't understand.</p>\n<p>Let me try still. Maybe this thread can help: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165</a></p>\n<p>For a given NN there are three NNs trained in the \"pretraining hase\", then they are assembled into a single NN in a second training phase. </p>\n<p>Incomplete sessions are only used in the pretraining phase.</p>\n<p>For instance, if a session has only data for the first level group (i.e. questions 1,2,3), then it is used to train the first of the three NNs, and it is not used for the other two.</p>",
          "rawMarkdown": "I am not sure I understand what you don't understand.\n\nLet me try still. Maybe this thread can help: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165\n\nFor a given NN there are three NNs trained in the \"pretraining hase\", then they are assembled into a single NN in a second training phase. \n\nIncomplete sessions are only used in the pretraining phase.\n\nFor instance, if a session has only data for the first level group (i.e. questions 1,2,3), then it is used to train the first of the three NNs, and it is not used for the other two.",
          "votes": 3,
          "replies": [
            {
              "id": 2341136,
              "postDate": "2023-07-11T21:56:24.593Z",
              "content": "<p>Thanks for the quick reply!</p>\n<p>Indeed the following <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165\" target=\"_blank\">thread</a> has helped to interpret the model training phase. Specially the following schematics:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F960b98da4ee0cefe936edf44d5fec9e9%2F1.png?generation=1688613004356860&amp;alt=media\" alt=\"\"><br>\nKudos to <a href=\"https://www.kaggle.com/dongyk\" target=\"_blank\">@dongyk</a> , for that.</p>\n<p>It seems that you have leveraged the full dataset (complete + incomplete sessions), to create specialized models for each level_group, concatenating the inputs. Which enabled the model to capture specific information of all the level_groups independently and all together, by concatenating the model inputs. Therefore you could optimize BCE with F1 score as metric, since it turned out to be a classification problem….</p>\n<p>Once again congratulations <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for the excellent work and thanks for contributing to the Kaggle Community🙏🚀</p>",
              "rawMarkdown": "Thanks for the quick reply!\n\nIndeed the following [thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165) has helped to interpret the model training phase. Specially the following schematics:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F960b98da4ee0cefe936edf44d5fec9e9%2F1.png?generation=1688613004356860&alt=media)\nKudos to @dongyk , for that.\n\nIt seems that you have leveraged the full dataset (complete + incomplete sessions), to create specialized models for each level_group, concatenating the inputs. Which enabled the model to capture specific information of all the level_groups independently and all together, by concatenating the model inputs. Therefore you could optimize BCE with F1 score as metric, since it turned out to be a classification problem....\n\nOnce again congratulations @pdnartreb and @cpmpml for the excellent work and thanks for contributing to the Kaggle Community🙏🚀\n\n\n",
              "votes": 1
            },
            {
              "id": 2341645,
              "postDate": "2023-07-12T07:59:22.233Z",
              "content": "<p>yes, that's what we did:</p>\n<blockquote>\n  <p>It seems that you have leveraged the full dataset (complete + incomplete sessions), to create specialized models for each level_group, concatenating the inputs. </p>\n</blockquote>",
              "rawMarkdown": "yes, that's what we did:\n\n> It seems that you have leveraged the full dataset (complete + incomplete sessions), to create specialized models for each level_group, concatenating the inputs. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2334085,
      "postDate": "2023-07-07T12:53:45.227Z",
      "content": "<p>Congratulations ! Very interesting !</p>",
      "rawMarkdown": "Congratulations ! Very interesting !",
      "votes": 1
    },
    {
      "id": 2333655,
      "postDate": "2023-07-07T05:31:07.173Z",
      "content": "<p>Congratulations! very helpful. </p>",
      "rawMarkdown": "Congratulations! very helpful. ",
      "votes": 1
    },
    {
      "id": 2332588,
      "postDate": "2023-07-06T09:47:28.037Z",
      "content": "<p>Very insightful </p>",
      "rawMarkdown": "Very insightful ",
      "votes": 1
    },
    {
      "id": 2331909,
      "postDate": "2023-07-05T21:38:04.753Z",
      "content": "<p>Congratulations! Really interesting work</p>",
      "rawMarkdown": "Congratulations! Really interesting work",
      "votes": 1
    },
    {
      "id": 2331165,
      "postDate": "2023-07-05T12:05:07.797Z",
      "content": "<p>Hi, I was wondering whether I understand it correctly.<br>\n<a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">PSPFGP 1st Place - NN Training</a></p>\n<pre><code>outputs = {}\noutputs = heads(convnet_outputs)\noutputs = heads(\n        tf()(, convnet_outputs])\n    )\noutputs = heads(\n        tf()(, convnet_outputs, convnet_outputs])\n    )\n</code></pre>\n<p>Does it describe like this??<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F5bd337c5017d3ab5606024c830391f40%2F.png?generation=1688556655026917&amp;alt=media\" alt=\"\"><br>\nBackpropagation is performed: <br>\nfrom outputs['0-4'] to convnets['0-4'],    (I will call this as submodel A)<br>\nfrom outputs['5-12'] to convnets['0-4'] and convnets['5-12'], and     (submodel B)<br>\nfrom outputs['13-22'] to convnets['0-4'], convnets['5-12'], and convnets['13-22'].     (submodel C)<br>\nHowever, this submodels are separated. So, Backpropagation which is performed from outputs['13-22'] to convnets['0-4'] (submodel C) does not affect convnets['0-4'] of submodel A.<br>\nDid I understand correctly??</p>\n<p>+) I was really surprised this approach to compute F1 score during training. Wow…</p>",
      "rawMarkdown": "Hi, I was wondering whether I understand it correctly.\n[PSPFGP 1st Place - NN Training](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training)\n```\noutputs = {}\noutputs['0-4'] = heads['0-4'](convnet_outputs['0-4'])\noutputs['5-12'] = heads['5-12'](\n        tf.keras.layers.Concatenate()([convnet_outputs['0-4'], convnet_outputs['5-12']])\n    )\noutputs['13-22'] = heads['13-22'](\n        tf.keras.layers.Concatenate()([convnet_outputs['0-4'], convnet_outputs['5-12'], convnet_outputs['13-22']])\n    )\n```\nDoes it describe like this??\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F5bd337c5017d3ab5606024c830391f40%2F.png?generation=1688556655026917&alt=media)\nBackpropagation is performed: \nfrom outputs['0-4'] to convnets['0-4'],    (I will call this as submodel A)\nfrom outputs['5-12'] to convnets['0-4'] and convnets['5-12'], and     (submodel B)\nfrom outputs['13-22'] to convnets['0-4'], convnets['5-12'], and convnets['13-22'].     (submodel C)\nHowever, this submodels are separated. So, Backpropagation which is performed from outputs['13-22'] to convnets['0-4'] (submodel C) does not affect convnets['0-4'] of submodel A.\nDid I understand correctly??\n\n+) I was really surprised this approach to compute F1 score during training. Wow...",
      "votes": 1,
      "replies": [
        {
          "id": 2331653,
          "postDate": "2023-07-05T17:22:41.923Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dongyk\" target=\"_blank\">@dongyk</a>!</p>\n<p>You are working hard. Great! This is really satisfying to see that our work is useful for you. With this mindset you are going to learn a lot.</p>\n<p>Your understanding and your schematic seem nearly OK.  <br>\nKeep in mind that in step 2 of the training the weights of the ConvNets are freezed. This means backpropagation only changes the heads weights.  <br>\nIn this 2nd step we have 3 sequences of data as inputs, 1 for each level_group. Each of these sequences feeds its own ConvNet/embedder that outputs a representation in the form of a 24 dim vector. The representations for level_group 0-4 feeds the SimpleHead for level_group 0-4 which outputs a vector of 3 values that are the probability of the 3 responses of this level_group. The representations for level_group 0-4 and 5-12 are concatenated in a 48 dim vector the feeds the SimpleHead for the 2nd level_group that outputs 10 values that correspond to the 10 questions of the level_group 5-12. Same for level_group 13-22.</p>\n<p>You correctly spotted that this had been made to train the whole ensemble to optimize the overall F1 score that is what we want.<br>\nWe tried every setup: training the whole ensemble from scratch, each part independently, … The setup we presented here corresponds to what worked best for us.</p>\n<p>Does this respond to your question(s)? If not feel free to ask.</p>",
          "rawMarkdown": "Hi @dongyk!\n\nYou are working hard. Great! This is really satisfying to see that our work is useful for you. With this mindset you are going to learn a lot.\n\nYour understanding and your schematic seem nearly OK.  \nKeep in mind that in step 2 of the training the weights of the ConvNets are freezed. This means backpropagation only changes the heads weights.  \nIn this 2nd step we have 3 sequences of data as inputs, 1 for each level_group. Each of these sequences feeds its own ConvNet/embedder that outputs a representation in the form of a 24 dim vector. The representations for level_group 0-4 feeds the SimpleHead for level_group 0-4 which outputs a vector of 3 values that are the probability of the 3 responses of this level_group. The representations for level_group 0-4 and 5-12 are concatenated in a 48 dim vector the feeds the SimpleHead for the 2nd level_group that outputs 10 values that correspond to the 10 questions of the level_group 5-12. Same for level_group 13-22.\n\nYou correctly spotted that this had been made to train the whole ensemble to optimize the overall F1 score that is what we want.\nWe tried every setup: training the whole ensemble from scratch, each part independently, ... The setup we presented here corresponds to what worked best for us.\n\nDoes this respond to your question(s)? If not feel free to ask.",
          "votes": 1,
          "replies": [
            {
              "id": 2332166,
              "postDate": "2023-07-06T03:25:15.107Z",
              "content": "<p>Thank you for your answering!!<br>\nI'm learning a lot from your code.</p>\n<blockquote>\n  <p>We tried every setup: training the whole ensemble from scratch, each part independently, … The setup we presented here corresponds to what worked best for us.</p>\n</blockquote>\n<p>I was thinking the same. Did they try the whole ensemble from scratch? Because it's more cumbersome to train seperately. Training separately can be helpful to train.<br>\nIs the reason why you used only convnet['5-12'] for outputs['5-12'] not contained convnet['0-4']?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F824ba97fb079de1424be6625e9e99f0b%2F2.png?generation=1688620218055724&amp;alt=media\" alt=\"\"></p>\n<p>And… I editted the figure.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F960b98da4ee0cefe936edf44d5fec9e9%2F1.png?generation=1688613004356860&amp;alt=media\" alt=\"\"><br>\nI found that you fed train_dataset only once and the same convnet['0-4'] is fed into simplehead['0-4'], simplehead['5-12'], and simplehead['13-22'].<br>\nIs it more accurate?<br>\nIf I didn't freeze the convnets, can convnet['0-4'] backpropagated(affected) by not only outputs['0-4'] but outputs['5-12'] and outputs['13-22']?</p>",
              "rawMarkdown": "Thank you for your answering!!\nI'm learning a lot from your code.\n\n>We tried every setup: training the whole ensemble from scratch, each part independently, … The setup we presented here corresponds to what worked best for us.\n\nI was thinking the same. Did they try the whole ensemble from scratch? Because it's more cumbersome to train seperately. Training separately can be helpful to train.\nIs the reason why you used only convnet['5-12'] for outputs['5-12'] not contained convnet['0-4']?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F824ba97fb079de1424be6625e9e99f0b%2F2.png?generation=1688620218055724&alt=media)\n\n\nAnd... I editted the figure.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F960b98da4ee0cefe936edf44d5fec9e9%2F1.png?generation=1688613004356860&alt=media)\nI found that you fed train_dataset only once and the same convnet['0-4'] is fed into simplehead['0-4'], simplehead['5-12'], and simplehead['13-22'].\nIs it more accurate?\nIf I didn't freeze the convnets, can convnet['0-4'] backpropagated(affected) by not only outputs['0-4'] but outputs['5-12'] and outputs['13-22']?\n",
              "votes": 2
            },
            {
              "id": 2332546,
              "postDate": "2023-07-06T09:08:12.617Z",
              "content": "<p>Wow! Thanks for your schematics that are 100% correct as I understand them! The write-up has been updated to link to your post.</p>\n<p>If the weights of the convnets are not freezed then they are updated by the backprop coming from each output they are linked to when training end-to-end.</p>\n<p>We first trained separatly which gave a baseline. Then we tried to train the end-to-end setup from scratch. As you might have noticed the output allow to monitor the loss which is composed by the losses of each level group that can also be monitored. We observed that the loss for 0-4 was the first to converge but also that, as the others go down, it tends to increase. We interpreted that as the optimization of previous level groups representation for next level groups. This is not what we wanted. We wanted a good representation for a level group that can be used by the next level groups. So we chose to first optimize each convnet separatly and to exploit its representation by heads in a second time.</p>\n<p>As F1 score cannot be optimized directy, the main goal of the end-2-end approach was to be able to monitor the F1 score to select the best weights.</p>",
              "rawMarkdown": "Wow! Thanks for your schematics that are 100% correct as I understand them! The write-up has been updated to link to your post.\n\nIf the weights of the convnets are not freezed then they are updated by the backprop coming from each output they are linked to when training end-to-end.\n\nWe first trained separatly which gave a baseline. Then we tried to train the end-to-end setup from scratch. As you might have noticed the output allow to monitor the loss which is composed by the losses of each level group that can also be monitored. We observed that the loss for 0-4 was the first to converge but also that, as the others go down, it tends to increase. We interpreted that as the optimization of previous level groups representation for next level groups. This is not what we wanted. We wanted a good representation for a level group that can be used by the next level groups. So we chose to first optimize each convnet separatly and to exploit its representation by heads in a second time.\n\nAs F1 score cannot be optimized directy, the main goal of the end-2-end approach was to be able to monitor the F1 score to select the best weights.",
              "votes": 1
            },
            {
              "id": 2332884,
              "postDate": "2023-07-06T13:36:42.110Z",
              "content": "<p>Cool.<br>\nI feel good to understand it.</p>\n<blockquote>\n  <p>We observed that the loss for 0-4 was the first to converge but also that, as the others go down, it tends to increase. We interpreted that as the optimization of previous level groups representation for next level groups. This is not what we wanted. We wanted a good representation for a level group that can be used by the next level groups. So we chose to first optimize each convnet separatly and to exploit its representation by heads in a second time.</p>\n</blockquote>\n<p>I also tried to come up with a way to change the architecture of the model, but I realized that it's the best idea considering the pros and cons of possible models.</p>\n<blockquote>\n  <p>As F1 score cannot be optimized directy, the main goal of the end-2-end approach was to be able to monitor the F1 score to select the best weights.</p>\n</blockquote>\n<p>The way I see it, it's the most important approach in the NN code.</p>\n<p>Again, thank you so much.</p>",
              "rawMarkdown": "Cool.\nI feel good to understand it.\n\n>We observed that the loss for 0-4 was the first to converge but also that, as the others go down, it tends to increase. We interpreted that as the optimization of previous level groups representation for next level groups. This is not what we wanted. We wanted a good representation for a level group that can be used by the next level groups. So we chose to first optimize each convnet separatly and to exploit its representation by heads in a second time.\n\nI also tried to come up with a way to change the architecture of the model, but I realized that it's the best idea considering the pros and cons of possible models.\n\n>As F1 score cannot be optimized directy, the main goal of the end-2-end approach was to be able to monitor the F1 score to select the best weights.\n\nThe way I see it, it's the most important approach in the NN code.\n\nAgain, thank you so much.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2326886,
      "postDate": "2023-07-02T13:22:02.613Z",
      "content": "<p>bravo 1St place</p>",
      "rawMarkdown": "bravo 1St place",
      "votes": 1
    },
    {
      "id": 2326179,
      "postDate": "2023-07-02T00:55:58.820Z",
      "content": "<p>Congratulations on your win🎉 And thanks you also for sharing your solution!</p>\n<p>If I may, may I ask a question about the third chapter, the beginning of the <strong>Data</strong>?</p>\n<p>Due to the small number of game sessions, the length of the sequence (i.e. the length of the game to be put into the NN) is not long enough. This as not ideal for the Deep Learning approach, I interpreted.<br>\nIf so, what does it mean by \"Moreover these are long sequences.\"?　Does it relevant that in some cases the game is played many times and is actually a long sequence, but what we have as data is a short sequence?</p>",
      "rawMarkdown": "Congratulations on your win🎉 And thanks you also for sharing your solution!\n\nIf I may, may I ask a question about the third chapter, the beginning of the **Data**?\n\nDue to the small number of game sessions, the length of the sequence (i.e. the length of the game to be put into the NN) is not long enough. This as not ideal for the Deep Learning approach, I interpreted.\nIf so, what does it mean by \"Moreover these are long sequences.\"?　Does it relevant that in some cases the game is played many times and is actually a long sequence, but what we have as data is a short sequence?",
      "votes": 1,
      "replies": [
        {
          "id": 2326740,
          "postDate": "2023-07-02T11:11:22.223Z",
          "content": "<p>There are too few sequences to enable complex NN model training. But sequence average length is rather long.</p>",
          "rawMarkdown": "There are too few sequences to enable complex NN model training. But sequence average length is rather long.",
          "votes": 2,
          "replies": [
            {
              "id": 2326787,
              "postDate": "2023-07-02T11:59:11.803Z",
              "content": "<p>Thank you.  I understood!</p>",
              "rawMarkdown": "Thank you.  I understood!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2325659,
      "postDate": "2023-07-01T14:02:09.720Z",
      "content": "<p><a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> 1st place bravo. Great Victory.</p>",
      "rawMarkdown": "@pdnartreb 1st place bravo. Great Victory.",
      "votes": 1
    },
    {
      "id": 2325346,
      "postDate": "2023-07-01T09:10:11.110Z",
      "content": "<p>Congratulations on your first place!</p>",
      "rawMarkdown": "Congratulations on your first place!",
      "votes": 1
    },
    {
      "id": 2325317,
      "postDate": "2023-07-01T08:49:29.843Z",
      "content": "<p>very useful, thank you for your fanscinating work.</p>",
      "rawMarkdown": "very useful, thank you for your fanscinating work.",
      "votes": 1
    },
    {
      "id": 2325268,
      "postDate": "2023-07-01T08:07:48.353Z",
      "content": "<p>Such an impressive combination of good ideas and hard hard work.</p>",
      "rawMarkdown": "Such an impressive combination of good ideas and hard hard work.",
      "votes": 1
    },
    {
      "id": 2325097,
      "postDate": "2023-07-01T05:51:09.790Z",
      "content": "<p>Congrats on winning first place! Thanks a lot for your detailed explanation. Impressive! 🔥🔥🔥🔥🔥🔥</p>",
      "rawMarkdown": "Congrats on winning first place! Thanks a lot for your detailed explanation. Impressive! 🔥🔥🔥🔥🔥🔥",
      "votes": 1
    },
    {
      "id": 2325053,
      "postDate": "2023-07-01T05:26:58.413Z",
      "content": "<p>Congrats on the win! Really love the NN design here. I has a very similar NN design, except that I just one-hot encode everything due to low cardinality of categorical features.</p>",
      "rawMarkdown": "Congrats on the win! Really love the NN design here. I has a very similar NN design, except that I just one-hot encode everything due to low cardinality of categorical features.",
      "votes": 1
    },
    {
      "id": 2324735,
      "postDate": "2023-06-30T20:45:31.277Z",
      "content": "<p>Wow, just wow! :D</p>\n<p>Such an impressive combination of good ideas and hard hard work. Very inspiring!</p>",
      "rawMarkdown": "Wow, just wow! :D\n\nSuch an impressive combination of good ideas and hard hard work. Very inspiring!",
      "votes": 1
    },
    {
      "id": 2324715,
      "postDate": "2023-06-30T20:21:55.020Z",
      "content": "<p>This is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lesson</p>",
      "rawMarkdown": "This is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lesson",
      "votes": 1
    },
    {
      "id": 2324446,
      "postDate": "2023-06-30T15:53:44.053Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> for winning the 1st place! Also, thanks a lot for the detailed explanation!</p>\n<p>May I ask how you decide to simply divide the numeric features (<em>e.g.,</em> <code>duration</code>) by a constant (60000 in the code you open source), instead of  using other normalization methods. I suppose that shrinking the scale of the raw feature could make the training process converge faster and other normalization techniques don't show superiority over the way you eventually adopted. If there's misunderstanding, please put me right. Thanks a lot!</p>",
      "rawMarkdown": "Congrats @cpmpml and @pdnartreb for winning the 1st place! Also, thanks a lot for the detailed explanation!\n\nMay I ask how you decide to simply divide the numeric features (*e.g.,* `duration`) by a constant (60000 in the code you open source), instead of  using other normalization methods. I suppose that shrinking the scale of the raw feature could make the training process converge faster and other normalization techniques don't show superiority over the way you eventually adopted. If there's misunderstanding, please put me right. Thanks a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 2324560,
          "postDate": "2023-06-30T17:20:55.463Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>!</p>\n<p>The writeup has been updated to mention your excellent notebook. We first saw the time-aware events and the Conv1D ideas in your work. Finding once more these ideas while exploring the litterature likely made us focus more on them.</p>\n<p>Continuous features are tricky to preprocess. We knew from our GBDT modeling work that time and especially <code>duration</code> was very important. We put a lot of careful efforts to try to model it correctly. Processing the <code>duration</code> seemed important so we tried a lot of things: scaling, normalization, standardization, clipping, … We explored the data to analyze the distribution of duration times and experimented with values around a baseline that seemed to us reasonnable because of the better distribution on the continuum. What worked best was clipping at 60 seconds + simple scaling.</p>",
          "rawMarkdown": "Thanks @abaojiang!\n\nThe writeup has been updated to mention your excellent notebook. We first saw the time-aware events and the Conv1D ideas in your work. Finding once more these ideas while exploring the litterature likely made us focus more on them.\n\nContinuous features are tricky to preprocess. We knew from our GBDT modeling work that time and especially `duration` was very important. We put a lot of careful efforts to try to model it correctly. Processing the `duration` seemed important so we tried a lot of things: scaling, normalization, standardization, clipping, ... We explored the data to analyze the distribution of duration times and experimented with values around a baseline that seemed to us reasonnable because of the better distribution on the continuum. What worked best was clipping at 60 seconds + simple scaling.",
          "votes": 2,
          "replies": [
            {
              "id": 2325122,
              "postDate": "2023-07-01T06:14:04.730Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a>,</p>\n<p>Thanks very much for the mention. I'm so glad that my inconspicuous sharing can inspire you and other teams to come up with far better solutions. And, I think that's the reason why I learn to share and gradually enjoy sharing!</p>\n<p>During the first month of the competition, I tried techniques like normalization, standardization, quantile transformation, <code>np.log1p</code>, etc., all of which didn't help improve the CV score. Hence, I sticked to the original processing method that clipped <code>duration</code> at <code>3.6e6</code> (<em>i.e.,</em> one hour). However, I didn't try to clip it with smaller values. I'll try them out and verify more ideas with late submissions!</p>\n<p>I learn so so much from your sharing. (1) Instead of just using beautiful visualization to interpret the data, you try hard to understand the mechanism behind raw data generation and clean the data to improve the data quality. (2) In addition to the model architecture design, robust training process and techniques are also crucial (<em>e.g.,</em> the loss criterion to use, whether to share \"embedders\" across different <code>level_group</code>, how to choose checkpoints). (3) Create a robust CV strategy and resist the temptation to select submission based on LB. (4) FE plays an important role all the time, to name a few.</p>\n<p>Thanks a lot for the clarification, and good luck with your next competition!!</p>",
              "rawMarkdown": "Hi @pdnartreb,\n\nThanks very much for the mention. I'm so glad that my inconspicuous sharing can inspire you and other teams to come up with far better solutions. And, I think that's the reason why I learn to share and gradually enjoy sharing!\n\nDuring the first month of the competition, I tried techniques like normalization, standardization, quantile transformation, `np.log1p`, etc., all of which didn't help improve the CV score. Hence, I sticked to the original processing method that clipped `duration` at `3.6e6` (*i.e.,* one hour). However, I didn't try to clip it with smaller values. I'll try them out and verify more ideas with late submissions!\n\nI learn so so much from your sharing. (1) Instead of just using beautiful visualization to interpret the data, you try hard to understand the mechanism behind raw data generation and clean the data to improve the data quality. (2) In addition to the model architecture design, robust training process and techniques are also crucial (*e.g.,* the loss criterion to use, whether to share \"embedders\" across different `level_group`, how to choose checkpoints). (3) Create a robust CV strategy and resist the temptation to select submission based on LB. (4) FE plays an important role all the time, to name a few.\n\nThanks a lot for the clarification, and good luck with your next competition!!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2324246,
      "postDate": "2023-06-30T13:42:44.707Z",
      "content": "<p>it is very useful content</p>",
      "rawMarkdown": "it is very useful content",
      "votes": 1
    },
    {
      "id": 2324241,
      "postDate": "2023-06-30T13:40:58.287Z",
      "content": "<p>Congrats and thank you for sharing 🎉 Impressive solution with many creative ideas!</p>",
      "rawMarkdown": "Congrats and thank you for sharing 🎉 Impressive solution with many creative ideas!",
      "votes": 1
    },
    {
      "id": 2324151,
      "postDate": "2023-06-30T12:21:00.883Z",
      "content": "<p>Congrats, mate!</p>",
      "rawMarkdown": "Congrats, mate!",
      "votes": 1
    },
    {
      "id": 2324107,
      "postDate": "2023-06-30T11:41:11.563Z",
      "content": "<p>Congratulations!  Keep it up!</p>",
      "rawMarkdown": "Congratulations!  Keep it up!",
      "votes": 1
    },
    {
      "id": 2324083,
      "postDate": "2023-06-30T11:20:05.807Z",
      "content": "<p>Congratulations on winning!</p>",
      "rawMarkdown": "Congratulations on winning!",
      "votes": 1
    },
    {
      "id": 2323256,
      "postDate": "2023-06-29T19:18:23.833Z",
      "content": "<p>Winner winner chicken dinner! :p</p>",
      "rawMarkdown": "Winner winner chicken dinner! :p",
      "votes": 1
    },
    {
      "id": 2323126,
      "postDate": "2023-06-29T17:21:02.953Z",
      "content": "<p>Congratulations on coming 1st <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> 🎉. Keep on achieving new milestones 👍</p>",
      "rawMarkdown": "Congratulations on coming 1st @pdnartreb 🎉. Keep on achieving new milestones 👍",
      "votes": 1,
      "replies": [
        {
          "id": 2324716,
          "postDate": "2023-06-30T20:22:44.147Z",
          "content": "<p>It's true he is very detailed in terms of giving knowledge I like </p>",
          "rawMarkdown": "It's true he is very detailed in terms of giving knowledge I like ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2327560,
      "postDate": "2023-07-03T03:34:57.963Z",
      "content": "<p>Congratulations! What a great work! You and your team deserve the 1st place.<br>\nI am a newbee. May I ask a question about probing the private test?<br>\nAs you mentioned here:</p>\n<blockquote>\n  <p>Probing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API.</p>\n</blockquote>\n<p>As far as what I know, the public test data can be probed, because the public learderboard give us the score on the public test dataset. But the private leaderboard seems won't give us feedback before the competition ends.<br>\nSo how do you probe the private test dataset?<br>\nThanks a lot.</p>",
      "rawMarkdown": "Congratulations! What a great work! You and your team deserve the 1st place.\nI am a newbee. May I ask a question about probing the private test?\nAs you mentioned here:\n>Probing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API.\n\nAs far as what I know, the public test data can be probed, because the public learderboard give us the score on the public test dataset. But the private leaderboard seems won't give us feedback before the competition ends.\nSo how do you probe the private test dataset?\nThanks a lot.",
      "votes": 2,
      "replies": [
        {
          "id": 2327778,
          "postDate": "2023-07-03T07:01:33.717Z",
          "content": "<p>You can switch some predictions and see if the public LB changes at all. If it does not change then the flipped predictions are on the private test set.</p>",
          "rawMarkdown": "You can switch some predictions and see if the public LB changes at all. If it does not change then the flipped predictions are on the private test set.",
          "votes": 2,
          "replies": [
            {
              "id": 2329321,
              "postDate": "2023-07-04T07:52:27.473Z",
              "content": "<p>Really appreciate for the reply.<br>\nYou mean the whole test dataset is provided in one \"for loop\",  after all predictions are made, some are used for pubic score calculation, some are used for private score calculation.<br>\nDid I understand correctly?</p>",
              "rawMarkdown": "Really appreciate for the reply.\nYou mean the whole test dataset is provided in one \"for loop\",  after all predictions are made, some are used for pubic score calculation, some are used for private score calculation.\nDid I understand correctly?",
              "votes": 2
            },
            {
              "id": 2329799,
              "postDate": "2023-07-04T14:10:58.863Z",
              "content": "<p>You understood correctly. Private LB score is computed when you submit, but it is only shown when the competition ends.</p>",
              "rawMarkdown": "You understood correctly. Private LB score is computed when you submit, but it is only shown when the competition ends.\n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2326751,
      "postDate": "2023-07-02T11:19:42.630Z",
      "content": "<p>Congratulations! Great work hope to learn from your excellent work！</p>",
      "rawMarkdown": "Congratulations! Great work hope to learn from your excellent work！",
      "votes": 2
    },
    {
      "id": 2325063,
      "postDate": "2023-07-01T05:31:38.557Z",
      "content": "<p>Thank you for your posting!</p>\n<p>I have a question about \"for example monitor all folds in a CV (and accept only on majority or consensus)\".</p>\n<p>Am I correct in understanding that if adding a certain feature improves accuracy for fold1~4 and decreases accuracy for fold5, and as a result there is no change in the overall average score, then the feature is adopted?</p>",
      "rawMarkdown": "Thank you for your posting!\n\nI have a question about \"for example monitor all folds in a CV (and accept only on majority or consensus)\".\n\nAm I correct in understanding that if adding a certain feature improves accuracy for fold1~4 and decreases accuracy for fold5, and as a result there is no change in the overall average score, then the feature is adopted?",
      "votes": 2,
      "replies": [
        {
          "id": 2325352,
          "postDate": "2023-07-01T09:14:42.890Z",
          "content": "<p>Thanks for your question <a href=\"https://www.kaggle.com/aesoptacit\" target=\"_blank\">@aesoptacit</a>!</p>\n<p>There is not a single response as there are dependancies with the data, the problem, the metric, the level of performance achieved, …<br>\nIt is unlikely that if 4 folds are improved except 1 the average won't be improved but it is possible. It probably would require to explore why such a dynamic.</p>\n<p>In this competition, for the GBDT approach, we mainly worked with 10 bags (maybe 5 would have been enough) and estimated the noise to be ~0.0003. Only &gt; 0.0003 overall improvements have been considered. See the code to look at the outputs we monitored: <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training\" target=\"_blank\">https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training</a>. For the NN as we needed to iterate quicker we only used a single bag and only incorporated &gt; 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.</p>",
          "rawMarkdown": "Thanks for your question @aesoptacit!\n\nThere is not a single response as there are dependancies with the data, the problem, the metric, the level of performance achieved, ...\nIt is unlikely that if 4 folds are improved except 1 the average won't be improved but it is possible. It probably would require to explore why such a dynamic.\n\nIn this competition, for the GBDT approach, we mainly worked with 10 bags (maybe 5 would have been enough) and estimated the noise to be ~0.0003. Only > 0.0003 overall improvements have been considered. See the code to look at the outputs we monitored: https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training. For the NN as we needed to iterate quicker we only used a single bag and only incorporated > 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.",
          "votes": 2,
          "replies": [
            {
              "id": 2325512,
              "postDate": "2023-07-01T12:26:03.133Z",
              "content": "<p>Thank you for your detailed explanation!</p>",
              "rawMarkdown": "Thank you for your detailed explanation!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2324109,
      "postDate": "2023-06-30T11:46:36.813Z",
      "content": "<p>Congrats, impressive solution! Thank you for your detailed explanation.💥</p>",
      "rawMarkdown": "Congrats, impressive solution! Thank you for your detailed explanation.💥",
      "votes": 2,
      "replies": [
        {
          "id": 2324133,
          "postDate": "2023-06-30T12:07:02.890Z",
          "content": "<p>+1  congrats and thank you for the explanation</p>",
          "rawMarkdown": "+1  congrats and thank you for the explanation",
          "votes": 1
        }
      ]
    },
    {
      "id": 2323791,
      "postDate": "2023-06-30T06:51:19.437Z",
      "content": "<p>Congrats! </p>",
      "rawMarkdown": "Congrats! ",
      "votes": 2
    },
    {
      "id": 2323653,
      "postDate": "2023-06-30T05:19:24.717Z",
      "content": "<p>Congrats!!<br>\nThank you for providing this wonderful solution.<br>\nI'm very pleased to be able to dig into this solution.</p>",
      "rawMarkdown": "Congrats!!\nThank you for providing this wonderful solution.\nI'm very pleased to be able to dig into this solution.",
      "votes": 2,
      "replies": [
        {
          "id": 2323681,
          "postDate": "2023-06-30T05:43:25.340Z",
          "content": "<p>Would I see your entire code??</p>",
          "rawMarkdown": "Would I see your entire code??",
          "votes": 2,
          "replies": [
            {
              "id": 2323759,
              "postDate": "2023-06-30T06:38:54.600Z",
              "content": "<p>We are considering to release the code but are not alone to decide. Stay tuned in the coming week.</p>",
              "rawMarkdown": "We are considering to release the code but are not alone to decide. Stay tuned in the coming week.",
              "votes": 2
            },
            {
              "id": 2324014,
              "postDate": "2023-06-30T10:23:02.350Z",
              "content": "<p>😃😃😃😃😃</p>",
              "rawMarkdown": "😃😃😃😃😃",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2323625,
      "postDate": "2023-06-30T04:56:34.523Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> congrats you with 1st place and gold medals!</p>",
      "rawMarkdown": "@cpmpml and @pdnartreb congrats you with 1st place and gold medals!",
      "votes": 2
    },
    {
      "id": 2323517,
      "postDate": "2023-06-30T02:47:35.463Z",
      "content": "<p>Hi, Congratulations! How much of an effort your team has put in this competition! Your really deserve to win this competition!</p>\n<p>I have a question which I hope you won't mind answering. In fact, this is a question I would like to ask every Kaggle competition winner…hope you won't mistake me..</p>\n<p>In the spirit of competition, the Kaggle community makes tremendous effort to win over others that leads to fighting for third/fourth  decimal accuracy, which I consider unavoidable, and may even be good for innovation and further progress in research..</p>\n<p>But as a competition host, suppose I ask you: \"For the competition sake, it's ok, but we are really not interested in third or fourth decimal accuracy for deployment in real world environment, but we want a much simpler but very robust model which can be extended for versions / deployed for long term..\" . What would be your recommended model architecture for a best second decimal accuracy for this specific problem? Given the amount of efforts and experimentations you have done, you must have such a solution in mind…Of course, there were thousands of submissions with 0.7 score, but we know all are not necessarily / equally that robust that can be extended to real world changes  or that would meet the client requirement..</p>",
      "rawMarkdown": "Hi, Congratulations! How much of an effort your team has put in this competition! Your really deserve to win this competition!\n\nI have a question which I hope you won't mind answering. In fact, this is a question I would like to ask every Kaggle competition winner...hope you won't mistake me..\n\nIn the spirit of competition, the Kaggle community makes tremendous effort to win over others that leads to fighting for third/fourth  decimal accuracy, which I consider unavoidable, and may even be good for innovation and further progress in research..\n\nBut as a competition host, suppose I ask you: \"For the competition sake, it's ok, but we are really not interested in third or fourth decimal accuracy for deployment in real world environment, but we want a much simpler but very robust model which can be extended for versions / deployed for long term..\" . What would be your recommended model architecture for a best second decimal accuracy for this specific problem? Given the amount of efforts and experimentations you have done, you must have such a solution in mind...Of course, there were thousands of submissions with 0.7 score, but we know all are not necessarily / equally that robust that can be extended to real world changes  or that would meet the client requirement..",
      "votes": 2,
      "replies": [
        {
          "id": 2323735,
          "postDate": "2023-06-30T06:33:36.117Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/murugesann\" target=\"_blank\">@murugesann</a>!</p>\n<p>Thanks!<br>\nNo problem with your question.</p>\n<h1>TL;DR</h1>\n<p>My recommendation is to use our solution that can be considered real world compliant.</p>\n<h1>Detailled reply</h1>\n<p>Given the fact that predicting the mean of each question lead to ~0.66 score and that the ML models do not learn that much (+0.04), 3rd decimal (+/- 0.002) is what we want to improve. The final LB shows that 3rd decimal is meaningful: for example even if 2nd place private is +101 positions, it is only 0.002 diff with public. 4th decimal is noise. I do not understand why you state that only the 2nd decimal is meaningful, implying that 3rd is just luck. I disagree.<br>\nMy first competition was <a href=\"https://www.kaggle.com/competitions/microsoft-malware-prediction\" target=\"_blank\">https://www.kaggle.com/competitions/microsoft-malware-prediction</a>. The private dataset was totally skewed because the host did not have all the targets so they were put to 0 (<a href=\"https://www.kaggle.com/competitions/microsoft-malware-prediction/discussion/83946)\" target=\"_blank\">https://www.kaggle.com/competitions/microsoft-malware-prediction/discussion/83946)</a>. A lot of competitors, including me, had worked very hard (ask to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>) and the shakeup was monstruous and very hard to accept. I invite you to have a look at <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> results: strong on the public LB and strong on the private LB! Production-ready models, right? Argue that it is just luck is biased. I learned a lot from my failure as well as from <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> success. These guys are true legends.<br>\nLast year I participated in <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating\" target=\"_blank\">https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating</a> and only entered in the last few days because of <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> feedback: <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating/discussion/304441#1671696\" target=\"_blank\">https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating/discussion/304441#1671696</a>. I only submitted twice. Just luck? I would rather say the private test set was pretty unfair and that robustness (my solution was significantly better than every benchmarks used in the field) ensured a good place that was in fact, for me and without modesty, a hugh disappointment. I let you search for the details if you want to understand why: <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating/discussion/306181\" target=\"_blank\">https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating/discussion/306181</a>.</p>\n<p>As the context is set, it is now possible to respond to your question:</p>\n<blockquote>\n  <p>What would be your recommended model architecture for a best second decimal accuracy for this specific problem?</p>\n</blockquote>\n<p>I would recommand to consider 3rd decimal and the model architecture we used. As it is explained (sorry for the english that is far from perfect) in the writeup, we focused on CV scores and just use the LB as a subsidiary information. It means that the private score corresponds to the real world environment you are referring at. Even if luck is present, keep in mind that our solution scores CV 0.705, public LB 0.705 (the winning solution is not our best LB because we thought our best LB overfitted) and private LB 0.705.<br>\nIf you want a stable/robust solution, I would recommand to blend GBDT and NN because as you need robustness, you need good but also diverse models. That being said proceeding like this you will limit variance so this choice is a tradeoff between ultimate performance and robustness. Real-world requires tradeoffs and robustness and it is the path we chose.<br>\nHope it helps.</p>",
          "rawMarkdown": "Hi @murugesann!\n\nThanks!\nNo problem with your question.\n\n# TL;DR\n\nMy recommendation is to use our solution that can be considered real world compliant.\n\n# Detailled reply\n\nGiven the fact that predicting the mean of each question lead to ~0.66 score and that the ML models do not learn that much (+0.04), 3rd decimal (+/- 0.002) is what we want to improve. The final LB shows that 3rd decimal is meaningful: for example even if 2nd place private is +101 positions, it is only 0.002 diff with public. 4th decimal is noise. I do not understand why you state that only the 2nd decimal is meaningful, implying that 3rd is just luck. I disagree.\nMy first competition was https://www.kaggle.com/competitions/microsoft-malware-prediction. The private dataset was totally skewed because the host did not have all the targets so they were put to 0 (https://www.kaggle.com/competitions/microsoft-malware-prediction/discussion/83946). A lot of competitors, including me, had worked very hard (ask to @cdeotte) and the shakeup was monstruous and very hard to accept. I invite you to have a look at @cpmpml results: strong on the public LB and strong on the private LB! Production-ready models, right? Argue that it is just luck is biased. I learned a lot from my failure as well as from @cpmpml and @titericz success. These guys are true legends.\nLast year I participated in https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating and only entered in the last few days because of @cpmpml feedback: https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating/discussion/304441#1671696. I only submitted twice. Just luck? I would rather say the private test set was pretty unfair and that robustness (my solution was significantly better than every benchmarks used in the field) ensured a good place that was in fact, for me and without modesty, a hugh disappointment. I let you search for the details if you want to understand why: https://www.kaggle.com/competitions/jigsaw-toxic-severity-rating/discussion/306181.\n\nAs the context is set, it is now possible to respond to your question:\n\n>What would be your recommended model architecture for a best second decimal accuracy for this specific problem?\n\nI would recommand to consider 3rd decimal and the model architecture we used. As it is explained (sorry for the english that is far from perfect) in the writeup, we focused on CV scores and just use the LB as a subsidiary information. It means that the private score corresponds to the real world environment you are referring at. Even if luck is present, keep in mind that our solution scores CV 0.705, public LB 0.705 (the winning solution is not our best LB because we thought our best LB overfitted) and private LB 0.705.\nIf you want a stable/robust solution, I would recommand to blend GBDT and NN because as you need robustness, you need good but also diverse models. That being said proceeding like this you will limit variance so this choice is a tradeoff between ultimate performance and robustness. Real-world requires tradeoffs and robustness and it is the path we chose.\nHope it helps.",
          "votes": 1
        },
        {
          "id": 2324098,
          "postDate": "2023-06-30T11:32:20.727Z",
          "content": "<p>Kaggle always ask winners to provide a simple mode that provides 90% of the solution quality. Exactly your question.</p>",
          "rawMarkdown": "Kaggle always ask winners to provide a simple mode that provides 90% of the solution quality. Exactly your question.",
          "votes": 1,
          "replies": [
            {
              "id": 2324279,
              "postDate": "2023-06-30T14:03:25.947Z",
              "content": "<p>How would you measure the quality?</p>\n<p>If that’s the private score, then it is 90% of 0.705, and this is less than the score of submission with the constant mean correctness for each question.</p>",
              "rawMarkdown": "How would you measure the quality?\n\nIf that’s the private score, then it is 90% of 0.705, and this is less than the score of submission with the constant mean correctness for each question."
            },
            {
              "id": 2327313,
              "postDate": "2023-07-02T20:48:01.177Z",
              "content": "<p>I guess we can provide a simple model that gives more thna 90% of the quality.</p>",
              "rawMarkdown": "I guess we can provide a simple model that gives more thna 90% of the quality.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2323230,
      "postDate": "2023-06-29T18:59:04.270Z",
      "content": "<p>I have been lucky to team with you!</p>",
      "rawMarkdown": "I have been lucky to team with you!",
      "votes": 2,
      "replies": [
        {
          "id": 2323264,
          "postDate": "2023-06-29T19:28:40.540Z",
          "content": "<p>I, too, think that I have been lucky to team with you!</p>",
          "rawMarkdown": "I, too, think that I have been lucky to team with you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2323912,
      "postDate": "2023-06-30T08:42:58.907Z",
      "content": "<p>I need help. i am beginner and do not know what to do I want to collect information about Talent acquisition data science fundamentals to learn about them what to write in parameters?<br>\nimport requests<br>\nimport time<br>\nimport json</p>\n<h1>Set up API request parameters</h1>\n<p>url = 'https://www.kaggle.com/'<br>\nheaders = {'Authorization': 'My API'}<br>\nparams = {'q': ''}</p>\n<h1>Set up rate limiting parameters</h1>\n<p>requests_per_minute = 10<br>\nseconds_per_request = 60 / requests_per_minute</p>\n<h1>Send API request and get response</h1>\n<p>response = requests.get(url, headers=headers, params=params)</p>\n<h1>Check if the request was successful</h1>\n<p>if response.status_code == 200:<br>\n    # Parse the response<br>\n    data = json.loads(response.content)</p>\n<pre><code>\n item  data[]:\n    (, item[])\n    (, item[])\n    (, item[])\n    ()\n</code></pre>\n<p>else:<br>\n    print('Error:', response.status_code)</p>\n<h1>Wait to ensure rate limiting compliance</h1>\n<p>time.sleep(seconds_per_request)</p>\n<h1>This script scrapes data from example.com for research purposes.</h1>\n<h1>All data is used with permission and proper attribution is given.</h1>",
      "rawMarkdown": "I need help. i am beginner and do not know what to do I want to collect information about Talent acquisition data science fundamentals to learn about them what to write in parameters?\nimport requests\nimport time\nimport json\n\n# Set up API request parameters\nurl = 'https://www.kaggle.com/'\nheaders = {'Authorization': 'My API'}\nparams = {'q': ''}\n\n# Set up rate limiting parameters\nrequests_per_minute = 10\nseconds_per_request = 60 / requests_per_minute\n\n# Send API request and get response\nresponse = requests.get(url, headers=headers, params=params)\n\n# Check if the request was successful\nif response.status_code == 200:\n    # Parse the response\n    data = json.loads(response.content)\n\n    # Extract relevant information\n    for item in data['items']:\n        print('Title:', item['title'])\n        print('Link:', item['link'])\n        print('Description:', item['snippet'])\n        print('---')\nelse:\n    print('Error:', response.status_code)\n\n# Wait to ensure rate limiting compliance\ntime.sleep(seconds_per_request)\n# This script scrapes data from example.com for research purposes.\n# All data is used with permission and proper attribution is given.",
      "votes": -1,
      "replies": [
        {
          "id": 2324096,
          "postDate": "2023-06-30T11:31:11.873Z",
          "content": "<p>How is this related to the competition solution we share?</p>\n<p>The first thing to learn is to ask questions in the right forum.</p>\n<p>The second is to ask clear questions . Your question is impossible to understand. </p>",
          "rawMarkdown": "How is this related to the competition solution we share?\n\nThe first thing to learn is to ask questions in the right forum.\n\nThe second is to ask clear questions . Your question is impossible to understand. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2539675,
      "postDate": "2023-11-27T08:58:44.050Z",
      "content": "<p>I still don't understand how \"duration\" in your article is derived? Can you explain to me, thank you</p>",
      "rawMarkdown": "I still don't understand how \"duration\" in your article is derived? Can you explain to me, thank you"
    },
    {
      "id": 2343891,
      "postDate": "2023-07-14T04:56:27.350Z",
      "content": "<p>Congratulations. And thanks for sharing the solution. The ML competition is great place to learn.</p>",
      "rawMarkdown": "Congratulations. And thanks for sharing the solution. The ML competition is great place to learn."
    },
    {
      "id": 2327283,
      "postDate": "2023-07-02T20:08:16.840Z",
      "content": "<p>Congratulations! Great work hope to learn from your excellent workCongratulations! Great work hope to learn from your excellent work SIP</p>",
      "rawMarkdown": "Congratulations! Great work hope to learn from your excellent workCongratulations! Great work hope to learn from your excellent work SIP"
    },
    {
      "id": 3153970,
      "postDate": "2025-03-19T11:44:35.700Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2327699,
      "postDate": "2023-07-03T05:49:09.270Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 2327556,
      "postDate": "2023-07-03T03:33:50.393Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2327659,
      "postDate": "2023-07-03T05:09:39.983Z",
      "content": "<p>Thanks for helping us learn!</p>",
      "rawMarkdown": "Thanks for helping us learn!",
      "votes": 3
    },
    {
      "id": 3322537,
      "postDate": "2025-11-13T15:09:47.883Z",
      "content": "<p>looks great!</p>",
      "rawMarkdown": "looks great!"
    }
  ],
  "comments": [
    {
      "id": 2323315,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2023-06-29T20:59:06.750000",
      "content": "<p>Huge congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a>! Thanks for sharing such a great solution. Double congrats for <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for finally break the 2nd place curse and get your first winning!</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2323974,
      "author_name": "Rongchu",
      "author_url": "",
      "post_date": "2023-06-30T09:41:15.223000",
      "content": "<p>Congrats, impressive solution! Thank you for your detailed explanation. </p>\n<p>Will you share the kernel? I can't wait to learn.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2324001,
          "author_name": "Bertrand P",
          "author_url": "",
          "post_date": "2023-06-30T10:01:06.463000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/zrongchu\" target=\"_blank\">@zrongchu</a>!<br>\nWe have decided to share our code: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332</a>.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2324042,
              "author_name": "Rongchu",
              "author_url": "",
              "post_date": "2023-06-30T10:49:01.050000",
              "content": "<p>Thanks! <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> , <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2339570,
      "author_name": "Vladimir Simões da Luz Junior",
      "author_url": "",
      "post_date": "2023-07-10T22:35:08.770000",
      "content": "<p>Greetings, <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> !</p>\n<p>I wanted to express my sincere appreciation and congratulations for your outstanding work in this recent competition. Your solution is truly impressive and it's evident that you put a tremendous amount of effort and expertise into developing it.</p>\n<p>Your focus on understanding the data model and extracting valuable insights is impressive. The combination of GBDT and NN models in your ensemble and search of external feature engineering demonstrates your deep understanding of the problem and your ability to leverage different techniques effectively.</p>\n<p>I have a question regarding your approach: In the training pipeline, you mentioned that the backbone, representing the events of a level group, is trained on all the available data for that group. Could you elaborate on how you handle the incomplete sessions during this training phase? How do you ensure that the backbone captures the most relevant information from both complete and incomplete sessions?</p>\n<p>Once again, congratulations on your remarkable achievement, and thank you for your contributions to the Kaggle community. Your solution is an inspiration to fellow data scientists and serves as a testament to your expertise.</p>\n<p>Best regards,keep up with the great work🦾🚀</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2340147,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2023-07-11T09:58:40.450000",
          "content": "<p>I am not sure I understand what you don't understand.</p>\n<p>Let me try still. Maybe this thread can help: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165</a></p>\n<p>For a given NN there are three NNs trained in the \"pretraining hase\", then they are assembled into a single NN in a second training phase. </p>\n<p>Incomplete sessions are only used in the pretraining phase.</p>\n<p>For instance, if a session has only data for the first level group (i.e. questions 1,2,3), then it is used to train the first of the three NNs, and it is not used for the other two.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2341136,
              "author_name": "Vladimir Simões da Luz Junior",
              "author_url": "",
              "post_date": "2023-07-11T21:56:24.593000",
              "content": "<p>Thanks for the quick reply!</p>\n<p>Indeed the following <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2331165\" target=\"_blank\">thread</a> has helped to interpret the model training phase. Specially the following schematics:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F960b98da4ee0cefe936edf44d5fec9e9%2F1.png?generation=1688613004356860&amp;alt=media\" alt=\"\"><br>\nKudos to <a href=\"https://www.kaggle.com/dongyk\" target=\"_blank\">@dongyk</a> , for that.</p>\n<p>It seems that you have leveraged the full dataset (complete + incomplete sessions), to create specialized models for each level_group, concatenating the inputs. Which enabled the model to capture specific information of all the level_groups independently and all together, by concatenating the model inputs. Therefore you could optimize BCE with F1 score as metric, since it turned out to be a classification problem….</p>\n<p>Once again congratulations <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> and <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for the excellent work and thanks for contributing to the Kaggle Community🙏🚀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2341645,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-07-12T07:59:22.233000",
              "content": "<p>yes, that's what we did:</p>\n<blockquote>\n  <p>It seems that you have leveraged the full dataset (complete + incomplete sessions), to create specialized models for each level_group, concatenating the inputs. </p>\n</blockquote>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2334085,
      "author_name": "Hamed JOORATI",
      "author_url": "",
      "post_date": "2023-07-07T12:53:45.227000",
      "content": "<p>Congratulations ! Very interesting !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2333655,
      "author_name": "UoM_268345H",
      "author_url": "",
      "post_date": "2023-07-07T05:31:07.173000",
      "content": "<p>Congratulations! very helpful. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2332588,
      "author_name": "Ansh Tanwar",
      "author_url": "",
      "post_date": "2023-07-06T09:47:28.037000",
      "content": "<p>Very insightful </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2331909,
      "author_name": "MG",
      "author_url": "",
      "post_date": "2023-07-05T21:38:04.753000",
      "content": "<p>Congratulations! Really interesting work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2331165,
      "author_name": "DongYK",
      "author_url": "",
      "post_date": "2023-07-05T12:05:07.797000",
      "content": "<p>Hi, I was wondering whether I understand it correctly.<br>\n<a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training\" target=\"_blank\">PSPFGP 1st Place - NN Training</a></p>\n<pre><code>outputs = {}\noutputs = heads(convnet_outputs)\noutputs = heads(\n        tf()(, convnet_outputs])\n    )\noutputs = heads(\n        tf()(, convnet_outputs, convnet_outputs])\n    )\n</code></pre>\n<p>Does it describe like this??<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F5bd337c5017d3ab5606024c830391f40%2F.png?generation=1688556655026917&amp;alt=media\" alt=\"\"><br>\nBackpropagation is performed: <br>\nfrom outputs['0-4'] to convnets['0-4'],    (I will call this as submodel A)<br>\nfrom outputs['5-12'] to convnets['0-4'] and convnets['5-12'], and     (submodel B)<br>\nfrom outputs['13-22'] to convnets['0-4'], convnets['5-12'], and convnets['13-22'].     (submodel C)<br>\nHowever, this submodels are separated. So, Backpropagation which is performed from outputs['13-22'] to convnets['0-4'] (submodel C) does not affect convnets['0-4'] of submodel A.<br>\nDid I understand correctly??</p>\n<p>+) I was really surprised this approach to compute F1 score during training. Wow…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2331653,
          "author_name": "Bertrand P",
          "author_url": "",
          "post_date": "2023-07-05T17:22:41.923000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dongyk\" target=\"_blank\">@dongyk</a>!</p>\n<p>You are working hard. Great! This is really satisfying to see that our work is useful for you. With this mindset you are going to learn a lot.</p>\n<p>Your understanding and your schematic seem nearly OK.  <br>\nKeep in mind that in step 2 of the training the weights of the ConvNets are freezed. This means backpropagation only changes the heads weights.  <br>\nIn this 2nd step we have 3 sequences of data as inputs, 1 for each level_group. Each of these sequences feeds its own ConvNet/embedder that outputs a representation in the form of a 24 dim vector. The representations for level_group 0-4 feeds the SimpleHead for level_group 0-4 which outputs a vector of 3 values that are the probability of the 3 responses of this level_group. The representations for level_group 0-4 and 5-12 are concatenated in a 48 dim vector the feeds the SimpleHead for the 2nd level_group that outputs 10 values that correspond to the 10 questions of the level_group 5-12. Same for level_group 13-22.</p>\n<p>You correctly spotted that this had been made to train the whole ensemble to optimize the overall F1 score that is what we want.<br>\nWe tried every setup: training the whole ensemble from scratch, each part independently, … The setup we presented here corresponds to what worked best for us.</p>\n<p>Does this respond to your question(s)? If not feel free to ask.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2332166,
              "author_name": "DongYK",
              "author_url": "",
              "post_date": "2023-07-06T03:25:15.107000",
              "content": "<p>Thank you for your answering!!<br>\nI'm learning a lot from your code.</p>\n<blockquote>\n  <p>We tried every setup: training the whole ensemble from scratch, each part independently, … The setup we presented here corresponds to what worked best for us.</p>\n</blockquote>\n<p>I was thinking the same. Did they try the whole ensemble from scratch? Because it's more cumbersome to train seperately. Training separately can be helpful to train.<br>\nIs the reason why you used only convnet['5-12'] for outputs['5-12'] not contained convnet['0-4']?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F824ba97fb079de1424be6625e9e99f0b%2F2.png?generation=1688620218055724&amp;alt=media\" alt=\"\"></p>\n<p>And… I editted the figure.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F960b98da4ee0cefe936edf44d5fec9e9%2F1.png?generation=1688613004356860&amp;alt=media\" alt=\"\"><br>\nI found that you fed train_dataset only once and the same convnet['0-4'] is fed into simplehead['0-4'], simplehead['5-12'], and simplehead['13-22'].<br>\nIs it more accurate?<br>\nIf I didn't freeze the convnets, can convnet['0-4'] backpropagated(affected) by not only outputs['0-4'] but outputs['5-12'] and outputs['13-22']?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2332546,
              "author_name": "Bertrand P",
              "author_url": "",
              "post_date": "2023-07-06T09:08:12.617000",
              "content": "<p>Wow! Thanks for your schematics that are 100% correct as I understand them! The write-up has been updated to link to your post.</p>\n<p>If the weights of the convnets are not freezed then they are updated by the backprop coming from each output they are linked to when training end-to-end.</p>\n<p>We first trained separatly which gave a baseline. Then we tried to train the end-to-end setup from scratch. As you might have noticed the output allow to monitor the loss which is composed by the losses of each level group that can also be monitored. We observed that the loss for 0-4 was the first to converge but also that, as the others go down, it tends to increase. We interpreted that as the optimization of previous level groups representation for next level groups. This is not what we wanted. We wanted a good representation for a level group that can be used by the next level groups. So we chose to first optimize each convnet separatly and to exploit its representation by heads in a second time.</p>\n<p>As F1 score cannot be optimized directy, the main goal of the end-2-end approach was to be able to monitor the F1 score to select the best weights.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2332884,
              "author_name": "DongYK",
              "author_url": "",
              "post_date": "2023-07-06T13:36:42.110000",
              "content": "<p>Cool.<br>\nI feel good to understand it.</p>\n<blockquote>\n  <p>We observed that the loss for 0-4 was the first to converge but also that, as the others go down, it tends to increase. We interpreted that as the optimization of previous level groups representation for next level groups. This is not what we wanted. We wanted a good representation for a level group that can be used by the next level groups. So we chose to first optimize each convnet separatly and to exploit its representation by heads in a second time.</p>\n</blockquote>\n<p>I also tried to come up with a way to change the architecture of the model, but I realized that it's the best idea considering the pros and cons of possible models.</p>\n<blockquote>\n  <p>As F1 score cannot be optimized directy, the main goal of the end-2-end approach was to be able to monitor the F1 score to select the best weights.</p>\n</blockquote>\n<p>The way I see it, it's the most important approach in the NN code.</p>\n<p>Again, thank you so much.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2326886,
      "author_name": "Mayur B. Ingole",
      "author_url": "",
      "post_date": "2023-07-02T13:22:02.613000",
      "content": "<p>bravo 1St place</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2326179,
      "author_name": "yCarbon",
      "author_url": "",
      "post_date": "2023-07-02T00:55:58.820000",
      "content": "<p>Congratulations on your win🎉 And thanks you also for sharing your solution!</p>\n<p>If I may, may I ask a question about the third chapter, the beginning of the <strong>Data</strong>?</p>\n<p>Due to the small number of game sessions, the length of the sequence (i.e. the length of the game to be put into the NN) is not long enough. This as not ideal for the Deep Learning approach, I interpreted.<br>\nIf so, what does it mean by \"Moreover these are long sequences.\"?　Does it relevant that in some cases the game is played many times and is actually a long sequence, but what we have as data is a short sequence?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2326740,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2023-07-02T11:11:22.223000",
          "content": "<p>There are too few sequences to enable complex NN model training. But sequence average length is rather long.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2326787,
              "author_name": "yCarbon",
              "author_url": "",
              "post_date": "2023-07-02T11:59:11.803000",
              "content": "<p>Thank you.  I understood!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2325659,
      "author_name": "Pranav Jadhav",
      "author_url": "",
      "post_date": "2023-07-01T14:02:09.720000",
      "content": "<p><a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> 1st place bravo. Great Victory.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325346,
      "author_name": "Sercan Yeşilöz",
      "author_url": "",
      "post_date": "2023-07-01T09:10:11.110000",
      "content": "<p>Congratulations on your first place!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325317,
      "author_name": "Minze Li",
      "author_url": "",
      "post_date": "2023-07-01T08:49:29.843000",
      "content": "<p>very useful, thank you for your fanscinating work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325268,
      "author_name": "PRIYANSHU YADAV",
      "author_url": "",
      "post_date": "2023-07-01T08:07:48.353000",
      "content": "<p>Such an impressive combination of good ideas and hard hard work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325097,
      "author_name": "Dimas Mufid",
      "author_url": "",
      "post_date": "2023-07-01T05:51:09.790000",
      "content": "<p>Congrats on winning first place! Thanks a lot for your detailed explanation. Impressive! 🔥🔥🔥🔥🔥🔥</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325053,
      "author_name": "Mohsin hasan",
      "author_url": "",
      "post_date": "2023-07-01T05:26:58.413000",
      "content": "<p>Congrats on the win! Really love the NN design here. I has a very similar NN design, except that I just one-hot encode everything due to low cardinality of categorical features.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324735,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-06-30T20:45:31.277000",
      "content": "<p>Wow, just wow! :D</p>\n<p>Such an impressive combination of good ideas and hard hard work. Very inspiring!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324715,
      "author_name": "Andriansyah1",
      "author_url": "",
      "post_date": "2023-06-30T20:21:55.020000",
      "content": "<p>This is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lesson</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324446,
      "author_name": "AbaoJiang",
      "author_url": "",
      "post_date": "2023-06-30T15:53:44.053000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> for winning the 1st place! Also, thanks a lot for the detailed explanation!</p>\n<p>May I ask how you decide to simply divide the numeric features (<em>e.g.,</em> <code>duration</code>) by a constant (60000 in the code you open source), instead of  using other normalization methods. I suppose that shrinking the scale of the raw feature could make the training process converge faster and other normalization techniques don't show superiority over the way you eventually adopted. If there's misunderstanding, please put me right. Thanks a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2324560,
          "author_name": "Bertrand P",
          "author_url": "",
          "post_date": "2023-06-30T17:20:55.463000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>!</p>\n<p>The writeup has been updated to mention your excellent notebook. We first saw the time-aware events and the Conv1D ideas in your work. Finding once more these ideas while exploring the litterature likely made us focus more on them.</p>\n<p>Continuous features are tricky to preprocess. We knew from our GBDT modeling work that time and especially <code>duration</code> was very important. We put a lot of careful efforts to try to model it correctly. Processing the <code>duration</code> seemed important so we tried a lot of things: scaling, normalization, standardization, clipping, … We explored the data to analyze the distribution of duration times and experimented with values around a baseline that seemed to us reasonnable because of the better distribution on the continuum. What worked best was clipping at 60 seconds + simple scaling.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2325122,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2023-07-01T06:14:04.730000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a>,</p>\n<p>Thanks very much for the mention. I'm so glad that my inconspicuous sharing can inspire you and other teams to come up with far better solutions. And, I think that's the reason why I learn to share and gradually enjoy sharing!</p>\n<p>During the first month of the competition, I tried techniques like normalization, standardization, quantile transformation, <code>np.log1p</code>, etc., all of which didn't help improve the CV score. Hence, I sticked to the original processing method that clipped <code>duration</code> at <code>3.6e6</code> (<em>i.e.,</em> one hour). However, I didn't try to clip it with smaller values. I'll try them out and verify more ideas with late submissions!</p>\n<p>I learn so so much from your sharing. (1) Instead of just using beautiful visualization to interpret the data, you try hard to understand the mechanism behind raw data generation and clean the data to improve the data quality. (2) In addition to the model architecture design, robust training process and techniques are also crucial (<em>e.g.,</em> the loss criterion to use, whether to share \"embedders\" across different <code>level_group</code>, how to choose checkpoints). (3) Create a robust CV strategy and resist the temptation to select submission based on LB. (4) FE plays an important role all the time, to name a few.</p>\n<p>Thanks a lot for the clarification, and good luck with your next competition!!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2324246,
      "author_name": "ANSH TIWARI Student",
      "author_url": "",
      "post_date": "2023-06-30T13:42:44.707000",
      "content": "<p>it is very useful content</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324241,
      "author_name": "Joel Erikanders",
      "author_url": "",
      "post_date": "2023-06-30T13:40:58.287000",
      "content": "<p>Congrats and thank you for sharing 🎉 Impressive solution with many creative ideas!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324151,
      "author_name": "Fahmi Rizal Kurnia",
      "author_url": "",
      "post_date": "2023-06-30T12:21:00.883000",
      "content": "<p>Congrats, mate!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324107,
      "author_name": "BELSONRAJA T",
      "author_url": "",
      "post_date": "2023-06-30T11:41:11.563000",
      "content": "<p>Congratulations!  Keep it up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324083,
      "author_name": "Piyush Sukhija",
      "author_url": "",
      "post_date": "2023-06-30T11:20:05.807000",
      "content": "<p>Congratulations on winning!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2323256,
      "author_name": "Patrick Gendotti",
      "author_url": "",
      "post_date": "2023-06-29T19:18:23.833000",
      "content": "<p>Winner winner chicken dinner! :p</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2323126,
      "author_name": "Swapnil Chowdhury",
      "author_url": "",
      "post_date": "2023-06-29T17:21:02.953000",
      "content": "<p>Congratulations on coming 1st <a href=\"https://www.kaggle.com/pdnartreb\" target=\"_blank\">@pdnartreb</a> 🎉. Keep on achieving new milestones 👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2324716,
          "author_name": "Andriansyah1",
          "author_url": "",
          "post_date": "2023-06-30T20:22:44.147000",
          "content": "<p>It's true he is very detailed in terms of giving knowledge I like </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2327560,
      "author_name": "huoyeqianxun",
      "author_url": "",
      "post_date": "2023-07-03T03:34:57.963000",
      "content": "<p>Congratulations! What a great work! You and your team deserve the 1st place.<br>\nI am a newbee. May I ask a question about probing the private test?<br>\nAs you mentioned here:</p>\n<blockquote>\n  <p>Probing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API.</p>\n</blockquote>\n<p>As far as what I know, the public test data can be probed, because the public learderboard give us the score on the public test dataset. But the private leaderboard seems won't give us feedback before the competition ends.<br>\nSo how do you probe the private test dataset?<br>\nThanks a lot.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2327778,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2023-07-03T07:01:33.717000",
          "content": "<p>You can switch some predictions and see if the public LB changes at all. If it does not change then the flipped predictions are on the private test set.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2329321,
              "author_name": "huoyeqianxun",
              "author_url": "",
              "post_date": "2023-07-04T07:52:27.473000",
              "content": "<p>Really appreciate for the reply.<br>\nYou mean the whole test dataset is provided in one \"for loop\",  after all predictions are made, some are used for pubic score calculation, some are used for private score calculation.<br>\nDid I understand correctly?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2329799,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-07-04T14:10:58.863000",
              "content": "<p>You understood correctly. Private LB score is computed when you submit, but it is only shown when the competition ends.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2326751,
      "author_name": "Hoping",
      "author_url": "",
      "post_date": "2023-07-02T11:19:42.630000",
      "content": "<p>Congratulations! Great work hope to learn from your excellent work！</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2325063,
      "author_name": "wakaka",
      "author_url": "",
      "post_date": "2023-07-01T05:31:38.557000",
      "content": "<p>Thank you for your posting!</p>\n<p>I have a question about \"for example monitor all folds in a CV (and accept only on majority or consensus)\".</p>\n<p>Am I correct in understanding that if adding a certain feature improves accuracy for fold1~4 and decreases accuracy for fold5, and as a result there is no change in the overall average score, then the feature is adopted?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2325352,
          "author_name": "Bertrand P",
          "author_url": "",
          "post_date": "2023-07-01T09:14:42.890000",
          "content": "<p>Thanks for your question <a href=\"https://www.kaggle.com/aesoptacit\" target=\"_blank\">@aesoptacit</a>!</p>\n<p>There is not a single response as there are dependancies with the data, the problem, the metric, the level of performance achieved, …<br>\nIt is unlikely that if 4 folds are improved except 1 the average won't be improved but it is possible. It probably would require to explore why such a dynamic.</p>\n<p>In this competition, for the GBDT approach, we mainly worked with 10 bags (maybe 5 would have been enough) and estimated the noise to be ~0.0003. Only &gt; 0.0003 overall improvements have been considered. See the code to look at the outputs we monitored: <a href=\"https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training\" target=\"_blank\">https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training</a>. For the NN as we needed to iterate quicker we only used a single bag and only incorporated &gt; 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2325512,
              "author_name": "wakaka",
              "author_url": "",
              "post_date": "2023-07-01T12:26:03.133000",
              "content": "<p>Thank you for your detailed explanation!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2324109,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T11:46:36.813000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2324133,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-30T12:07:02.890000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2323791,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T06:51:19.437000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2323653,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T05:19:24.717000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2323681,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-30T05:43:25.340000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2323759,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-30T06:38:54.600000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2324014,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-30T10:23:02.350000",
              "content": "",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2323625,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T04:56:34.523000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2323517,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T02:47:35.463000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2323735,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-30T06:33:36.117000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2324098,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-30T11:32:20.727000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2324279,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-30T14:03:25.947000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2327313,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-07-02T20:48:01.177000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2323230,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-29T18:59:04.270000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2323264,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-29T19:28:40.540000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2323912,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T08:42:58.907000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 2324096,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-30T11:31:11.873000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2539675,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-27T08:58:44.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2343891,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-14T04:56:27.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2327283,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-02T20:08:16.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3153970,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-19T11:44:35.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2327699,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-03T05:49:09.270000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2327556,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-03T03:33:50.393000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2327659,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-03T05:09:39.983000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3322537,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-11-13T15:09:47.883000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2323109": "Unbelievable to write this!\n\n# Thanks!\n\nAs it is the usage, we first **thank the host and Kaggle**. These are special thanks because you and us have had a special link in this competition as we gave you more work by reporting data leaks. No doubt you tried to do your best. You are right to animate this community and to trust in it. You are part of it. Please take care of this community that is able to build so much together by sharing. As all of us you have made mistakes and we hope you will learn from them.\n\nWe also want to **thank all of you**, Kagglers. We love and are grateful to be part of our group/community. Thanks for sharing and for the collective learning experience.\n\n# Context\n\n- Business context: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview,\n- Data context: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data.\n\n# Overview of the Approach\n\nOur solution is essentially a blend of a XGBoost and a NN models. Both heavily rely on duration that appeared to be a powerful leverage. Time was aggregated in different ways and combined with counts for the GBDT while it is transformed via a custom TimeEmbedding block based on 1D convolutions that produce a representation combined with user event representations for the NN.\nRobustness and efficiency founded our work. XGBoost models were validated on 10 bags of 5 folds and features incorporated only if the mean of the CV of these 10 bags was greater than the level of noise we quantified while we opted for a majority/consensus strategy to build the NN, i.e. validate choices only if 4 of 5 folds were improved. The 3rd place of the efficiency LB was achieved with a lightweight NN accelerated via TF Lite.\n\n# Details of the submission\n\n## Code\n\nAfter publishing this write-up we decided to open our code: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332.\nIt is composed by several parts: [how to train the XGBoost models](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-gbdt-training), how to [pretrain](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-pretraining) and [train](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training) the NN models and the [inference notebook](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-inference) used to win this competition.\n\n## Data\n\nLooking at the 1st data released showed that there aren't a lot of sessions so not a lot of sequences. Moreover these are long sequences. This is not ideal for a deep learning approach.  \nExploring the Field Day Lab research instructed that the Jo Wilder application was built to help learning to read and that way more than 11,500 learners had played this game.  \nThese 2 ideas led to search for a bigger dataset. In 1 Google search and 3 clicks we came up to the open data portal (https://fielddaylab.wisc.edu/opengamedata/) which contains a lot of sessions. 1 hour and 3 bash commands latter we knew that the train set was in part in the open data. So we took a week to **build a pipeline that extracts 98 % of the sessions of the train set perfectly and with minor errors for the last 2 %**. Our data are even better than the comp data because we knew before the host confirmation that for the sessions with 2 games the target was skewed (0 if wrong in 1 of the 2 games when we aim at predicting the responses for the 1st game). It seems that fixing these targets can bring a significant boost up to +0.002.\n\nWe took 1 more week to build a GBDT/XGBoost baseline that would have scored top 10 given the CV score, with the use of the supplemental data (~20,000 sessions) that gave +0.003/0.004 at that time. As we simulated the API locally (see after), we used some training sessions to infer and noticed that it scored 0.718. We were hoping that the LB sessions were not part of the open data portal but our 1st submission, LB 0.708, immediately showed to us that we had rebuilt about a half of the data and especially the targets in the public LB, because 0.708 = (0.698 + 0.718) / 2. The host and Kaggle have been immediately informed. You know what happened next (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/415820).  \nAfter the release of the LB data we measured that we perfectly rebuilt ~7000 sessions over the ~11,500 of the LB data.\n\nWe spent the first month exploring the data until we understood/knew it pretty well. For example we even reconstituted sessions for what might be schools (several games on 1 IP session), extracted every single session with at least 1 answer, ...\n\nAfter the update we made a first submission that scored 0.72. This was shocking because this meant that some leaked data were remaining. A few days later we noticed that the open data was not totally similar with the state we found it 1 month before. A file was missing. So we returned to the host and Kaggle to give them more work.\n\n**This process/work led us to perfectly understand the data model** (that changed since the 1st release of the game). This also allowed us to deeply understand the data itself.\n\nNote that we only used the sessions for which we had responses to all questions of the 2 1st level groups. 1) This is more consistant with the sessions we want to predict (game from the beginning to the end) and 2) this approach preserves performance (vs all data) while reducing training time.\n\nOur dataset is constituted by **37323 complete sessions (23562 comp + logs) in a total of 66376 sessions**.\n\nThe supplemental data (that we fully added 1 month ago) gave us consistently **CV +0.002**.\n\n## Model\n\nOur solution is mainly an ensemble of GBDT + NN models.\n\n### Trust your validation\n\nWe think that **the main reason of the robustness of our solution is that we only relied on CV** for decision making. No choice had been made on LB.\n\nProbing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API. This is a small set. In our experiments 5,000 sessions is the minimum to guarantee a stable CV/LB alignment. A set less than 2,000 is very noisy so **robustness was the way to go**.\n\nWe **only added features that improved the CV for sure**. This is not easy to delete features that you believe in but this is needed as science is not a matter of belief. There are several ways to do so: for example monitor all folds in a CV (and accept only on majority or consensus), monitor several bags (composition of CV to not overfit validation), ...\n\nFor the GBDT approach, we mainly validated on the mean of 10 bags (we defined a bag as a composition of the folds). As we estimated the noise to be ~0.0003, only improvements greater than the noise have been considered. For the NN as we needed to iterate quicker we only used a single bag and only incorporated > 0.0003 overall improvements with at least 3 or 4 (over 5) folds improved.\n\n### Metric\n\nWe experimented a lot on finding a threshold by question but found that this approach is less robust than a single threshold. We mainly used 0.625 as global threshold despite our highest LB scores that were obtained with a threshold per question.\n\n### GBDT\n\nWe prototyped a baseline with **XGBoost because of the structured/tabular nature of the data**. The feature engineering process is interesting to understand what is predictive and to understand the causation, i.e. how the features or decision criteria that enable to predict correctly.\n\nGenerally speaking we followed 3 ways to build features: **business knowledge**, our **intuition** playing the game and a meticulous **exploration of the data**.\nBusiness knowledge refers to using expert knowledge. Reading the papers of the researchers that built this game allow to understand the game beyond usage. For example, Jo Wilder has been built to improve the players reading skills. So this means that the text duration should be important. These are like killer features.\n\nWe exclusively made use of Polars because of the CPU constraints and to simply learn it.\nOur features (663, 1993, 3734 for each level\\_group) are mainly **durations and counts for different aggregations**: how much time in a level, in a room, reading a text, interacting in some way (event type), how many events in a level_group, how many events of each type, how many events of each type in a room or a level, ...  \nWe also built a few notebook dedicated features: how many type of events on the notebook in a level, ...  \nDespite our efforts we weren't able to extract useful information from the coordinates, the only few features of this type had been mean and std for some events in the activities (journal interactions for example).\n\nWe considered that injecting targets predicted in the previous level groups was a compression of the signal, meaning a loss of information, so we used, for each session, **all interactions from the beginning of the game/session**. This led to a +0.002 at the time of this choice.\n\nAfter the API needed to order the data, we noticed that **models trained both on original order and on index order** but validated on index order (inference order) improved our scores. This leads to more variety that was needed to **improve stability and robustness**. The same goes for the composition of the validation sets: usage of several bags (composition of validation sets) based on the comp data but also on the extracted data improved our scores. We detected late that increasing the number of folds from 5 to 10 could also be leveraged.\n\nThe code for GBDT allows to switch from XGBoost to LightGBM and CatBoost with a simple variable parameter but despite the good scores (~0.001 less than XGBoost), this did not bring to ensemble so we sticked to only XGBoost.\n\nWe experimented a lot around feature selection but were unable to build a stable strategy. So instead of a top-down approach consisting in deleting useless features, we adopted a bottom-up approach choosing carefully each group of features.\n\nOur **XGBoost models score CV ~0.7025 +/-0.0003** and blending 5 of them (the only XGBoost we still have with correct score) scores **LB 0.704**.\n\n### NN\n\nAfter achieving a good score with gradient boosting and having understood well the data we focused on deep learning.\n\nThe **first attempt was with Transformers**. The 1st results were disappointed: CV 0.685 with 2 hours / fold (as far as we can remember). Transformers are very computationally intensive. Resources: https://arxiv.org/pdf/1912.09363.pdf, https://arxiv.org/pdf/2001.08317.pdf, https://arxiv.org/pdf/1711.03905.pdf, https://arxiv.org/pdf/1907.00235.pdf, ...\n\nWe then gave a try to **Conv1D**. In one day we had a very simple model that scored as Transformers but **10x faster** allowing to iterate quicker. So we pushed this approach and could seamlessly scaled it beyond our expectations.\n\nDifficult to share the **tens or hundreds of experimentations** needed to achieve the final solution which is both based on a simple architecture and a slightly complex training pipeline.\n\n#### Architecture roots\n\nWe browsed the literature based on the question: how to model time in deep learning?  \nThis research made us come to the idea of **time-aware events** (i.e. https://proceedings.mlr.press/v126/zhang20c/zhang20c.pdf) and back to **WaveNet** (https://arxiv.org/pdf/1609.03499.pdf) because it uses **Conv1D to model long sequences with considerations on causation**.  \nOther papers also inspired us: https://arxiv.org/pdf/1703.04691.pdf build on top of WaveNet paper for time series, https://idus.us.es/bitstream/handle/11441/114701/Short-Term%20Load%20Forecasting%20Using%20Encoder-Decoder%20WaveNet.pdf?sequence=1&isAllowed=y also build on top of WaveNet.  \nWe also have to mention the excellent work that @abaojiang shared (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/398565 and https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part). It inspired our research and maybe successfully biased it.\n\nLet's focus on the model of our efficiency submission that is also one of our final ensemble and which performance is nearly the same as models with a few more features.\n\n#### Feature representations\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Ff692423c63172494ea4df97b01d65d33%2Fdata.png?generation=1702113769960411&alt=media)\n\n**5 features as inputs: duration, text\\_fqid, room\\_fqid, fqid, event\\_name + name** (this is the event type from the original data model as far as we remember). Each of these information is encoded/embedded into a vector representation (d_model = 24) to be the merged. The 4 **categorical features feed a classical Embedding layer and the duration a TimeEmbedding** which is a custom block.\n\nDeveloping the GBDT solution showed that the **duration** was crucial, so we put a crucial amount of time trying to model it greatly. The TimeEmbedding layer is a composition of 4x ConvBlock which is inspired by the Transformer main block: Conv1D -> skip connection -> layer norm -> dropout.\n\n```\nclass TimeEmbedding(tf.keras.layers.Layer):\n    def __init__(self, n_blocks, d_model, dropout_rate):\n        super(TimeEmbedding, self).__init__()\n        self.conv_blocks = [ConvBlock(d_model, dropout_rate=dropout_rate) for _ in range(n_blocks)]\n        \n    def call(self, inputs):\n        x = tf.expand_dims(inputs, axis=-1)\n        for conv_block in self.conv_blocks:\n            x = conv_block(x)\n        return x\n```\n\n```\nclass ConvBlock(tf.keras.layers.Layer):\n    def __init__(self, d_model, dropout_rate):\n        super(ConvBlock, self).__init__()\n        self.conv1d = tf.keras.layers.Conv1D(d_model, kernel_size=5, padding='same', activation='gelu')\n        self.layer_norm = tf.keras.layers.LayerNormalization()\n        self.dropout = tf.keras.layers.Dropout(rate=dropout_rate)\n        \n    def call(self, inputs):\n        x = self.conv1d(inputs)\n        x = x + inputs\n        x = self.layer_norm(x)\n        outputs = self.dropout(x)\n        return outputs\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F9407390ffd7f226b7be0f57c37051f77%2Ftime_embedding.png?generation=1702113795144186&alt=media)\n\n#### Time-aware events\n\nAs said, the goal of building these representations was to model time-aware events. We considered the **categorical features as events** because they represent the user interactions with business entities of the game. We then tried to incorporate duration to make them time-awared. Our main intuition showed to be the best. It is a **simple solution based on operation priority to represent that the duration should be associated to each event before associated them together**: duration * event_1 + duration * event_2 + ... which had been factorized to duration * (event_1 + event_2 + ...).\n\n```\nclass ConvNet(tf.keras.Model):\n    def __init__(self, input_dims, n_outputs, d_model, n_blocks=4, name=None):\n        super(ConvNet, self).__init__(name=name)\n        self.input_dims = input_dims\n        self.n_outputs = n_outputs\n        self.d_model = d_model\n        self.n_blocks = n_blocks\n        self.event_embedding = tf.keras.layers.Embedding(input_dims['event_name_name'], d_model, mask_zero=True)\n        self.room_embedding = tf.keras.layers.Embedding(input_dims['room_fqid'], d_model, mask_zero=True)\n        self.text_embedding = tf.keras.layers.Embedding(input_dims['text'], d_model, mask_zero=True)\n        self.fqid_embedding = tf.keras.layers.Embedding(input_dims['fqid'], d_model, mask_zero=True)\n        self.duration_embedding = TimeEmbedding(n_blocks=n_blocks, d_model=d_model, dropout_rate=0.2)\n        self.gap = tf.keras.layers.GlobalAveragePooling1D()\n        \n    def call(self, inputs):\n        event = self.event_embedding(inputs['event_name_name'])\n        room = self.room_embedding(inputs['room_fqid'])\n        text = self.text_embedding(inputs['text'])\n        fqid = self.fqid_embedding(inputs['fqid'])\n        duration = self.duration_embedding(inputs['duration'])\n        x = duration * (event + room + text + fqid)\n        outputs = self.gap(x)\n        return outputs\n\n    def get_config(self):\n        config = super().get_config().copy()\n        config.update({\n            'input_dims': self.input_dims,\n            'n_outputs': self.n_outputs,\n            'd_model': self.d_model,\n            'n_blocks': self.n_blocks,\n            'name': self._name,\n        })\n        return config\n\n    @classmethod\n    def from_config(cls, config):\n        return cls(**config)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2Fca563daa4f9ef7bf8d8b20e0c85c7951%2Ftime_aware_events_1.png?generation=1702113822015438&alt=media)\nThe 2 representations are equivalent: either you can think time-aware events as a combination of time and sub-events or as a combination of sub-events and time.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F1f65d399dfd93f128ca6c887d0f0be89%2Ftime_aware_events_2.png?generation=1702113853644045&alt=media)\n \n#### Training pipeline\n\nThe training pipeline is not totally straight forward.\n\n@dongyk published great schematics that can be useful to illustrate what is explained bellow: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420217#2332166.\n\n##### 1st step (pre-training?)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F970baa66fc37b301bb8960cc92cf8ab4%2Fpre_training.png?generation=1702113879947155&alt=media)\n\nThe best approach for us consists in a kind of **backbone that represents the events of a level_group**.\n\nThis backbone is trained on all the data available for this level_group (i.e. on complete + incomplete sessions). It is associated with a temporary SimpleHead optimizing BCE loss.\n\n```\nclass SimpleHead(tf.keras.Model):\n    def __init__(self, n_units, n_outputs, name=None):\n        super(SimpleHead, self).__init__(name=name)\n        self.ffs = [tf.keras.layers.Dense(units, activation='gelu') for units in n_units]\n        self.out = tf.keras.layers.Dense(n_outputs, activation='sigmoid')\n        \n    def call(self, inputs):\n        x = inputs\n        for ff in self.ffs:\n            x = ff(x)\n        outputs = self.out(x)\n        return outputs\n```\n\nThis approach allows to score **CV 0.70025 +/- 0.0005**.\n\n##### 2nd step (training?)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2208184%2F98e32800c54ffab63d98d90d220fbbcd%2Fend_2_end.png?generation=1702113900174294&alt=media)\n\n**The weights of each of the 3 backbones (1 by level_group) are freezed** for the 2nd level of training to speedup training but also because it is more stable and efficient. These backbones can be thought as \"embedders\".\n\nDuring this 2nd step, **all the submodels that composed the solution were trained on all complete sessions in an end-to-end setup**. The input data are 3 sequences of the 5 features, 1 for each of the 3 level groups. Each \"embedders\" outputs a 24 dim-vector representation. These outputs are the inputs of a head in which enters the representation of level\\_group '0-4' to predict the 3 first questions and the concatenation of the previous and the current representations for level_groups '5-12' and '13-22' to make use of all information.\n\nProceeding like this allows to optimize the overall performance and to monitor it based on the F1 score that is the score of the competition. This means we optimized BCE with F1 score as a metric.\n\nOur winning submission uses a simple **MLP head** but also a **skip head** (512 -> 512 -> 512 allow it for example). **MMoE** did not improve the simplest approaches.\n\nThis approach allows to score **CV 0.70175 +/- 0.0003** which is **comparable to the GBDT solution**.\n\n### Inference\n\n#### Build a simulator\n\nEarly in the competition we built a simulator of the API. Doing so we never experimented any submission error. Maybe trying to keep ideas and code as simple as possible was also key to debug easily.\n\n#### Efficiency\n\nWe invested the efficiency part of the challenge for GBDT as well as NNs.  \nUsing **Treelite** for XGBoost allow us to divide by 2 the execution time.  \nOur deep learning models were lights: **400,000 weights** for the end-to-end model which combines every parts/sub-models. Having already used **TF Lite** we knew it could be a game changer. Converting our models led to a significant boost in inference time without any performance loss (we do not remember exactly but we think it is at least **6x faster** on our local inference simulator).  \nBeginning to explore pruning as well as hard quantization showed that the performance loss would be significant (which is OK in production but not in a competition) so we sticked to a simple TF Lite conversion.\n\nWe have not leveraged what seems to be a problem in the efficiency metric. As we identified the private test sessions to be the 1450/1500 first served by the API we tried to just predict the others to check which time was used (public for public and private for private). Doing so we gain a place but choose to not use this.\n\nOur **efficiency submission is a NN that scores public LB 0.702 and private LB 0.699 in less than 5 minutes**.\n\n#### Ensemble\n\nWe experimented a lot of ensembling alternatives. In the end we sticked to a simple average 50/50 GBDT/NN with:\n\n  * 2 kinds of GBDT: trained on original order + trained on index order (validated on index order that is the inference case),\n  * 3 kinds of NNs: trained on original order + trained on index order with 5 or all features.\n  \nAs our models are lightweight we were able to build a hugh ensemble: **2 x 4 x 10 folds XGBoost + 3 x 4 x 5 folds NNs**. The bottleneck for us is the 8 Go RAM constraint.\n\nThe winning submission scores **CV 0.705, public LB 0.705 and private LB 0.705**.\n\n## Conclusion\n\nThe main achievement of our work is that it is a good solution for the researchers, learners and children that can benefit of it and we hope it will contributes to progresses for a better learning experience. Up to you guys!\n\nThanks if you read until here!\nIf you have any question do not hesitate to ask. We will do our best to respond.\n\n## Presentation to the host\n\nA video presentation to the host has been recorded and can be available on demand. Feel free to ask via PM.\n\n# Sources\n\nBelow are the main sources that we used. More sources can be found in section *Details of the submission* above.\n\n- https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420332,\n- https://fielddaylab.wisc.edu/opengamedata/,\n- https://arxiv.org/pdf/1609.03499.pdf,\n- https://www.tensorflow.org/lite/guide",
    "2323315": "Huge congrats @cpmpml and @pdnartreb! Thanks for sharing such a great solution. Double congrats for @cpmpml for finally break the 2nd place curse and get your first winning!",
    "2323974": "Congrats, impressive solution! Thank you for your detailed explanation. \n\nWill you share the kernel? I can't wait to learn.",
    "2339570": "Greetings, @pdnartreb and @cpmpml !\n\nI wanted to express my sincere appreciation and congratulations for your outstanding work in this recent competition. Your solution is truly impressive and it's evident that you put a tremendous amount of effort and expertise into developing it.\n\nYour focus on understanding the data model and extracting valuable insights is impressive. The combination of GBDT and NN models in your ensemble and search of external feature engineering demonstrates your deep understanding of the problem and your ability to leverage different techniques effectively.\n\nI have a question regarding your approach: In the training pipeline, you mentioned that the backbone, representing the events of a level group, is trained on all the available data for that group. Could you elaborate on how you handle the incomplete sessions during this training phase? How do you ensure that the backbone captures the most relevant information from both complete and incomplete sessions?\n\nOnce again, congratulations on your remarkable achievement, and thank you for your contributions to the Kaggle community. Your solution is an inspiration to fellow data scientists and serves as a testament to your expertise.\n\nBest regards,keep up with the great work🦾🚀",
    "2334085": "Congratulations ! Very interesting !",
    "2333655": "Congratulations! very helpful. ",
    "2332588": "Very insightful ",
    "2331909": "Congratulations! Really interesting work",
    "2331165": "Hi, I was wondering whether I understand it correctly.\n[PSPFGP 1st Place - NN Training](https://www.kaggle.com/code/pdnartreb/pspfgp-1st-place-nn-training)\n```\noutputs = {}\noutputs['0-4'] = heads['0-4'](convnet_outputs['0-4'])\noutputs['5-12'] = heads['5-12'](\n        tf.keras.layers.Concatenate()([convnet_outputs['0-4'], convnet_outputs['5-12']])\n    )\noutputs['13-22'] = heads['13-22'](\n        tf.keras.layers.Concatenate()([convnet_outputs['0-4'], convnet_outputs['5-12'], convnet_outputs['13-22']])\n    )\n```\nDoes it describe like this??\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8366173%2F5bd337c5017d3ab5606024c830391f40%2F.png?generation=1688556655026917&alt=media)\nBackpropagation is performed: \nfrom outputs['0-4'] to convnets['0-4'],    (I will call this as submodel A)\nfrom outputs['5-12'] to convnets['0-4'] and convnets['5-12'], and     (submodel B)\nfrom outputs['13-22'] to convnets['0-4'], convnets['5-12'], and convnets['13-22'].     (submodel C)\nHowever, this submodels are separated. So, Backpropagation which is performed from outputs['13-22'] to convnets['0-4'] (submodel C) does not affect convnets['0-4'] of submodel A.\nDid I understand correctly??\n\n+) I was really surprised this approach to compute F1 score during training. Wow...",
    "2326886": "bravo 1St place",
    "2326179": "Congratulations on your win🎉 And thanks you also for sharing your solution!\n\nIf I may, may I ask a question about the third chapter, the beginning of the **Data**?\n\nDue to the small number of game sessions, the length of the sequence (i.e. the length of the game to be put into the NN) is not long enough. This as not ideal for the Deep Learning approach, I interpreted.\nIf so, what does it mean by \"Moreover these are long sequences.\"?　Does it relevant that in some cases the game is played many times and is actually a long sequence, but what we have as data is a short sequence?",
    "2325659": "@pdnartreb 1st place bravo. Great Victory.",
    "2325346": "Congratulations on your first place!",
    "2325317": "very useful, thank you for your fanscinating work.",
    "2325268": "Such an impressive combination of good ideas and hard hard work.",
    "2325097": "Congrats on winning first place! Thanks a lot for your detailed explanation. Impressive! 🔥🔥🔥🔥🔥🔥",
    "2325053": "Congrats on the win! Really love the NN design here. I has a very similar NN design, except that I just one-hot encode everything due to low cardinality of categorical features.",
    "2324735": "Wow, just wow! :D\n\nSuch an impressive combination of good ideas and hard hard work. Very inspiring!",
    "2324715": "This is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lessonThis is very cool the understanding is very good and I like this lesson",
    "2324446": "Congrats @cpmpml and @pdnartreb for winning the 1st place! Also, thanks a lot for the detailed explanation!\n\nMay I ask how you decide to simply divide the numeric features (*e.g.,* `duration`) by a constant (60000 in the code you open source), instead of  using other normalization methods. I suppose that shrinking the scale of the raw feature could make the training process converge faster and other normalization techniques don't show superiority over the way you eventually adopted. If there's misunderstanding, please put me right. Thanks a lot!",
    "2324246": "it is very useful content",
    "2324241": "Congrats and thank you for sharing 🎉 Impressive solution with many creative ideas!",
    "2324151": "Congrats, mate!",
    "2324107": "Congratulations!  Keep it up!",
    "2324083": "Congratulations on winning!",
    "2323256": "Winner winner chicken dinner! :p",
    "2323126": "Congratulations on coming 1st @pdnartreb 🎉. Keep on achieving new milestones 👍",
    "2327560": "Congratulations! What a great work! You and your team deserve the 1st place.\nI am a newbee. May I ask a question about probing the private test?\nAs you mentioned here:\n>Probing showed to us that the private test consists in the 1st 1450/1500 sessions served by the API.\n\nAs far as what I know, the public test data can be probed, because the public learderboard give us the score on the public test dataset. But the private leaderboard seems won't give us feedback before the competition ends.\nSo how do you probe the private test dataset?\nThanks a lot.",
    "2326751": "Congratulations! Great work hope to learn from your excellent work！",
    "2325063": "Thank you for your posting!\n\nI have a question about \"for example monitor all folds in a CV (and accept only on majority or consensus)\".\n\nAm I correct in understanding that if adding a certain feature improves accuracy for fold1~4 and decreases accuracy for fold5, and as a result there is no change in the overall average score, then the feature is adopted?",
    "2324109": "Congrats, impressive solution! Thank you for your detailed explanation.💥",
    "2323791": "Congrats! ",
    "2323653": "Congrats!!\nThank you for providing this wonderful solution.\nI'm very pleased to be able to dig into this solution.",
    "2323625": "@cpmpml and @pdnartreb congrats you with 1st place and gold medals!",
    "2323517": "Hi, Congratulations! How much of an effort your team has put in this competition! Your really deserve to win this competition!\n\nI have a question which I hope you won't mind answering. In fact, this is a question I would like to ask every Kaggle competition winner...hope you won't mistake me..\n\nIn the spirit of competition, the Kaggle community makes tremendous effort to win over others that leads to fighting for third/fourth  decimal accuracy, which I consider unavoidable, and may even be good for innovation and further progress in research..\n\nBut as a competition host, suppose I ask you: \"For the competition sake, it's ok, but we are really not interested in third or fourth decimal accuracy for deployment in real world environment, but we want a much simpler but very robust model which can be extended for versions / deployed for long term..\" . What would be your recommended model architecture for a best second decimal accuracy for this specific problem? Given the amount of efforts and experimentations you have done, you must have such a solution in mind...Of course, there were thousands of submissions with 0.7 score, but we know all are not necessarily / equally that robust that can be extended to real world changes  or that would meet the client requirement..",
    "2323230": "I have been lucky to team with you!",
    "2323912": "I need help. i am beginner and do not know what to do I want to collect information about Talent acquisition data science fundamentals to learn about them what to write in parameters?\nimport requests\nimport time\nimport json\n\n# Set up API request parameters\nurl = 'https://www.kaggle.com/'\nheaders = {'Authorization': 'My API'}\nparams = {'q': ''}\n\n# Set up rate limiting parameters\nrequests_per_minute = 10\nseconds_per_request = 60 / requests_per_minute\n\n# Send API request and get response\nresponse = requests.get(url, headers=headers, params=params)\n\n# Check if the request was successful\nif response.status_code == 200:\n    # Parse the response\n    data = json.loads(response.content)\n\n    # Extract relevant information\n    for item in data['items']:\n        print('Title:', item['title'])\n        print('Link:', item['link'])\n        print('Description:', item['snippet'])\n        print('---')\nelse:\n    print('Error:', response.status_code)\n\n# Wait to ensure rate limiting compliance\ntime.sleep(seconds_per_request)\n# This script scrapes data from example.com for research purposes.\n# All data is used with permission and proper attribution is given.",
    "2539675": "I still don't understand how \"duration\" in your article is derived? Can you explain to me, thank you",
    "2343891": "Congratulations. And thanks for sharing the solution. The ML competition is great place to learn.",
    "2327283": "Congratulations! Great work hope to learn from your excellent workCongratulations! Great work hope to learn from your excellent work SIP",
    "3153970": "",
    "2327699": "",
    "2327556": "",
    "2327659": "Thanks for helping us learn!",
    "3322537": "looks great!"
  }
}