{
  "id": 209585,
  "title": "3rd place solution (0.818 LB): The Transformer",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209585",
  "author_name": "Javier Martín",
  "post_date": "2021-01-08T00:24:31.391000",
  "votes": 162,
  "comment_count": 37,
  "views": 0,
  "content": "<ul>\n<li><strong>2021-01-16 edit 2</strong>: added missing linear layers after categorical embeddings + tag features to diagram</li>\n<li><strong>2021-01-16 edit 1</strong>: the <a href=\"https://github.com/jamarju/riiid-acp-pub\" target=\"_blank\">source code</a> is up.</li>\n</ul>\n<p>Wow, what a ride this has been. First off, thanks <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/hoonpyotimjeon\" target=\"_blank\">@hoonpyotimjeon</a> Kaggle, Riiid and everyone involved in setting up this challenging competition.</p>\n<p>Congratulations to #1 and #2 <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> and <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>!!! We truly look forward to reading about your solutions!</p>\n<p>Also huge thanks to my teammate <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> who joined me in the last stretch and without whose ideas, intuition and hardware I wouldn’t have gotten this far.</p>\n<p>I was attracted to this competition by the relatively small dataset footprint compared to my other two previous competitions (deep fakes and RSNA pulmonary embolisms) but this ended up being much more resource intensive than I anticipated.</p>\n<p>Our solution is a mixture of two Transformer models with carefully crafted attention, engineered features and a time-aware adaptive ensembling mechanism we nick-named “The Blindfolded Gunslinger”.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F59b490f2f38b6d8b2412e5fd83585d2c%2Fblindfolded.jpg?generation=1610064421822590&amp;alt=media\" alt=\"\"></p>\n<p>(“Blindfolded gunslinger” hand-drawn by <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> inspired by Red Dead Redemption 2)</p>\n<h1>The Transformer</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F2bba4e4717ba0cc0f7e6af00dfe4a709%2FEsquemaRiid%202.jpg?generation=1610837878777203&amp;alt=media\" alt=\"\"></p>\n<p>We use two transformers trained separately with 2.5% of the users held out for validation and sequences of 500 interactions. At train we simply split user stories in 500 non-overlapping interaction chunks and sample the chunks randomly.</p>\n<ul>\n<li>Transformer 1: 3+3 layers (encoder+decoder), no LayerNorms, <a href=\"http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf\" target=\"_blank\">T-Fixup init</a> see paper for reasons why we used it, ReLU activations, d_model=512</li>\n<li>Transformer 2: 4+4 layers, no LayerNorms, T-Fixup init, GELU activations, d_model=512</li>\n</ul>\n<p>We feed both the encoder and the decoder ALL the features. We use learned features for continuous variables (simply projecting them to d_model) and for categorical variables we first map it to embeddings with low dimensionality and then project it also to d_model=512 (to avoid potential overfitting). We use an embedding bag for question tags.</p>\n<p>To prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right, hide all the answer-specific features from the decoder (user_answer, answered_correctly, qhe, qet), ie. those that are not immediately available upon inferring the interaction and use the appropriate attention masks in all 3 attentions.</p>\n<p>To our surprise the 3+3 model outperformed its bigger 4+4 brother even if we tried to finetune the latter at the final hours of the competition.</p>\n<h1>Engineered features</h1>\n<p>This was a very rich dataset  but we found the following derived features helped the transformer converge faster and reach a higher AUROC score. A lot of them have been discussed in the forum:</p>\n<ul>\n<li><code>qet</code>, <code>qhe</code>: these are the <code>prior_qet</code> and <code>prior_qhe</code> counterparts shifted upwards one container. This was probably discussed first by <a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194184\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194184</a></li>\n<li><code>tsli</code>: time since list interaction, AKA timestamp delta. Discussed in many threads.</li>\n<li><code>clipped_tsli</code>: <code>tsli</code> clipped to 20 minutes. <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> hinted at this in the Saint benchmark mega-thread: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632</a></li>\n<li><code>ts_mod_1day</code>: <code>timestamp</code> modulus 1 day. This may reveal daily patterns such as the user being more / less attentive / tired in the mornings / after work, etc.</li>\n<li><code>ts_mod_1week</code>: timestamp modulus 1 week. We similarly hope this will reveal weekly patterns (are “mondays” a bad day? etc.)</li>\n<li><code>attempt_num</code>, <code>attempts_correct</code>, <code>attempts_correct_avg</code>: about 11% of the questions were <strong>repeated</strong> questions, so it made a lot of sense to keep a record of which question had been answered by whom and how many times it was answered correctly. This was revealed by <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a> in his great thread: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266</a> and it was an extremely demanding feature to code since it takes a total of 11.2 GB of space at inference all by itself.</li>\n</ul>\n<h1>Attention</h1>\n<p>Our architecture follows the auto-regressive application of sequence to sequence transformers, so a causal attention mask is needed to make sure a given interaction cannot attend to interactions in the future; however there is an exception to this which we believe is critical: interaction grouped by the same task_container_id.</p>\n<p>We compute separate attention masks for the encoder and decoder, preventing the encoder self-attention from attending past interactions if they belong to the same task_container_id, and conversely modifying the decoder self-attention to allow to attend to all interaction belonging to the same task_container_id, we further restrain the output of the encoder to the decoder with the encoder attention to prevent a leakage of information from the residual connections in the encoder.</p>\n<pre><code>causal_mask  = ~torch.tril(torch.ones(1,sl, sl,dtype=torch.bool,device=x_cat.device)).expand(b,-1,-1)\nx_tci   = x_cat[...,Cats.task_container_id]\nx_tci_s = torch.zeros_like(x_tci)\nx_tci_s[...,1:] = x_tci[...,:-1]\nenc_container_aware_mask =  (x_tci.unsqueeze(-1) == x_tci_s.unsqueeze(-1).permute(0,2,1)) | causal_mask\ndec_container_aware_mask = ~(x_tci.unsqueeze(-1) == x_tci.unsqueeze(-1).permute(0,2,1))   &amp; causal_mask\n</code></pre>\n<h1>The Blindfolded Gunslinger</h1>\n<p>We made a joke in the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356#1137339\" target=\"_blank\">meme thread</a> about us wanting to ensemble multiple models, but the competition having only 9 hours to run full inference…</p>\n<p>It was already reported that the public test set was sitting on <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203383#1115680\" target=\"_blank\">the first 20% of the test set</a> so it is possible to maximize the allotted time for the private test set by skipping model inference in the first 20% and predicting only the last 80%. </p>\n<p>We implemented dynamic ensembling that attempts to perform as much ensembling as it is possible in the allotted time for the last 80%.</p>\n<p>We dubbed this idea “The Blindfolded Gunslinger” because it fires two guns (models) as much as it can (after a while it will only fire one) but it is blindfolded in the sense that the public LB will be ~0.5 so we cannot be sure if it worked or not until now…</p>\n<h1>Hardware</h1>\n<ul>\n<li>1 computer with Ryzen 3950x (16c32t) + 64 Gb RAM + 1x3090</li>\n<li>1 computer with Threadripper 1950x (16c32t) + 256 Gb RAM + 6x3090</li>\n</ul>\n<p>We set up the big computer during the competition which was a project on its own:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Ff1186277b6c2045ba9c00053192d6be2%2Fpasted%20image%200.png?generation=1610065136682440&amp;alt=media\" alt=\"\"></p>\n<p>Also in the last 8 hours of the competition we rented a 190 Gb RAM + 6x3090 but it did not help us much.</p>\n<h1>Software</h1>\n<p>We used pytorch 1.7.1, fastai and we trained using distributed training and mixed precision (both as implemented in fastai). </p>\n<p>For inference we included the last 500 interactions and summaries in both pickle files and memory-mapped numpy matrices.</p>\n<h1>Source code</h1>\n<p>Available at: <a href=\"https://github.com/jamarju/riiid-acp-pub\" target=\"_blank\">https://github.com/jamarju/riiid-acp-pub</a></p>",
  "messages": [
    {
      "id": 1143555,
      "postDate": "2021-01-08T00:24:31.390Z",
      "content": "<ul>\n<li><strong>2021-01-16 edit 2</strong>: added missing linear layers after categorical embeddings + tag features to diagram</li>\n<li><strong>2021-01-16 edit 1</strong>: the <a href=\"https://github.com/jamarju/riiid-acp-pub\" target=\"_blank\">source code</a> is up.</li>\n</ul>\n<p>Wow, what a ride this has been. First off, thanks <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/hoonpyotimjeon\" target=\"_blank\">@hoonpyotimjeon</a> Kaggle, Riiid and everyone involved in setting up this challenging competition.</p>\n<p>Congratulations to #1 and #2 <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> and <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>!!! We truly look forward to reading about your solutions!</p>\n<p>Also huge thanks to my teammate <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> who joined me in the last stretch and without whose ideas, intuition and hardware I wouldn’t have gotten this far.</p>\n<p>I was attracted to this competition by the relatively small dataset footprint compared to my other two previous competitions (deep fakes and RSNA pulmonary embolisms) but this ended up being much more resource intensive than I anticipated.</p>\n<p>Our solution is a mixture of two Transformer models with carefully crafted attention, engineered features and a time-aware adaptive ensembling mechanism we nick-named “The Blindfolded Gunslinger”.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F59b490f2f38b6d8b2412e5fd83585d2c%2Fblindfolded.jpg?generation=1610064421822590&amp;alt=media\" alt=\"\"></p>\n<p>(“Blindfolded gunslinger” hand-drawn by <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> inspired by Red Dead Redemption 2)</p>\n<h1>The Transformer</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F2bba4e4717ba0cc0f7e6af00dfe4a709%2FEsquemaRiid%202.jpg?generation=1610837878777203&amp;alt=media\" alt=\"\"></p>\n<p>We use two transformers trained separately with 2.5% of the users held out for validation and sequences of 500 interactions. At train we simply split user stories in 500 non-overlapping interaction chunks and sample the chunks randomly.</p>\n<ul>\n<li>Transformer 1: 3+3 layers (encoder+decoder), no LayerNorms, <a href=\"http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf\" target=\"_blank\">T-Fixup init</a> see paper for reasons why we used it, ReLU activations, d_model=512</li>\n<li>Transformer 2: 4+4 layers, no LayerNorms, T-Fixup init, GELU activations, d_model=512</li>\n</ul>\n<p>We feed both the encoder and the decoder ALL the features. We use learned features for continuous variables (simply projecting them to d_model) and for categorical variables we first map it to embeddings with low dimensionality and then project it also to d_model=512 (to avoid potential overfitting). We use an embedding bag for question tags.</p>\n<p>To prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right, hide all the answer-specific features from the decoder (user_answer, answered_correctly, qhe, qet), ie. those that are not immediately available upon inferring the interaction and use the appropriate attention masks in all 3 attentions.</p>\n<p>To our surprise the 3+3 model outperformed its bigger 4+4 brother even if we tried to finetune the latter at the final hours of the competition.</p>\n<h1>Engineered features</h1>\n<p>This was a very rich dataset  but we found the following derived features helped the transformer converge faster and reach a higher AUROC score. A lot of them have been discussed in the forum:</p>\n<ul>\n<li><code>qet</code>, <code>qhe</code>: these are the <code>prior_qet</code> and <code>prior_qhe</code> counterparts shifted upwards one container. This was probably discussed first by <a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194184\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194184</a></li>\n<li><code>tsli</code>: time since list interaction, AKA timestamp delta. Discussed in many threads.</li>\n<li><code>clipped_tsli</code>: <code>tsli</code> clipped to 20 minutes. <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> hinted at this in the Saint benchmark mega-thread: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632</a></li>\n<li><code>ts_mod_1day</code>: <code>timestamp</code> modulus 1 day. This may reveal daily patterns such as the user being more / less attentive / tired in the mornings / after work, etc.</li>\n<li><code>ts_mod_1week</code>: timestamp modulus 1 week. We similarly hope this will reveal weekly patterns (are “mondays” a bad day? etc.)</li>\n<li><code>attempt_num</code>, <code>attempts_correct</code>, <code>attempts_correct_avg</code>: about 11% of the questions were <strong>repeated</strong> questions, so it made a lot of sense to keep a record of which question had been answered by whom and how many times it was answered correctly. This was revealed by <a href=\"https://www.kaggle.com/aravindpadman\" target=\"_blank\">@aravindpadman</a> in his great thread: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266</a> and it was an extremely demanding feature to code since it takes a total of 11.2 GB of space at inference all by itself.</li>\n</ul>\n<h1>Attention</h1>\n<p>Our architecture follows the auto-regressive application of sequence to sequence transformers, so a causal attention mask is needed to make sure a given interaction cannot attend to interactions in the future; however there is an exception to this which we believe is critical: interaction grouped by the same task_container_id.</p>\n<p>We compute separate attention masks for the encoder and decoder, preventing the encoder self-attention from attending past interactions if they belong to the same task_container_id, and conversely modifying the decoder self-attention to allow to attend to all interaction belonging to the same task_container_id, we further restrain the output of the encoder to the decoder with the encoder attention to prevent a leakage of information from the residual connections in the encoder.</p>\n<pre><code>causal_mask  = ~torch.tril(torch.ones(1,sl, sl,dtype=torch.bool,device=x_cat.device)).expand(b,-1,-1)\nx_tci   = x_cat[...,Cats.task_container_id]\nx_tci_s = torch.zeros_like(x_tci)\nx_tci_s[...,1:] = x_tci[...,:-1]\nenc_container_aware_mask =  (x_tci.unsqueeze(-1) == x_tci_s.unsqueeze(-1).permute(0,2,1)) | causal_mask\ndec_container_aware_mask = ~(x_tci.unsqueeze(-1) == x_tci.unsqueeze(-1).permute(0,2,1))   &amp; causal_mask\n</code></pre>\n<h1>The Blindfolded Gunslinger</h1>\n<p>We made a joke in the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356#1137339\" target=\"_blank\">meme thread</a> about us wanting to ensemble multiple models, but the competition having only 9 hours to run full inference…</p>\n<p>It was already reported that the public test set was sitting on <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203383#1115680\" target=\"_blank\">the first 20% of the test set</a> so it is possible to maximize the allotted time for the private test set by skipping model inference in the first 20% and predicting only the last 80%. </p>\n<p>We implemented dynamic ensembling that attempts to perform as much ensembling as it is possible in the allotted time for the last 80%.</p>\n<p>We dubbed this idea “The Blindfolded Gunslinger” because it fires two guns (models) as much as it can (after a while it will only fire one) but it is blindfolded in the sense that the public LB will be ~0.5 so we cannot be sure if it worked or not until now…</p>\n<h1>Hardware</h1>\n<ul>\n<li>1 computer with Ryzen 3950x (16c32t) + 64 Gb RAM + 1x3090</li>\n<li>1 computer with Threadripper 1950x (16c32t) + 256 Gb RAM + 6x3090</li>\n</ul>\n<p>We set up the big computer during the competition which was a project on its own:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Ff1186277b6c2045ba9c00053192d6be2%2Fpasted%20image%200.png?generation=1610065136682440&amp;alt=media\" alt=\"\"></p>\n<p>Also in the last 8 hours of the competition we rented a 190 Gb RAM + 6x3090 but it did not help us much.</p>\n<h1>Software</h1>\n<p>We used pytorch 1.7.1, fastai and we trained using distributed training and mixed precision (both as implemented in fastai). </p>\n<p>For inference we included the last 500 interactions and summaries in both pickle files and memory-mapped numpy matrices.</p>\n<h1>Source code</h1>\n<p>Available at: <a href=\"https://github.com/jamarju/riiid-acp-pub\" target=\"_blank\">https://github.com/jamarju/riiid-acp-pub</a></p>",
      "rawMarkdown": "- **2021-01-16 edit 2**: added missing linear layers after categorical embeddings + tag features to diagram\n- **2021-01-16 edit 1**: the [source code](https://github.com/jamarju/riiid-acp-pub) is up.\n\nWow, what a ride this has been. First off, thanks @sohier @hoonpyotimjeon Kaggle, Riiid and everyone involved in setting up this challenging competition.\n\nCongratulations to #1 and #2 @keetar and @mamasinkgs!!! We truly look forward to reading about your solutions!\n\nAlso huge thanks to my teammate @antorsae who joined me in the last stretch and without whose ideas, intuition and hardware I wouldn’t have gotten this far.\n\nI was attracted to this competition by the relatively small dataset footprint compared to my other two previous competitions (deep fakes and RSNA pulmonary embolisms) but this ended up being much more resource intensive than I anticipated.\n\nOur solution is a mixture of two Transformer models with carefully crafted attention, engineered features and a time-aware adaptive ensembling mechanism we nick-named “The Blindfolded Gunslinger”.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F59b490f2f38b6d8b2412e5fd83585d2c%2Fblindfolded.jpg?generation=1610064421822590&alt=media)\n\n(“Blindfolded gunslinger” hand-drawn by @antorsae inspired by Red Dead Redemption 2)\n\n# The Transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F2bba4e4717ba0cc0f7e6af00dfe4a709%2FEsquemaRiid%202.jpg?generation=1610837878777203&alt=media)\n\nWe use two transformers trained separately with 2.5% of the users held out for validation and sequences of 500 interactions. At train we simply split user stories in 500 non-overlapping interaction chunks and sample the chunks randomly.\n\n* Transformer 1: 3+3 layers (encoder+decoder), no LayerNorms, [T-Fixup init] (http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf) see paper for reasons why we used it, ReLU activations, d_model=512\n* Transformer 2: 4+4 layers, no LayerNorms, T-Fixup init, GELU activations, d_model=512\n\nWe feed both the encoder and the decoder ALL the features. We use learned features for continuous variables (simply projecting them to d_model) and for categorical variables we first map it to embeddings with low dimensionality and then project it also to d_model=512 (to avoid potential overfitting). We use an embedding bag for question tags.\n\nTo prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right, hide all the answer-specific features from the decoder (user_answer, answered_correctly, qhe, qet), ie. those that are not immediately available upon inferring the interaction and use the appropriate attention masks in all 3 attentions.\n\nTo our surprise the 3+3 model outperformed its bigger 4+4 brother even if we tried to finetune the latter at the final hours of the competition.\n\n# Engineered features\n\nThis was a very rich dataset  but we found the following derived features helped the transformer converge faster and reach a higher AUROC score. A lot of them have been discussed in the forum:\n\n* `qet`, `qhe`: these are the `prior_qet` and `prior_qhe` counterparts shifted upwards one container. This was probably discussed first by @doctorkael here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194184\n* `tsli`: time since list interaction, AKA timestamp delta. Discussed in many threads.\n* `clipped_tsli`: `tsli` clipped to 20 minutes. @claverru hinted at this in the Saint benchmark mega-thread: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\n* `ts_mod_1day`: `timestamp` modulus 1 day. This may reveal daily patterns such as the user being more / less attentive / tired in the mornings / after work, etc.\n* `ts_mod_1week`: timestamp modulus 1 week. We similarly hope this will reveal weekly patterns (are “mondays” a bad day? etc.)\n* `attempt_num`, `attempts_correct`, `attempts_correct_avg`: about 11% of the questions were **repeated** questions, so it made a lot of sense to keep a record of which question had been answered by whom and how many times it was answered correctly. This was revealed by @aravindpadman in his great thread: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266 and it was an extremely demanding feature to code since it takes a total of 11.2 GB of space at inference all by itself.\n\n# Attention\n\nOur architecture follows the auto-regressive application of sequence to sequence transformers, so a causal attention mask is needed to make sure a given interaction cannot attend to interactions in the future; however there is an exception to this which we believe is critical: interaction grouped by the same task_container_id.\n\nWe compute separate attention masks for the encoder and decoder, preventing the encoder self-attention from attending past interactions if they belong to the same task_container_id, and conversely modifying the decoder self-attention to allow to attend to all interaction belonging to the same task_container_id, we further restrain the output of the encoder to the decoder with the encoder attention to prevent a leakage of information from the residual connections in the encoder.\n\n```\ncausal_mask  = ~torch.tril(torch.ones(1,sl, sl,dtype=torch.bool,device=x_cat.device)).expand(b,-1,-1)\nx_tci   = x_cat[...,Cats.task_container_id]\nx_tci_s = torch.zeros_like(x_tci)\nx_tci_s[...,1:] = x_tci[...,:-1]\nenc_container_aware_mask =  (x_tci.unsqueeze(-1) == x_tci_s.unsqueeze(-1).permute(0,2,1)) | causal_mask\ndec_container_aware_mask = ~(x_tci.unsqueeze(-1) == x_tci.unsqueeze(-1).permute(0,2,1))   & causal_mask\n```\n\n# The Blindfolded Gunslinger\n\nWe made a joke in the [meme thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356#1137339) about us wanting to ensemble multiple models, but the competition having only 9 hours to run full inference...\n\nIt was already reported that the public test set was sitting on [the first 20% of the test set](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203383#1115680) so it is possible to maximize the allotted time for the private test set by skipping model inference in the first 20% and predicting only the last 80%. \n\nWe implemented dynamic ensembling that attempts to perform as much ensembling as it is possible in the allotted time for the last 80%.\n\nWe dubbed this idea “The Blindfolded Gunslinger” because it fires two guns (models) as much as it can (after a while it will only fire one) but it is blindfolded in the sense that the public LB will be ~0.5 so we cannot be sure if it worked or not until now…\n\n# Hardware\n\n* 1 computer with Ryzen 3950x (16c32t) + 64 Gb RAM + 1x3090\n* 1 computer with Threadripper 1950x (16c32t) + 256 Gb RAM + 6x3090\n\nWe set up the big computer during the competition which was a project on its own:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Ff1186277b6c2045ba9c00053192d6be2%2Fpasted%20image%200.png?generation=1610065136682440&alt=media)\n\nAlso in the last 8 hours of the competition we rented a 190 Gb RAM + 6x3090 but it did not help us much.\n\n# Software\n\nWe used pytorch 1.7.1, fastai and we trained using distributed training and mixed precision (both as implemented in fastai). \n\nFor inference we included the last 500 interactions and summaries in both pickle files and memory-mapped numpy matrices.\n\n# Source code\n\nAvailable at: https://github.com/jamarju/riiid-acp-pub\n",
      "votes": 162
    },
    {
      "id": 1155737,
      "postDate": "2021-01-16T16:16:05.027Z",
      "content": "<p>The <a href=\"https://github.com/jamarju/riiid-acp-pub\" target=\"_blank\">source code</a> is up. <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> </p>",
      "rawMarkdown": "The [source code](https://github.com/jamarju/riiid-acp-pub) is up. @imeintanis ",
      "votes": 7,
      "replies": [
        {
          "id": 1159375,
          "postDate": "2021-01-19T08:15:07.190Z",
          "content": "<p>The code is so clean that it makes my eyes wet. Learning a lot in many aspects. Thanks a lot!</p>",
          "rawMarkdown": "The code is so clean that it makes my eyes wet. Learning a lot in many aspects. Thanks a lot!",
          "votes": 4
        },
        {
          "id": 1159724,
          "postDate": "2021-01-19T12:30:57.540Z",
          "content": "<p>LOL I've never had anyone say anything this beautiful about something I've written, thanks 😄</p>",
          "rawMarkdown": "LOL I've never had anyone say anything this beautiful about something I've written, thanks 😄",
          "votes": 2
        }
      ]
    },
    {
      "id": 1144532,
      "postDate": "2021-01-08T14:09:40.320Z",
      "content": "<p>Congrats! No doubt why you got your well deserved 3rd position. Keep it up!</p>",
      "rawMarkdown": "Congrats! No doubt why you got your well deserved 3rd position. Keep it up!",
      "votes": 5
    },
    {
      "id": 1143683,
      "postDate": "2021-01-08T02:08:11.437Z",
      "content": "<p>Wow very impressive indeed!! Congratulations - are you planning to share the code or at least your Blindfolded Guns approach ?</p>",
      "rawMarkdown": "Wow very impressive indeed!! Congratulations - are you planning to share the code or at least your Blindfolded Guns approach ?",
      "votes": 5,
      "replies": [
        {
          "id": 1145138,
          "postDate": "2021-01-08T23:33:45.597Z",
          "content": "<p>Yes, we will. Allow me a couple of days to recover mentally and clean up.</p>",
          "rawMarkdown": "Yes, we will. Allow me a couple of days to recover mentally and clean up.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1144472,
      "postDate": "2021-01-08T13:33:42.117Z",
      "content": "<p>Congratulations! TIL the T-fixup trick. I think the right link to the paper is <a href=\"http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf\" target=\"_blank\">http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf</a> </p>",
      "rawMarkdown": "Congratulations! TIL the T-fixup trick. I think the right link to the paper is http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf ",
      "votes": 3,
      "replies": [
        {
          "id": 1144493,
          "postDate": "2021-01-08T13:41:49.540Z",
          "content": "<p>You are right, I fixed the link. Thanks for noticing!</p>",
          "rawMarkdown": "You are right, I fixed the link. Thanks for noticing!"
        }
      ]
    },
    {
      "id": 1144360,
      "postDate": "2021-01-08T12:20:30.213Z",
      "content": "<p>Thanks for the very interesting read, and congratulations on the 3rd place !</p>",
      "rawMarkdown": "Thanks for the very interesting read, and congratulations on the 3rd place !",
      "votes": 3
    },
    {
      "id": 1143990,
      "postDate": "2021-01-08T07:02:01.270Z",
      "content": "<p>Congratulations and amazing sulotion !<br>\nOne quick question: why you choose an encoder-decoder framework? Since the inputs and outputs are always equal length, is sequence-tagging method more suitable?</p>",
      "rawMarkdown": "Congratulations and amazing sulotion !\nOne quick question: why you choose an encoder-decoder framework? Since the inputs and outputs are always equal length, is sequence-tagging method more suitable?",
      "votes": 3,
      "replies": [
        {
          "id": 1144214,
          "postDate": "2021-01-08T10:03:43.853Z",
          "content": "<p>Thank you! We tried a variety of encoder+decoder layouts (0+4, 4+0, 2+4, 4+2, etc.) and those two (3+3, 4+4) turned out the best. I believe we are actually doing sequence tagging, I just didn't know it had such name.</p>",
          "rawMarkdown": "Thank you! We tried a variety of encoder+decoder layouts (0+4, 4+0, 2+4, 4+2, etc.) and those two (3+3, 4+4) turned out the best. I believe we are actually doing sequence tagging, I just didn't know it had such name.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1143560,
      "postDate": "2021-01-08T00:27:15.923Z",
      "content": "<p>Amazing solution. Congratulations on results <a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> and team. </p>",
      "rawMarkdown": "Amazing solution. Congratulations on results @bacterio and team. ",
      "votes": 3
    },
    {
      "id": 1151076,
      "postDate": "2021-01-13T05:47:04.790Z",
      "content": "<p>Congratulations and thank you for sharing <br>\nthis computer is amazing.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing \nthis computer is amazing.",
      "votes": 2
    },
    {
      "id": 1145755,
      "postDate": "2021-01-09T10:14:56.123Z",
      "content": "<p>That computer is impressive. 😄<br>\nCongratulations on the third place and thanks for sharing your solution. 👌</p>",
      "rawMarkdown": "That computer is impressive. 😄\nCongratulations on the third place and thanks for sharing your solution. 👌",
      "votes": 2
    },
    {
      "id": 1144631,
      "postDate": "2021-01-08T15:19:56.953Z",
      "content": "<p>Thanks for sharing and congratulations!</p>\n<p>I have a question about the hardware: was this used only to train a model? Or was it used in the submission process somehow (this would confuse me since I thought submission had to be made through the kernel)?</p>",
      "rawMarkdown": "Thanks for sharing and congratulations!\n\nI have a question about the hardware: was this used only to train a model? Or was it used in the submission process somehow (this would confuse me since I thought submission had to be made through the kernel)?",
      "votes": 2,
      "replies": [
        {
          "id": 1144658,
          "postDate": "2021-01-08T15:44:36.117Z",
          "content": "<p>Train only.</p>",
          "rawMarkdown": "Train only.",
          "votes": 1
        },
        {
          "id": 1144664,
          "postDate": "2021-01-08T15:46:18.667Z",
          "content": "<p>Thanks very much!</p>",
          "rawMarkdown": "Thanks very much!"
        }
      ]
    },
    {
      "id": 1144564,
      "postDate": "2021-01-08T14:33:40.537Z",
      "content": "<p>Congrats you and your team for the 3rd position!!! Thank you for your detailed solutions, so much to learn</p>",
      "rawMarkdown": "Congrats you and your team for the 3rd position!!! Thank you for your detailed solutions, so much to learn",
      "votes": 2
    },
    {
      "id": 1143896,
      "postDate": "2021-01-08T05:53:50.517Z",
      "content": "<p>Congrats, learn a lot  from your team.</p>",
      "rawMarkdown": "Congrats, learn a lot  from your team.",
      "votes": 2
    },
    {
      "id": 1143754,
      "postDate": "2021-01-08T03:37:22.960Z",
      "content": "<p>Congratulations! Amazing work! 💥🎉</p>",
      "rawMarkdown": "Congratulations! Amazing work! 💥🎉",
      "votes": 2
    },
    {
      "id": 1143588,
      "postDate": "2021-01-08T00:47:40.723Z",
      "content": "<p>Such graphics :D</p>",
      "rawMarkdown": "Such graphics :D",
      "votes": 2
    },
    {
      "id": 1143583,
      "postDate": "2021-01-08T00:43:46.463Z",
      "content": "<p>Wow impresionante. Enhorabuena por ese tercer puesto  <a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> !  ¿Cuánto habéis tardado en montar esa bestia de ordenador?</p>",
      "rawMarkdown": "Wow impresionante. Enhorabuena por ese tercer puesto  @bacterio @antorsae !  ¿Cuánto habéis tardado en montar esa bestia de ordenador?",
      "votes": 2,
      "replies": [
        {
          "id": 1143591,
          "postDate": "2021-01-08T00:48:52.010Z",
          "content": "<p>Question was how long did it took to set up the computer.</p>\n<p>Well, I had to build my own chassis using 20x20 aluminum profiles, dual 1600W PSU, some 3d-printed supports, PCIe switches (there's 4 PCIe slots only), etc. I had a lot of issues POSTing w/ 5+ GPUs which I fixed by disabling non-critical stuff in BIOS (e.g. HD audio, etc.). Xataka (Spanish website) will soon release a small post about the build. Stay tuned.</p>",
          "rawMarkdown": "Question was how long did it took to set up the computer.\n\nWell, I had to build my own chassis using 20x20 aluminum profiles, dual 1600W PSU, some 3d-printed supports, PCIe switches (there's 4 PCIe slots only), etc. I had a lot of issues POSTing w/ 5+ GPUs which I fixed by disabling non-critical stuff in BIOS (e.g. HD audio, etc.). Xataka (Spanish website) will soon release a small post about the build. Stay tuned.",
          "votes": 5
        },
        {
          "id": 1143717,
          "postDate": "2021-01-08T02:58:12.317Z",
          "content": "<p>I did not know the separate PCIe switches exist, do you have a link? Would be very interested to read about your build.</p>",
          "rawMarkdown": "I did not know the separate PCIe switches exist, do you have a link? Would be very interested to read about your build.",
          "votes": 2
        },
        {
          "id": 1144383,
          "postDate": "2021-01-08T12:37:09.680Z",
          "content": "<p>It's actually a PCIe bifurcator and your bios has to support the actual PCIe switching. Mine lucklily had it (Asrock X399), then it's this: <a href=\"https://peine-braun.net/shop/index.php?route=product/category&amp;path=65\" target=\"_blank\">https://peine-braun.net/shop/index.php?route=product/category&amp;path=65</a> alternatively other PCIe 4.0 solutions exist <a href=\"http://www.ioi.com.tw/products/proddetail.aspx?CatID=106&amp;DeviceID=3050&amp;HostID=2108&amp;ProdID=1060249\" target=\"_blank\">http://www.ioi.com.tw/products/proddetail.aspx?CatID=106&amp;DeviceID=3050&amp;HostID=2108&amp;ProdID=1060249</a></p>",
          "rawMarkdown": "It's actually a PCIe bifurcator and your bios has to support the actual PCIe switching. Mine lucklily had it (Asrock X399), then it's this: https://peine-braun.net/shop/index.php?route=product/category&path=65 alternatively other PCIe 4.0 solutions exist http://www.ioi.com.tw/products/proddetail.aspx?CatID=106&DeviceID=3050&HostID=2108&ProdID=1060249",
          "votes": 1
        },
        {
          "id": 1145137,
          "postDate": "2021-01-08T23:32:59.957Z",
          "content": "<p>Thanks, and congratulations with the great solution and 3rd place!</p>",
          "rawMarkdown": "Thanks, and congratulations with the great solution and 3rd place!",
          "votes": 1
        },
        {
          "id": 1145773,
          "postDate": "2021-01-09T10:25:26.957Z",
          "content": "<p>Would love to read the blog post. That's an impressive build on its own. 👌<br>\nMaybe there will be a new Kaggle category: builder GM. :p</p>",
          "rawMarkdown": "Would love to read the blog post. That's an impressive build on its own. 👌\nMaybe there will be a new Kaggle category: builder GM. :p"
        }
      ]
    },
    {
      "id": 1143572,
      "postDate": "2021-01-08T00:36:54.963Z",
      "content": "<p>Congrats for the 3rd place and thank you for sharing. I was very impressed by the elaborate NN structure. </p>",
      "rawMarkdown": "Congrats for the 3rd place and thank you for sharing. I was very impressed by the elaborate NN structure. ",
      "votes": 2
    },
    {
      "id": 1143566,
      "postDate": "2021-01-08T00:30:22.387Z",
      "content": "<p>Wow very nice solution! And insane compute, this comp is very compute hungry!</p>",
      "rawMarkdown": "Wow very nice solution! And insane compute, this comp is very compute hungry!",
      "votes": 2
    },
    {
      "id": 3424427,
      "postDate": "2026-03-19T17:14:30.423Z",
      "content": "<p>Impressive work dude!</p>",
      "rawMarkdown": "Impressive work dude!"
    },
    {
      "id": 1606342,
      "postDate": "2021-12-05T00:06:09.627Z",
      "content": "<p>Very late congratulations! and thank you very much for the solution write-up. </p>\n<p>I have a question and just wondering if you could kindly clarify:</p>\n<p>\"To prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right\"</p>\n<p>Since you have already used a causal mask to prevent attention from the future to the past, why is the above operation (shifting the encode input to the right) still needed?</p>\n<p>many thanks!</p>",
      "rawMarkdown": "Very late congratulations! and thank you very much for the solution write-up. \n\nI have a question and just wondering if you could kindly clarify:\n\n\"To prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right\"\n\nSince you have already used a causal mask to prevent attention from the future to the past, why is the above operation (shifting the encode input to the right) still needed?\n\nmany thanks!\n\n",
      "replies": [
        {
          "id": 1607276,
          "postDate": "2021-12-05T19:36:40.047Z",
          "content": "<p>Let's say the encoder input is: $$QA_1, QA_2,  QA_3, …$$ and the decoder input is: $$Q_1, Q_2, Q_3, …$$ where \\(QA_x\\) is question+answer \\(x\\)'s data and \\(Q_x\\) is question \\(x\\)'s data.</p>\n<p>Now suppose we don't shift the encoder input: For question \\(Q_x\\) we MUST mask all QA from \\(QA_1\\) up to (and including) \\(QA_x\\), otherwise question 1 would be able to see its own answer and cheat. This means that for \\(Q_1\\) all QAs should be masked but we can't do that because in the attention mechanism, masking means we set the masked values to -inf after calculating \\(QK^T\\). In the case of \\(Q_1\\), the first row would be <strong>entirely</strong> set to -inf, and that's a problem because the softmax of all -inf is nan:</p>\n<pre><code>a = torch.triu(torch.full((5, 5), -float('inf')))\ntorch.nn.functional.softmax(a, 1)\ntensor([[   nan,    nan,    nan,    nan,    nan],\n        [1.0000, 0.0000, 0.0000, 0.0000, 0.0000],\n        [0.5000, 0.5000, 0.0000, 0.0000, 0.0000],\n        [0.3333, 0.3333, 0.3333, 0.0000, 0.0000],\n        [0.2500, 0.2500, 0.2500, 0.2500, 0.0000]])\n</code></pre>\n<p>And that's why we always need to shift the encoder to the right in causal applications: there always has to be at least one token to pay attention to so that softmax can be calculated.</p>",
          "rawMarkdown": "Let's say the encoder input is: $$QA_1, QA_2,  QA_3, ...$$ and the decoder input is: $$Q_1, Q_2, Q_3, ...$$ where \\\\(QA_x\\\\) is question+answer \\\\(x\\\\)'s data and \\\\(Q_x\\\\) is question \\\\(x\\\\)'s data.\n\nNow suppose we don't shift the encoder input: For question \\\\(Q_x\\\\) we MUST mask all QA from \\\\(QA_1\\\\) up to (and including) \\\\(QA_x\\\\), otherwise question 1 would be able to see its own answer and cheat. This means that for \\\\(Q_1\\\\) all QAs should be masked but we can't do that because in the attention mechanism, masking means we set the masked values to -inf after calculating \\\\(QK^T\\\\). In the case of \\\\(Q_1\\\\), the first row would be **entirely** set to -inf, and that's a problem because the softmax of all -inf is nan:\n\n```\na = torch.triu(torch.full((5, 5), -float('inf')))\ntorch.nn.functional.softmax(a, 1)\ntensor([[   nan,    nan,    nan,    nan,    nan],\n        [1.0000, 0.0000, 0.0000, 0.0000, 0.0000],\n        [0.5000, 0.5000, 0.0000, 0.0000, 0.0000],\n        [0.3333, 0.3333, 0.3333, 0.0000, 0.0000],\n        [0.2500, 0.2500, 0.2500, 0.2500, 0.0000]])\n```\n\nAnd that's why we always need to shift the encoder to the right in causal applications: there always has to be at least one token to pay attention to so that softmax can be calculated.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1312135,
      "postDate": "2021-05-17T20:47:10.857Z",
      "content": "<p>Great solution!  Tried to submit the 05_inference.ipynb to Kaggle, but the notebook ran out of the memory.  So I ran this notebook and generated submission.csv locally, and copied submission.csv to the directory \"/kaggle/working/\" and submitted to Kaggle again, and the notebook was stopped after a few minutes with error, but without any score.</p>\n<p>Have anyone tried to submit the 05_inference.ipynb to Kaggle, and successfully got good score?</p>",
      "rawMarkdown": "Great solution!  Tried to submit the 05_inference.ipynb to Kaggle, but the notebook ran out of the memory.  So I ran this notebook and generated submission.csv locally, and copied submission.csv to the directory \"/kaggle/working/\" and submitted to Kaggle again, and the notebook was stopped after a few minutes with error, but without any score.\n\n\nHave anyone tried to submit the 05_inference.ipynb to Kaggle, and successfully got good score?",
      "replies": [
        {
          "id": 1312815,
          "postDate": "2021-05-18T09:04:53.970Z",
          "content": "<p>Hello, <a href=\"https://www.kaggle.com/julia5\" target=\"_blank\">@julia5</a> and thanks for the kind words.</p>\n<p>You can't run 05 off-line because 80% of the test data is hidden.</p>\n<p>Try to convert the 05.ipynb nb to .py and submit that instead, it should reduce RAM footprint and that's actually how we did it. The first cell in 05 explains how to do such conversion with the <code>ipynb-py-convert</code> tool.</p>",
          "rawMarkdown": "Hello, @julia5 and thanks for the kind words.\n\nYou can't run 05 off-line because 80% of the test data is hidden.\n\nTry to convert the 05.ipynb nb to .py and submit that instead, it should reduce RAM footprint and that's actually how we did it. The first cell in 05 explains how to do such conversion with the `ipynb-py-convert` tool.\n",
          "votes": 1
        },
        {
          "id": 1323103,
          "postDate": "2021-05-26T02:12:14.030Z",
          "content": "<p>Thank you for your kind reply, it's really helpful.  I did what you said, and submitted the notebook to Kaggle successfully. </p>",
          "rawMarkdown": "Thank you for your kind reply, it's really helpful.  I did what you said, and submitted the notebook to Kaggle successfully. "
        }
      ]
    },
    {
      "id": 1217134,
      "postDate": "2021-02-24T19:25:51.620Z",
      "content": "<p>Haha, truly embodying the \"python should be fun\" ethos, nice work :)</p>",
      "rawMarkdown": "Haha, truly embodying the \"python should be fun\" ethos, nice work :)"
    },
    {
      "id": 1659819,
      "postDate": "2022-01-22T06:14:04.487Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1155737,
      "author_name": "Javier Martín",
      "author_url": "",
      "post_date": "2021-01-16T16:16:05.027000",
      "content": "<p>The <a href=\"https://github.com/jamarju/riiid-acp-pub\" target=\"_blank\">source code</a> is up. <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> </p>",
      "votes": 7,
      "replies": [
        {
          "id": 1159375,
          "author_name": "kobi2000",
          "author_url": "",
          "post_date": "2021-01-19T08:15:07.190000",
          "content": "<p>The code is so clean that it makes my eyes wet. Learning a lot in many aspects. Thanks a lot!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1159724,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2021-01-19T12:30:57.540000",
          "content": "<p>LOL I've never had anyone say anything this beautiful about something I've written, thanks 😄</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1144532,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2021-01-08T14:09:40.320000",
      "content": "<p>Congrats! No doubt why you got your well deserved 3rd position. Keep it up!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1143683,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-01-08T02:08:11.437000",
      "content": "<p>Wow very impressive indeed!! Congratulations - are you planning to share the code or at least your Blindfolded Guns approach ?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1145138,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2021-01-08T23:33:45.597000",
          "content": "<p>Yes, we will. Allow me a couple of days to recover mentally and clean up.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1144472,
      "author_name": "Iván de Prado",
      "author_url": "",
      "post_date": "2021-01-08T13:33:42.117000",
      "content": "<p>Congratulations! TIL the T-fixup trick. I think the right link to the paper is <a href=\"http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf\" target=\"_blank\">http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf</a> </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1144493,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2021-01-08T13:41:49.540000",
          "content": "<p>You are right, I fixed the link. Thanks for noticing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1144360,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-01-08T12:20:30.213000",
      "content": "<p>Thanks for the very interesting read, and congratulations on the 3rd place !</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1143990,
      "author_name": "Yu Kang",
      "author_url": "",
      "post_date": "2021-01-08T07:02:01.270000",
      "content": "<p>Congratulations and amazing sulotion !<br>\nOne quick question: why you choose an encoder-decoder framework? Since the inputs and outputs are always equal length, is sequence-tagging method more suitable?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1144214,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2021-01-08T10:03:43.853000",
          "content": "<p>Thank you! We tried a variety of encoder+decoder layouts (0+4, 4+0, 2+4, 4+2, etc.) and those two (3+3, 4+4) turned out the best. I believe we are actually doing sequence tagging, I just didn't know it had such name.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1143560,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-01-08T00:27:15.923000",
      "content": "<p>Amazing solution. Congratulations on results <a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> and team. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1151076,
      "author_name": "LeiWang66808",
      "author_url": "",
      "post_date": "2021-01-13T05:47:04.790000",
      "content": "<p>Congratulations and thank you for sharing <br>\nthis computer is amazing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1145755,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-01-09T10:14:56.123000",
      "content": "<p>That computer is impressive. 😄<br>\nCongratulations on the third place and thanks for sharing your solution. 👌</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1144631,
      "author_name": "Neil Gibbons",
      "author_url": "",
      "post_date": "2021-01-08T15:19:56.953000",
      "content": "<p>Thanks for sharing and congratulations!</p>\n<p>I have a question about the hardware: was this used only to train a model? Or was it used in the submission process somehow (this would confuse me since I thought submission had to be made through the kernel)?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1144658,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2021-01-08T15:44:36.117000",
          "content": "<p>Train only.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1144664,
          "author_name": "Neil Gibbons",
          "author_url": "",
          "post_date": "2021-01-08T15:46:18.667000",
          "content": "<p>Thanks very much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1144564,
      "author_name": "Minh Tri Phan",
      "author_url": "",
      "post_date": "2021-01-08T14:33:40.537000",
      "content": "<p>Congrats you and your team for the 3rd position!!! Thank you for your detailed solutions, so much to learn</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1143896,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2021-01-08T05:53:50.517000",
      "content": "<p>Congrats, learn a lot  from your team.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1143754,
      "author_name": "Naresh Jagadeesan",
      "author_url": "",
      "post_date": "2021-01-08T03:37:22.960000",
      "content": "<p>Congratulations! Amazing work! 💥🎉</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1143588,
      "author_name": "AmorfEvo",
      "author_url": "",
      "post_date": "2021-01-08T00:47:40.723000",
      "content": "<p>Such graphics :D</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1143583,
      "author_name": "Maher el Ouahabi",
      "author_url": "",
      "post_date": "2021-01-08T00:43:46.463000",
      "content": "<p>Wow impresionante. Enhorabuena por ese tercer puesto  <a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> !  ¿Cuánto habéis tardado en montar esa bestia de ordenador?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1143591,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2021-01-08T00:48:52.010000",
          "content": "<p>Question was how long did it took to set up the computer.</p>\n<p>Well, I had to build my own chassis using 20x20 aluminum profiles, dual 1600W PSU, some 3d-printed supports, PCIe switches (there's 4 PCIe slots only), etc. I had a lot of issues POSTing w/ 5+ GPUs which I fixed by disabling non-critical stuff in BIOS (e.g. HD audio, etc.). Xataka (Spanish website) will soon release a small post about the build. Stay tuned.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1143717,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2021-01-08T02:58:12.317000",
          "content": "<p>I did not know the separate PCIe switches exist, do you have a link? Would be very interested to read about your build.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1144383,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2021-01-08T12:37:09.680000",
          "content": "<p>It's actually a PCIe bifurcator and your bios has to support the actual PCIe switching. Mine lucklily had it (Asrock X399), then it's this: <a href=\"https://peine-braun.net/shop/index.php?route=product/category&amp;path=65\" target=\"_blank\">https://peine-braun.net/shop/index.php?route=product/category&amp;path=65</a> alternatively other PCIe 4.0 solutions exist <a href=\"http://www.ioi.com.tw/products/proddetail.aspx?CatID=106&amp;DeviceID=3050&amp;HostID=2108&amp;ProdID=1060249\" target=\"_blank\">http://www.ioi.com.tw/products/proddetail.aspx?CatID=106&amp;DeviceID=3050&amp;HostID=2108&amp;ProdID=1060249</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145137,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2021-01-08T23:32:59.957000",
          "content": "<p>Thanks, and congratulations with the great solution and 3rd place!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145773,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2021-01-09T10:25:26.957000",
          "content": "<p>Would love to read the blog post. That's an impressive build on its own. 👌<br>\nMaybe there will be a new Kaggle category: builder GM. :p</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1143572,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2021-01-08T00:36:54.963000",
      "content": "<p>Congrats for the 3rd place and thank you for sharing. I was very impressed by the elaborate NN structure. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1143566,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2021-01-08T00:30:22.387000",
      "content": "<p>Wow very nice solution! And insane compute, this comp is very compute hungry!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3424427,
      "author_name": "Zayyam Wani",
      "author_url": "",
      "post_date": "2026-03-19T17:14:30.423000",
      "content": "<p>Impressive work dude!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1606342,
      "author_name": "YL",
      "author_url": "",
      "post_date": "2021-12-05T00:06:09.627000",
      "content": "<p>Very late congratulations! and thank you very much for the solution write-up. </p>\n<p>I have a question and just wondering if you could kindly clarify:</p>\n<p>\"To prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right\"</p>\n<p>Since you have already used a causal mask to prevent attention from the future to the past, why is the above operation (shifting the encode input to the right) still needed?</p>\n<p>many thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1607276,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2021-12-05T19:36:40.047000",
          "content": "<p>Let's say the encoder input is: $$QA_1, QA_2,  QA_3, …$$ and the decoder input is: $$Q_1, Q_2, Q_3, …$$ where \\(QA_x\\) is question+answer \\(x\\)'s data and \\(Q_x\\) is question \\(x\\)'s data.</p>\n<p>Now suppose we don't shift the encoder input: For question \\(Q_x\\) we MUST mask all QA from \\(QA_1\\) up to (and including) \\(QA_x\\), otherwise question 1 would be able to see its own answer and cheat. This means that for \\(Q_1\\) all QAs should be masked but we can't do that because in the attention mechanism, masking means we set the masked values to -inf after calculating \\(QK^T\\). In the case of \\(Q_1\\), the first row would be <strong>entirely</strong> set to -inf, and that's a problem because the softmax of all -inf is nan:</p>\n<pre><code>a = torch.triu(torch.full((5, 5), -float('inf')))\ntorch.nn.functional.softmax(a, 1)\ntensor([[   nan,    nan,    nan,    nan,    nan],\n        [1.0000, 0.0000, 0.0000, 0.0000, 0.0000],\n        [0.5000, 0.5000, 0.0000, 0.0000, 0.0000],\n        [0.3333, 0.3333, 0.3333, 0.0000, 0.0000],\n        [0.2500, 0.2500, 0.2500, 0.2500, 0.0000]])\n</code></pre>\n<p>And that's why we always need to shift the encoder to the right in causal applications: there always has to be at least one token to pay attention to so that softmax can be calculated.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1312135,
      "author_name": "Julia",
      "author_url": "",
      "post_date": "2021-05-17T20:47:10.857000",
      "content": "<p>Great solution!  Tried to submit the 05_inference.ipynb to Kaggle, but the notebook ran out of the memory.  So I ran this notebook and generated submission.csv locally, and copied submission.csv to the directory \"/kaggle/working/\" and submitted to Kaggle again, and the notebook was stopped after a few minutes with error, but without any score.</p>\n<p>Have anyone tried to submit the 05_inference.ipynb to Kaggle, and successfully got good score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1312815,
          "author_name": "Javier Martín",
          "author_url": "",
          "post_date": "2021-05-18T09:04:53.970000",
          "content": "<p>Hello, <a href=\"https://www.kaggle.com/julia5\" target=\"_blank\">@julia5</a> and thanks for the kind words.</p>\n<p>You can't run 05 off-line because 80% of the test data is hidden.</p>\n<p>Try to convert the 05.ipynb nb to .py and submit that instead, it should reduce RAM footprint and that's actually how we did it. The first cell in 05 explains how to do such conversion with the <code>ipynb-py-convert</code> tool.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1323103,
          "author_name": "Julia",
          "author_url": "",
          "post_date": "2021-05-26T02:12:14.030000",
          "content": "<p>Thank you for your kind reply, it's really helpful.  I did what you said, and submitted the notebook to Kaggle successfully. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1217134,
      "author_name": "Jonathan Bowden",
      "author_url": "",
      "post_date": "2021-02-24T19:25:51.620000",
      "content": "<p>Haha, truly embodying the \"python should be fun\" ethos, nice work :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1659819,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-22T06:14:04.487000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143555": "- **2021-01-16 edit 2**: added missing linear layers after categorical embeddings + tag features to diagram\n- **2021-01-16 edit 1**: the [source code](https://github.com/jamarju/riiid-acp-pub) is up.\n\nWow, what a ride this has been. First off, thanks @sohier @hoonpyotimjeon Kaggle, Riiid and everyone involved in setting up this challenging competition.\n\nCongratulations to #1 and #2 @keetar and @mamasinkgs!!! We truly look forward to reading about your solutions!\n\nAlso huge thanks to my teammate @antorsae who joined me in the last stretch and without whose ideas, intuition and hardware I wouldn’t have gotten this far.\n\nI was attracted to this competition by the relatively small dataset footprint compared to my other two previous competitions (deep fakes and RSNA pulmonary embolisms) but this ended up being much more resource intensive than I anticipated.\n\nOur solution is a mixture of two Transformer models with carefully crafted attention, engineered features and a time-aware adaptive ensembling mechanism we nick-named “The Blindfolded Gunslinger”.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F59b490f2f38b6d8b2412e5fd83585d2c%2Fblindfolded.jpg?generation=1610064421822590&alt=media)\n\n(“Blindfolded gunslinger” hand-drawn by @antorsae inspired by Red Dead Redemption 2)\n\n# The Transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2F2bba4e4717ba0cc0f7e6af00dfe4a709%2FEsquemaRiid%202.jpg?generation=1610837878777203&alt=media)\n\nWe use two transformers trained separately with 2.5% of the users held out for validation and sequences of 500 interactions. At train we simply split user stories in 500 non-overlapping interaction chunks and sample the chunks randomly.\n\n* Transformer 1: 3+3 layers (encoder+decoder), no LayerNorms, [T-Fixup init] (http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf) see paper for reasons why we used it, ReLU activations, d_model=512\n* Transformer 2: 4+4 layers, no LayerNorms, T-Fixup init, GELU activations, d_model=512\n\nWe feed both the encoder and the decoder ALL the features. We use learned features for continuous variables (simply projecting them to d_model) and for categorical variables we first map it to embeddings with low dimensionality and then project it also to d_model=512 (to avoid potential overfitting). We use an embedding bag for question tags.\n\nTo prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right, hide all the answer-specific features from the decoder (user_answer, answered_correctly, qhe, qet), ie. those that are not immediately available upon inferring the interaction and use the appropriate attention masks in all 3 attentions.\n\nTo our surprise the 3+3 model outperformed its bigger 4+4 brother even if we tried to finetune the latter at the final hours of the competition.\n\n# Engineered features\n\nThis was a very rich dataset  but we found the following derived features helped the transformer converge faster and reach a higher AUROC score. A lot of them have been discussed in the forum:\n\n* `qet`, `qhe`: these are the `prior_qet` and `prior_qhe` counterparts shifted upwards one container. This was probably discussed first by @doctorkael here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194184\n* `tsli`: time since list interaction, AKA timestamp delta. Discussed in many threads.\n* `clipped_tsli`: `tsli` clipped to 20 minutes. @claverru hinted at this in the Saint benchmark mega-thread: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\n* `ts_mod_1day`: `timestamp` modulus 1 day. This may reveal daily patterns such as the user being more / less attentive / tired in the mornings / after work, etc.\n* `ts_mod_1week`: timestamp modulus 1 week. We similarly hope this will reveal weekly patterns (are “mondays” a bad day? etc.)\n* `attempt_num`, `attempts_correct`, `attempts_correct_avg`: about 11% of the questions were **repeated** questions, so it made a lot of sense to keep a record of which question had been answered by whom and how many times it was answered correctly. This was revealed by @aravindpadman in his great thread: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266 and it was an extremely demanding feature to code since it takes a total of 11.2 GB of space at inference all by itself.\n\n# Attention\n\nOur architecture follows the auto-regressive application of sequence to sequence transformers, so a causal attention mask is needed to make sure a given interaction cannot attend to interactions in the future; however there is an exception to this which we believe is critical: interaction grouped by the same task_container_id.\n\nWe compute separate attention masks for the encoder and decoder, preventing the encoder self-attention from attending past interactions if they belong to the same task_container_id, and conversely modifying the decoder self-attention to allow to attend to all interaction belonging to the same task_container_id, we further restrain the output of the encoder to the decoder with the encoder attention to prevent a leakage of information from the residual connections in the encoder.\n\n```\ncausal_mask  = ~torch.tril(torch.ones(1,sl, sl,dtype=torch.bool,device=x_cat.device)).expand(b,-1,-1)\nx_tci   = x_cat[...,Cats.task_container_id]\nx_tci_s = torch.zeros_like(x_tci)\nx_tci_s[...,1:] = x_tci[...,:-1]\nenc_container_aware_mask =  (x_tci.unsqueeze(-1) == x_tci_s.unsqueeze(-1).permute(0,2,1)) | causal_mask\ndec_container_aware_mask = ~(x_tci.unsqueeze(-1) == x_tci.unsqueeze(-1).permute(0,2,1))   & causal_mask\n```\n\n# The Blindfolded Gunslinger\n\nWe made a joke in the [meme thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/208356#1137339) about us wanting to ensemble multiple models, but the competition having only 9 hours to run full inference...\n\nIt was already reported that the public test set was sitting on [the first 20% of the test set](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203383#1115680) so it is possible to maximize the allotted time for the private test set by skipping model inference in the first 20% and predicting only the last 80%. \n\nWe implemented dynamic ensembling that attempts to perform as much ensembling as it is possible in the allotted time for the last 80%.\n\nWe dubbed this idea “The Blindfolded Gunslinger” because it fires two guns (models) as much as it can (after a while it will only fire one) but it is blindfolded in the sense that the public LB will be ~0.5 so we cannot be sure if it worked or not until now…\n\n# Hardware\n\n* 1 computer with Ryzen 3950x (16c32t) + 64 Gb RAM + 1x3090\n* 1 computer with Threadripper 1950x (16c32t) + 256 Gb RAM + 6x3090\n\nWe set up the big computer during the competition which was a project on its own:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3076212%2Ff1186277b6c2045ba9c00053192d6be2%2Fpasted%20image%200.png?generation=1610065136682440&alt=media)\n\nAlso in the last 8 hours of the competition we rented a 190 Gb RAM + 6x3090 but it did not help us much.\n\n# Software\n\nWe used pytorch 1.7.1, fastai and we trained using distributed training and mixed precision (both as implemented in fastai). \n\nFor inference we included the last 500 interactions and summaries in both pickle files and memory-mapped numpy matrices.\n\n# Source code\n\nAvailable at: https://github.com/jamarju/riiid-acp-pub\n",
    "1155737": "The [source code](https://github.com/jamarju/riiid-acp-pub) is up. @imeintanis ",
    "1144532": "Congrats! No doubt why you got your well deserved 3rd position. Keep it up!",
    "1143683": "Wow very impressive indeed!! Congratulations - are you planning to share the code or at least your Blindfolded Guns approach ?",
    "1144472": "Congratulations! TIL the T-fixup trick. I think the right link to the paper is http://www.cs.toronto.edu/~mvolkovs/ICML2020_tfixup.pdf ",
    "1144360": "Thanks for the very interesting read, and congratulations on the 3rd place !",
    "1143990": "Congratulations and amazing sulotion !\nOne quick question: why you choose an encoder-decoder framework? Since the inputs and outputs are always equal length, is sequence-tagging method more suitable?",
    "1143560": "Amazing solution. Congratulations on results @bacterio and team. ",
    "1151076": "Congratulations and thank you for sharing \nthis computer is amazing.",
    "1145755": "That computer is impressive. 😄\nCongratulations on the third place and thanks for sharing your solution. 👌",
    "1144631": "Thanks for sharing and congratulations!\n\nI have a question about the hardware: was this used only to train a model? Or was it used in the submission process somehow (this would confuse me since I thought submission had to be made through the kernel)?",
    "1144564": "Congrats you and your team for the 3rd position!!! Thank you for your detailed solutions, so much to learn",
    "1143896": "Congrats, learn a lot  from your team.",
    "1143754": "Congratulations! Amazing work! 💥🎉",
    "1143588": "Such graphics :D",
    "1143583": "Wow impresionante. Enhorabuena por ese tercer puesto  @bacterio @antorsae !  ¿Cuánto habéis tardado en montar esa bestia de ordenador?",
    "1143572": "Congrats for the 3rd place and thank you for sharing. I was very impressed by the elaborate NN structure. ",
    "1143566": "Wow very nice solution! And insane compute, this comp is very compute hungry!",
    "3424427": "Impressive work dude!",
    "1606342": "Very late congratulations! and thank you very much for the solution write-up. \n\nI have a question and just wondering if you could kindly clarify:\n\n\"To prevent the transformer from looking into the future we shift the encoder input including both questions + answers to the right\"\n\nSince you have already used a causal mask to prevent attention from the future to the past, why is the above operation (shifting the encode input to the right) still needed?\n\nmany thanks!\n\n",
    "1312135": "Great solution!  Tried to submit the 05_inference.ipynb to Kaggle, but the notebook ran out of the memory.  So I ran this notebook and generated submission.csv locally, and copied submission.csv to the directory \"/kaggle/working/\" and submitted to Kaggle again, and the notebook was stopped after a few minutes with error, but without any score.\n\n\nHave anyone tried to submit the 05_inference.ipynb to Kaggle, and successfully got good score?",
    "1217134": "Haha, truly embodying the \"python should be fun\" ethos, nice work :)",
    "1659819": ""
  }
}