{
  "id": 210171,
  "title": "4th place solution : Single Transformer Model",
  "url": "/competitions/riiid-test-answer-prediction/writeups/emmy-4th-place-solution-single-transformer-model",
  "author_name": "",
  "post_date": "2021-10-30T09:36:25.517Z",
  "votes": 124,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Hi everyone,<br>\nIt has been a real pleasure competing in this wonderful challenge.  Thank you Kaggle and the Competition Host for making it possible.</p>\n<p>I'm happy to share here my solution which got me to the 4th place. It is a single transformer model inspired from previous works (like <em>SAINT, SAKT</em>) very much discussed in this competition. </p>\n<h1>The model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F2ec8f0682b82bb47954f2816f221093b%2Friiid-model%20(1).png?generation=1610229212586102&amp;alt=media\" alt=\"Riiid 4th solution model architecture\"><br>\nI hope my figure is straightforward. Below are key features of the model I'd like to explain more:</p>\n<h3>Input sequences</h3>\n<p>I tried to include all data available from the train table and metadata tables. I also add <em>time lag</em>, which is the delta time from the previous interaction (questions in the same container have the same timestamp so they share the same <em>time lag</em>). <br>\nAlso, on the question table, I added 2 features : difficulty level (<em>correct response rate</em> of each questions), popularity (<em>number of appearances</em>), which are computed from the whole training data.</p>\n<h3>Input embeddings</h3>\n<p>Same size of embeddings for all inputs, the embeddings are then concatenated and go through linear transform to feed to the first encoder and decoder of the transformer.</p>\n<p>Embeddings of continuous features (<em>time lag</em>, <em>question elapsed time</em>,  <em>question difficulty</em>, <em>question popularity</em>) are computed using a <em>ContinuousEmbedding</em> layer.  The idea of <em>ContinuousEmbedding</em> is to sum up <code>(</code>weighted sum<code>)</code> a number of consecutive embedding vectors (from the embedding <em>weight matrix</em>). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F13b93a94023bb1ee627b220204f34272%2FContinuousEmbeddings.png?generation=1610279445750633&amp;alt=media\" alt=\"\"><br>\nThis way we have a \"smooth\" version for the embeddings of the continuous variable: 2 values very close together should have similar embeddings. </p>\n<h3>The transformer</h3>\n<p>Input of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements.<br>\nEncoder and decoder layers are almost the same as in the original paper (<a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">Attention is all you need</a>).<br>\nOne key difference is of course the causal masks to prevent the current position from seeing the future. The other is a feature that I add to improve the performance and convergence speed : a kind of time aware weighted attention. The idea is to decay the attention coefficient by a factor of \\(dt^{-w}\\) where \\(dt\\) is the difference in timestamp of a position and the position it attends to and \\(w\\) is a trainable parameter constrained to be non-negative <code>(</code>one parameter per attention head<code>)</code>. This is pretty easy to implement: compute the timestamp difference matrix in log scale, multiply it with the parameters \\(w\\) and subtract it from the attention logits (<em>scaled dot product</em> output of the attention layer).</p>\n<h1>Training</h1>\n<p>I use the cv method <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a> (thanks <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>). The model was implemented in Tensorflow and trained on TPU with Colab Pro.<br>\nSequences are randomly cut and padded to have the same length and all parts are kept for training.<br>\nThe final version of my model has embeddings size of 128, model size of 512, 4 encoder layers and 4 decoder layers. <br>\nIt was trained with the sequence length of 1024 for about 36000 steps (<em>warmup</em> 4000 steps then cosine decay) and with batch size 64. Training took about 4-5 hours.<br>\nOn the submission kernel I had to reduce sequence length to 512 due to resource limit.</p>\n<h1>Some observations</h1>\n<ul>\n<li>Input embeddings: concatenation is better than sum</li>\n<li>Longer sequence <code>(</code>for both training and inference<code>)</code> improves the performance</li>\n<li>Model size also matters: bigger model size generally improves the performance but going beyond size of 512 and 4 layers does not improve much. </li>\n</ul>\n<h1>Why not ensemble</h1>\n<p>I didn't have much time toward the end of the competition. When I still made improvement on my single model I made the choice of staying on that rather than spending time on making ensemble of smaller models. I'm not sure it was a good choice,  but it got me this far so I'm still happy. </p>\n<h1>Code</h1>\n<p><a href=\"https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-place-public-version\" target=\"_blank\">Here</a> is the submission kernel that I made public. You should find all my code source training log in the kernel. <br>\nAs you can see it scores 0.8180 on valid set, 0.815 on public LB and 0.817 on private LB.<br>\nMy top scored submission has the same model and training configuration but was trained on the whole training set, which did not improve much.</p>\n<p><a href=\"https://github.com/dkletran/riiid-challenge-4th-place\" target=\"_blank\">Here</a> is the source code on github including all steps to reproduce the solution.</p>\n<h3>Best regards to all</h3>",
  "messages": [
    {
      "id": "1146645",
      "postDate": "01/09/2021 22:54:54",
      "content": "<p>Hi everyone,<br>\nIt has been a real pleasure competing in this wonderful challenge.  Thank you Kaggle and the Competition Host for making it possible.</p>\n<p>I'm happy to share here my solution which got me to the 4th place. It is a single transformer model inspired from previous works (like <em>SAINT, SAKT</em>) very much discussed in this competition. </p>\n<h1>The model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F2ec8f0682b82bb47954f2816f221093b%2Friiid-model%20(1).png?generation=1610229212586102&amp;alt=media\" alt=\"Riiid 4th solution model architecture\"><br>\nI hope my figure is straightforward. Below are key features of the model I'd like to explain more:</p>\n<h3>Input sequences</h3>\n<p>I tried to include all data available from the train table and metadata tables. I also add <em>time lag</em>, which is the delta time from the previous interaction (questions in the same container have the same timestamp so they share the same <em>time lag</em>). <br>\nAlso, on the question table, I added 2 features : difficulty level (<em>correct response rate</em> of each questions), popularity (<em>number of appearances</em>), which are computed from the whole training data.</p>\n<h3>Input embeddings</h3>\n<p>Same size of embeddings for all inputs, the embeddings are then concatenated and go through linear transform to feed to the first encoder and decoder of the transformer.</p>\n<p>Embeddings of continuous features (<em>time lag</em>, <em>question elapsed time</em>,  <em>question difficulty</em>, <em>question popularity</em>) are computed using a <em>ContinuousEmbedding</em> layer.  The idea of <em>ContinuousEmbedding</em> is to sum up <code>(</code>weighted sum<code>)</code> a number of consecutive embedding vectors (from the embedding <em>weight matrix</em>). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F13b93a94023bb1ee627b220204f34272%2FContinuousEmbeddings.png?generation=1610279445750633&amp;alt=media\" alt=\"\"><br>\nThis way we have a \"smooth\" version for the embeddings of the continuous variable: 2 values very close together should have similar embeddings. </p>\n<h3>The transformer</h3>\n<p>Input of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements.<br>\nEncoder and decoder layers are almost the same as in the original paper (<a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">Attention is all you need</a>).<br>\nOne key difference is of course the causal masks to prevent the current position from seeing the future. The other is a feature that I add to improve the performance and convergence speed : a kind of time aware weighted attention. The idea is to decay the attention coefficient by a factor of \\(dt^{-w}\\) where \\(dt\\) is the difference in timestamp of a position and the position it attends to and \\(w\\) is a trainable parameter constrained to be non-negative <code>(</code>one parameter per attention head<code>)</code>. This is pretty easy to implement: compute the timestamp difference matrix in log scale, multiply it with the parameters \\(w\\) and subtract it from the attention logits (<em>scaled dot product</em> output of the attention layer).</p>\n<h1>Training</h1>\n<p>I use the cv method <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a> (thanks <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>). The model was implemented in Tensorflow and trained on TPU with Colab Pro.<br>\nSequences are randomly cut and padded to have the same length and all parts are kept for training.<br>\nThe final version of my model has embeddings size of 128, model size of 512, 4 encoder layers and 4 decoder layers. <br>\nIt was trained with the sequence length of 1024 for about 36000 steps (<em>warmup</em> 4000 steps then cosine decay) and with batch size 64. Training took about 4-5 hours.<br>\nOn the submission kernel I had to reduce sequence length to 512 due to resource limit.</p>\n<h1>Some observations</h1>\n<ul>\n<li>Input embeddings: concatenation is better than sum</li>\n<li>Longer sequence <code>(</code>for both training and inference<code>)</code> improves the performance</li>\n<li>Model size also matters: bigger model size generally improves the performance but going beyond size of 512 and 4 layers does not improve much. </li>\n</ul>\n<h1>Why not ensemble</h1>\n<p>I didn't have much time toward the end of the competition. When I still made improvement on my single model I made the choice of staying on that rather than spending time on making ensemble of smaller models. I'm not sure it was a good choice,  but it got me this far so I'm still happy. </p>\n<h1>Code</h1>\n<p><a href=\"https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-place-public-version\" target=\"_blank\">Here</a> is the submission kernel that I made public. You should find all my code source training log in the kernel. <br>\nAs you can see it scores 0.8180 on valid set, 0.815 on public LB and 0.817 on private LB.<br>\nMy top scored submission has the same model and training configuration but was trained on the whole training set, which did not improve much.</p>\n<p><a href=\"https://github.com/dkletran/riiid-challenge-4th-place\" target=\"_blank\">Here</a> is the source code on github including all steps to reproduce the solution.</p>\n<h3>Best regards to all</h3>",
      "rawMarkdown": "Hi everyone,\nIt has been a real pleasure competing in this wonderful challenge.  Thank you Kaggle and the Competition Host for making it possible.\n\nI'm happy to share here my solution which got me to the 4th place. It is a single transformer model inspired from previous works (like *SAINT, SAKT*) very much discussed in this competition. \n\n# The model \n![Riiid 4th solution model architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F2ec8f0682b82bb47954f2816f221093b%2Friiid-model%20(1).png?generation=1610229212586102&alt=media)\nI hope my figure is straightforward. Below are key features of the model I'd like to explain more:\n\n### Input sequences\nI tried to include all data available from the train table and metadata tables. I also add *time lag*, which is the delta time from the previous interaction (questions in the same container have the same timestamp so they share the same *time lag*). \nAlso, on the question table, I added 2 features : difficulty level (*correct response rate* of each questions), popularity (*number of appearances*), which are computed from the whole training data.\n### Input embeddings\nSame size of embeddings for all inputs, the embeddings are then concatenated and go through linear transform to feed to the first encoder and decoder of the transformer.\n\nEmbeddings of continuous features (*time lag*, *question elapsed time*,  *question difficulty*, *question popularity*) are computed using a *ContinuousEmbedding* layer.  The idea of *ContinuousEmbedding* is to sum up `(`weighted sum`)` a number of consecutive embedding vectors (from the embedding *weight matrix*). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F13b93a94023bb1ee627b220204f34272%2FContinuousEmbeddings.png?generation=1610279445750633&alt=media)\nThis way we have a \"smooth\" version for the embeddings of the continuous variable: 2 values very close together should have similar embeddings. \n\n### The transformer\nInput of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements.\nEncoder and decoder layers are almost the same as in the original paper ([Attention is all you need](https://arxiv.org/abs/1706.03762)).\nOne key difference is of course the causal masks to prevent the current position from seeing the future. The other is a feature that I add to improve the performance and convergence speed : a kind of time aware weighted attention. The idea is to decay the attention coefficient by a factor of \\\\(dt^{-w}\\\\) where \\\\(dt\\\\) is the difference in timestamp of a position and the position it attends to and \\\\(w\\\\) is a trainable parameter constrained to be non-negative ` (`one parameter per attention head`)`. This is pretty easy to implement: compute the timestamp difference matrix in log scale, multiply it with the parameters \\\\(w\\\\) and subtract it from the attention logits (*scaled dot product* output of the attention layer).\n  \t\n# Training\nI use the cv method https://www.kaggle.com/its7171/cv-strategy (thanks @its7171). The model was implemented in Tensorflow and trained on TPU with Colab Pro.\nSequences are randomly cut and padded to have the same length and all parts are kept for training.\nThe final version of my model has embeddings size of 128, model size of 512, 4 encoder layers and 4 decoder layers. \nIt was trained with the sequence length of 1024 for about 36000 steps (*warmup* 4000 steps then cosine decay) and with batch size 64. Training took about 4-5 hours.\nOn the submission kernel I had to reduce sequence length to 512 due to resource limit.\n\n#Some observations\n- Input embeddings: concatenation is better than sum\n- Longer sequence `(`for both training and inference`)` improves the performance\n- Model size also matters: bigger model size generally improves the performance but going beyond size of 512 and 4 layers does not improve much. \n\n#Why not ensemble\nI didn't have much time toward the end of the competition. When I still made improvement on my single model I made the choice of staying on that rather than spending time on making ensemble of smaller models. I'm not sure it was a good choice,  but it got me this far so I'm still happy. \n\n# Code\n[Here](https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-place-public-version) is the submission kernel that I made public. You should find all my code source training log in the kernel. \nAs you can see it scores 0.8180 on valid set, 0.815 on public LB and 0.817 on private LB.\nMy top scored submission has the same model and training configuration but was trained on the whole training set, which did not improve much.\n\n[Here](https://github.com/dkletran/riiid-challenge-4th-place) is the source code on github including all steps to reproduce the solution.\n\n###Best regards to all",
      "votes": null
    },
    {
      "id": "1146650",
      "postDate": "01/09/2021 23:04:31",
      "content": "<p>Congrats 4th place! I was very surprised that your model is trained using only Colab TPU. <br>\nThe idea of time-weighted multihead attention is great! I tried similar idea, but I gave up it because of too long training time.</p>",
      "rawMarkdown": "Congrats 4th place! I was very surprised that your model is trained using only Colab TPU. \nThe idea of time-weighted multihead attention is great! I tried similar idea, but I gave up it because of too long training time.",
      "votes": null
    },
    {
      "id": "1146687",
      "postDate": "01/10/2021 00:17:06",
      "content": "<p>Thank you for sharing and congrats on your solo gold &amp; 4th place🎉</p>",
      "rawMarkdown": "Thank you for sharing and congrats on your solo gold & 4th place🎉",
      "votes": null
    },
    {
      "id": "1146769",
      "postDate": "01/10/2021 03:15:47",
      "content": "<p>From Viet nam, big congratulations to you on 4th place <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>, amazing solution. Thanks for sharing! </p>",
      "rawMarkdown": "From Viet nam, big congratulations to you on 4th place @letranduckinh, amazing solution. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1146791",
      "postDate": "01/10/2021 03:57:29",
      "content": "<p>What's the difference between a Continuous Embedding and a Linear layer? Ty, and congratulations.</p>",
      "rawMarkdown": "What's the difference between a Continuous Embedding and a Linear layer? Ty, and congratulations.",
      "votes": null
    },
    {
      "id": "1146977",
      "postDate": "01/10/2021 08:13:52",
      "content": "<p>Congratulations on your solution and very clean code.</p>\n<p>Is ContinuousEmbedding similar to <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html</a> where you want to compute the sum or average of multiple embeddings but you are not interested in each one individually? (We used the latter for question tags)</p>",
      "rawMarkdown": "Congratulations on your solution and very clean code.\n\nIs ContinuousEmbedding similar to https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html where you want to compute the sum or average of multiple embeddings but you are not interested in each one individually? (We used the latter for question tags)",
      "votes": null
    },
    {
      "id": "1147007",
      "postDate": "01/10/2021 08:37:24",
      "content": "<p>4th place with a single Transformer, amazing job <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a> and thanks for sharing the code!<br>\nRegarding this sentence: </p>\n<blockquote>\n  <p>Input of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements</p>\n</blockquote>\n<p>Could you give insights on how you made the decision? <br>\nDid it work better for CV to include features in both encoder &amp; decoder?</p>",
      "rawMarkdown": "4th place with a single Transformer, amazing job @letranduckinh and thanks for sharing the code!\nRegarding this sentence: \n> Input of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements\n\nCould you give insights on how you made the decision? \nDid it work better for CV to include features in both encoder & decoder?",
      "votes": null
    },
    {
      "id": "1147057",
      "postDate": "01/10/2021 09:02:07",
      "content": "<p>Congratulations, another single model solution!</p>",
      "rawMarkdown": "Congratulations, another single model solution!",
      "votes": null
    },
    {
      "id": "1147064",
      "postDate": "01/10/2021 09:07:06",
      "content": "<p>This is a very clean code an straight forward solution, thanks for sharing</p>",
      "rawMarkdown": "This is a very clean code an straight forward solution, thanks for sharing",
      "votes": null
    },
    {
      "id": "1147145",
      "postDate": "01/10/2021 10:05:45",
      "content": "<p><a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a> Congratulations! I have a question about your inputs: It seems to me that put the content and lag time into both encoder / decoder, and user answer/ user correctness / had_explanation into the encoder only.</p>\n<p>Could you explain a bit your choice? Usually, the content is put in the encoder, and the properties of the answering questions are put into the decoder. Of course, in your case, it works. But I am still wondering the reason behind your design.</p>\n<p>Maybe you use questions as query, and the historical answer correctness as the key?</p>",
      "rawMarkdown": "letranduckinh Congratulations! I have a question about your inputs: It seems to me that put the content and lag time into both encoder / decoder, and user answer/ user correctness / had_explanation into the encoder only.\n\nCould you explain a bit your choice? Usually, the content is put in the encoder, and the properties of the answering questions are put into the decoder. Of course, in your case, it works. But I am still wondering the reason behind your design.\n\nMaybe you use questions as query, and the historical answer correctness as the key?",
      "votes": null
    },
    {
      "id": "1147174",
      "postDate": "01/10/2021 10:35:58",
      "content": "<p>I'm also looking forward to <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>'s answer. We did it the same way and that was exactly our intuition.</p>\n<p>By putting Q+A in the encoder and Q in the decoder you make the transformer's encoder-decoder attention ask itself the following question:</p>\n<p>\"Hey, I'm a question about grammar (query), let's match other grammar questions in history (key) and let me know <strong>how the user did</strong>, ie. answers (value)\"</p>",
      "rawMarkdown": "I'm also looking forward to @letranduckinh's answer. We did it the same way and that was exactly our intuition.\n\nBy putting Q+A in the encoder and Q in the decoder you make the transformer's encoder-decoder attention ask itself the following question:\n\n\"Hey, I'm a question about grammar (query), let's match other grammar questions in history (key) and let me know **how the user did**, ie. answers (value)\"",
      "votes": null
    },
    {
      "id": "1147186",
      "postDate": "01/10/2021 10:46:28",
      "content": "<p><a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> , Yes, but I am curious why we need the content (questions + lecture if we want) in both side? Isn't the self-attention on the decoder side (here the content) enough to see the <code>other grammar questions in history</code> already? Of course, adding it to both side might get stronger information from the network.</p>\n<p>Have your team tried the usual SAINT(+), where content on encoder and answer info on the decoder? If so, do you know which approach gets better results?</p>",
      "rawMarkdown": "bacterio , Yes, but I am curious why we need the content (questions + lecture if we want) in both side? Isn't the self-attention on the decoder side (here the content) enough to see the `other grammar questions in history ` already? Of course, adding it to both side might get stronger information from the network.\n\nHave your team tried the usual SAINT(+), where content on encoder and answer info on the decoder? If so, do you know which approach gets better results?",
      "votes": null
    },
    {
      "id": "1147189",
      "postDate": "01/10/2021 10:50:14",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>  , why you are surprised <code>training using only Colab TPU.</code> 😄</p>\n<p>It is very powerful, in my case, 1M sequences of length 128 takes 5 minutes. A model is trained will be finished in 4-5 hours.</p>",
      "rawMarkdown": "mamasinkgs  , why you are surprised `training using only Colab TPU. ` 😄\n\nIt is very powerful, in my case, 1M sequences of length 128 takes 5 minutes. A model is trained will be finished in 4-5 hours.",
      "votes": null
    },
    {
      "id": "1147200",
      "postDate": "01/10/2021 10:57:19",
      "content": "<p>Creating a difficult model is easy. Creating an easy model is difficult! Creating an easy model that breaks into top 5 is most difficult. Hearty congratulations! These are the kind of solutions that have strong potential for the hosts..</p>",
      "rawMarkdown": "Creating a difficult model is easy. Creating an easy model is difficult! Creating an easy model that breaks into top 5 is most difficult. Hearty congratulations! These are the kind of solutions that have strong potential for the hosts..",
      "votes": null
    },
    {
      "id": "1147270",
      "postDate": "01/10/2021 11:54:54",
      "content": "<p>Just added a new figure to explain the embeddings of continuous variables.</p>",
      "rawMarkdown": "Just added a new figure to explain the embeddings of continuous variables.",
      "votes": null
    },
    {
      "id": "1147273",
      "postDate": "01/10/2021 11:56:26",
      "content": "<p>I added a schema to explain this, hope it helps.</p>",
      "rawMarkdown": "I added a schema to explain this, hope it helps.",
      "votes": null
    },
    {
      "id": "1147275",
      "postDate": "01/10/2021 11:57:40",
      "content": "<p>I updated my post with with a schema to explain this.</p>",
      "rawMarkdown": "I updated my post with with a schema to explain this.",
      "votes": null
    },
    {
      "id": "1147293",
      "postDate": "01/10/2021 12:23:03",
      "content": "<p>Here is the intuition behind my choice:</p>\n<ul>\n<li>The encoder input value at a position represents the knowledge of the student up to that position. We need all information (content, time lag, his answer). </li>\n<li>The decoder input (query) at a position represents the exercise (question) at that position.  We only have information about the content itself and time lag from the previous exercises. </li>\n<li>To predict the user answer at a position we need to know his \"knowledge\" (the encoder output) and the current question (the decoder query). Self attention and attention helps to look back the  history  in the past.</li>\n</ul>\n<p>Like <a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> said, it is an intuitive choice. We cannot guarantee it is the best choice. I did try some other ways (like in Saint++ put user answer in the decoder query) but the the did not work out well so I stayed with this. </p>",
      "rawMarkdown": "Here is the intuition behind my choice:\n- The encoder input value at a position represents the knowledge of the student up to that position. We need all information (content, time lag, his answer). \n- The decoder input (query) at a position represents the exercise (question) at that position.  We only have information about the content itself and time lag from the previous exercises. \n- To predict the user answer at a position we need to know his \"knowledge\" (the encoder output) and the current question (the decoder query). Self attention and attention helps to look back the  history  in the past.\n\nLike @bacterio said, it is an intuitive choice. We cannot guarantee it is the best choice. I did try some other ways (like in Saint++ put user answer in the decoder query) but the the did not work out well so I stayed with this.",
      "votes": null
    },
    {
      "id": "1147297",
      "postDate": "01/10/2021 12:28:35",
      "content": "<p>Thank you for the reply</p>",
      "rawMarkdown": "Thank you for the reply",
      "votes": null
    },
    {
      "id": "1147591",
      "postDate": "01/10/2021 15:44:26",
      "content": "<p>The amazing solution, I love the way it is relatively straightforward to understand, also a single model and especially, trainable on the Google Colab TPU, which is totally approachable to everyone. Congratulations on your winning!</p>",
      "rawMarkdown": "The amazing solution, I love the way it is relatively straightforward to understand, also a single model and especially, trainable on the Google Colab TPU, which is totally approachable to everyone. Congratulations on your winning!",
      "votes": null
    },
    {
      "id": "1147921",
      "postDate": "01/10/2021 20:01:17",
      "content": "<p>4th with a single model is impressive ! Congratz !</p>",
      "rawMarkdown": "4th with a single model is impressive ! Congratz !",
      "votes": null
    },
    {
      "id": "1148299",
      "postDate": "01/11/2021 04:19:15",
      "content": "<p>Congrats ! and its a great solution.. <br>\nQuestion: How did you sample the dataset? I see that you have a batch size of 64 and only 2048 steps per epoch.. is that just 64*2048 samples per epoch ?</p>",
      "rawMarkdown": "Congrats ! and its a great solution.. \nQuestion: How did you sample the dataset? I see that you have a batch size of 64 and only 2048 steps per epoch.. is that just 64*2048 samples per epoch ?",
      "votes": null
    },
    {
      "id": "1148498",
      "postDate": "01/11/2021 07:19:51",
      "content": "<p>2048 is not the number of steps for a full traversal of the train dataset. Really sorry the naming in my code is a little misleading. Go through the whole dataset needs around 5000-6000 steps so the model was trained for around 6-7 epochs in the good sense of this word.</p>",
      "rawMarkdown": "2048 is not the number of steps for a full traversal of the train dataset. Really sorry the naming in my code is a little misleading. Go through the whole dataset needs around 5000-6000 steps so the model was trained for around 6-7 epochs in the good sense of this word.",
      "votes": null
    },
    {
      "id": "1149511",
      "postDate": "01/11/2021 22:31:19",
      "content": "<p>Thanks for sharing the code, <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>. Congrats on your gold &amp; 4th position. </p>\n<p>How was your experience with TPU in Google Colab Pro, compared to its GPU, and other Cloud platforms?</p>",
      "rawMarkdown": "Thanks for sharing the code, @letranduckinh. Congrats on your gold & 4th position. \n\nHow was your experience with TPU in Google Colab Pro, compared to its GPU, and other Cloud platforms?",
      "votes": null
    },
    {
      "id": "1150445",
      "postDate": "01/12/2021 15:41:00",
      "content": "<p>Thanks.<br>\nI didn't have much experience with TPU, only started using Colab TPU recently for Kaggle competitions because it is cheaper than cloud GPU. </p>",
      "rawMarkdown": "Thanks.\nI didn't have much experience with TPU, only started using Colab TPU recently for Kaggle competitions because it is cheaper than cloud GPU.",
      "votes": null
    },
    {
      "id": "1150956",
      "postDate": "01/13/2021 02:11:10",
      "content": "<p>Nice solution,  i also find increase the sequence length will give big boost, and also the model training time will increase by O(N^2)</p>",
      "rawMarkdown": "Nice solution,  i also find increase the sequence length will give big boost, and also the model training time will increase by O(N^2)",
      "votes": null
    },
    {
      "id": "1150958",
      "postDate": "01/13/2021 02:12:19",
      "content": "<p>In saint paper, continuous is just a linear layer without bias <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> </p>",
      "rawMarkdown": "In saint paper, continuous is just a linear layer without bias @returnofsputnik",
      "votes": null
    },
    {
      "id": "1151305",
      "postDate": "01/13/2021 08:25:31",
      "content": "<p>Ahh got it!<br>\nSo the training set has approx 6000*64 =  380K samples (I'm guessing one per user).<br>\nWhy not train the model on each point in the training data (100 Mi) …. use the lookback from that point backwards as input?</p>",
      "rawMarkdown": "Ahh got it!\nSo the training set has approx 6000*64 =  380K samples (I'm guessing one per user).\nWhy not train the model on each point in the training data (100 Mi) .... use the lookback from that point backwards as input?",
      "votes": null
    },
    {
      "id": "1151480",
      "postDate": "01/13/2021 11:03:12",
      "content": "<p>Basically one sample per user except for users having more than <em>sequence_length</em> (1024) interactions. Training the model on each point in the training data would need a lot more time because in this case the gradient is propagated from only one position of a long sequence. I did try partial label masking as a kind of regularization (compute loss function only for some final positions of the sequence) but the model converges slowly without any improvement on performance.</p>",
      "rawMarkdown": "Basically one sample per user except for users having more than *sequence_length* (1024) interactions. Training the model on each point in the training data would need a lot more time because in this case the gradient is propagated from only one position of a long sequence. I did try partial label masking as a kind of regularization (compute loss function only for some final positions of the sequence) but the model converges slowly without any improvement on performance.",
      "votes": null
    },
    {
      "id": "1156562",
      "postDate": "01/17/2021 08:53:24",
      "content": "<p>Congratulations !<br>\nI have spent some time reading all the codes of your great solution.<br>\nBut I'm confused why there is a first_token_embedding before each sequence for the encoder.<br>\nThanks !</p>",
      "rawMarkdown": "Congratulations !\nI have spent some time reading all the codes of your great solution.\nBut I'm confused why there is a first_token_embedding before each sequence for the encoder.\nThanks !",
      "votes": null
    },
    {
      "id": "1158225",
      "postDate": "01/18/2021 13:12:24",
      "content": "<p>Thanks.<br>\nHere's the deal with <em>first_token_embedding</em>: because of the causal mask of the encoder - decoder attention (a position only attends to future position, excluding positions with the same timestamp), no position (of the encoder output) attends to the first position of the decoder. In the multi-head attention layer, attention masks are achieved by subtracting 1e9 from the attention logits; which doesn't work when no position attends to a position (because in that case, all attentions logits at the target position are subtracted by 1e9; these subtracted terms are canceled out when computing the softmax and as a result the mask has no effect) . To avoid this situation the encoder output is pre-padded with a <em>first_token_embedding</em> which attends to all positions of the decoder.</p>",
      "rawMarkdown": "Thanks.\nHere's the deal with *first_token_embedding*: because of the causal mask of the encoder - decoder attention (a position only attends to future position, excluding positions with the same timestamp), no position (of the encoder output) attends to the first position of the decoder. In the multi-head attention layer, attention masks are achieved by subtracting 1e9 from the attention logits; which doesn't work when no position attends to a position (because in that case, all attentions logits at the target position are subtracted by 1e9; these subtracted terms are canceled out when computing the softmax and as a result the mask has no effect) . To avoid this situation the encoder output is pre-padded with a *first_token_embedding* which attends to all positions of the decoder.",
      "votes": null
    },
    {
      "id": "1160251",
      "postDate": "01/19/2021 18:57:24",
      "content": "<p><a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a> Amongst all the notebooks I've seen in this competition, I think that yours was the most well-organized and readable. Thank you for sharing your solution 🙂</p>\n<p>A question: Do you have a comparison of the importance of each feature used by your model?</p>",
      "rawMarkdown": "letranduckinh Amongst all the notebooks I've seen in this competition, I think that yours was the most well-organized and readable. Thank you for sharing your solution 🙂\n\nA question: Do you have a comparison of the importance of each feature used by your model?",
      "votes": null
    },
    {
      "id": "1160708",
      "postDate": "01/20/2021 04:45:10",
      "content": "<p>Congratulations! Thank you for sharing!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1162778",
      "postDate": "01/21/2021 10:39:05",
      "content": "<p>Thank you for your kind words.<br>\nUnfortunately I don't have a quantitative analysis of feature importance of my model. All I can say is that I already tried to remove some features from the model, it performed worse or sometimes no significant changes on valid set.  In the early stage of the competition, when I added <em>time lag</em>  it gave an important boost, <em>answered correctly, user answer</em> are also important.</p>",
      "rawMarkdown": "Thank you for your kind words.\nUnfortunately I don't have a quantitative analysis of feature importance of my model. All I can say is that I already tried to remove some features from the model, it performed worse or sometimes no significant changes on valid set.  In the early stage of the competition, when I added *time lag*  it gave an important boost, *answered correctly, user answer* are also important.",
      "votes": null
    },
    {
      "id": "1171514",
      "postDate": "01/26/2021 22:32:27",
      "content": "<p>Thanks for the share! Sorry if this is a newb question, but is it possible to make the TF Records dataset public or provide a link to the one you used?</p>",
      "rawMarkdown": "Thanks for the share! Sorry if this is a newb question, but is it possible to make the TF Records dataset public or provide a link to the one you used?",
      "votes": null
    },
    {
      "id": "1179829",
      "postDate": "01/31/2021 21:37:20",
      "content": "<p>I'm also having trouble understanding what <code>data_map_512.pickle</code> is. I'm able to see it being loaded, and we have the file, but I don't know how it's generated.</p>",
      "rawMarkdown": "I'm also having trouble understanding what `data_map_512.pickle` is. I'm able to see it being loaded, and we have the file, but I don't know how it's generated.",
      "votes": null
    },
    {
      "id": "1180721",
      "postDate": "02/01/2021 12:38:46",
      "content": "<p>Amazing, awesome and very clean solution! Thanks for sharing!</p>",
      "rawMarkdown": "Amazing, awesome and very clean solution! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1186381",
      "postDate": "02/04/2021 19:39:05",
      "content": "<p>Here is the github link of my solution: <a href=\"https://github.com/dkletran/riiid-challenge-4th-place\" target=\"_blank\">https://github.com/dkletran/riiid-challenge-4th-place</a>. You should find all scripts and instruction to regenerate all files you need to train &amp; submit the model. Sorry for the late response.</p>",
      "rawMarkdown": "Here is the github link of my solution: https://github.com/dkletran/riiid-challenge-4th-place. You should find all scripts and instruction to regenerate all files you need to train & submit the model. Sorry for the late response.",
      "votes": null
    },
    {
      "id": "1186551",
      "postDate": "02/04/2021 22:16:12",
      "content": "<p>Thank you so much, Duc-Kinh, this is a life-saver.</p>",
      "rawMarkdown": "Thank you so much, Duc-Kinh, this is a life-saver.",
      "votes": null
    },
    {
      "id": "1195644",
      "postDate": "02/11/2021 01:02:05",
      "content": "<p>I noticed that the submission notebook linked to in the GitHub README doesn't work with the weights file generated from the training. The notebook gives a dimension-mismatch error. However, the weights work with your second notebook:</p>\n<p><a href=\"https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-solution\" target=\"_blank\">https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-solution</a></p>\n<p>I've made a \"Late Submission\" with this notebook, and it matches your 0.817 private. Thanks again.</p>",
      "rawMarkdown": "I noticed that the submission notebook linked to in the GitHub README doesn't work with the weights file generated from the training. The notebook gives a dimension-mismatch error. However, the weights work with your second notebook:\n\nhttps://www.kaggle.com/letranduckinh/riiid-model-submission-4th-solution\n\nI've made a \"Late Submission\" with this notebook, and it matches your 0.817 private. Thanks again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1146650,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "01/09/2021 23:04:31",
      "content": "<p>Congrats 4th place! I was very surprised that your model is trained using only Colab TPU. <br>\nThe idea of time-weighted multihead attention is great! I tried similar idea, but I gave up it because of too long training time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147189,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/10/2021 10:50:14",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>  , why you are surprised <code>training using only Colab TPU.</code> 😄</p>\n<p>It is very powerful, in my case, 1M sequences of length 128 takes 5 minutes. A model is trained will be finished in 4-5 hours.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146687,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/10/2021 00:17:06",
      "content": "<p>Thank you for sharing and congrats on your solo gold &amp; 4th place🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146769,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/10/2021 03:15:47",
      "content": "<p>From Viet nam, big congratulations to you on 4th place <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>, amazing solution. Thanks for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146791,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "01/10/2021 03:57:29",
      "content": "<p>What's the difference between a Continuous Embedding and a Linear layer? Ty, and congratulations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147275,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/10/2021 11:57:40",
          "content": "<p>I updated my post with with a schema to explain this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150958,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "01/13/2021 02:12:19",
          "content": "<p>In saint paper, continuous is just a linear layer without bias <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146977,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/10/2021 08:13:52",
      "content": "<p>Congratulations on your solution and very clean code.</p>\n<p>Is ContinuousEmbedding similar to <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html</a> where you want to compute the sum or average of multiple embeddings but you are not interested in each one individually? (We used the latter for question tags)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147273,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/10/2021 11:56:26",
          "content": "<p>I added a schema to explain this, hope it helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147007,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "01/10/2021 08:37:24",
      "content": "<p>4th place with a single Transformer, amazing job <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a> and thanks for sharing the code!<br>\nRegarding this sentence: </p>\n<blockquote>\n  <p>Input of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements</p>\n</blockquote>\n<p>Could you give insights on how you made the decision? <br>\nDid it work better for CV to include features in both encoder &amp; decoder?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147057,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "01/10/2021 09:02:07",
      "content": "<p>Congratulations, another single model solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147064,
      "author_name": "bacterio",
      "author_url": "",
      "post_date": "01/10/2021 09:07:06",
      "content": "<p>This is a very clean code an straight forward solution, thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147145,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/10/2021 10:05:45",
      "content": "<p><a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a> Congratulations! I have a question about your inputs: It seems to me that put the content and lag time into both encoder / decoder, and user answer/ user correctness / had_explanation into the encoder only.</p>\n<p>Could you explain a bit your choice? Usually, the content is put in the encoder, and the properties of the answering questions are put into the decoder. Of course, in your case, it works. But I am still wondering the reason behind your design.</p>\n<p>Maybe you use questions as query, and the historical answer correctness as the key?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147174,
          "author_name": "bacterio",
          "author_url": "",
          "post_date": "01/10/2021 10:35:58",
          "content": "<p>I'm also looking forward to <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>'s answer. We did it the same way and that was exactly our intuition.</p>\n<p>By putting Q+A in the encoder and Q in the decoder you make the transformer's encoder-decoder attention ask itself the following question:</p>\n<p>\"Hey, I'm a question about grammar (query), let's match other grammar questions in history (key) and let me know <strong>how the user did</strong>, ie. answers (value)\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147186,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/10/2021 10:46:28",
          "content": "<p><a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> , Yes, but I am curious why we need the content (questions + lecture if we want) in both side? Isn't the self-attention on the decoder side (here the content) enough to see the <code>other grammar questions in history</code> already? Of course, adding it to both side might get stronger information from the network.</p>\n<p>Have your team tried the usual SAINT(+), where content on encoder and answer info on the decoder? If so, do you know which approach gets better results?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147293,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/10/2021 12:23:03",
          "content": "<p>Here is the intuition behind my choice:</p>\n<ul>\n<li>The encoder input value at a position represents the knowledge of the student up to that position. We need all information (content, time lag, his answer). </li>\n<li>The decoder input (query) at a position represents the exercise (question) at that position.  We only have information about the content itself and time lag from the previous exercises. </li>\n<li>To predict the user answer at a position we need to know his \"knowledge\" (the encoder output) and the current question (the decoder query). Self attention and attention helps to look back the  history  in the past.</li>\n</ul>\n<p>Like <a href=\"https://www.kaggle.com/bacterio\" target=\"_blank\">@bacterio</a> said, it is an intuitive choice. We cannot guarantee it is the best choice. I did try some other ways (like in Saint++ put user answer in the decoder query) but the the did not work out well so I stayed with this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147297,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/10/2021 12:28:35",
          "content": "<p>Thank you for the reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147200,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "01/10/2021 10:57:19",
      "content": "<p>Creating a difficult model is easy. Creating an easy model is difficult! Creating an easy model that breaks into top 5 is most difficult. Hearty congratulations! These are the kind of solutions that have strong potential for the hosts..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147270,
      "author_name": "letranduckinh",
      "author_url": "",
      "post_date": "01/10/2021 11:54:54",
      "content": "<p>Just added a new figure to explain the embeddings of continuous variables.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147591,
      "author_name": "shinomoriaoshi",
      "author_url": "",
      "post_date": "01/10/2021 15:44:26",
      "content": "<p>The amazing solution, I love the way it is relatively straightforward to understand, also a single model and especially, trainable on the Google Colab TPU, which is totally approachable to everyone. Congratulations on your winning!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1147921,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/10/2021 20:01:17",
      "content": "<p>4th with a single model is impressive ! Congratz !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148299,
      "author_name": "rarun2596",
      "author_url": "",
      "post_date": "01/11/2021 04:19:15",
      "content": "<p>Congrats ! and its a great solution.. <br>\nQuestion: How did you sample the dataset? I see that you have a batch size of 64 and only 2048 steps per epoch.. is that just 64*2048 samples per epoch ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148498,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/11/2021 07:19:51",
          "content": "<p>2048 is not the number of steps for a full traversal of the train dataset. Really sorry the naming in my code is a little misleading. Go through the whole dataset needs around 5000-6000 steps so the model was trained for around 6-7 epochs in the good sense of this word.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151305,
          "author_name": "rarun2596",
          "author_url": "",
          "post_date": "01/13/2021 08:25:31",
          "content": "<p>Ahh got it!<br>\nSo the training set has approx 6000*64 =  380K samples (I'm guessing one per user).<br>\nWhy not train the model on each point in the training data (100 Mi) …. use the lookback from that point backwards as input?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151480,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/13/2021 11:03:12",
          "content": "<p>Basically one sample per user except for users having more than <em>sequence_length</em> (1024) interactions. Training the model on each point in the training data would need a lot more time because in this case the gradient is propagated from only one position of a long sequence. I did try partial label masking as a kind of regularization (compute loss function only for some final positions of the sequence) but the model converges slowly without any improvement on performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1149511,
      "author_name": "typhoonasian",
      "author_url": "",
      "post_date": "01/11/2021 22:31:19",
      "content": "<p>Thanks for sharing the code, <a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a>. Congrats on your gold &amp; 4th position. </p>\n<p>How was your experience with TPU in Google Colab Pro, compared to its GPU, and other Cloud platforms?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1150445,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/12/2021 15:41:00",
          "content": "<p>Thanks.<br>\nI didn't have much experience with TPU, only started using Colab TPU recently for Kaggle competitions because it is cheaper than cloud GPU. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1150956,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "01/13/2021 02:11:10",
      "content": "<p>Nice solution,  i also find increase the sequence length will give big boost, and also the model training time will increase by O(N^2)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156562,
      "author_name": "dylanchen",
      "author_url": "",
      "post_date": "01/17/2021 08:53:24",
      "content": "<p>Congratulations !<br>\nI have spent some time reading all the codes of your great solution.<br>\nBut I'm confused why there is a first_token_embedding before each sequence for the encoder.<br>\nThanks !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1158225,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/18/2021 13:12:24",
          "content": "<p>Thanks.<br>\nHere's the deal with <em>first_token_embedding</em>: because of the causal mask of the encoder - decoder attention (a position only attends to future position, excluding positions with the same timestamp), no position (of the encoder output) attends to the first position of the decoder. In the multi-head attention layer, attention masks are achieved by subtracting 1e9 from the attention logits; which doesn't work when no position attends to a position (because in that case, all attentions logits at the target position are subtracted by 1e9; these subtracted terms are canceled out when computing the softmax and as a result the mask has no effect) . To avoid this situation the encoder output is pre-padded with a <em>first_token_embedding</em> which attends to all positions of the decoder.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1160251,
      "author_name": "nyakaggle",
      "author_url": "",
      "post_date": "01/19/2021 18:57:24",
      "content": "<p><a href=\"https://www.kaggle.com/letranduckinh\" target=\"_blank\">@letranduckinh</a> Amongst all the notebooks I've seen in this competition, I think that yours was the most well-organized and readable. Thank you for sharing your solution 🙂</p>\n<p>A question: Do you have a comparison of the importance of each feature used by your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1162778,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "01/21/2021 10:39:05",
          "content": "<p>Thank you for your kind words.<br>\nUnfortunately I don't have a quantitative analysis of feature importance of my model. All I can say is that I already tried to remove some features from the model, it performed worse or sometimes no significant changes on valid set.  In the early stage of the competition, when I added <em>time lag</em>  it gave an important boost, <em>answered correctly, user answer</em> are also important.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1160708,
      "author_name": "junjijiang",
      "author_url": "",
      "post_date": "01/20/2021 04:45:10",
      "content": "<p>Congratulations! Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1171514,
      "author_name": "philipkd",
      "author_url": "",
      "post_date": "01/26/2021 22:32:27",
      "content": "<p>Thanks for the share! Sorry if this is a newb question, but is it possible to make the TF Records dataset public or provide a link to the one you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186381,
          "author_name": "letranduckinh",
          "author_url": "",
          "post_date": "02/04/2021 19:39:05",
          "content": "<p>Here is the github link of my solution: <a href=\"https://github.com/dkletran/riiid-challenge-4th-place\" target=\"_blank\">https://github.com/dkletran/riiid-challenge-4th-place</a>. You should find all scripts and instruction to regenerate all files you need to train &amp; submit the model. Sorry for the late response.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1186551,
          "author_name": "philipkd",
          "author_url": "",
          "post_date": "02/04/2021 22:16:12",
          "content": "<p>Thank you so much, Duc-Kinh, this is a life-saver.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1195644,
          "author_name": "philipkd",
          "author_url": "",
          "post_date": "02/11/2021 01:02:05",
          "content": "<p>I noticed that the submission notebook linked to in the GitHub README doesn't work with the weights file generated from the training. The notebook gives a dimension-mismatch error. However, the weights work with your second notebook:</p>\n<p><a href=\"https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-solution\" target=\"_blank\">https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-solution</a></p>\n<p>I've made a \"Late Submission\" with this notebook, and it matches your 0.817 private. Thanks again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1179829,
      "author_name": "philipkd",
      "author_url": "",
      "post_date": "01/31/2021 21:37:20",
      "content": "<p>I'm also having trouble understanding what <code>data_map_512.pickle</code> is. I'm able to see it being loaded, and we have the file, but I don't know how it's generated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1180721,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "02/01/2021 12:38:46",
      "content": "<p>Amazing, awesome and very clean solution! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1146645": "Hi everyone,\nIt has been a real pleasure competing in this wonderful challenge.  Thank you Kaggle and the Competition Host for making it possible.\n\nI'm happy to share here my solution which got me to the 4th place. It is a single transformer model inspired from previous works (like *SAINT, SAKT*) very much discussed in this competition. \n\n# The model \n![Riiid 4th solution model architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F2ec8f0682b82bb47954f2816f221093b%2Friiid-model%20(1).png?generation=1610229212586102&alt=media)\nI hope my figure is straightforward. Below are key features of the model I'd like to explain more:\n\n### Input sequences\nI tried to include all data available from the train table and metadata tables. I also add *time lag*, which is the delta time from the previous interaction (questions in the same container have the same timestamp so they share the same *time lag*). \nAlso, on the question table, I added 2 features : difficulty level (*correct response rate* of each questions), popularity (*number of appearances*), which are computed from the whole training data.\n### Input embeddings\nSame size of embeddings for all inputs, the embeddings are then concatenated and go through linear transform to feed to the first encoder and decoder of the transformer.\n\nEmbeddings of continuous features (*time lag*, *question elapsed time*,  *question difficulty*, *question popularity*) are computed using a *ContinuousEmbedding* layer.  The idea of *ContinuousEmbedding* is to sum up `(`weighted sum`)` a number of consecutive embedding vectors (from the embedding *weight matrix*). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F37949%2F13b93a94023bb1ee627b220204f34272%2FContinuousEmbeddings.png?generation=1610279445750633&alt=media)\nThis way we have a \"smooth\" version for the embeddings of the continuous variable: 2 values very close together should have similar embeddings. \n\n### The transformer\nInput of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements.\nEncoder and decoder layers are almost the same as in the original paper ([Attention is all you need](https://arxiv.org/abs/1706.03762)).\nOne key difference is of course the causal masks to prevent the current position from seeing the future. The other is a feature that I add to improve the performance and convergence speed : a kind of time aware weighted attention. The idea is to decay the attention coefficient by a factor of \\\\(dt^{-w}\\\\) where \\\\(dt\\\\) is the difference in timestamp of a position and the position it attends to and \\\\(w\\\\) is a trainable parameter constrained to be non-negative ` (`one parameter per attention head`)`. This is pretty easy to implement: compute the timestamp difference matrix in log scale, multiply it with the parameters \\\\(w\\\\) and subtract it from the attention logits (*scaled dot product* output of the attention layer).\n  \t\n# Training\nI use the cv method https://www.kaggle.com/its7171/cv-strategy (thanks @its7171). The model was implemented in Tensorflow and trained on TPU with Colab Pro.\nSequences are randomly cut and padded to have the same length and all parts are kept for training.\nThe final version of my model has embeddings size of 128, model size of 512, 4 encoder layers and 4 decoder layers. \nIt was trained with the sequence length of 1024 for about 36000 steps (*warmup* 4000 steps then cosine decay) and with batch size 64. Training took about 4-5 hours.\nOn the submission kernel I had to reduce sequence length to 512 due to resource limit.\n\n#Some observations\n- Input embeddings: concatenation is better than sum\n- Longer sequence `(`for both training and inference`)` improves the performance\n- Model size also matters: bigger model size generally improves the performance but going beyond size of 512 and 4 layers does not improve much. \n\n#Why not ensemble\nI didn't have much time toward the end of the competition. When I still made improvement on my single model I made the choice of staying on that rather than spending time on making ensemble of smaller models. I'm not sure it was a good choice,  but it got me this far so I'm still happy. \n\n# Code\n[Here](https://www.kaggle.com/letranduckinh/riiid-model-submission-4th-place-public-version) is the submission kernel that I made public. You should find all my code source training log in the kernel. \nAs you can see it scores 0.8180 on valid set, 0.815 on public LB and 0.817 on private LB.\nMy top scored submission has the same model and training configuration but was trained on the whole training set, which did not improve much.\n\n[Here](https://github.com/dkletran/riiid-challenge-4th-place) is the source code on github including all steps to reproduce the solution.\n\n###Best regards to all",
    "1146650": "Congrats 4th place! I was very surprised that your model is trained using only Colab TPU. \nThe idea of time-weighted multihead attention is great! I tried similar idea, but I gave up it because of too long training time.",
    "1146687": "Thank you for sharing and congrats on your solo gold & 4th place🎉",
    "1146769": "From Viet nam, big congratulations to you on 4th place @letranduckinh, amazing solution. Thanks for sharing!",
    "1146791": "What's the difference between a Continuous Embedding and a Linear layer? Ty, and congratulations.",
    "1146977": "Congratulations on your solution and very clean code.\n\nIs ContinuousEmbedding similar to https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html where you want to compute the sum or average of multiple embeddings but you are not interested in each one individually? (We used the latter for question tags)",
    "1147007": "4th place with a single Transformer, amazing job @letranduckinh and thanks for sharing the code!\nRegarding this sentence: \n> Input of the encoder are embeddings of all input elements. Input of the decoder doesn't not contain user answer related elements\n\nCould you give insights on how you made the decision? \nDid it work better for CV to include features in both encoder & decoder?",
    "1147057": "Congratulations, another single model solution!",
    "1147064": "This is a very clean code an straight forward solution, thanks for sharing",
    "1147145": "letranduckinh Congratulations! I have a question about your inputs: It seems to me that put the content and lag time into both encoder / decoder, and user answer/ user correctness / had_explanation into the encoder only.\n\nCould you explain a bit your choice? Usually, the content is put in the encoder, and the properties of the answering questions are put into the decoder. Of course, in your case, it works. But I am still wondering the reason behind your design.\n\nMaybe you use questions as query, and the historical answer correctness as the key?",
    "1147174": "I'm also looking forward to @letranduckinh's answer. We did it the same way and that was exactly our intuition.\n\nBy putting Q+A in the encoder and Q in the decoder you make the transformer's encoder-decoder attention ask itself the following question:\n\n\"Hey, I'm a question about grammar (query), let's match other grammar questions in history (key) and let me know **how the user did**, ie. answers (value)\"",
    "1147186": "bacterio , Yes, but I am curious why we need the content (questions + lecture if we want) in both side? Isn't the self-attention on the decoder side (here the content) enough to see the `other grammar questions in history ` already? Of course, adding it to both side might get stronger information from the network.\n\nHave your team tried the usual SAINT(+), where content on encoder and answer info on the decoder? If so, do you know which approach gets better results?",
    "1147189": "mamasinkgs  , why you are surprised `training using only Colab TPU. ` 😄\n\nIt is very powerful, in my case, 1M sequences of length 128 takes 5 minutes. A model is trained will be finished in 4-5 hours.",
    "1147200": "Creating a difficult model is easy. Creating an easy model is difficult! Creating an easy model that breaks into top 5 is most difficult. Hearty congratulations! These are the kind of solutions that have strong potential for the hosts..",
    "1147270": "Just added a new figure to explain the embeddings of continuous variables.",
    "1147273": "I added a schema to explain this, hope it helps.",
    "1147275": "I updated my post with with a schema to explain this.",
    "1147293": "Here is the intuition behind my choice:\n- The encoder input value at a position represents the knowledge of the student up to that position. We need all information (content, time lag, his answer). \n- The decoder input (query) at a position represents the exercise (question) at that position.  We only have information about the content itself and time lag from the previous exercises. \n- To predict the user answer at a position we need to know his \"knowledge\" (the encoder output) and the current question (the decoder query). Self attention and attention helps to look back the  history  in the past.\n\nLike @bacterio said, it is an intuitive choice. We cannot guarantee it is the best choice. I did try some other ways (like in Saint++ put user answer in the decoder query) but the the did not work out well so I stayed with this.",
    "1147297": "Thank you for the reply",
    "1147591": "The amazing solution, I love the way it is relatively straightforward to understand, also a single model and especially, trainable on the Google Colab TPU, which is totally approachable to everyone. Congratulations on your winning!",
    "1147921": "4th with a single model is impressive ! Congratz !",
    "1148299": "Congrats ! and its a great solution.. \nQuestion: How did you sample the dataset? I see that you have a batch size of 64 and only 2048 steps per epoch.. is that just 64*2048 samples per epoch ?",
    "1148498": "2048 is not the number of steps for a full traversal of the train dataset. Really sorry the naming in my code is a little misleading. Go through the whole dataset needs around 5000-6000 steps so the model was trained for around 6-7 epochs in the good sense of this word.",
    "1149511": "Thanks for sharing the code, @letranduckinh. Congrats on your gold & 4th position. \n\nHow was your experience with TPU in Google Colab Pro, compared to its GPU, and other Cloud platforms?",
    "1150445": "Thanks.\nI didn't have much experience with TPU, only started using Colab TPU recently for Kaggle competitions because it is cheaper than cloud GPU.",
    "1150956": "Nice solution,  i also find increase the sequence length will give big boost, and also the model training time will increase by O(N^2)",
    "1150958": "In saint paper, continuous is just a linear layer without bias @returnofsputnik",
    "1151305": "Ahh got it!\nSo the training set has approx 6000*64 =  380K samples (I'm guessing one per user).\nWhy not train the model on each point in the training data (100 Mi) .... use the lookback from that point backwards as input?",
    "1151480": "Basically one sample per user except for users having more than *sequence_length* (1024) interactions. Training the model on each point in the training data would need a lot more time because in this case the gradient is propagated from only one position of a long sequence. I did try partial label masking as a kind of regularization (compute loss function only for some final positions of the sequence) but the model converges slowly without any improvement on performance.",
    "1156562": "Congratulations !\nI have spent some time reading all the codes of your great solution.\nBut I'm confused why there is a first_token_embedding before each sequence for the encoder.\nThanks !",
    "1158225": "Thanks.\nHere's the deal with *first_token_embedding*: because of the causal mask of the encoder - decoder attention (a position only attends to future position, excluding positions with the same timestamp), no position (of the encoder output) attends to the first position of the decoder. In the multi-head attention layer, attention masks are achieved by subtracting 1e9 from the attention logits; which doesn't work when no position attends to a position (because in that case, all attentions logits at the target position are subtracted by 1e9; these subtracted terms are canceled out when computing the softmax and as a result the mask has no effect) . To avoid this situation the encoder output is pre-padded with a *first_token_embedding* which attends to all positions of the decoder.",
    "1160251": "letranduckinh Amongst all the notebooks I've seen in this competition, I think that yours was the most well-organized and readable. Thank you for sharing your solution 🙂\n\nA question: Do you have a comparison of the importance of each feature used by your model?",
    "1160708": "Congratulations! Thank you for sharing!",
    "1162778": "Thank you for your kind words.\nUnfortunately I don't have a quantitative analysis of feature importance of my model. All I can say is that I already tried to remove some features from the model, it performed worse or sometimes no significant changes on valid set.  In the early stage of the competition, when I added *time lag*  it gave an important boost, *answered correctly, user answer* are also important.",
    "1171514": "Thanks for the share! Sorry if this is a newb question, but is it possible to make the TF Records dataset public or provide a link to the one you used?",
    "1179829": "I'm also having trouble understanding what `data_map_512.pickle` is. I'm able to see it being loaded, and we have the file, but I don't know how it's generated.",
    "1180721": "Amazing, awesome and very clean solution! Thanks for sharing!",
    "1186381": "Here is the github link of my solution: https://github.com/dkletran/riiid-challenge-4th-place. You should find all scripts and instruction to regenerate all files you need to train & submit the model. Sorry for the late response.",
    "1186551": "Thank you so much, Duc-Kinh, this is a life-saver.",
    "1195644": "I noticed that the submission notebook linked to in the GitHub README doesn't work with the weights file generated from the training. The notebook gives a dimension-mismatch error. However, the weights work with your second notebook:\n\nhttps://www.kaggle.com/letranduckinh/riiid-model-submission-4th-solution\n\nI've made a \"Late Submission\" with this notebook, and it matches your 0.817 private. Thanks again."
  },
  "source": "meta"
}