{
  "id": 218318,
  "title": "[1st place solution] Last Query Transformer RNN",
  "url": "/competitions/riiid-test-answer-prediction/discussion/218318",
  "author_name": "keetar",
  "post_date": "2021-02-10T05:40:56.623000",
  "votes": 170,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hi Kagglers, here is the paper link.<br>\nPaper Link : <a href=\"https://arxiv.org/abs/2102.05038\" target=\"_blank\">https://arxiv.org/abs/2102.05038</a><br>\nPlease refer the paper for details.</p>\n<h1>Summary</h1>\n<p>I wanted to use Transformer but I could not input long history to the model because QK matrix multiplication in Transformer has O(L^2) time complexity when input length is L. </p>\n<h1> </h1>\n<p>My approach is to use only last input as Query, because I only predict last question's answer correctness per history inputs. It means I will only compare between last question(query) and other questions(key), and not between other questions. It makes QK matrix multiplication in Transformer to have O(L) time complexity(because len(Q)=1, len(K)=L), which allows me to input much longer history. </p>\n<h1> </h1>\n<p>In final submission, I ensemble 5 models with 1728 length history inputs.<br>\nI didn't do feature engineering much, since I can use extremely long history, I wanted the model to learn it by itself. 5 input features I used are question id, question part, answer correctness, current question elapsed time, and timestamp difference.</p>\n<h1>Acknowledgement</h1>\n<p>Thank you <a href=\"https://www.kaggle.com/limerobot\" target=\"_blank\">@limerobot</a>, to share amazing transformer model to approach table data problem in the 3rd place solution of 2019 Data Science Bowl. It was a good motivation to work on transformer encoder. <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891</a> </p>\n<h1> </h1>\n<p>Thanks to competition sponsers and organizers for hosting this fun competition.</p>\n<h1> </h1>\n<p>Thank you for reading.</p>",
  "messages": [
    {
      "id": 1194242,
      "postDate": "2021-02-10T05:40:56.623Z",
      "content": "<p>Hi Kagglers, here is the paper link.<br>\nPaper Link : <a href=\"https://arxiv.org/abs/2102.05038\" target=\"_blank\">https://arxiv.org/abs/2102.05038</a><br>\nPlease refer the paper for details.</p>\n<h1>Summary</h1>\n<p>I wanted to use Transformer but I could not input long history to the model because QK matrix multiplication in Transformer has O(L^2) time complexity when input length is L. </p>\n<h1> </h1>\n<p>My approach is to use only last input as Query, because I only predict last question's answer correctness per history inputs. It means I will only compare between last question(query) and other questions(key), and not between other questions. It makes QK matrix multiplication in Transformer to have O(L) time complexity(because len(Q)=1, len(K)=L), which allows me to input much longer history. </p>\n<h1> </h1>\n<p>In final submission, I ensemble 5 models with 1728 length history inputs.<br>\nI didn't do feature engineering much, since I can use extremely long history, I wanted the model to learn it by itself. 5 input features I used are question id, question part, answer correctness, current question elapsed time, and timestamp difference.</p>\n<h1>Acknowledgement</h1>\n<p>Thank you <a href=\"https://www.kaggle.com/limerobot\" target=\"_blank\">@limerobot</a>, to share amazing transformer model to approach table data problem in the 3rd place solution of 2019 Data Science Bowl. It was a good motivation to work on transformer encoder. <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891</a> </p>\n<h1> </h1>\n<p>Thanks to competition sponsers and organizers for hosting this fun competition.</p>\n<h1> </h1>\n<p>Thank you for reading.</p>",
      "rawMarkdown": " Hi Kagglers, here is the paper link.\nPaper Link : https://arxiv.org/abs/2102.05038\nPlease refer the paper for details.\n# Summary\nI wanted to use Transformer but I could not input long history to the model because QK matrix multiplication in Transformer has O(L^2) time complexity when input length is L. \n# \nMy approach is to use only last input as Query, because I only predict last question's answer correctness per history inputs. It means I will only compare between last question(query) and other questions(key), and not between other questions. It makes QK matrix multiplication in Transformer to have O(L) time complexity(because len(Q)=1, len(K)=L), which allows me to input much longer history. \n# \nIn final submission, I ensemble 5 models with 1728 length history inputs.\nI didn't do feature engineering much, since I can use extremely long history, I wanted the model to learn it by itself. 5 input features I used are question id, question part, answer correctness, current question elapsed time, and timestamp difference.\n# Acknowledgement\nThank you @limerobot, to share amazing transformer model to approach table data problem in the 3rd place solution of 2019 Data Science Bowl. It was a good motivation to work on transformer encoder. https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891 \n# \nThanks to competition sponsers and organizers for hosting this fun competition.\n\n# \nThank you for reading.",
      "votes": 170
    },
    {
      "id": 1195053,
      "postDate": "2021-02-10T14:28:02.963Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> for sharing this approach. Very very nice, and great to see such long histories used. If you have published code somewhere, it would be great to see the encoder layer set up - if it is not public, no worries. Well done again, well deserved.  </p>",
      "rawMarkdown": "Thank you @keetar for sharing this approach. Very very nice, and great to see such long histories used. If you have published code somewhere, it would be great to see the encoder layer set up - if it is not public, no worries. Well done again, well deserved.  ",
      "votes": 3,
      "replies": [
        {
          "id": 1195096,
          "postDate": "2021-02-10T14:56:00.353Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>, thank you. <br>\nActually, it's really simple to implement the model.<br>\nI only changed the one line from normal transformer encoder.<br>\nIn encoder layer,<br>\nI changed<br>\n     <code>attn_output, _ = self.mha(x, x, x[:,:,:], mask) #mha : multi head attention</code><br>\nto<br>\n     <code>attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)</code><br>\nHope it helps!</p>",
          "rawMarkdown": "Hi @darraghdog, thank you. \nActually, it's really simple to implement the model.\nI only changed the one line from normal transformer encoder.\nIn encoder layer,\nI changed\n     `attn_output, _ = self.mha(x, x, x[:,:,:], mask) #mha : multi head attention`\nto\n     `attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)`\nHope it helps!",
          "votes": 9
        }
      ]
    },
    {
      "id": 1200641,
      "postDate": "2021-02-14T20:14:52.433Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> Congrats. <br>\nI had a couple of doubts, can you please answer them-</p>\n<p>Assume that I represent Query, Key and Value as Q, K, and V respectively, and let's assume that I have an input sequence with-  <br>\n<strong>D</strong>- Dimension of model  <br>\n<strong>S</strong> - Sequence Length of input  </p>\n<p>According to the <a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">paper</a>  first we need to multiply the Q and K transpose so -  <br>\n<strong>result1 = Q * K    ---&gt;      (S , D) * ( D ,S)  ---&gt; (S,S)</strong>  dimension  </p>\n<p>Now, we divide by sq root of a scaling factor and then apply softmax and get result2 with dim (S,S).  <br>\n<strong>result2 = softmax(  result1  /  sqRoot(ScalingFactor) )</strong>  </p>\n<p>Now we multiply the result2 with our Values V--&gt;  <br>\n<strong>FinalResult  =   result2 *  V       ---&gt;  i.e.  (S,S) * (S,D) ---&gt; (S,D)</strong>  dimension</p>\n<p>My doubts-  <br>\n<strong>1. If you keep seq len = 1  in the Query Q, then the final output will be having dim (1, D). \nThis is just a single feature with dim D, then how are you passing it through the LSTM? \n( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )</strong>    </p>\n<p><strong>2. In one of the comments you mentioned the following -</strong><br>\n<code>attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)</code>  <br>\n<strong>Can you please elaborate more on the code, What is x and what is x[:,-1:,:] ?\nWhich one of those 3 is Query Q ?\n( If the code above is used from <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html\" target=\"_blank\">torch.nn.MultiheadAttention</a> then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )</strong>    </p>\n<p>Thanks</p>",
      "rawMarkdown": "Hey @keetar Congrats. \nI had a couple of doubts, can you please answer them-\n\nAssume that I represent Query, Key and Value as Q, K, and V respectively, and let's assume that I have an input sequence with-  \n**D**- Dimension of model  \n**S** - Sequence Length of input  \n\nAccording to the [paper](https://arxiv.org/abs/1706.03762)  first we need to multiply the Q and K transpose so -  \n**result1 = Q * K    --->      (S , D) * ( D ,S)  ---> (S,S)**  dimension  \n\nNow, we divide by sq root of a scaling factor and then apply softmax and get result2 with dim (S,S).  \n**result2 = softmax(  result1  /  sqRoot(ScalingFactor) )**  \n\nNow we multiply the result2 with our Values V-->  \n**FinalResult  =   result2 *  V       --->  i.e.  (S,S) * (S,D) ---> (S,D)**  dimension\n\nMy doubts-  \n**1. If you keep seq len = 1  in the Query Q, then the final output will be having dim (1, D). \nThis is just a single feature with dim D, then how are you passing it through the LSTM? \n( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )**    \n\n**2. In one of the comments you mentioned the following -**\n`attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)`  \n**Can you please elaborate more on the code, What is x and what is x[:,-1:,:] ?\nWhich one of those 3 is Query Q ?\n( If the code above is used from [torch.nn.MultiheadAttention](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html) then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )**    \n\nThanks\n\n \n \n\n",
      "votes": 4,
      "replies": [
        {
          "id": 1201657,
          "postDate": "2021-02-15T14:51:42.267Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1201901,
          "postDate": "2021-02-15T18:19:58.613Z",
          "content": "<p>Based on my understanding of transformers and what <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> said, it looks like <code>x[:,-1:,:]</code> is the query. I don't think he uses torch's MHA. I've seen other MHAs where the prototype is KVQ instead of QKV.</p>\n<p>In one of the <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">reference implementations</a> for Transformers, <code>x</code> represents Q and has a shape of <code>(batch_size, sequence_len, embedding_dim_size)</code>. By tweaking <code>x</code> to <code>x[:,-1,:]</code> you change the <code>sequence_length</code> to 1.</p>",
          "rawMarkdown": "Based on my understanding of transformers and what @keetar said, it looks like `x[:,-1:,:]` is the query. I don't think he uses torch's MHA. I've seen other MHAs where the prototype is KVQ instead of QKV.\n\nIn one of the [reference implementations](https://www.tensorflow.org/tutorials/text/transformer) for Transformers, `x` represents Q and has a shape of `(batch_size, sequence_len, embedding_dim_size)`. By tweaking `x` to `x[:,-1,:]` you change the `sequence_length` to 1.",
          "votes": 2
        },
        {
          "id": 1201917,
          "postDate": "2021-02-15T18:36:17.650Z",
          "content": "<p>Same doubts, especial the first doubt</p>",
          "rawMarkdown": "Same doubts, especial the first doubt"
        },
        {
          "id": 1202115,
          "postDate": "2021-02-15T21:10:03.073Z",
          "content": "<p>Hi,</p>\n<pre><code>1. If you keep seq len = 1 in the Query Q, then the final output will be having dim (1, D). This is just a single feature with dim D, then how are you passing it through the LSTM? ( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )\n</code></pre>\n<h1> </h1>\n<p>In transformer encoder, there's a residual addition, which adds 'multi head attention output' to original input. In multi head attention, I use Query with len=1 to get some speed advantages in QK matrix multiplication. In result, the output of multi head attention becomes (batch size, 1, d_model) instead of (batch size, sequence length, d_model). So, for residual addition in transformer encoder, operands are original input with shape (batch size,  sequence length, d_model) and multi head attention output with shape (batch size, 1, d_model). Because two shapes are different, I add multi head attention output to every point of original input.(like broadcasting) Therefore, transformer encoder output is (batch size, sequence length, d_model).</p>\n<pre><code>2. In one of the comments you mentioned the following -\nattn_output, _ = self.mha(x, x, x[:,-1:,:], mask)\nCan you please elaborate more on the code, What is x and what is x[:,-1:,:] ? Which one of those 3 is Query Q ? ( If the code above is used from torch.nn.MultiheadAttention then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )\n</code></pre>\n<p>x is an original input to encoder. x's shape is (batch size, seq_len, d_model). The order of mha(multihead attention) parameters is v,k,q,mask.  </p>",
          "rawMarkdown": "Hi,\n```\n1. If you keep seq len = 1 in the Query Q, then the final output will be having dim (1, D). This is just a single feature with dim D, then how are you passing it through the LSTM? ( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )\n```\n# \nIn transformer encoder, there's a residual addition, which adds 'multi head attention output' to original input. In multi head attention, I use Query with len=1 to get some speed advantages in QK matrix multiplication. In result, the output of multi head attention becomes (batch size, 1, d_model) instead of (batch size, sequence length, d_model). So, for residual addition in transformer encoder, operands are original input with shape (batch size,  sequence length, d_model) and multi head attention output with shape (batch size, 1, d_model). Because two shapes are different, I add multi head attention output to every point of original input.(like broadcasting) Therefore, transformer encoder output is (batch size, sequence length, d_model).\n\n```\n2. In one of the comments you mentioned the following -\nattn_output, _ = self.mha(x, x, x[:,-1:,:], mask)\nCan you please elaborate more on the code, What is x and what is x[:,-1:,:] ? Which one of those 3 is Query Q ? ( If the code above is used from torch.nn.MultiheadAttention then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )\n```\nx is an original input to encoder. x's shape is (batch size, seq_len, d_model). The order of mha(multihead attention) parameters is v,k,q,mask.  ",
          "votes": 12
        },
        {
          "id": 1204329,
          "postDate": "2021-02-16T04:16:51.840Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a>  for the reply. Once again congrats, its a great solution.</p>",
          "rawMarkdown": "Thanks @keetar  for the reply. Once again congrats, its a great solution.",
          "votes": 1
        },
        {
          "id": 1204407,
          "postDate": "2021-02-16T05:53:12.453Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)",
          "votes": 1
        },
        {
          "id": 1332599,
          "postDate": "2021-06-02T07:52:51.703Z",
          "content": "<p>Thanks for sharing this amazing solution! One question I wanna ask is you add 'last query embedding' to the original input(which contains the last query), which seems like adding a 'scaler' to all sequences, does this make sense?</p>",
          "rawMarkdown": "Thanks for sharing this amazing solution! One question I wanna ask is you add 'last query embedding' to the original input(which contains the last query), which seems like adding a 'scaler' to all sequences, does this make sense?"
        }
      ]
    },
    {
      "id": 1485709,
      "postDate": "2021-08-22T10:56:56.997Z",
      "content": "<p>Pytorch Implementation of 1st place solution - <a href=\"https://github.com/arshadshk/Last_Query_Transformer_RNN-PyTorch\" target=\"_blank\">https://github.com/arshadshk/Last_Query_Transformer_RNN-PyTorch</a> <br>\n( Let me know if any changes are needed or any bugs present )</p>",
      "rawMarkdown": "Pytorch Implementation of 1st place solution - https://github.com/arshadshk/Last_Query_Transformer_RNN-PyTorch \n( Let me know if any changes are needed or any bugs present )",
      "votes": 1
    },
    {
      "id": 1195536,
      "postDate": "2021-02-10T22:09:36.580Z",
      "content": "<p>Thank you keetar for your great solution, it was really fun to compete with you. Finally I lost, though.<br>\nIt is not surprising to use only the last input as query, because I heard <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>, <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> and <a href=\"https://www.kaggle.com/nadare\" target=\"_blank\">@nadare</a> did the same thing and I suppose many people did this, as it's easy to implement.<br>\nHowever, they said using the last input as query is very time-consuming (I remember <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> said training takes 1 week). How long is your training time, why is it so fast and what is your trick? Did you use TPU? </p>\n<p>My another question is about the sequence length and the input features. I was very surprised that your sequence length is 1728 and you used only 5 features. How did you notice sequence length is so important? Did you throw away other useful features to make the sequence length longer? I would like to know your journey to L = 1728 :)</p>",
      "rawMarkdown": "Thank you keetar for your great solution, it was really fun to compete with you. Finally I lost, though.\nIt is not surprising to use only the last input as query, because I heard @its7171, @takoihiraokazu and @nadare did the same thing and I suppose many people did this, as it's easy to implement.\nHowever, they said using the last input as query is very time-consuming (I remember @its7171 said training takes 1 week). How long is your training time, why is it so fast and what is your trick? Did you use TPU? \n\nMy another question is about the sequence length and the input features. I was very surprised that your sequence length is 1728 and you used only 5 features. How did you notice sequence length is so important? Did you throw away other useful features to make the sequence length longer? I would like to know your journey to L = 1728 :)",
      "votes": 2,
      "replies": [
        {
          "id": 1195672,
          "postDate": "2021-02-11T01:51:18.893Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, thank you.<br>\nI'm surprised to hear many people tried this, as I saw none. Could you provide some links?</p>",
          "rawMarkdown": "Hi @mamasinkgs, thank you.\nI'm surprised to hear many people tried this, as I saw none. Could you provide some links?",
          "votes": 2
        },
        {
          "id": 1195685,
          "postDate": "2021-02-11T02:20:35.783Z",
          "content": "<p>Yes, it's not clearly written in many people's solution, but I believe this was done by many people. I noticed my japanese friends did this in a meeting after the competition ends. For example, please look at the takoi's figure of <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209622\" target=\"_blank\">their solution</a>. it uses only the last query for training. tito also said <code>only trained and predicted for answered_correctly of last question of the sequence.</code> in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/210354\" target=\"_blank\">their solution</a>, which is based on the <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">public kernel</a>. Moreover, I didn't use the last query technique in training phase (instead I used masking), but I used last query technique in the inference phase, as can be seen in <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">my kernel</a>. The corresponding part of my code is here.</p>\n<pre><code>            #key and value process 2\n            query_clone = query.clone()\n            query_clone[token_idx, -2] = 0\n            memory_cat = torch.cat([query_clone, ohe_explanation, ohe_correctness, \n                                    normed_elapsed.unsqueeze(2), ohe_user_answer], dim = 2)\n            memory = self.norm4(self.linear8(self.dropout(F.relu(self.linear7(memory_cat)))))\n            memory = memory[:, :-1]\n\n            #transpose\n            query = query.transpose(0, 1)\n            memory = memory.transpose(0, 1)\n            spc_query = query[-1:]\n\n            #MyEncoder\n            out = self.MyEncoderLayer1(spc_query, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer2(out, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer3(out, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer4(out, memory, memory, None, padding_mask).squeeze(0)\n</code></pre>",
          "rawMarkdown": "Yes, it's not clearly written in many people's solution, but I believe this was done by many people. I noticed my japanese friends did this in a meeting after the competition ends. For example, please look at the takoi's figure of [their solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209622). it uses only the last query for training. tito also said `only trained and predicted for answered_correctly of last question of the sequence.` in [their solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/210354), which is based on the [public kernel](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public). Moreover, I didn't use the last query technique in training phase (instead I used masking), but I used last query technique in the inference phase, as can be seen in [my kernel](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold). The corresponding part of my code is here.\n\n```\n            #key and value process 2\n            query_clone = query.clone()\n            query_clone[token_idx, -2] = 0\n            memory_cat = torch.cat([query_clone, ohe_explanation, ohe_correctness, \n                                    normed_elapsed.unsqueeze(2), ohe_user_answer], dim = 2)\n            memory = self.norm4(self.linear8(self.dropout(F.relu(self.linear7(memory_cat)))))\n            memory = memory[:, :-1]\n\n            #transpose\n            query = query.transpose(0, 1)\n            memory = memory.transpose(0, 1)\n            spc_query = query[-1:]\n\n            #MyEncoder\n            out = self.MyEncoderLayer1(spc_query, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer2(out, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer3(out, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer4(out, memory, memory, None, padding_mask).squeeze(0)\n```\n\n",
          "votes": 1
        },
        {
          "id": 1195704,
          "postDate": "2021-02-11T02:49:42.963Z",
          "content": "<p>I think they're totally different.<br>\nNormally, even if you train and predict for only last input, you don't change query vector in transformer to have length 1. In usual case, all query and value and key are set to have length L.(when input length is L). Because that's how transformer works. I believe the cases you mentioned have the same configurations. <br>\nChanging query vector in transformer to have length 1 changes the inner workings of transformer significantly. When multi head attention output is added to each inputs, last query's attention output is added to each inputs like constant, instead of each query's attention output. I found that it does not hurt the performance much, while making QK matrix multiplication in transformer to have O(L) instead of O(L^2).<br>\nTo answer why my model is faster than others who also used last query for training, it's simply because my QK matrix multiplication is O(L) and others is O(L^2).</p>",
          "rawMarkdown": "I think they're totally different.\nNormally, even if you train and predict for only last input, you don't change query vector in transformer to have length 1. In usual case, all query and value and key are set to have length L.(when input length is L). Because that's how transformer works. I believe the cases you mentioned have the same configurations. \nChanging query vector in transformer to have length 1 changes the inner workings of transformer significantly. When multi head attention output is added to each inputs, last query's attention output is added to each inputs like constant, instead of each query's attention output. I found that it does not hurt the performance much, while making QK matrix multiplication in transformer to have O(L) instead of O(L^2).\nTo answer why my model is faster than others who also used last query for training, it's simply because my QK matrix multiplication is O(L) and others is O(L^2).",
          "votes": 4
        },
        {
          "id": 1195741,
          "postDate": "2021-02-11T03:29:10.940Z",
          "content": "<p>Yes, as you say, in usual all query and key have length L and thus the matrix multiplication is O(L^2). But I think my inference code above is O(L), because the size of the query is 1 as can be seen from             <code>spc_query = query[-1:]</code> for the input of my custom transformer, so I just guessed other competitors who used the last input for training did the same thing like my inference phase (i.e. O(L) matrix multiplication). But it may be wrong, as I don't perfectly understand others' solution. Anyway, you are the kaggler I respect the most and your solution is really great without any doubt :) Thank you again!</p>",
          "rawMarkdown": "Yes, as you say, in usual all query and key have length L and thus the matrix multiplication is O(L^2). But I think my inference code above is O(L), because the size of the query is 1 as can be seen from             `spc_query = query[-1:]` for the input of my custom transformer, so I just guessed other competitors who used the last input for training did the same thing like my inference phase (i.e. O(L) matrix multiplication). But it may be wrong, as I don't perfectly understand others' solution. Anyway, you are the kaggler I respect the most and your solution is really great without any doubt :) Thank you again!",
          "votes": 1
        },
        {
          "id": 1195759,
          "postDate": "2021-02-11T04:00:23.103Z",
          "content": "<p>Thank you :) And congratulations on your successful finish.</p>",
          "rawMarkdown": "Thank you :) And congratulations on your successful finish.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1233189,
      "postDate": "2021-03-10T08:19:17.753Z",
      "content": "<p>Hi, I had some doubts about your inputs, can you please anwer it.</p>\n<ol>\n<li><p><strong>Categorical embedding is used for first three features, and continuous embedding is used for last two continuous features.</strong><br>\nbut how you make continuous embedding?</p></li>\n<li><p><strong>'Timestamp difference' feature indicates the difference from the past question timestamp to the current question timestamp, and it is clipped by maximum value, 3 day.</strong><br>\nhow you fill the first 'Timestamp difference' in a sequence?</p></li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi, I had some doubts about your inputs, can you please anwer it.\n\n1.  **Categorical embedding is used for first three features, and continuous embedding is used for last two continuous features.**\nbut how you make continuous embedding?\n\n2.  **'Timestamp difference' feature indicates the difference from the past question timestamp to the current question timestamp, and it is clipped by maximum value, 3 day.**\nhow you fill the first 'Timestamp difference' in a sequence?\n\nThanks!\n\n",
      "votes": 1
    },
    {
      "id": 1195385,
      "postDate": "2021-02-10T18:51:53.747Z",
      "content": "<p>Cong! What a simple but sweet solution:)</p>",
      "rawMarkdown": "Cong! What a simple but sweet solution:)",
      "votes": 1
    },
    {
      "id": 1194577,
      "postDate": "2021-02-10T08:53:02.880Z",
      "content": "<p>Congrats for this really elegant and original solution !</p>",
      "rawMarkdown": "Congrats for this really elegant and original solution !",
      "votes": 1,
      "replies": [
        {
          "id": 1195111,
          "postDate": "2021-02-10T15:04:36.273Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1349555,
      "postDate": "2021-06-14T22:25:27.257Z",
      "content": "<p>Hi,<br>\ndoes the LSTM use a mask as well? It looks like Tensorflow's LSTM accepts mask.<br>\nthe residual addition in the transformer includes the padding tokens which we may want to ignore in the lstm.</p>",
      "rawMarkdown": "Hi,\ndoes the LSTM use a mask as well? It looks like Tensorflow's LSTM accepts mask.\nthe residual addition in the transformer includes the padding tokens which we may want to ignore in the lstm."
    },
    {
      "id": 1230528,
      "postDate": "2021-03-08T07:58:16.633Z",
      "content": "<p>great point</p>",
      "rawMarkdown": "great point"
    },
    {
      "id": 1219606,
      "postDate": "2021-02-27T01:55:10.553Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1216721,
      "postDate": "2021-02-24T12:53:15.533Z",
      "content": "<p>Good work! Thanks for sharing.</p>",
      "rawMarkdown": "Good work! Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 1195053,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2021-02-10T14:28:02.963000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> for sharing this approach. Very very nice, and great to see such long histories used. If you have published code somewhere, it would be great to see the encoder layer set up - if it is not public, no worries. Well done again, well deserved.  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1195096,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-10T14:56:00.353000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>, thank you. <br>\nActually, it's really simple to implement the model.<br>\nI only changed the one line from normal transformer encoder.<br>\nIn encoder layer,<br>\nI changed<br>\n     <code>attn_output, _ = self.mha(x, x, x[:,:,:], mask) #mha : multi head attention</code><br>\nto<br>\n     <code>attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)</code><br>\nHope it helps!</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 1200641,
      "author_name": "Arshad Shaikh",
      "author_url": "",
      "post_date": "2021-02-14T20:14:52.433000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> Congrats. <br>\nI had a couple of doubts, can you please answer them-</p>\n<p>Assume that I represent Query, Key and Value as Q, K, and V respectively, and let's assume that I have an input sequence with-  <br>\n<strong>D</strong>- Dimension of model  <br>\n<strong>S</strong> - Sequence Length of input  </p>\n<p>According to the <a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">paper</a>  first we need to multiply the Q and K transpose so -  <br>\n<strong>result1 = Q * K    ---&gt;      (S , D) * ( D ,S)  ---&gt; (S,S)</strong>  dimension  </p>\n<p>Now, we divide by sq root of a scaling factor and then apply softmax and get result2 with dim (S,S).  <br>\n<strong>result2 = softmax(  result1  /  sqRoot(ScalingFactor) )</strong>  </p>\n<p>Now we multiply the result2 with our Values V--&gt;  <br>\n<strong>FinalResult  =   result2 *  V       ---&gt;  i.e.  (S,S) * (S,D) ---&gt; (S,D)</strong>  dimension</p>\n<p>My doubts-  <br>\n<strong>1. If you keep seq len = 1  in the Query Q, then the final output will be having dim (1, D). \nThis is just a single feature with dim D, then how are you passing it through the LSTM? \n( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )</strong>    </p>\n<p><strong>2. In one of the comments you mentioned the following -</strong><br>\n<code>attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)</code>  <br>\n<strong>Can you please elaborate more on the code, What is x and what is x[:,-1:,:] ?\nWhich one of those 3 is Query Q ?\n( If the code above is used from <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html\" target=\"_blank\">torch.nn.MultiheadAttention</a> then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )</strong>    </p>\n<p>Thanks</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1201657,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-15T14:51:42.267000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1201901,
          "author_name": "Philip Dhingra",
          "author_url": "",
          "post_date": "2021-02-15T18:19:58.613000",
          "content": "<p>Based on my understanding of transformers and what <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> said, it looks like <code>x[:,-1:,:]</code> is the query. I don't think he uses torch's MHA. I've seen other MHAs where the prototype is KVQ instead of QKV.</p>\n<p>In one of the <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">reference implementations</a> for Transformers, <code>x</code> represents Q and has a shape of <code>(batch_size, sequence_len, embedding_dim_size)</code>. By tweaking <code>x</code> to <code>x[:,-1,:]</code> you change the <code>sequence_length</code> to 1.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1201917,
          "author_name": "HenryHZY",
          "author_url": "",
          "post_date": "2021-02-15T18:36:17.650000",
          "content": "<p>Same doubts, especial the first doubt</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1202115,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-15T21:10:03.073000",
          "content": "<p>Hi,</p>\n<pre><code>1. If you keep seq len = 1 in the Query Q, then the final output will be having dim (1, D). This is just a single feature with dim D, then how are you passing it through the LSTM? ( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )\n</code></pre>\n<h1> </h1>\n<p>In transformer encoder, there's a residual addition, which adds 'multi head attention output' to original input. In multi head attention, I use Query with len=1 to get some speed advantages in QK matrix multiplication. In result, the output of multi head attention becomes (batch size, 1, d_model) instead of (batch size, sequence length, d_model). So, for residual addition in transformer encoder, operands are original input with shape (batch size,  sequence length, d_model) and multi head attention output with shape (batch size, 1, d_model). Because two shapes are different, I add multi head attention output to every point of original input.(like broadcasting) Therefore, transformer encoder output is (batch size, sequence length, d_model).</p>\n<pre><code>2. In one of the comments you mentioned the following -\nattn_output, _ = self.mha(x, x, x[:,-1:,:], mask)\nCan you please elaborate more on the code, What is x and what is x[:,-1:,:] ? Which one of those 3 is Query Q ? ( If the code above is used from torch.nn.MultiheadAttention then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )\n</code></pre>\n<p>x is an original input to encoder. x's shape is (batch size, seq_len, d_model). The order of mha(multihead attention) parameters is v,k,q,mask.  </p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 1204329,
          "author_name": "Arshad Shaikh",
          "author_url": "",
          "post_date": "2021-02-16T04:16:51.840000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a>  for the reply. Once again congrats, its a great solution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1204407,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-16T05:53:12.453000",
          "content": "<p>Thank you :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332599,
          "author_name": "Gerryl",
          "author_url": "",
          "post_date": "2021-06-02T07:52:51.703000",
          "content": "<p>Thanks for sharing this amazing solution! One question I wanna ask is you add 'last query embedding' to the original input(which contains the last query), which seems like adding a 'scaler' to all sequences, does this make sense?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1485709,
      "author_name": "Arshad Shaikh",
      "author_url": "",
      "post_date": "2021-08-22T10:56:56.997000",
      "content": "<p>Pytorch Implementation of 1st place solution - <a href=\"https://github.com/arshadshk/Last_Query_Transformer_RNN-PyTorch\" target=\"_blank\">https://github.com/arshadshk/Last_Query_Transformer_RNN-PyTorch</a> <br>\n( Let me know if any changes are needed or any bugs present )</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1195536,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2021-02-10T22:09:36.580000",
      "content": "<p>Thank you keetar for your great solution, it was really fun to compete with you. Finally I lost, though.<br>\nIt is not surprising to use only the last input as query, because I heard <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>, <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> and <a href=\"https://www.kaggle.com/nadare\" target=\"_blank\">@nadare</a> did the same thing and I suppose many people did this, as it's easy to implement.<br>\nHowever, they said using the last input as query is very time-consuming (I remember <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> said training takes 1 week). How long is your training time, why is it so fast and what is your trick? Did you use TPU? </p>\n<p>My another question is about the sequence length and the input features. I was very surprised that your sequence length is 1728 and you used only 5 features. How did you notice sequence length is so important? Did you throw away other useful features to make the sequence length longer? I would like to know your journey to L = 1728 :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1195672,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-11T01:51:18.893000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>, thank you.<br>\nI'm surprised to hear many people tried this, as I saw none. Could you provide some links?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1195685,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-02-11T02:20:35.783000",
          "content": "<p>Yes, it's not clearly written in many people's solution, but I believe this was done by many people. I noticed my japanese friends did this in a meeting after the competition ends. For example, please look at the takoi's figure of <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209622\" target=\"_blank\">their solution</a>. it uses only the last query for training. tito also said <code>only trained and predicted for answered_correctly of last question of the sequence.</code> in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/210354\" target=\"_blank\">their solution</a>, which is based on the <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">public kernel</a>. Moreover, I didn't use the last query technique in training phase (instead I used masking), but I used last query technique in the inference phase, as can be seen in <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">my kernel</a>. The corresponding part of my code is here.</p>\n<pre><code>            #key and value process 2\n            query_clone = query.clone()\n            query_clone[token_idx, -2] = 0\n            memory_cat = torch.cat([query_clone, ohe_explanation, ohe_correctness, \n                                    normed_elapsed.unsqueeze(2), ohe_user_answer], dim = 2)\n            memory = self.norm4(self.linear8(self.dropout(F.relu(self.linear7(memory_cat)))))\n            memory = memory[:, :-1]\n\n            #transpose\n            query = query.transpose(0, 1)\n            memory = memory.transpose(0, 1)\n            spc_query = query[-1:]\n\n            #MyEncoder\n            out = self.MyEncoderLayer1(spc_query, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer2(out, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer3(out, memory, memory, None, padding_mask)\n            out = self.MyEncoderLayer4(out, memory, memory, None, padding_mask).squeeze(0)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1195704,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-11T02:49:42.963000",
          "content": "<p>I think they're totally different.<br>\nNormally, even if you train and predict for only last input, you don't change query vector in transformer to have length 1. In usual case, all query and value and key are set to have length L.(when input length is L). Because that's how transformer works. I believe the cases you mentioned have the same configurations. <br>\nChanging query vector in transformer to have length 1 changes the inner workings of transformer significantly. When multi head attention output is added to each inputs, last query's attention output is added to each inputs like constant, instead of each query's attention output. I found that it does not hurt the performance much, while making QK matrix multiplication in transformer to have O(L) instead of O(L^2).<br>\nTo answer why my model is faster than others who also used last query for training, it's simply because my QK matrix multiplication is O(L) and others is O(L^2).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1195741,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-02-11T03:29:10.940000",
          "content": "<p>Yes, as you say, in usual all query and key have length L and thus the matrix multiplication is O(L^2). But I think my inference code above is O(L), because the size of the query is 1 as can be seen from             <code>spc_query = query[-1:]</code> for the input of my custom transformer, so I just guessed other competitors who used the last input for training did the same thing like my inference phase (i.e. O(L) matrix multiplication). But it may be wrong, as I don't perfectly understand others' solution. Anyway, you are the kaggler I respect the most and your solution is really great without any doubt :) Thank you again!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1195759,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-11T04:00:23.103000",
          "content": "<p>Thank you :) And congratulations on your successful finish.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1233189,
      "author_name": "hlf",
      "author_url": "",
      "post_date": "2021-03-10T08:19:17.753000",
      "content": "<p>Hi, I had some doubts about your inputs, can you please anwer it.</p>\n<ol>\n<li><p><strong>Categorical embedding is used for first three features, and continuous embedding is used for last two continuous features.</strong><br>\nbut how you make continuous embedding?</p></li>\n<li><p><strong>'Timestamp difference' feature indicates the difference from the past question timestamp to the current question timestamp, and it is clipped by maximum value, 3 day.</strong><br>\nhow you fill the first 'Timestamp difference' in a sequence?</p></li>\n</ol>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1195385,
      "author_name": "HenryHZY",
      "author_url": "",
      "post_date": "2021-02-10T18:51:53.747000",
      "content": "<p>Cong! What a simple but sweet solution:)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1194577,
      "author_name": "TitiTest",
      "author_url": "",
      "post_date": "2021-02-10T08:53:02.880000",
      "content": "<p>Congrats for this really elegant and original solution !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1195111,
          "author_name": "keetar",
          "author_url": "",
          "post_date": "2021-02-10T15:04:36.273000",
          "content": "<p>Thank you :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1349555,
      "author_name": "Pierre",
      "author_url": "",
      "post_date": "2021-06-14T22:25:27.257000",
      "content": "<p>Hi,<br>\ndoes the LSTM use a mask as well? It looks like Tensorflow's LSTM accepts mask.<br>\nthe residual addition in the transformer includes the padding tokens which we may want to ignore in the lstm.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1230528,
      "author_name": "Neo Zhao",
      "author_url": "",
      "post_date": "2021-03-08T07:58:16.633000",
      "content": "<p>great point</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1219606,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-27T01:55:10.553000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1216721,
      "author_name": "ZavodRobotov",
      "author_url": "",
      "post_date": "2021-02-24T12:53:15.533000",
      "content": "<p>Good work! Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1194242": " Hi Kagglers, here is the paper link.\nPaper Link : https://arxiv.org/abs/2102.05038\nPlease refer the paper for details.\n# Summary\nI wanted to use Transformer but I could not input long history to the model because QK matrix multiplication in Transformer has O(L^2) time complexity when input length is L. \n# \nMy approach is to use only last input as Query, because I only predict last question's answer correctness per history inputs. It means I will only compare between last question(query) and other questions(key), and not between other questions. It makes QK matrix multiplication in Transformer to have O(L) time complexity(because len(Q)=1, len(K)=L), which allows me to input much longer history. \n# \nIn final submission, I ensemble 5 models with 1728 length history inputs.\nI didn't do feature engineering much, since I can use extremely long history, I wanted the model to learn it by itself. 5 input features I used are question id, question part, answer correctness, current question elapsed time, and timestamp difference.\n# Acknowledgement\nThank you @limerobot, to share amazing transformer model to approach table data problem in the 3rd place solution of 2019 Data Science Bowl. It was a good motivation to work on transformer encoder. https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891 \n# \nThanks to competition sponsers and organizers for hosting this fun competition.\n\n# \nThank you for reading.",
    "1195053": "Thank you @keetar for sharing this approach. Very very nice, and great to see such long histories used. If you have published code somewhere, it would be great to see the encoder layer set up - if it is not public, no worries. Well done again, well deserved.  ",
    "1200641": "Hey @keetar Congrats. \nI had a couple of doubts, can you please answer them-\n\nAssume that I represent Query, Key and Value as Q, K, and V respectively, and let's assume that I have an input sequence with-  \n**D**- Dimension of model  \n**S** - Sequence Length of input  \n\nAccording to the [paper](https://arxiv.org/abs/1706.03762)  first we need to multiply the Q and K transpose so -  \n**result1 = Q * K    --->      (S , D) * ( D ,S)  ---> (S,S)**  dimension  \n\nNow, we divide by sq root of a scaling factor and then apply softmax and get result2 with dim (S,S).  \n**result2 = softmax(  result1  /  sqRoot(ScalingFactor) )**  \n\nNow we multiply the result2 with our Values V-->  \n**FinalResult  =   result2 *  V       --->  i.e.  (S,S) * (S,D) ---> (S,D)**  dimension\n\nMy doubts-  \n**1. If you keep seq len = 1  in the Query Q, then the final output will be having dim (1, D). \nThis is just a single feature with dim D, then how are you passing it through the LSTM? \n( Or you are assuming the single feature with dim D as a sequence and passing it through the LSTM ? If so then can you tell the intuition behind this please )**    \n\n**2. In one of the comments you mentioned the following -**\n`attn_output, _ = self.mha(x, x, x[:,-1:,:], mask)`  \n**Can you please elaborate more on the code, What is x and what is x[:,-1:,:] ?\nWhich one of those 3 is Query Q ?\n( If the code above is used from [torch.nn.MultiheadAttention](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html) then in that case you are passing the values V as x[:,-1:,:], which will raise an error. )**    \n\nThanks\n\n \n \n\n",
    "1485709": "Pytorch Implementation of 1st place solution - https://github.com/arshadshk/Last_Query_Transformer_RNN-PyTorch \n( Let me know if any changes are needed or any bugs present )",
    "1195536": "Thank you keetar for your great solution, it was really fun to compete with you. Finally I lost, though.\nIt is not surprising to use only the last input as query, because I heard @its7171, @takoihiraokazu and @nadare did the same thing and I suppose many people did this, as it's easy to implement.\nHowever, they said using the last input as query is very time-consuming (I remember @its7171 said training takes 1 week). How long is your training time, why is it so fast and what is your trick? Did you use TPU? \n\nMy another question is about the sequence length and the input features. I was very surprised that your sequence length is 1728 and you used only 5 features. How did you notice sequence length is so important? Did you throw away other useful features to make the sequence length longer? I would like to know your journey to L = 1728 :)",
    "1233189": "Hi, I had some doubts about your inputs, can you please anwer it.\n\n1.  **Categorical embedding is used for first three features, and continuous embedding is used for last two continuous features.**\nbut how you make continuous embedding?\n\n2.  **'Timestamp difference' feature indicates the difference from the past question timestamp to the current question timestamp, and it is clipped by maximum value, 3 day.**\nhow you fill the first 'Timestamp difference' in a sequence?\n\nThanks!\n\n",
    "1195385": "Cong! What a simple but sweet solution:)",
    "1194577": "Congrats for this really elegant and original solution !",
    "1349555": "Hi,\ndoes the LSTM use a mask as well? It looks like Tensorflow's LSTM accepts mask.\nthe residual addition in the transformer includes the padding tokens which we may want to ignore in the lstm.",
    "1230528": "great point",
    "1219606": "",
    "1216721": "Good work! Thanks for sharing."
  }
}