{
  "id": 206584,
  "title": "Query on position encoding (specific to Riiid as well as generic)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206584",
  "author_name": "",
  "post_date": "2020-12-25T10:39:03.373351300Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Why do we use same dimensions for position embedding and for the actual input data? Even in BERT they have used 768 dimensions to encode the input word as well as the position. There is no comparison between the number of words (runs into 10s of 1000’s)  and the number of position (I think it is 0-512 in BERT). Do we need that high dimension embedding for position?<br>\nOne logical reason is that we want to add up the input embedding with the position embedding and they need to be in same shape. I have a question around that also? Why do they need to be in same shape? Why can’t we concatenate(matrix join) them instead of (matrix)adding?<br>\nInterestingly I observed in some implementations of transformers that the input embedding is multiplied by a scalar before adding to the position embedding. This sort of makes sense if we want to ensure that the word embedding (which is a stronger signal) has a larger say in determining the output compared to position embeddings. But neither Vaswani’s paper nor some of the subsequent ones talk about this..</p>\n<p>Coming to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?</p>\n<p><strong>Edit</strong>: Couple of hours &amp; a small nap later. <br>\nI am guessing that addition could be conveniently simpler than concatenation. There is no need to play around with shapes. So addition is preferred. Coming to the question of why the position embeddings would not distort the signals from the main input data is because the model in all its wisdom learns to assign lower values to position embeddings during training itself. So there is no need to scale the main data embeddings in any way. <br>\nThis explanation seems intuitive but would love to hear if someone has a differing view point. </p>\n<p>Of course there still remains the question of why the scalar multiplication is done in some transformer implementations like this one:<br>\nrefer to: x *= tf.math.sqrt(tf.cast(self.d_model, tf.float32)) in <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/transformer</a></p>\n<p>I can only imagine that this is some (harmless) unwanted relic from the past which nobody bothered to correct</p>\n<p>Thankfully the reference implementation of the transformer - the <a href=\"https://nlp.seas.harvard.edu/2018/04/03/attention.html#positional-encoding\" target=\"_blank\">https://nlp.seas.harvard.edu/2018/04/03/attention.html#positional-encoding</a> does not have this.</p>\n<p>BTW, I am guessing there are not too many takers for theoretical discussions at the fag end of the competition :)</p>",
  "messages": [
    {
      "id": "1126104",
      "postDate": "12/25/2020 10:39:03",
      "content": "<p>Why do we use same dimensions for position embedding and for the actual input data? Even in BERT they have used 768 dimensions to encode the input word as well as the position. There is no comparison between the number of words (runs into 10s of 1000’s)  and the number of position (I think it is 0-512 in BERT). Do we need that high dimension embedding for position?<br>\nOne logical reason is that we want to add up the input embedding with the position embedding and they need to be in same shape. I have a question around that also? Why do they need to be in same shape? Why can’t we concatenate(matrix join) them instead of (matrix)adding?<br>\nInterestingly I observed in some implementations of transformers that the input embedding is multiplied by a scalar before adding to the position embedding. This sort of makes sense if we want to ensure that the word embedding (which is a stronger signal) has a larger say in determining the output compared to position embeddings. But neither Vaswani’s paper nor some of the subsequent ones talk about this..</p>\n<p>Coming to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?</p>\n<p><strong>Edit</strong>: Couple of hours &amp; a small nap later. <br>\nI am guessing that addition could be conveniently simpler than concatenation. There is no need to play around with shapes. So addition is preferred. Coming to the question of why the position embeddings would not distort the signals from the main input data is because the model in all its wisdom learns to assign lower values to position embeddings during training itself. So there is no need to scale the main data embeddings in any way. <br>\nThis explanation seems intuitive but would love to hear if someone has a differing view point. </p>\n<p>Of course there still remains the question of why the scalar multiplication is done in some transformer implementations like this one:<br>\nrefer to: x *= tf.math.sqrt(tf.cast(self.d_model, tf.float32)) in <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/transformer</a></p>\n<p>I can only imagine that this is some (harmless) unwanted relic from the past which nobody bothered to correct</p>\n<p>Thankfully the reference implementation of the transformer - the <a href=\"https://nlp.seas.harvard.edu/2018/04/03/attention.html#positional-encoding\" target=\"_blank\">https://nlp.seas.harvard.edu/2018/04/03/attention.html#positional-encoding</a> does not have this.</p>\n<p>BTW, I am guessing there are not too many takers for theoretical discussions at the fag end of the competition :)</p>",
      "rawMarkdown": "Why do we use same dimensions for position embedding and for the actual input data? Even in BERT they have used 768 dimensions to encode the input word as well as the position. There is no comparison between the number of words (runs into 10s of 1000’s)  and the number of position (I think it is 0-512 in BERT). Do we need that high dimension embedding for position?\nOne logical reason is that we want to add up the input embedding with the position embedding and they need to be in same shape. I have a question around that also? Why do they need to be in same shape? Why can’t we concatenate(matrix join) them instead of (matrix)adding?\nInterestingly I observed in some implementations of transformers that the input embedding is multiplied by a scalar before adding to the position embedding. This sort of makes sense if we want to ensure that the word embedding (which is a stronger signal) has a larger say in determining the output compared to position embeddings. But neither Vaswani’s paper nor some of the subsequent ones talk about this..\n\nComing to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?\n\n**Edit**: Couple of hours & a small nap later. \nI am guessing that addition could be conveniently simpler than concatenation. There is no need to play around with shapes. So addition is preferred. Coming to the question of why the position embeddings would not distort the signals from the main input data is because the model in all its wisdom learns to assign lower values to position embeddings during training itself. So there is no need to scale the main data embeddings in any way. \nThis explanation seems intuitive but would love to hear if someone has a differing view point. \n\nOf course there still remains the question of why the scalar multiplication is done in some transformer implementations like this one:\nrefer to: x *= tf.math.sqrt(tf.cast(self.d_model, tf.float32)) in https://www.tensorflow.org/tutorials/text/transformer\n\nI can only imagine that this is some (harmless) unwanted relic from the past which nobody bothered to correct\n\nThankfully the reference implementation of the transformer - the https://nlp.seas.harvard.edu/2018/04/03/attention.html#positional-encoding does not have this.\n\nBTW, I am guessing there are not too many takers for theoretical discussions at the fag end of the competition :)",
      "votes": null
    },
    {
      "id": "1126516",
      "postDate": "12/25/2020 17:12:29",
      "content": "<p>I forget where I saw about the postion encoding,one paper say sum position encoding compare to concat position  encoding ,the result didn't change.maybe SAKT paper or <a href=\"https://nlp.seas.harvard.edu/2018/04/03/attention.html#applications-of-attention-in-our-model\" target=\"_blank\">https://nlp.seas.harvard.edu/2018/04/03/attention.html#applications-of-attention-in-our-model</a><br>\nI just read these papers a few hours ago because I'm a  neural network beginner<br>\nI'll update if I find again. <br>\n<a href=\"https://arxiv.org/pdf/1705.03122.pdf\" target=\"_blank\">this</a> may has a more detailed explanation of position encoding </p>\n<p>BTW Theoretical discussion can help me learn more thoroughly.😊😊</p>",
      "rawMarkdown": "I forget where I saw about the postion encoding,one paper say sum position encoding compare to concat position  encoding ,the result didn't change.maybe SAKT paper or https://nlp.seas.harvard.edu/2018/04/03/attention.html#applications-of-attention-in-our-model\nI just read these papers a few hours ago because I'm a  neural network beginner\nI'll update if I find again. \n[this](https://arxiv.org/pdf/1705.03122.pdf) may has a more detailed explanation of position encoding \n\nBTW Theoretical discussion can help me learn more thoroughly.😊😊",
      "votes": null
    },
    {
      "id": "1126550",
      "postDate": "12/25/2020 17:32:10",
      "content": "<p>cool. that somewhat confirms my understanding. 'sum' is just more convenient than 'concat'. There is no difference because the model trains itself and outputs lesser values in the position embedding compared to actual input embeddings. </p>\n<p>BTW, positional encoding is slightly incorrect in the public kernels. hope u have taken care in your model</p>",
      "rawMarkdown": "cool. that somewhat confirms my understanding. 'sum' is just more convenient than 'concat'. There is no difference because the model trains itself and outputs lesser values in the position embedding compared to actual input embeddings. \n\nBTW, positional encoding is slightly incorrect in the public kernels. hope u have taken care in your model",
      "votes": null
    },
    {
      "id": "1126718",
      "postDate": "12/25/2020 21:09:31",
      "content": "<p>I've thought about this, not just in the context of pos embeddings, but other embeddings as well. Traditionally before these seq2seq transformer networks, we would use a heuristic to determine embedding size as a function of the # of elements, concat our embeddings, then drop them into an RNN. In SAKT/SAINT/+ all embeddings share the same dimension and are embedded into the same space.</p>\n<p>What does this mean?</p>\n<ul>\n<li>A correctness embedding that only has 2 options (0/1) is embedded at the same dimensionality as exercise embeddings that have almost 100k values(!)</li>\n<li>The two (or more) embeddings actually <em>share</em> an embedding space; if you were to put this into NLP terms, it would be like taking two different \"languages\" and then force them to share the same embedding. That might seem just plan WRONG. But take into account those language models which are trained without the use of a language-language corpus. In a sense, it might act as a regularizer because it's bound to add a ton of noise, but it also can lead to overfitting if the embeddings with only a few members aren't properly taken care of. For example—</li>\n<li>If the embeddings are multiplied what is that doing? One embedding's value literally knock out the contribution of that dimension for your entire stack. That's pretty intense. Better use a slow LR and hope you don't have dead neurons. Or a ton of dropout.</li>\n</ul>\n<blockquote>\n  <p>Coming to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?</p>\n</blockquote>\n<p>This has been on my TODO list for testing but haven't gotten around to it because I haven't found a magic signal to bank on for taking me over 800AUC yet. If you end up experimenting locally, please do share results. We still 6 days before the soft deadline for the stopping of sharing nice details.</p>",
      "rawMarkdown": "I've thought about this, not just in the context of pos embeddings, but other embeddings as well. Traditionally before these seq2seq transformer networks, we would use a heuristic to determine embedding size as a function of the # of elements, concat our embeddings, then drop them into an RNN. In SAKT/SAINT/+ all embeddings share the same dimension and are embedded into the same space.\n\nWhat does this mean?\n\n- A correctness embedding that only has 2 options (0/1) is embedded at the same dimensionality as exercise embeddings that have almost 100k values(!)\n- The two (or more) embeddings actually *share* an embedding space; if you were to put this into NLP terms, it would be like taking two different \"languages\" and then force them to share the same embedding. That might seem just plan WRONG. But take into account those language models which are trained without the use of a language-language corpus. In a sense, it might act as a regularizer because it's bound to add a ton of noise, but it also can lead to overfitting if the embeddings with only a few members aren't properly taken care of. For example—\n- If the embeddings are multiplied what is that doing? One embedding's value literally knock out the contribution of that dimension for your entire stack. That's pretty intense. Better use a slow LR and hope you don't have dead neurons. Or a ton of dropout.\n\n> Coming to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?\n\nThis has been on my TODO list for testing but haven't gotten around to it because I haven't found a magic signal to bank on for taking me over 800AUC yet. If you end up experimenting locally, please do share results. We still 6 days before the soft deadline for the stopping of sharing nice details.",
      "votes": null
    },
    {
      "id": "1127026",
      "postDate": "12/26/2020 07:21:27",
      "content": "<p>Will let u know authman if I do something locally. Given that my problem is not yet fixed by Kaggle support, I am thinking to move onto other competitions…</p>",
      "rawMarkdown": "Will let u know authman if I do something locally. Given that my problem is not yet fixed by Kaggle support, I am thinking to move onto other competitions...",
      "votes": null
    },
    {
      "id": "1127034",
      "postDate": "12/26/2020 07:33:11",
      "content": "<p>Riiid team did what was done earlier almost; So that's why they are hosting the comp here to find other better alternative ideas/features! </p>\n<blockquote>\n  <p>the input embedding is multiplied by a scalar before adding to the position embedding.</p>\n</blockquote>\n<p>I have tried this but it didn't help at all; I don't know why; Though there's a bug in my code, so i am not sure what works or not; But i am interested to know it all once the comp ends for good (or maybe before an hour the comp ends) !</p>",
      "rawMarkdown": "Riiid team did what was done earlier almost; So that's why they are hosting the comp here to find other better alternative ideas/features! \n\n>the input embedding is multiplied by a scalar before adding to the position embedding.\n\nI have tried this but it didn't help at all; I don't know why; Though there's a bug in my code, so i am not sure what works or not; But i am interested to know it all once the comp ends for good (or maybe before an hour the comp ends) !",
      "votes": null
    },
    {
      "id": "1127295",
      "postDate": "12/26/2020 12:14:30",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>, sorry didn’t understand your third point initially, No I didn’t intend to say multiply instead of sum…I meant concatenate(join) instead of sum, so basically the main input could have 256 dimensions and the position (or any other less significant input) could have lesser dimensions..say 56 and we just join the embeddings to take it to 312 and then use this one.</p>\n<p>I later learnt that there is no harm in this approach but it will not improve the results in any way. </p>\n<p>So the original question remains - why should all inputs share the same dimension? Especially since some of them (like position) dont need to be of that size at all. More importantly wouldn’t they distort the main input data. Is this why the main input data is multiplied by a scalar (to enhance its value)?</p>\n<p>Though I couldn’t find the answers anywhere, what I believe</p>\n<ul>\n<li>they share same dimension so that they can be added to one other (matter of convenience)</li>\n<li>Do they need that many dimensions - No. But apparently there is no damage done in having higher dimensions represent same data. As I sad BERT for e.g. uses 768 dimensions for words (all words in the language) and same 768 dimensions for position (I think their seq length is 512)</li>\n<li>Would this distort the signal from the main input data - The answer is NO else BERT would not have been the emperor of all embeddings</li>\n<li>Why does it not distort - Because the model trains to assign lower numbers to the position encodings and higher numbers to the main input. When a matrix sum is done (element to element addition), the effect of position embedding on the sum is very very low (my opinion)</li>\n<li>Why is in some transformer implementations there is a scalar multiplication before the addition - this is some old practice and does NOT add value (my opinion)</li>\n</ul>\n<p>It is nice to see we have been thinking of similar issues, These apply beyond RiiId and will definitely help me in future as well.</p>\n<p>PS - My last contribution to RiiiD. I believe that positional encoding is NOT needed at all and instead we just need a simple memory decay module. In case you are interested - please check if this makes sense: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a></p>",
      "rawMarkdown": "authman, sorry didn’t understand your third point initially, No I didn’t intend to say multiply instead of sum…I meant concatenate(join) instead of sum, so basically the main input could have 256 dimensions and the position (or any other less significant input) could have lesser dimensions..say 56 and we just join the embeddings to take it to 312 and then use this one.\n\nI later learnt that there is no harm in this approach but it will not improve the results in any way. \n\nSo the original question remains - why should all inputs share the same dimension? Especially since some of them (like position) dont need to be of that size at all. More importantly wouldn’t they distort the main input data. Is this why the main input data is multiplied by a scalar (to enhance its value)?\n\nThough I couldn’t find the answers anywhere, what I believe\n- they share same dimension so that they can be added to one other (matter of convenience)\n- Do they need that many dimensions - No. But apparently there is no damage done in having higher dimensions represent same data. As I sad BERT for e.g. uses 768 dimensions for words (all words in the language) and same 768 dimensions for position (I think their seq length is 512)\n- Would this distort the signal from the main input data - The answer is NO else BERT would not have been the emperor of all embeddings\n- Why does it not distort - Because the model trains to assign lower numbers to the position encodings and higher numbers to the main input. When a matrix sum is done (element to element addition), the effect of position embedding on the sum is very very low (my opinion)\n- Why is in some transformer implementations there is a scalar multiplication before the addition - this is some old practice and does NOT add value (my opinion)\n\nIt is nice to see we have been thinking of similar issues, These apply beyond RiiId and will definitely help me in future as well.\n\nPS - My last contribution to RiiiD. I believe that positional encoding is NOT needed at all and instead we just need a simple memory decay module. In case you are interested - please check if this makes sense: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719",
      "votes": null
    },
    {
      "id": "1127546",
      "postDate": "12/26/2020 16:21:42",
      "content": "<p>use 512 dimension for 7 part ….It does look strange,Maybe after I read more papers, I can solve this puzzle😃</p>",
      "rawMarkdown": "use 512 dimension for 7 part ....It does look strange,Maybe after I read more papers, I can solve this puzzle😃",
      "votes": null
    },
    {
      "id": "1127997",
      "postDate": "12/27/2020 05:04:20",
      "content": "<p>yes <a href=\"https://www.kaggle.com/yangxiaoshuai\" target=\"_blank\">@yangxiaoshuai</a> :) Authman in fact talks of a scenario where there could be a binary…so a 0 and 1 encoded by 512 dimensions..</p>\n<p>I did not see any papers covering this, but if you are interested, you could try generating these embeddings, then take a similarity measure and see how similar the combined embedding is to the original input embedding versus the position or part embedding. I am guessing it would be several scales of difference which shows that the model pays attention to the main signal and gives minimal-needed attention to the position and other weaker signals</p>\n<p>that would be quite a blog </p>",
      "rawMarkdown": "yes @yangxiaoshuai :) Authman in fact talks of a scenario where there could be a binary…so a 0 and 1 encoded by 512 dimensions..\n\nI did not see any papers covering this, but if you are interested, you could try generating these embeddings, then take a similarity measure and see how similar the combined embedding is to the original input embedding versus the position or part embedding. I am guessing it would be several scales of difference which shows that the model pays attention to the main signal and gives minimal-needed attention to the position and other weaker signals\n\nthat would be quite a blog",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1126516,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/25/2020 17:12:29",
      "content": "<p>I forget where I saw about the postion encoding,one paper say sum position encoding compare to concat position  encoding ,the result didn't change.maybe SAKT paper or <a href=\"https://nlp.seas.harvard.edu/2018/04/03/attention.html#applications-of-attention-in-our-model\" target=\"_blank\">https://nlp.seas.harvard.edu/2018/04/03/attention.html#applications-of-attention-in-our-model</a><br>\nI just read these papers a few hours ago because I'm a  neural network beginner<br>\nI'll update if I find again. <br>\n<a href=\"https://arxiv.org/pdf/1705.03122.pdf\" target=\"_blank\">this</a> may has a more detailed explanation of position encoding </p>\n<p>BTW Theoretical discussion can help me learn more thoroughly.😊😊</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126550,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/25/2020 17:32:10",
          "content": "<p>cool. that somewhat confirms my understanding. 'sum' is just more convenient than 'concat'. There is no difference because the model trains itself and outputs lesser values in the position embedding compared to actual input embeddings. </p>\n<p>BTW, positional encoding is slightly incorrect in the public kernels. hope u have taken care in your model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1126718,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/25/2020 21:09:31",
      "content": "<p>I've thought about this, not just in the context of pos embeddings, but other embeddings as well. Traditionally before these seq2seq transformer networks, we would use a heuristic to determine embedding size as a function of the # of elements, concat our embeddings, then drop them into an RNN. In SAKT/SAINT/+ all embeddings share the same dimension and are embedded into the same space.</p>\n<p>What does this mean?</p>\n<ul>\n<li>A correctness embedding that only has 2 options (0/1) is embedded at the same dimensionality as exercise embeddings that have almost 100k values(!)</li>\n<li>The two (or more) embeddings actually <em>share</em> an embedding space; if you were to put this into NLP terms, it would be like taking two different \"languages\" and then force them to share the same embedding. That might seem just plan WRONG. But take into account those language models which are trained without the use of a language-language corpus. In a sense, it might act as a regularizer because it's bound to add a ton of noise, but it also can lead to overfitting if the embeddings with only a few members aren't properly taken care of. For example—</li>\n<li>If the embeddings are multiplied what is that doing? One embedding's value literally knock out the contribution of that dimension for your entire stack. That's pretty intense. Better use a slow LR and hope you don't have dead neurons. Or a ton of dropout.</li>\n</ul>\n<blockquote>\n  <p>Coming to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?</p>\n</blockquote>\n<p>This has been on my TODO list for testing but haven't gotten around to it because I haven't found a magic signal to bank on for taking me over 800AUC yet. If you end up experimenting locally, please do share results. We still 6 days before the soft deadline for the stopping of sharing nice details.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1127026,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/26/2020 07:21:27",
          "content": "<p>Will let u know authman if I do something locally. Given that my problem is not yet fixed by Kaggle support, I am thinking to move onto other competitions…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127295,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/26/2020 12:14:30",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>, sorry didn’t understand your third point initially, No I didn’t intend to say multiply instead of sum…I meant concatenate(join) instead of sum, so basically the main input could have 256 dimensions and the position (or any other less significant input) could have lesser dimensions..say 56 and we just join the embeddings to take it to 312 and then use this one.</p>\n<p>I later learnt that there is no harm in this approach but it will not improve the results in any way. </p>\n<p>So the original question remains - why should all inputs share the same dimension? Especially since some of them (like position) dont need to be of that size at all. More importantly wouldn’t they distort the main input data. Is this why the main input data is multiplied by a scalar (to enhance its value)?</p>\n<p>Though I couldn’t find the answers anywhere, what I believe</p>\n<ul>\n<li>they share same dimension so that they can be added to one other (matter of convenience)</li>\n<li>Do they need that many dimensions - No. But apparently there is no damage done in having higher dimensions represent same data. As I sad BERT for e.g. uses 768 dimensions for words (all words in the language) and same 768 dimensions for position (I think their seq length is 512)</li>\n<li>Would this distort the signal from the main input data - The answer is NO else BERT would not have been the emperor of all embeddings</li>\n<li>Why does it not distort - Because the model trains to assign lower numbers to the position encodings and higher numbers to the main input. When a matrix sum is done (element to element addition), the effect of position embedding on the sum is very very low (my opinion)</li>\n<li>Why is in some transformer implementations there is a scalar multiplication before the addition - this is some old practice and does NOT add value (my opinion)</li>\n</ul>\n<p>It is nice to see we have been thinking of similar issues, These apply beyond RiiId and will definitely help me in future as well.</p>\n<p>PS - My last contribution to RiiiD. I believe that positional encoding is NOT needed at all and instead we just need a simple memory decay module. In case you are interested - please check if this makes sense: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127546,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/26/2020 16:21:42",
          "content": "<p>use 512 dimension for 7 part ….It does look strange,Maybe after I read more papers, I can solve this puzzle😃</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127997,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/27/2020 05:04:20",
          "content": "<p>yes <a href=\"https://www.kaggle.com/yangxiaoshuai\" target=\"_blank\">@yangxiaoshuai</a> :) Authman in fact talks of a scenario where there could be a binary…so a 0 and 1 encoded by 512 dimensions..</p>\n<p>I did not see any papers covering this, but if you are interested, you could try generating these embeddings, then take a similarity measure and see how similar the combined embedding is to the original input embedding versus the position or part embedding. I am guessing it would be several scales of difference which shows that the model pays attention to the main signal and gives minimal-needed attention to the position and other weaker signals</p>\n<p>that would be quite a blog </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127034,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "12/26/2020 07:33:11",
      "content": "<p>Riiid team did what was done earlier almost; So that's why they are hosting the comp here to find other better alternative ideas/features! </p>\n<blockquote>\n  <p>the input embedding is multiplied by a scalar before adding to the position embedding.</p>\n</blockquote>\n<p>I have tried this but it didn't help at all; I don't know why; Though there's a bug in my code, so i am not sure what works or not; But i am interested to know it all once the comp ends for good (or maybe before an hour the comp ends) !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1126104": "Why do we use same dimensions for position embedding and for the actual input data? Even in BERT they have used 768 dimensions to encode the input word as well as the position. There is no comparison between the number of words (runs into 10s of 1000’s)  and the number of position (I think it is 0-512 in BERT). Do we need that high dimension embedding for position?\nOne logical reason is that we want to add up the input embedding with the position embedding and they need to be in same shape. I have a question around that also? Why do they need to be in same shape? Why can’t we concatenate(matrix join) them instead of (matrix)adding?\nInterestingly I observed in some implementations of transformers that the input embedding is multiplied by a scalar before adding to the position embedding. This sort of makes sense if we want to ensure that the word embedding (which is a stronger signal) has a larger say in determining the output compared to position embeddings. But neither Vaswani’s paper nor some of the subsequent ones talk about this..\n\nComing to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?\n\n**Edit**: Couple of hours & a small nap later. \nI am guessing that addition could be conveniently simpler than concatenation. There is no need to play around with shapes. So addition is preferred. Coming to the question of why the position embeddings would not distort the signals from the main input data is because the model in all its wisdom learns to assign lower values to position embeddings during training itself. So there is no need to scale the main data embeddings in any way. \nThis explanation seems intuitive but would love to hear if someone has a differing view point. \n\nOf course there still remains the question of why the scalar multiplication is done in some transformer implementations like this one:\nrefer to: x *= tf.math.sqrt(tf.cast(self.d_model, tf.float32)) in https://www.tensorflow.org/tutorials/text/transformer\n\nI can only imagine that this is some (harmless) unwanted relic from the past which nobody bothered to correct\n\nThankfully the reference implementation of the transformer - the https://nlp.seas.harvard.edu/2018/04/03/attention.html#positional-encoding does not have this.\n\nBTW, I am guessing there are not too many takers for theoretical discussions at the fag end of the competition :)",
    "1126516": "I forget where I saw about the postion encoding,one paper say sum position encoding compare to concat position  encoding ,the result didn't change.maybe SAKT paper or https://nlp.seas.harvard.edu/2018/04/03/attention.html#applications-of-attention-in-our-model\nI just read these papers a few hours ago because I'm a  neural network beginner\nI'll update if I find again. \n[this](https://arxiv.org/pdf/1705.03122.pdf) may has a more detailed explanation of position encoding \n\nBTW Theoretical discussion can help me learn more thoroughly.😊😊",
    "1126550": "cool. that somewhat confirms my understanding. 'sum' is just more convenient than 'concat'. There is no difference because the model trains itself and outputs lesser values in the position embedding compared to actual input embeddings. \n\nBTW, positional encoding is slightly incorrect in the public kernels. hope u have taken care in your model",
    "1126718": "I've thought about this, not just in the context of pos embeddings, but other embeddings as well. Traditionally before these seq2seq transformer networks, we would use a heuristic to determine embedding size as a function of the # of elements, concat our embeddings, then drop them into an RNN. In SAKT/SAINT/+ all embeddings share the same dimension and are embedded into the same space.\n\nWhat does this mean?\n\n- A correctness embedding that only has 2 options (0/1) is embedded at the same dimensionality as exercise embeddings that have almost 100k values(!)\n- The two (or more) embeddings actually *share* an embedding space; if you were to put this into NLP terms, it would be like taking two different \"languages\" and then force them to share the same embedding. That might seem just plan WRONG. But take into account those language models which are trained without the use of a language-language corpus. In a sense, it might act as a regularizer because it's bound to add a ton of noise, but it also can lead to overfitting if the embeddings with only a few members aren't properly taken care of. For example—\n- If the embeddings are multiplied what is that doing? One embedding's value literally knock out the contribution of that dimension for your entire stack. That's pretty intense. Better use a slow LR and hope you don't have dead neurons. Or a ton of dropout.\n\n> Coming to Riiid, given the computation constraints, does it make sense to use smaller dimension embeddings for position (and any of the other weaker signals we want to use) and concatenate (instead of adding) to the input data embedding?\n\nThis has been on my TODO list for testing but haven't gotten around to it because I haven't found a magic signal to bank on for taking me over 800AUC yet. If you end up experimenting locally, please do share results. We still 6 days before the soft deadline for the stopping of sharing nice details.",
    "1127026": "Will let u know authman if I do something locally. Given that my problem is not yet fixed by Kaggle support, I am thinking to move onto other competitions...",
    "1127034": "Riiid team did what was done earlier almost; So that's why they are hosting the comp here to find other better alternative ideas/features! \n\n>the input embedding is multiplied by a scalar before adding to the position embedding.\n\nI have tried this but it didn't help at all; I don't know why; Though there's a bug in my code, so i am not sure what works or not; But i am interested to know it all once the comp ends for good (or maybe before an hour the comp ends) !",
    "1127295": "authman, sorry didn’t understand your third point initially, No I didn’t intend to say multiply instead of sum…I meant concatenate(join) instead of sum, so basically the main input could have 256 dimensions and the position (or any other less significant input) could have lesser dimensions..say 56 and we just join the embeddings to take it to 312 and then use this one.\n\nI later learnt that there is no harm in this approach but it will not improve the results in any way. \n\nSo the original question remains - why should all inputs share the same dimension? Especially since some of them (like position) dont need to be of that size at all. More importantly wouldn’t they distort the main input data. Is this why the main input data is multiplied by a scalar (to enhance its value)?\n\nThough I couldn’t find the answers anywhere, what I believe\n- they share same dimension so that they can be added to one other (matter of convenience)\n- Do they need that many dimensions - No. But apparently there is no damage done in having higher dimensions represent same data. As I sad BERT for e.g. uses 768 dimensions for words (all words in the language) and same 768 dimensions for position (I think their seq length is 512)\n- Would this distort the signal from the main input data - The answer is NO else BERT would not have been the emperor of all embeddings\n- Why does it not distort - Because the model trains to assign lower numbers to the position encodings and higher numbers to the main input. When a matrix sum is done (element to element addition), the effect of position embedding on the sum is very very low (my opinion)\n- Why is in some transformer implementations there is a scalar multiplication before the addition - this is some old practice and does NOT add value (my opinion)\n\nIt is nice to see we have been thinking of similar issues, These apply beyond RiiId and will definitely help me in future as well.\n\nPS - My last contribution to RiiiD. I believe that positional encoding is NOT needed at all and instead we just need a simple memory decay module. In case you are interested - please check if this makes sense: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719",
    "1127546": "use 512 dimension for 7 part ....It does look strange,Maybe after I read more papers, I can solve this puzzle😃",
    "1127997": "yes @yangxiaoshuai :) Authman in fact talks of a scenario where there could be a binary…so a 0 and 1 encoded by 512 dimensions..\n\nI did not see any papers covering this, but if you are interested, you could try generating these embeddings, then take a similarity measure and see how similar the combined embedding is to the original input embedding versus the position or part embedding. I am guessing it would be several scales of difference which shows that the model pays attention to the main signal and gives minimal-needed attention to the position and other weaker signals\n\nthat would be quite a blog"
  },
  "source": "meta"
}