{
  "id": 201798,
  "title": "Transformers: Add vs Concat for Positional Embedding",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201798",
  "author_name": "",
  "post_date": "2020-12-06T20:56:29.657627600Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>So in my experiments, Concatenating different embeddings works either similar or slightly better vs Adding the before feeding into Transformers.</p>\n<p>Though I understand the notion of creating Positional Embedding, I've not yet found any concrete and sound reasoning on why they are added directly to the other embeddings.</p>\n<p>Wouldn't it be more reasonable to do a concatenation layer and then add a dense layer to reduce the dimensions back to d_model.</p>\n<p>Any links/sources on the intuition of adding them would be highly appreciated. </p>",
  "messages": [
    {
      "id": "1104354",
      "postDate": "12/06/2020 20:56:29",
      "content": "<p>Hi,</p>\n<p>So in my experiments, Concatenating different embeddings works either similar or slightly better vs Adding the before feeding into Transformers.</p>\n<p>Though I understand the notion of creating Positional Embedding, I've not yet found any concrete and sound reasoning on why they are added directly to the other embeddings.</p>\n<p>Wouldn't it be more reasonable to do a concatenation layer and then add a dense layer to reduce the dimensions back to d_model.</p>\n<p>Any links/sources on the intuition of adding them would be highly appreciated. </p>",
      "rawMarkdown": "Hi,\n\nSo in my experiments, Concatenating different embeddings works either similar or slightly better vs Adding the before feeding into Transformers.\n\nThough I understand the notion of creating Positional Embedding, I've not yet found any concrete and sound reasoning on why they are added directly to the other embeddings.\n\nWouldn't it be more reasonable to do a concatenation layer and then add a dense layer to reduce the dimensions back to d_model.\n\nAny links/sources on the intuition of adding them would be highly appreciated.",
      "votes": null
    },
    {
      "id": "1105453",
      "postDate": "12/07/2020 22:12:53",
      "content": "<p>If you concatenate them, you'll still send both to the latent space and add them together when you apply the matrix multiplication because $M \\times X$ can be seen as $[M1,M2] \\times [X1, X2]^t = M1\\times X1 + M2 \\times X2$. But, by concatenating, you decrease the number of dimensions allowed for each <br>\nof the concatenated parts which limits the expressivity of the model. So, overall, it seems best to add instead of concatenating.<br>\nI tried both and it didn't change much though.</p>",
      "rawMarkdown": "If you concatenate them, you'll still send both to the latent space and add them together when you apply the matrix multiplication because $M \\times X$ can be seen as $[M1,M2] \\times [X1, X2]^t = M1\\times X1 + M2 \\times X2$. But, by concatenating, you decrease the number of dimensions allowed for each \nof the concatenated parts which limits the expressivity of the model. So, overall, it seems best to add instead of concatenating.\nI tried both and it didn't change much though.",
      "votes": null
    },
    {
      "id": "1105687",
      "postDate": "12/08/2020 05:03:38",
      "content": "<p>Thanks <br>\nFor me concat leads to an earlier convergence vs addition. Not sure if that's a good thing or an issue in my model :/</p>",
      "rawMarkdown": "Thanks \nFor me concat leads to an earlier convergence vs addition. Not sure if that's a good thing or an issue in my model :/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105453,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "12/07/2020 22:12:53",
      "content": "<p>If you concatenate them, you'll still send both to the latent space and add them together when you apply the matrix multiplication because $M \\times X$ can be seen as $[M1,M2] \\times [X1, X2]^t = M1\\times X1 + M2 \\times X2$. But, by concatenating, you decrease the number of dimensions allowed for each <br>\nof the concatenated parts which limits the expressivity of the model. So, overall, it seems best to add instead of concatenating.<br>\nI tried both and it didn't change much though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105687,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/08/2020 05:03:38",
          "content": "<p>Thanks <br>\nFor me concat leads to an earlier convergence vs addition. Not sure if that's a good thing or an issue in my model :/</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1104354": "Hi,\n\nSo in my experiments, Concatenating different embeddings works either similar or slightly better vs Adding the before feeding into Transformers.\n\nThough I understand the notion of creating Positional Embedding, I've not yet found any concrete and sound reasoning on why they are added directly to the other embeddings.\n\nWouldn't it be more reasonable to do a concatenation layer and then add a dense layer to reduce the dimensions back to d_model.\n\nAny links/sources on the intuition of adding them would be highly appreciated.",
    "1105453": "If you concatenate them, you'll still send both to the latent space and add them together when you apply the matrix multiplication because $M \\times X$ can be seen as $[M1,M2] \\times [X1, X2]^t = M1\\times X1 + M2 \\times X2$. But, by concatenating, you decrease the number of dimensions allowed for each \nof the concatenated parts which limits the expressivity of the model. So, overall, it seems best to add instead of concatenating.\nI tried both and it didn't change much though.",
    "1105687": "Thanks \nFor me concat leads to an earlier convergence vs addition. Not sure if that's a good thing or an issue in my model :/"
  },
  "source": "meta"
}