{
  "id": 235509,
  "title": "Understanding Vocabulary Embedding Dimension",
  "url": "/competitions/bms-molecular-translation/discussion/235509",
  "author_name": "Michael Wolff",
  "post_date": "2021-04-29T22:16:23.866000",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi guys,<br>\nI was wondering what a reasonable vocabulary embedding dimension, depending on the vocabulary size, is:</p>\n<p>To be specific, I'm referring to this line in a typical LSTM or GRU decoder:<br>\n<code>self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)</code><br>\nwhich is then later used to embed the integer representation of the vocabulary.</p>\n<p>The <a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">kernel</a> I am adapting had an embedding_dim equal to the decoder dimension (meaning the dimension of the cell/hidden unit) of 512, now reduced to 256 in the latest version. </p>\n<p>I have been using 512 until now, however, from my mathematical intuition, an embedding_dim larger than the vocabulary size should give no benefit, because in the 'equal' case all elements of the vocabulary could already be embedded orthogonally if needed, and all extra dimensions should have no impact. </p>\n<p>A test with embedding_dim = vocab_size over 2 epochs seems to support this, also giving a slight speed up and less decoder parameters.</p>\n<p>I read that in a 'typical' text generation problem the embedding dim is much smaller than the vocabulary, with vocabularies containing thousands of words, so much larger than a typical InChI vocabulary.</p>\n<p>Can you guys give any insight on this topic?</p>",
  "messages": [
    {
      "id": 1288770,
      "postDate": "2021-04-30T10:15:05.180Z",
      "content": "<p>The embeddings are not only there to separate tokens from each other. They are feature vectors. While these features describe the relations to other tokens, the benefit of a higher embedding size is influenced by (at least) five things: Amount of data, sequence length, complexity of relations between tokens, vocabulary size, model architecture. In about that order of decreasing importance  (sequence length also induces complexity).</p>\n<p>So you can benefit from an embedding size far higher than your vocabulary. A good example is the <a href=\"https://paperswithcode.com/sota/language-modelling-on-enwiki8\" target=\"_blank\">enwik8 benchmark</a> which has a vocabulary size of ~200 and usually uses sequence lengths of 8000. You will find, that the top models (which are not pretrained on other data) often use a embedding size of 1024.</p>",
      "rawMarkdown": "The embeddings are not only there to separate tokens from each other. They are feature vectors. While these features describe the relations to other tokens, the benefit of a higher embedding size is influenced by (at least) five things: Amount of data, sequence length, complexity of relations between tokens, vocabulary size, model architecture. In about that order of decreasing importance  (sequence length also induces complexity).\n\nSo you can benefit from an embedding size far higher than your vocabulary. A good example is the [enwik8 benchmark](https://paperswithcode.com/sota/language-modelling-on-enwiki8) which has a vocabulary size of ~200 and usually uses sequence lengths of 8000. You will find, that the top models (which are not pretrained on other data) often use a embedding size of 1024.",
      "votes": 6,
      "replies": [
        {
          "id": 1288810,
          "postDate": "2021-04-30T11:04:00.617Z",
          "content": "<p>Thanks a lot for your explanation!</p>",
          "rawMarkdown": "Thanks a lot for your explanation!"
        }
      ]
    },
    {
      "id": 1288342,
      "postDate": "2021-04-29T22:16:23.867Z",
      "content": "<p>Hi guys,<br>\nI was wondering what a reasonable vocabulary embedding dimension, depending on the vocabulary size, is:</p>\n<p>To be specific, I'm referring to this line in a typical LSTM or GRU decoder:<br>\n<code>self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)</code><br>\nwhich is then later used to embed the integer representation of the vocabulary.</p>\n<p>The <a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">kernel</a> I am adapting had an embedding_dim equal to the decoder dimension (meaning the dimension of the cell/hidden unit) of 512, now reduced to 256 in the latest version. </p>\n<p>I have been using 512 until now, however, from my mathematical intuition, an embedding_dim larger than the vocabulary size should give no benefit, because in the 'equal' case all elements of the vocabulary could already be embedded orthogonally if needed, and all extra dimensions should have no impact. </p>\n<p>A test with embedding_dim = vocab_size over 2 epochs seems to support this, also giving a slight speed up and less decoder parameters.</p>\n<p>I read that in a 'typical' text generation problem the embedding dim is much smaller than the vocabulary, with vocabularies containing thousands of words, so much larger than a typical InChI vocabulary.</p>\n<p>Can you guys give any insight on this topic?</p>",
      "rawMarkdown": "Hi guys,\nI was wondering what a reasonable vocabulary embedding dimension, depending on the vocabulary size, is:\n\nTo be specific, I'm referring to this line in a typical LSTM or GRU decoder:\n`self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)`\nwhich is then later used to embed the integer representation of the vocabulary.\n\nThe [kernel](https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92) I am adapting had an embedding_dim equal to the decoder dimension (meaning the dimension of the cell/hidden unit) of 512, now reduced to 256 in the latest version. \n\nI have been using 512 until now, however, from my mathematical intuition, an embedding_dim larger than the vocabulary size should give no benefit, because in the 'equal' case all elements of the vocabulary could already be embedded orthogonally if needed, and all extra dimensions should have no impact. \n\nA test with embedding_dim = vocab_size over 2 epochs seems to support this, also giving a slight speed up and less decoder parameters.\n\nI read that in a 'typical' text generation problem the embedding dim is much smaller than the vocabulary, with vocabularies containing thousands of words, so much larger than a typical InChI vocabulary.\n\nCan you guys give any insight on this topic?\n",
      "votes": 6
    },
    {
      "id": 1289313,
      "postDate": "2021-04-30T20:51:40.483Z",
      "content": "<p>\"4.1 Embedding Dimensionality<br>\nWith a large vocabulary, the embedding layer may<br>\naccount for a significant fraction of the model<br>\nparameters. Historically, researchers have used<br>\n620-dimensional (Bahdanau et al., 2015) or 1024-<br>\ndimensional (Luong et al., 2015a) embeddings.<br>\nWe expected larger embeddings to result in better BLEU scores, or at least lower perplexities,<br>\nbut this wasn’t always the case. While table 1<br>\nshows that 2048-dimensional embeddings yielded<br>\nthe overall best result, they only outperformed the<br>\nsmallest 128-dimensional embeddings by a narrow yet statistically significant margin (p = 0.01),<br>\nbut took nearly twice as long to converge.\"</p>\n<p>Massive Exploration of Neural Machine Translation Architectures<br>\n<a href=\"https://www.aclweb.org/anthology/D17-1151.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/D17-1151.pdf</a></p>\n<p>google for \"exploration encoder decoder seq model experiments\" for many more papers</p>",
      "rawMarkdown": "\"4.1 Embedding Dimensionality\nWith a large vocabulary, the embedding layer may\naccount for a significant fraction of the model\nparameters. Historically, researchers have used\n620-dimensional (Bahdanau et al., 2015) or 1024-\ndimensional (Luong et al., 2015a) embeddings.\nWe expected larger embeddings to result in better BLEU scores, or at least lower perplexities,\nbut this wasn’t always the case. While table 1\nshows that 2048-dimensional embeddings yielded\nthe overall best result, they only outperformed the\nsmallest 128-dimensional embeddings by a narrow yet statistically significant margin (p = 0.01),\nbut took nearly twice as long to converge.\"\n\nMassive Exploration of Neural Machine Translation Architectures\nhttps://www.aclweb.org/anthology/D17-1151.pdf\n\ngoogle for \"exploration encoder decoder seq model experiments\" for many more papers",
      "votes": 1
    },
    {
      "id": 1288389,
      "postDate": "2021-04-30T00:52:31.860Z",
      "content": "<p>a embedding plot like tSNE (over different embedding dim used) will show the differences</p>",
      "rawMarkdown": "a embedding plot like tSNE (over different embedding dim used) will show the differences",
      "votes": 1
    },
    {
      "id": 1310638,
      "postDate": "2021-05-16T19:23:18.953Z",
      "content": "<p>For those using transformers, what dimension are you using for the vocab?</p>",
      "rawMarkdown": "For those using transformers, what dimension are you using for the vocab?"
    },
    {
      "id": 1288425,
      "postDate": "2021-04-30T02:01:14.223Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1288770,
      "author_name": "Gabriel Lindenmaier",
      "author_url": "",
      "post_date": "2021-04-30T10:15:05.180000",
      "content": "<p>The embeddings are not only there to separate tokens from each other. They are feature vectors. While these features describe the relations to other tokens, the benefit of a higher embedding size is influenced by (at least) five things: Amount of data, sequence length, complexity of relations between tokens, vocabulary size, model architecture. In about that order of decreasing importance  (sequence length also induces complexity).</p>\n<p>So you can benefit from an embedding size far higher than your vocabulary. A good example is the <a href=\"https://paperswithcode.com/sota/language-modelling-on-enwiki8\" target=\"_blank\">enwik8 benchmark</a> which has a vocabulary size of ~200 and usually uses sequence lengths of 8000. You will find, that the top models (which are not pretrained on other data) often use a embedding size of 1024.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1288810,
          "author_name": "Michael Wolff",
          "author_url": "",
          "post_date": "2021-04-30T11:04:00.617000",
          "content": "<p>Thanks a lot for your explanation!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1289313,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-30T20:51:40.483000",
      "content": "<p>\"4.1 Embedding Dimensionality<br>\nWith a large vocabulary, the embedding layer may<br>\naccount for a significant fraction of the model<br>\nparameters. Historically, researchers have used<br>\n620-dimensional (Bahdanau et al., 2015) or 1024-<br>\ndimensional (Luong et al., 2015a) embeddings.<br>\nWe expected larger embeddings to result in better BLEU scores, or at least lower perplexities,<br>\nbut this wasn’t always the case. While table 1<br>\nshows that 2048-dimensional embeddings yielded<br>\nthe overall best result, they only outperformed the<br>\nsmallest 128-dimensional embeddings by a narrow yet statistically significant margin (p = 0.01),<br>\nbut took nearly twice as long to converge.\"</p>\n<p>Massive Exploration of Neural Machine Translation Architectures<br>\n<a href=\"https://www.aclweb.org/anthology/D17-1151.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/D17-1151.pdf</a></p>\n<p>google for \"exploration encoder decoder seq model experiments\" for many more papers</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1288389,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-30T00:52:31.860000",
      "content": "<p>a embedding plot like tSNE (over different embedding dim used) will show the differences</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1310638,
      "author_name": "Andrew Shao",
      "author_url": "",
      "post_date": "2021-05-16T19:23:18.953000",
      "content": "<p>For those using transformers, what dimension are you using for the vocab?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1288425,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-30T02:01:14.223000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1288770": "The embeddings are not only there to separate tokens from each other. They are feature vectors. While these features describe the relations to other tokens, the benefit of a higher embedding size is influenced by (at least) five things: Amount of data, sequence length, complexity of relations between tokens, vocabulary size, model architecture. In about that order of decreasing importance  (sequence length also induces complexity).\n\nSo you can benefit from an embedding size far higher than your vocabulary. A good example is the [enwik8 benchmark](https://paperswithcode.com/sota/language-modelling-on-enwiki8) which has a vocabulary size of ~200 and usually uses sequence lengths of 8000. You will find, that the top models (which are not pretrained on other data) often use a embedding size of 1024.",
    "1288342": "Hi guys,\nI was wondering what a reasonable vocabulary embedding dimension, depending on the vocabulary size, is:\n\nTo be specific, I'm referring to this line in a typical LSTM or GRU decoder:\n`self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)`\nwhich is then later used to embed the integer representation of the vocabulary.\n\nThe [kernel](https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92) I am adapting had an embedding_dim equal to the decoder dimension (meaning the dimension of the cell/hidden unit) of 512, now reduced to 256 in the latest version. \n\nI have been using 512 until now, however, from my mathematical intuition, an embedding_dim larger than the vocabulary size should give no benefit, because in the 'equal' case all elements of the vocabulary could already be embedded orthogonally if needed, and all extra dimensions should have no impact. \n\nA test with embedding_dim = vocab_size over 2 epochs seems to support this, also giving a slight speed up and less decoder parameters.\n\nI read that in a 'typical' text generation problem the embedding dim is much smaller than the vocabulary, with vocabularies containing thousands of words, so much larger than a typical InChI vocabulary.\n\nCan you guys give any insight on this topic?\n",
    "1289313": "\"4.1 Embedding Dimensionality\nWith a large vocabulary, the embedding layer may\naccount for a significant fraction of the model\nparameters. Historically, researchers have used\n620-dimensional (Bahdanau et al., 2015) or 1024-\ndimensional (Luong et al., 2015a) embeddings.\nWe expected larger embeddings to result in better BLEU scores, or at least lower perplexities,\nbut this wasn’t always the case. While table 1\nshows that 2048-dimensional embeddings yielded\nthe overall best result, they only outperformed the\nsmallest 128-dimensional embeddings by a narrow yet statistically significant margin (p = 0.01),\nbut took nearly twice as long to converge.\"\n\nMassive Exploration of Neural Machine Translation Architectures\nhttps://www.aclweb.org/anthology/D17-1151.pdf\n\ngoogle for \"exploration encoder decoder seq model experiments\" for many more papers",
    "1288389": "a embedding plot like tSNE (over different embedding dim used) will show the differences",
    "1310638": "For those using transformers, what dimension are you using for the vocab?",
    "1288425": ""
  }
}