{
  "id": 77092,
  "title": "Question regarding embedding matrix",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77092",
  "author_name": "",
  "post_date": "2019-01-09T12:31:08.601278200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,\nWanted to ask a question regarding embedding matrix. I have seen on a few kernels that the EM is constructed using following code:</p>\n\n<pre><code>all_embs = np.stack(embeddings_index.values())\n\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\nword_index = word_index\nembedding_matrix = np.random.normal(emb_mean, emb_std, (len_voc, embed_size))\n\nfor word, i in word_index.items():\n    if i &gt;= len_voc:\n        continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: \n        embedding_matrix[i] = embedding_vector\n</code></pre>\n\n<p>The thing that's bugging me is that why is the 0 index of the matrix not given a specific value like all  zeros:</p>\n\n<pre><code>embedding_matrix[0] = np.zeros(shape=(1,embed_size))\n</code></pre>\n\n<p>Since word embedding's is a mapping to embed_size dimension of the original word. Wouldn't mapping it to 0 (instead of random normally distributed values) be better.</p>",
  "messages": [
    {
      "id": "452973",
      "postDate": "01/09/2019 12:31:08",
      "content": "<p>Hi,\nWanted to ask a question regarding embedding matrix. I have seen on a few kernels that the EM is constructed using following code:</p>\n\n<pre><code>all_embs = np.stack(embeddings_index.values())\n\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\nword_index = word_index\nembedding_matrix = np.random.normal(emb_mean, emb_std, (len_voc, embed_size))\n\nfor word, i in word_index.items():\n    if i &gt;= len_voc:\n        continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: \n        embedding_matrix[i] = embedding_vector\n</code></pre>\n\n<p>The thing that's bugging me is that why is the 0 index of the matrix not given a specific value like all  zeros:</p>\n\n<pre><code>embedding_matrix[0] = np.zeros(shape=(1,embed_size))\n</code></pre>\n\n<p>Since word embedding's is a mapping to embed_size dimension of the original word. Wouldn't mapping it to 0 (instead of random normally distributed values) be better.</p>",
      "rawMarkdown": "Hi,\nWanted to ask a question regarding embedding matrix. I have seen on a few kernels that the EM is constructed using following code:\n\n    all_embs = np.stack(embeddings_index.values())\n    \n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n    word_index = word_index\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (len_voc, embed_size))\n    \n    for word, i in word_index.items():\n        if i &gt;= len_voc:\n            continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n\n\nThe thing that's bugging me is that why is the 0 index of the matrix not given a specific value like all  zeros:\n\n    embedding_matrix[0] = np.zeros(shape=(1,embed_size))\n\nSince word embedding's is a mapping to embed_size dimension of the original word. Wouldn't mapping it to 0 (instead of random normally distributed values) be better.",
      "votes": null
    },
    {
      "id": "453016",
      "postDate": "01/09/2019 14:21:57",
      "content": "<p>Keras tokenizer doesn't use 0th index.Unless you specify it to use oov_index.So its useless.</p>",
      "rawMarkdown": "Keras tokenizer doesn't use 0th index.Unless you specify it to use oov_index.So its useless.",
      "votes": null
    },
    {
      "id": "453024",
      "postDate": "01/09/2019 14:29:58",
      "content": "<p>Acknowledged, but we do pad the sequence with 0's, so the mapping of 0 to random normally distributed values is still there.</p>",
      "rawMarkdown": "Acknowledged, but we do pad the sequence with 0's, so the mapping of 0 to random normally distributed values is still there.",
      "votes": null
    },
    {
      "id": "453033",
      "postDate": "01/09/2019 14:53:29",
      "content": "<p>I don't see any reason for it.Many are using zeros.And if masking used it will be hidden.</p>",
      "rawMarkdown": "I don't see any reason for it.Many are using zeros.And if masking used it will be hidden.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 453016,
      "author_name": "mayurnewase",
      "author_url": "",
      "post_date": "01/09/2019 14:21:57",
      "content": "<p>Keras tokenizer doesn't use 0th index.Unless you specify it to use oov_index.So its useless.</p>",
      "votes": null,
      "replies": [
        {
          "id": 453024,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "01/09/2019 14:29:58",
          "content": "<p>Acknowledged, but we do pad the sequence with 0's, so the mapping of 0 to random normally distributed values is still there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 453033,
      "author_name": "mayurnewase",
      "author_url": "",
      "post_date": "01/09/2019 14:53:29",
      "content": "<p>I don't see any reason for it.Many are using zeros.And if masking used it will be hidden.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "452973": "Hi,\nWanted to ask a question regarding embedding matrix. I have seen on a few kernels that the EM is constructed using following code:\n\n    all_embs = np.stack(embeddings_index.values())\n    \n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n    word_index = word_index\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (len_voc, embed_size))\n    \n    for word, i in word_index.items():\n        if i &gt;= len_voc:\n            continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n\n\nThe thing that's bugging me is that why is the 0 index of the matrix not given a specific value like all  zeros:\n\n    embedding_matrix[0] = np.zeros(shape=(1,embed_size))\n\nSince word embedding's is a mapping to embed_size dimension of the original word. Wouldn't mapping it to 0 (instead of random normally distributed values) be better.",
    "453016": "Keras tokenizer doesn't use 0th index.Unless you specify it to use oov_index.So its useless.",
    "453024": "Acknowledged, but we do pad the sequence with 0's, so the mapping of 0 to random normally distributed values is still there.",
    "453033": "I don't see any reason for it.Many are using zeros.And if masking used it will be hidden."
  },
  "source": "meta"
}