{
  "id": 72893,
  "title": "Should pre-trained Embedding Matrix be normalized before using it for seq-2-seq attention learning based sentiment analysis?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72893",
  "author_name": "",
  "post_date": "2018-11-28T05:47:42.087591900Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I appreciate its background... <a href=\"https://arxiv.org/pdf/1808.06305.pdf\">https://arxiv.org/pdf/1808.06305.pdf</a> In the word embedding field, it is observed that learned word vectors usually share a large mean and several dominant principal components, which prevents word embedding from being isotropic. Word vectors that are isotropically distributed (or uniformly distributed in spatial angles) can be differentiated from each other more easily.</p>\n\n<p>First, how to get its mean &amp; std? Next, should we always PVN (post processing through variance normalization) pre-trained embeddings, i.e., is it likely to help better representation &amp; impact in all kind of NLP problems?</p>",
  "messages": [
    {
      "id": "428946",
      "postDate": "11/28/2018 05:47:42",
      "content": "<p>I appreciate its background... <a href=\"https://arxiv.org/pdf/1808.06305.pdf\">https://arxiv.org/pdf/1808.06305.pdf</a> In the word embedding field, it is observed that learned word vectors usually share a large mean and several dominant principal components, which prevents word embedding from being isotropic. Word vectors that are isotropically distributed (or uniformly distributed in spatial angles) can be differentiated from each other more easily.</p>\n\n<p>First, how to get its mean &amp; std? Next, should we always PVN (post processing through variance normalization) pre-trained embeddings, i.e., is it likely to help better representation &amp; impact in all kind of NLP problems?</p>",
      "rawMarkdown": "I appreciate its background... https://arxiv.org/pdf/1808.06305.pdf In the word embedding field, it is observed that learned word vectors usually share a large mean and several dominant principal components, which prevents word embedding from being isotropic. Word vectors that are isotropically distributed (or uniformly distributed in spatial angles) can be differentiated from each other more easily.\n\nFirst, how to get its mean &amp; std? Next, should we always PVN (post processing through variance normalization) pre-trained embeddings, i.e., is it likely to help better representation &amp; impact in all kind of NLP problems?",
      "votes": null
    },
    {
      "id": "429074",
      "postDate": "11/28/2018 09:50:28",
      "content": "<p>Just try it out:</p>\n\n<pre><code>emb_mean = np.mean(embedding_matrix,axis = 0)\nemb_std = np.std(embedding_matrix, axis = 0)\nembedding_matrix = (embedding_matrix-emb_mean)/emb_std\n</code></pre>",
      "rawMarkdown": "Just try it out:\n\n    emb_mean = np.mean(embedding_matrix,axis = 0)\n    emb_std = np.std(embedding_matrix, axis = 0)\n    embedding_matrix = (embedding_matrix-emb_mean)/emb_std",
      "votes": null
    },
    {
      "id": "430008",
      "postDate": "11/29/2018 16:48:54",
      "content": "<p>Thanks Dieter!</p>\n\n<p>This doubt had come up due to the way embedding matrix was being initialized with some empirical values as in <a href=\"https://www.kaggle.com/suicaokhoailang/magic-numbers-is-all-you-need-0-692-lb\">https://www.kaggle.com/suicaokhoailang/magic-numbers-is-all-you-need-0-692-lb</a> !</p>",
      "rawMarkdown": "Thanks Dieter!\n\nThis doubt had come up due to the way embedding matrix was being initialized with some empirical values as in https://www.kaggle.com/suicaokhoailang/magic-numbers-is-all-you-need-0-692-lb !",
      "votes": null
    },
    {
      "id": "430589",
      "postDate": "11/30/2018 16:38:13",
      "content": "<p>My guess is that that is hard-coded in order to speed it up? But then it is weird that all_embs is loaded.</p>",
      "rawMarkdown": "My guess is that that is hard-coded in order to speed it up? But then it is weird that all_embs is loaded.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 429074,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "11/28/2018 09:50:28",
      "content": "<p>Just try it out:</p>\n\n<pre><code>emb_mean = np.mean(embedding_matrix,axis = 0)\nemb_std = np.std(embedding_matrix, axis = 0)\nembedding_matrix = (embedding_matrix-emb_mean)/emb_std\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 430008,
          "author_name": "suniliitb96",
          "author_url": "",
          "post_date": "11/29/2018 16:48:54",
          "content": "<p>Thanks Dieter!</p>\n\n<p>This doubt had come up due to the way embedding matrix was being initialized with some empirical values as in <a href=\"https://www.kaggle.com/suicaokhoailang/magic-numbers-is-all-you-need-0-692-lb\">https://www.kaggle.com/suicaokhoailang/magic-numbers-is-all-you-need-0-692-lb</a> !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 430589,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/30/2018 16:38:13",
          "content": "<p>My guess is that that is hard-coded in order to speed it up? But then it is weird that all_embs is loaded.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "428946": "I appreciate its background... https://arxiv.org/pdf/1808.06305.pdf In the word embedding field, it is observed that learned word vectors usually share a large mean and several dominant principal components, which prevents word embedding from being isotropic. Word vectors that are isotropically distributed (or uniformly distributed in spatial angles) can be differentiated from each other more easily.\n\nFirst, how to get its mean &amp; std? Next, should we always PVN (post processing through variance normalization) pre-trained embeddings, i.e., is it likely to help better representation &amp; impact in all kind of NLP problems?",
    "429074": "Just try it out:\n\n    emb_mean = np.mean(embedding_matrix,axis = 0)\n    emb_std = np.std(embedding_matrix, axis = 0)\n    embedding_matrix = (embedding_matrix-emb_mean)/emb_std",
    "430008": "Thanks Dieter!\n\nThis doubt had come up due to the way embedding matrix was being initialized with some empirical values as in https://www.kaggle.com/suicaokhoailang/magic-numbers-is-all-you-need-0-692-lb !",
    "430589": "My guess is that that is hard-coded in order to speed it up? But then it is weird that all_embs is loaded."
  },
  "source": "meta"
}