{
  "id": 57407,
  "title": "Why the range of embedding make so much difference ?",
  "url": "/competitions/avito-demand-prediction/discussion/57407",
  "author_name": "",
  "post_date": "2018-05-23T12:15:40.898582500Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have tried many pre trained word embedding for my NN model. I find that when I initialize the embedding randomly with zero mean and 0.15 std, the training of NN goes fine. And if I initialize the embedding with a pre-trained model whose values all fall into [-1, 1], it also goes fine. But when I load the embedding where some values go beyond [-1, 1], like -12, -4, 8, 9 etc...  the training of my NN model will get stuck within one epoch. The loss is swing between a value and it is stuck.\nIf I use the model to predict, all the output will be zeros. \nSo I am confused what's wrong with the embedding initialization ? Any solutions ?</p>",
  "messages": [
    {
      "id": "332568",
      "postDate": "05/23/2018 12:15:40",
      "content": "<p>I have tried many pre trained word embedding for my NN model. I find that when I initialize the embedding randomly with zero mean and 0.15 std, the training of NN goes fine. And if I initialize the embedding with a pre-trained model whose values all fall into [-1, 1], it also goes fine. But when I load the embedding where some values go beyond [-1, 1], like -12, -4, 8, 9 etc...  the training of my NN model will get stuck within one epoch. The loss is swing between a value and it is stuck.\nIf I use the model to predict, all the output will be zeros. \nSo I am confused what's wrong with the embedding initialization ? Any solutions ?</p>",
      "rawMarkdown": "I have tried many pre trained word embedding for my NN model. I find that when I initialize the embedding randomly with zero mean and 0.15 std, the training of NN goes fine. And if I initialize the embedding with a pre-trained model whose values all fall into [-1, 1], it also goes fine. But when I load the embedding where some values go beyond [-1, 1], like -12, -4, 8, 9 etc...  the training of my NN model will get stuck within one epoch. The loss is swing between a value and it is stuck.\nIf I use the model to predict, all the output will be zeros. \nSo I am confused what's wrong with the embedding initialization ? Any solutions ?",
      "votes": null
    },
    {
      "id": "332789",
      "postDate": "05/23/2018 18:43:20",
      "content": "<p>Sounds like exploding gradient problem. Without knowing any details about the model, few ideas come to my mind: \nI would recommend scaling the initial values to between -1 and 1, if there is no any special reason to keep the original values. Generally values with zero mean and unit variance work well. \nTry different loss function. \nTry weight and/or gradient regularization.\nBatch normalization layer might help. \nSimplify the model.</p>",
      "rawMarkdown": "Sounds like exploding gradient problem. Without knowing any details about the model, few ideas come to my mind: \nI would recommend scaling the initial values to between -1 and 1, if there is no any special reason to keep the original values. Generally values with zero mean and unit variance work well. \nTry different loss function. \nTry weight and/or gradient regularization.\nBatch normalization layer might help. \nSimplify the model.",
      "votes": null
    },
    {
      "id": "332847",
      "postDate": "05/23/2018 22:11:47",
      "content": "<p>Neural nets are not very good at dealing with data input that change in scale. This can be the case for an unormalized embedding. Think of it this way. You have one word vector that has -0.3, -.9 etc... and then for the same sentence position of the next observation it is -39 56 70. How is the network supposed to handle the weights of the neurons involved in that part of the calculation ? Also depending on architecture, too high numbers in an exponential and you risk having overflow issues.</p>\n\n<p>The strategy is pretty simple. Normalize your word embeddings:</p>\n\n<pre><code>Vect = word2vec[word]\nVect = Vect/np.linalg.norm(vect)\n</code></pre>\n\n<p>In essence it is like putting every word on the same scale so that their dot product is the same as their cosine. In other words, If you had only 2 dimensions, it is the same as differencing all words through their position on the unit circle rather than the whole 2D plane which makes them a lot more comparable between them than just between neighbors.</p>\n\n<p>Pretrained embeddings are often normalized.</p>",
      "rawMarkdown": "Neural nets are not very good at dealing with data input that change in scale. This can be the case for an unormalized embedding. Think of it this way. You have one word vector that has -0.3, -.9 etc... and then for the same sentence position of the next observation it is -39 56 70. How is the network supposed to handle the weights of the neurons involved in that part of the calculation ? Also depending on architecture, too high numbers in an exponential and you risk having overflow issues.\n\nThe strategy is pretty simple. Normalize your word embeddings:\n\n    Vect = word2vec[word]\n    Vect = Vect/np.linalg.norm(vect)\n\nIn essence it is like putting every word on the same scale so that their dot product is the same as their cosine. In other words, If you had only 2 dimensions, it is the same as differencing all words through their position on the unit circle rather than the whole 2D plane which makes them a lot more comparable between them than just between neighbors.\n\nPretrained embeddings are often normalized.",
      "votes": null
    },
    {
      "id": "332877",
      "postDate": "05/24/2018 01:19:38",
      "content": "<p>the model is just textCNN and bi-GRU without any expection. </p>",
      "rawMarkdown": "the model is just textCNN and bi-GRU without any expection.",
      "votes": null
    },
    {
      "id": "332878",
      "postDate": "05/24/2018 01:19:47",
      "content": "<p>thanks.</p>",
      "rawMarkdown": "thanks.",
      "votes": null
    },
    {
      "id": "332882",
      "postDate": "05/24/2018 01:26:12",
      "content": "<p>thanks. Well , actually there are three pre trained embedding files: russian, english glove and fasttext embedding trained on the dataset whose value range cause the problem.    I am not sure whether I should norm them at the same time as you mentioned because there may be three different  types of distribution and normalizating  them at the same time is not a proper way. </p>",
      "rawMarkdown": "thanks. Well , actually there are three pre trained embedding files: russian, english glove and fasttext embedding trained on the dataset whose value range cause the problem.    I am not sure whether I should norm them at the same time as you mentioned because there may be three different  types of distribution and normalizating  them at the same time is not a proper way.",
      "votes": null
    },
    {
      "id": "333092",
      "postDate": "05/24/2018 11:37:48",
      "content": "<p>But are you using all three embeddings to train your model. As in concatenating the three embedding vectors into a single vector?  </p>\n\n<p>If you are using a single value say your third option, you can use a simple min max scaler on the vectors and try. It might be worth a shot. </p>\n\n<p>Also,  Check out Professor Andrews Deep learning specialization courses in Coursera. It has lectures that deals specifically with exploding gradients and such problems. </p>",
      "rawMarkdown": "But are you using all three embeddings to train your model. As in concatenating the three embedding vectors into a single vector?  \n\nIf you are using a single value say your third option, you can use a simple min max scaler on the vectors and try. It might be worth a shot. \n\nAlso,  Check out Professor Andrews Deep learning specialization courses in Coursera. It has lectures that deals specifically with exploding gradients and such problems.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332789,
      "author_name": "mkursula",
      "author_url": "",
      "post_date": "05/23/2018 18:43:20",
      "content": "<p>Sounds like exploding gradient problem. Without knowing any details about the model, few ideas come to my mind: \nI would recommend scaling the initial values to between -1 and 1, if there is no any special reason to keep the original values. Generally values with zero mean and unit variance work well. \nTry different loss function. \nTry weight and/or gradient regularization.\nBatch normalization layer might help. \nSimplify the model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332877,
          "author_name": "",
          "author_url": "",
          "post_date": "05/24/2018 01:19:38",
          "content": "<p>the model is just textCNN and bi-GRU without any expection. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332878,
          "author_name": "",
          "author_url": "",
          "post_date": "05/24/2018 01:19:47",
          "content": "<p>thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332847,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "05/23/2018 22:11:47",
      "content": "<p>Neural nets are not very good at dealing with data input that change in scale. This can be the case for an unormalized embedding. Think of it this way. You have one word vector that has -0.3, -.9 etc... and then for the same sentence position of the next observation it is -39 56 70. How is the network supposed to handle the weights of the neurons involved in that part of the calculation ? Also depending on architecture, too high numbers in an exponential and you risk having overflow issues.</p>\n\n<p>The strategy is pretty simple. Normalize your word embeddings:</p>\n\n<pre><code>Vect = word2vec[word]\nVect = Vect/np.linalg.norm(vect)\n</code></pre>\n\n<p>In essence it is like putting every word on the same scale so that their dot product is the same as their cosine. In other words, If you had only 2 dimensions, it is the same as differencing all words through their position on the unit circle rather than the whole 2D plane which makes them a lot more comparable between them than just between neighbors.</p>\n\n<p>Pretrained embeddings are often normalized.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332882,
          "author_name": "",
          "author_url": "",
          "post_date": "05/24/2018 01:26:12",
          "content": "<p>thanks. Well , actually there are three pre trained embedding files: russian, english glove and fasttext embedding trained on the dataset whose value range cause the problem.    I am not sure whether I should norm them at the same time as you mentioned because there may be three different  types of distribution and normalizating  them at the same time is not a proper way. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 333092,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/24/2018 11:37:48",
          "content": "<p>But are you using all three embeddings to train your model. As in concatenating the three embedding vectors into a single vector?  </p>\n\n<p>If you are using a single value say your third option, you can use a simple min max scaler on the vectors and try. It might be worth a shot. </p>\n\n<p>Also,  Check out Professor Andrews Deep learning specialization courses in Coursera. It has lectures that deals specifically with exploding gradients and such problems. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "332568": "I have tried many pre trained word embedding for my NN model. I find that when I initialize the embedding randomly with zero mean and 0.15 std, the training of NN goes fine. And if I initialize the embedding with a pre-trained model whose values all fall into [-1, 1], it also goes fine. But when I load the embedding where some values go beyond [-1, 1], like -12, -4, 8, 9 etc...  the training of my NN model will get stuck within one epoch. The loss is swing between a value and it is stuck.\nIf I use the model to predict, all the output will be zeros. \nSo I am confused what's wrong with the embedding initialization ? Any solutions ?",
    "332789": "Sounds like exploding gradient problem. Without knowing any details about the model, few ideas come to my mind: \nI would recommend scaling the initial values to between -1 and 1, if there is no any special reason to keep the original values. Generally values with zero mean and unit variance work well. \nTry different loss function. \nTry weight and/or gradient regularization.\nBatch normalization layer might help. \nSimplify the model.",
    "332847": "Neural nets are not very good at dealing with data input that change in scale. This can be the case for an unormalized embedding. Think of it this way. You have one word vector that has -0.3, -.9 etc... and then for the same sentence position of the next observation it is -39 56 70. How is the network supposed to handle the weights of the neurons involved in that part of the calculation ? Also depending on architecture, too high numbers in an exponential and you risk having overflow issues.\n\nThe strategy is pretty simple. Normalize your word embeddings:\n\n    Vect = word2vec[word]\n    Vect = Vect/np.linalg.norm(vect)\n\nIn essence it is like putting every word on the same scale so that their dot product is the same as their cosine. In other words, If you had only 2 dimensions, it is the same as differencing all words through their position on the unit circle rather than the whole 2D plane which makes them a lot more comparable between them than just between neighbors.\n\nPretrained embeddings are often normalized.",
    "332877": "the model is just textCNN and bi-GRU without any expection.",
    "332878": "thanks.",
    "332882": "thanks. Well , actually there are three pre trained embedding files: russian, english glove and fasttext embedding trained on the dataset whose value range cause the problem.    I am not sure whether I should norm them at the same time as you mentioned because there may be three different  types of distribution and normalizating  them at the same time is not a proper way.",
    "333092": "But are you using all three embeddings to train your model. As in concatenating the three embedding vectors into a single vector?  \n\nIf you are using a single value say your third option, you can use a simple min max scaler on the vectors and try. It might be worth a shot. \n\nAlso,  Check out Professor Andrews Deep learning specialization courses in Coursera. It has lectures that deals specifically with exploding gradients and such problems."
  },
  "source": "meta"
}