{
  "id": 76103,
  "title": "Why you shouldn't use Word2Vec",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76103",
  "author_name": "",
  "post_date": "2018-12-29T10:47:58.425419300Z",
  "votes": 32,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Stolen from <a href=\"https://www.kaggle.com/ryches\">@ryches</a> who stole from an old post from <a href=\"https://www.kaggle.com/tunguz\">@tunguz</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/173459/6298/f81e92c0-472b-4770-b808-1abdd9376edf-original.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": "447202",
      "postDate": "12/29/2018 10:47:58",
      "content": "<p>Stolen from <a href=\"https://www.kaggle.com/ryches\">@ryches</a> who stole from an old post from <a href=\"https://www.kaggle.com/tunguz\">@tunguz</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/173459/6298/f81e92c0-472b-4770-b808-1abdd9376edf-original.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Stolen from [@ryches][1] who stole from an old post from [@tunguz][2]\n\n![enter image description here][3]\n\n\n  [1]: https://www.kaggle.com/ryches\n  [2]: https://www.kaggle.com/tunguz\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/173459/6298/f81e92c0-472b-4770-b808-1abdd9376edf-original.png",
      "votes": null
    },
    {
      "id": "447210",
      "postDate": "12/29/2018 11:21:31",
      "content": "<p>Anyway while we are on the matter of word to vec in this thread, What embedding preprocessing is working best for everyone? We have many options :</p>\n\n<ol>\n<li>Concat of all 3 embeds</li>\n<li>Mean of all 3 embeds </li>\n<li>Concat/Mean of any 2 embeds</li>\n<li>Single embed</li>\n<li>train from scratch an embedding matrix as part of the network? Time consuming?</li>\n</ol>",
      "rawMarkdown": "Anyway while we are on the matter of word to vec in this thread, What embedding preprocessing is working best for everyone? We have many options :\n\n1. Concat of all 3 embeds\n2. Mean of all 3 embeds \n3. Concat/Mean of any 2 embeds\n4. Single embed\n5. train from scratch an embedding matrix as part of the network? Time consuming?",
      "votes": null
    },
    {
      "id": "447212",
      "postDate": "12/29/2018 11:25:45",
      "content": "<p>6 . Weighted average of 2+ embeds</p>\n\n<p>7 . Attention mechanism which embed to use for each word individually (computationally expensive)</p>",
      "rawMarkdown": "6 . Weighted average of 2+ embeds\n\n7 . Attention mechanism which embed to use for each word individually (computationally expensive)",
      "votes": null
    },
    {
      "id": "447215",
      "postDate": "12/29/2018 11:31:17",
      "content": "<p>weighted embed is a nice Idea. Any idea why all the kernels are not using Fasttext embedding? </p>",
      "rawMarkdown": "weighted embed is a nice Idea. Any idea why all the kernels are not using Fasttext embedding?",
      "votes": null
    },
    {
      "id": "447221",
      "postDate": "12/29/2018 11:44:01",
      "content": "<p>@Dieter How do you decide on weights for weighted average? <a href=\"/mlwhiz\">@mlwhiz</a>: Trial and error I guess and mismatch is larger.</p>",
      "rawMarkdown": "Dieter How do you decide on weights for weighted average? @mlwhiz: Trial and error I guess and mismatch is larger.",
      "votes": null
    },
    {
      "id": "447223",
      "postDate": "12/29/2018 11:47:31",
      "content": "<p>In my experience, if one has to do weighted average between Glove and Paragram, more weightage to be given to glove. Somehow my model CV is higher with only using single glove embedding. Tanks on LB though. </p>",
      "rawMarkdown": "In my experience, if one has to do weighted average between Glove and Paragram, more weightage to be given to glove. Somehow my model CV is higher with only using single glove embedding. Tanks on LB though.",
      "votes": null
    },
    {
      "id": "447224",
      "postDate": "12/29/2018 11:51:19",
      "content": "<p>Haha! That's why I prefer Tea at new places. Less chances of overfitting and tea generalizes well :) </p>",
      "rawMarkdown": "Haha! That's why I prefer Tea at new places. Less chances of overfitting and tea generalizes well :)",
      "votes": null
    },
    {
      "id": "447226",
      "postDate": "12/29/2018 11:57:22",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I train a model which has a Dense(4) with softmax activation layer to combine embeddings and then check weights</p>",
      "rawMarkdown": "philippsinger I train a model which has a Dense(4) with softmax activation layer to combine embeddings and then check weights",
      "votes": null
    },
    {
      "id": "447229",
      "postDate": "12/29/2018 12:03:20",
      "content": "<p>@Dieter neat! Do you have that as part of your overall model or seperately? I guess one could also just do hyperparameter tuning on the weights. I have observed similar things as <a href=\"/mlwhiz\">@mlwhiz</a> though that the LB score is quite sensitive on that choice.</p>",
      "rawMarkdown": "Dieter neat! Do you have that as part of your overall model or seperately? I guess one could also just do hyperparameter tuning on the weights. I have observed similar things as @mlwhiz though that the LB score is quite sensitive on that choice.",
      "votes": null
    },
    {
      "id": "447826",
      "postDate": "12/30/2018 16:52:12",
      "content": "<p>So if you don't mind which one are you using in your current model - mean/concat/single glove. :)</p>",
      "rawMarkdown": "So if you don't mind which one are you using in your current model - mean/concat/single glove. :)",
      "votes": null
    },
    {
      "id": "447945",
      "postDate": "12/30/2018 21:59:53",
      "content": "<p>It's a tradeoff between precision and recall.</p>",
      "rawMarkdown": "It's a tradeoff between precision and recall.",
      "votes": null
    },
    {
      "id": "447948",
      "postDate": "12/30/2018 22:11:28",
      "content": "<p>dieter,\nCan I know what are the inputs to dense layer?\nMaking a dense layer 4(activation='softmax') means that the output of dense layer will contain weights of each embedding matrix?</p>\n\n<p>Thanks  </p>",
      "rawMarkdown": "dieter,\nCan I know what are the inputs to dense layer?\nMaking a dense layer 4(activation='softmax') means that the output of dense layer will contain weights of each embedding matrix?\n\nThanks",
      "votes": null
    },
    {
      "id": "448102",
      "postDate": "12/31/2018 08:46:33",
      "content": "<p>haha</p>",
      "rawMarkdown": "haha",
      "votes": null
    },
    {
      "id": "448273",
      "postDate": "12/31/2018 16:45:18",
      "content": "<p>-8. Use cosine distance between embeddings for words as a feature (how are ambiguity and insincerity related?)\n-9. Train all embeddings for a couple of epochs then look at which word vectors changed the most and use this to help create other features\n-10. Train only the embeddings for words that are out of vocab (tricky to implement)\n-11. Train all embeddings for an epoch but don't actually update them - instead measure the magnitude of the attempted gradient updates for each word. Then train all words the model most wanted to update (i.e. above some gradient threshold) and freeze everything else.\n-12. Use the average of any synonyms word vectors for words that are out of vocab.</p>",
      "rawMarkdown": "8. Use cosine distance between embeddings for words as a feature (how are ambiguity and insincerity related?)\n-9. Train all embeddings for a couple of epochs then look at which word vectors changed the most and use this to help create other features\n-10. Train only the embeddings for words that are out of vocab (tricky to implement)\n-11. Train all embeddings for an epoch but don't actually update them - instead measure the magnitude of the attempted gradient updates for each word. Then train all words the model most wanted to update (i.e. above some gradient threshold) and freeze everything else.\n-12. Use the average of any synonyms word vectors for words that are out of vocab.",
      "votes": null
    },
    {
      "id": "448563",
      "postDate": "01/01/2019 13:58:10",
      "content": "<p>In response to the comic, I think espresso and cappuccino should be the considered same thing.</p>\n\n<p>\"Why don't liberals realise that ....\" and \"Why don't feminists realise that ...\" tend to be insincere questions.</p>\n\n<p>We should focus on question form. We can do it by decontextualising every question. Replace \"liberals\" with any other group of people and the question should be judged as insincere.</p>\n\n<p>Quora should use this approach as well since they do not want to be seen as biased and their dataset may be biased.</p>",
      "rawMarkdown": "In response to the comic, I think espresso and cappuccino should be the considered same thing.\n\n\"Why don't liberals realise that ....\" and \"Why don't feminists realise that ...\" tend to be insincere questions.\n\nWe should focus on question form. We can do it by decontextualising every question. Replace \"liberals\" with any other group of people and the question should be judged as insincere.\n\nQuora should use this approach as well since they do not want to be seen as biased and their dataset may be biased.",
      "votes": null
    },
    {
      "id": "461711",
      "postDate": "01/26/2019 19:58:23",
      "content": "<p>If you take mean of 3 embeddings, are they in equal position?</p>\n\n<p>I suppose if we build embedding several times, with different random,  we will not get same result, but each time different.</p>\n\n<p>One time word x1 might locate high positive value are and x2 other side of origo on negative. On other trained embedding they could be opposite: x1 negative and x2 positive - sort of rotated \"180 degrees\". Kind of symmetric, but whole embedding (the words) rotated.</p>\n\n<p>So if 2 embeddings would place word x1 on positive and 3d as negative, how does that work for taking a mean?</p>\n\n<p>Or have people who pretrained glove, paragraph, fasttext somehow tried to make em compatible by trying to rotate words to similar areas in the space?</p>",
      "rawMarkdown": "If you take mean of 3 embeddings, are they in equal position?\n\nI suppose if we build embedding several times, with different random,  we will not get same result, but each time different.\n\nOne time word x1 might locate high positive value are and x2 other side of origo on negative. On other trained embedding they could be opposite: x1 negative and x2 positive - sort of rotated \"180 degrees\". Kind of symmetric, but whole embedding (the words) rotated.\n\nSo if 2 embeddings would place word x1 on positive and 3d as negative, how does that work for taking a mean?\n\nOr have people who pretrained glove, paragraph, fasttext somehow tried to make em compatible by trying to rotate words to similar areas in the space?",
      "votes": null
    },
    {
      "id": "687240",
      "postDate": "12/04/2019 05:34:58",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍",
      "votes": null
    },
    {
      "id": "1077490",
      "postDate": "11/13/2020 16:33:46",
      "content": "<p>Notebook : <a href=\"https://www.kaggle.com/naim99/text-clustering-with-word2vec\" target=\"_blank\">https://www.kaggle.com/naim99/text-clustering-with-word2vec</a> <br>\nMy dataset : <a href=\"https://www.kaggle.com/naim99/ts-naim-mhedhbi\" target=\"_blank\">https://www.kaggle.com/naim99/ts-naim-mhedhbi</a></p>\n<p>Check this kernel for Arabic Text Analysis . I hope you find it useful and informative.</p>",
      "rawMarkdown": "Notebook : https://www.kaggle.com/naim99/text-clustering-with-word2vec \nMy dataset : https://www.kaggle.com/naim99/ts-naim-mhedhbi\n\nCheck this kernel for Arabic Text Analysis . I hope you find it useful and informative.",
      "votes": null
    },
    {
      "id": "2055628",
      "postDate": "12/05/2022 08:46:23",
      "content": "<p>hahahahahahahah😂This is so fun!</p>",
      "rawMarkdown": "hahahahahahahah😂This is so fun!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1077490,
      "author_name": "naim99",
      "author_url": "",
      "post_date": "11/13/2020 16:33:46",
      "content": "<p>Notebook : <a href=\"https://www.kaggle.com/naim99/text-clustering-with-word2vec\" target=\"_blank\">https://www.kaggle.com/naim99/text-clustering-with-word2vec</a> <br>\nMy dataset : <a href=\"https://www.kaggle.com/naim99/ts-naim-mhedhbi\" target=\"_blank\">https://www.kaggle.com/naim99/ts-naim-mhedhbi</a></p>\n<p>Check this kernel for Arabic Text Analysis . I hope you find it useful and informative.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2055628,
      "author_name": "jiangyunyao",
      "author_url": "",
      "post_date": "12/05/2022 08:46:23",
      "content": "<p>hahahahahahahah😂This is so fun!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 447210,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "12/29/2018 11:21:31",
      "content": "<p>Anyway while we are on the matter of word to vec in this thread, What embedding preprocessing is working best for everyone? We have many options :</p>\n\n<ol>\n<li>Concat of all 3 embeds</li>\n<li>Mean of all 3 embeds </li>\n<li>Concat/Mean of any 2 embeds</li>\n<li>Single embed</li>\n<li>train from scratch an embedding matrix as part of the network? Time consuming?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 447212,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/29/2018 11:25:45",
          "content": "<p>6 . Weighted average of 2+ embeds</p>\n\n<p>7 . Attention mechanism which embed to use for each word individually (computationally expensive)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447215,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/29/2018 11:31:17",
          "content": "<p>weighted embed is a nice Idea. Any idea why all the kernels are not using Fasttext embedding? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447221,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/29/2018 11:44:01",
          "content": "<p>@Dieter How do you decide on weights for weighted average? <a href=\"/mlwhiz\">@mlwhiz</a>: Trial and error I guess and mismatch is larger.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447223,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/29/2018 11:47:31",
          "content": "<p>In my experience, if one has to do weighted average between Glove and Paragram, more weightage to be given to glove. Somehow my model CV is higher with only using single glove embedding. Tanks on LB though. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447226,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/29/2018 11:57:22",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I train a model which has a Dense(4) with softmax activation layer to combine embeddings and then check weights</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447229,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/29/2018 12:03:20",
          "content": "<p>@Dieter neat! Do you have that as part of your overall model or seperately? I guess one could also just do hyperparameter tuning on the weights. I have observed similar things as <a href=\"/mlwhiz\">@mlwhiz</a> though that the LB score is quite sensitive on that choice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447826,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/30/2018 16:52:12",
          "content": "<p>So if you don't mind which one are you using in your current model - mean/concat/single glove. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447948,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "12/30/2018 22:11:28",
          "content": "<p>dieter,\nCan I know what are the inputs to dense layer?\nMaking a dense layer 4(activation='softmax') means that the output of dense layer will contain weights of each embedding matrix?</p>\n\n<p>Thanks  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448273,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "12/31/2018 16:45:18",
          "content": "<p>-8. Use cosine distance between embeddings for words as a feature (how are ambiguity and insincerity related?)\n-9. Train all embeddings for a couple of epochs then look at which word vectors changed the most and use this to help create other features\n-10. Train only the embeddings for words that are out of vocab (tricky to implement)\n-11. Train all embeddings for an epoch but don't actually update them - instead measure the magnitude of the attempted gradient updates for each word. Then train all words the model most wanted to update (i.e. above some gradient threshold) and freeze everything else.\n-12. Use the average of any synonyms word vectors for words that are out of vocab.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461711,
          "author_name": "jannen",
          "author_url": "",
          "post_date": "01/26/2019 19:58:23",
          "content": "<p>If you take mean of 3 embeddings, are they in equal position?</p>\n\n<p>I suppose if we build embedding several times, with different random,  we will not get same result, but each time different.</p>\n\n<p>One time word x1 might locate high positive value are and x2 other side of origo on negative. On other trained embedding they could be opposite: x1 negative and x2 positive - sort of rotated \"180 degrees\". Kind of symmetric, but whole embedding (the words) rotated.</p>\n\n<p>So if 2 embeddings would place word x1 on positive and 3d as negative, how does that work for taking a mean?</p>\n\n<p>Or have people who pretrained glove, paragraph, fasttext somehow tried to make em compatible by trying to rotate words to similar areas in the space?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447224,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "12/29/2018 11:51:19",
      "content": "<p>Haha! That's why I prefer Tea at new places. Less chances of overfitting and tea generalizes well :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 447945,
      "author_name": "ma2rten",
      "author_url": "",
      "post_date": "12/30/2018 21:59:53",
      "content": "<p>It's a tradeoff between precision and recall.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 448102,
      "author_name": "kailashnath1998",
      "author_url": "",
      "post_date": "12/31/2018 08:46:33",
      "content": "<p>haha</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 448563,
      "author_name": "huikang",
      "author_url": "",
      "post_date": "01/01/2019 13:58:10",
      "content": "<p>In response to the comic, I think espresso and cappuccino should be the considered same thing.</p>\n\n<p>\"Why don't liberals realise that ....\" and \"Why don't feminists realise that ...\" tend to be insincere questions.</p>\n\n<p>We should focus on question form. We can do it by decontextualising every question. Replace \"liberals\" with any other group of people and the question should be judged as insincere.</p>\n\n<p>Quora should use this approach as well since they do not want to be seen as biased and their dataset may be biased.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 687240,
      "author_name": "kaledonec",
      "author_url": "",
      "post_date": "12/04/2019 05:34:58",
      "content": "<p>👍 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "447202": "Stolen from [@ryches][1] who stole from an old post from [@tunguz][2]\n\n![enter image description here][3]\n\n\n  [1]: https://www.kaggle.com/ryches\n  [2]: https://www.kaggle.com/tunguz\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/173459/6298/f81e92c0-472b-4770-b808-1abdd9376edf-original.png",
    "447210": "Anyway while we are on the matter of word to vec in this thread, What embedding preprocessing is working best for everyone? We have many options :\n\n1. Concat of all 3 embeds\n2. Mean of all 3 embeds \n3. Concat/Mean of any 2 embeds\n4. Single embed\n5. train from scratch an embedding matrix as part of the network? Time consuming?",
    "447212": "6 . Weighted average of 2+ embeds\n\n7 . Attention mechanism which embed to use for each word individually (computationally expensive)",
    "447215": "weighted embed is a nice Idea. Any idea why all the kernels are not using Fasttext embedding?",
    "447221": "Dieter How do you decide on weights for weighted average? @mlwhiz: Trial and error I guess and mismatch is larger.",
    "447223": "In my experience, if one has to do weighted average between Glove and Paragram, more weightage to be given to glove. Somehow my model CV is higher with only using single glove embedding. Tanks on LB though.",
    "447224": "Haha! That's why I prefer Tea at new places. Less chances of overfitting and tea generalizes well :)",
    "447226": "philippsinger I train a model which has a Dense(4) with softmax activation layer to combine embeddings and then check weights",
    "447229": "Dieter neat! Do you have that as part of your overall model or seperately? I guess one could also just do hyperparameter tuning on the weights. I have observed similar things as @mlwhiz though that the LB score is quite sensitive on that choice.",
    "447826": "So if you don't mind which one are you using in your current model - mean/concat/single glove. :)",
    "447945": "It's a tradeoff between precision and recall.",
    "447948": "dieter,\nCan I know what are the inputs to dense layer?\nMaking a dense layer 4(activation='softmax') means that the output of dense layer will contain weights of each embedding matrix?\n\nThanks",
    "448102": "haha",
    "448273": "8. Use cosine distance between embeddings for words as a feature (how are ambiguity and insincerity related?)\n-9. Train all embeddings for a couple of epochs then look at which word vectors changed the most and use this to help create other features\n-10. Train only the embeddings for words that are out of vocab (tricky to implement)\n-11. Train all embeddings for an epoch but don't actually update them - instead measure the magnitude of the attempted gradient updates for each word. Then train all words the model most wanted to update (i.e. above some gradient threshold) and freeze everything else.\n-12. Use the average of any synonyms word vectors for words that are out of vocab.",
    "448563": "In response to the comic, I think espresso and cappuccino should be the considered same thing.\n\n\"Why don't liberals realise that ....\" and \"Why don't feminists realise that ...\" tend to be insincere questions.\n\nWe should focus on question form. We can do it by decontextualising every question. Replace \"liberals\" with any other group of people and the question should be judged as insincere.\n\nQuora should use this approach as well since they do not want to be seen as biased and their dataset may be biased.",
    "461711": "If you take mean of 3 embeddings, are they in equal position?\n\nI suppose if we build embedding several times, with different random,  we will not get same result, but each time different.\n\nOne time word x1 might locate high positive value are and x2 other side of origo on negative. On other trained embedding they could be opposite: x1 negative and x2 positive - sort of rotated \"180 degrees\". Kind of symmetric, but whole embedding (the words) rotated.\n\nSo if 2 embeddings would place word x1 on positive and 3d as negative, how does that work for taking a mean?\n\nOr have people who pretrained glove, paragraph, fasttext somehow tried to make em compatible by trying to rotate words to similar areas in the space?",
    "687240": "👍",
    "1077490": "Notebook : https://www.kaggle.com/naim99/text-clustering-with-word2vec \nMy dataset : https://www.kaggle.com/naim99/ts-naim-mhedhbi\n\nCheck this kernel for Arabic Text Analysis . I hope you find it useful and informative.",
    "2055628": "hahahahahahahah😂This is so fun!"
  },
  "source": "meta"
}