{
  "id": 72152,
  "title": "Are you using the right tokenizer for your embeddings?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72152",
  "author_name": "",
  "post_date": "2018-11-20T21:35:29.819632600Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The Quora Insincere Question Classification competition allows us to use the four embeddings: glove.840B.300d (GloVe), paragram_300_sl999 (paragram), wiki-news-300d-1M (wiki) and GoogleNews-vectors-negative300 (GoogleNews). In the kernel \"How to: Preprocessing when Using Embeddings\", the author raises the issue of tokenization and its effect on how much of the training vocabulary is covered by words in an embedding. The author uses Google news embeddings to illustrate this point. What about the other embeddings? Which tokenization techniques maximize the number of training vocabulary words covered by the embedding? </p>\n\n<p>Why is this important? Our classifier will have a hard time learning proper classification features if   our embedding misses a significant portion of the vocabulary.</p>\n\n<p>I explore the remaining three embeddings in this kernel:  <a href=\"https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility\">Tokenization and Word Embedding Compatibility</a></p>",
  "messages": [
    {
      "id": "424908",
      "postDate": "11/20/2018 21:35:29",
      "content": "<p>The Quora Insincere Question Classification competition allows us to use the four embeddings: glove.840B.300d (GloVe), paragram_300_sl999 (paragram), wiki-news-300d-1M (wiki) and GoogleNews-vectors-negative300 (GoogleNews). In the kernel \"How to: Preprocessing when Using Embeddings\", the author raises the issue of tokenization and its effect on how much of the training vocabulary is covered by words in an embedding. The author uses Google news embeddings to illustrate this point. What about the other embeddings? Which tokenization techniques maximize the number of training vocabulary words covered by the embedding? </p>\n\n<p>Why is this important? Our classifier will have a hard time learning proper classification features if   our embedding misses a significant portion of the vocabulary.</p>\n\n<p>I explore the remaining three embeddings in this kernel:  <a href=\"https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility\">Tokenization and Word Embedding Compatibility</a></p>",
      "rawMarkdown": "The Quora Insincere Question Classification competition allows us to use the four embeddings: glove.840B.300d (GloVe), paragram_300_sl999 (paragram), wiki-news-300d-1M (wiki) and GoogleNews-vectors-negative300 (GoogleNews). In the kernel \"How to: Preprocessing when Using Embeddings\", the author raises the issue of tokenization and its effect on how much of the training vocabulary is covered by words in an embedding. The author uses Google news embeddings to illustrate this point. What about the other embeddings? Which tokenization techniques maximize the number of training vocabulary words covered by the embedding? \n\nWhy is this important? Our classifier will have a hard time learning proper classification features if   our embedding misses a significant portion of the vocabulary.\n\nI explore the remaining three embeddings in this kernel:  [Tokenization and Word Embedding Compatibility][1]\n\n\n  [1]: https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility \"Tokenization and Word Embedding Compatibility\"",
      "votes": null
    },
    {
      "id": "425050",
      "postDate": "11/21/2018 04:01:48",
      "content": "<p>can't access you link~</p>",
      "rawMarkdown": "can't access you link~",
      "votes": null
    },
    {
      "id": "425393",
      "postDate": "11/21/2018 14:51:39",
      "content": "<p>Thanks for letting me know. Not sure what happened there!  It should work now. But just to be sure here is the link: <a href=\"https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility\">https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility</a></p>",
      "rawMarkdown": "Thanks for letting me know. Not sure what happened there!  It should work now. But just to be sure here is the link: https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility",
      "votes": null
    },
    {
      "id": "425655",
      "postDate": "11/21/2018 23:01:29",
      "content": "<p>Thanks for sharing.  I tried this but didn't find it increase accuracy. It is same in this kernel: <a href=\"https://www.kaggle.com/theoviel/should-you-clean-your-data\">https://www.kaggle.com/theoviel/should-you-clean-your-data</a></p>\n\n<p>Have you tried it into the model?</p>",
      "rawMarkdown": "Thanks for sharing.  I tried this but didn't find it increase accuracy. It is same in this kernel: https://www.kaggle.com/theoviel/should-you-clean-your-data\n\nHave you tried it into the model?",
      "votes": null
    },
    {
      "id": "425704",
      "postDate": "11/22/2018 01:52:43",
      "content": "<p>I haven't seen that post so I'm not sure how it differs from mine. However, I have used my findings from the Tokenization and Word Embedding Compatibility post in my classification model and found it to affect the final test score.  Paragram embeddings, for example, missed many words due to contractions. I found that expanding contractions improved the model's accuracy. And GloVe does not require converting all capitals to lower case. The embedding has both capitals and lower cases in the index. I found that converting all my training text to lower case decreased the test score. Capitalization seems to carry some context and the GloVe embeddings capture that. </p>",
      "rawMarkdown": "I haven't seen that post so I'm not sure how it differs from mine. However, I have used my findings from the Tokenization and Word Embedding Compatibility post in my classification model and found it to affect the final test score.  Paragram embeddings, for example, missed many words due to contractions. I found that expanding contractions improved the model's accuracy. And GloVe does not require converting all capitals to lower case. The embedding has both capitals and lower cases in the index. I found that converting all my training text to lower case decreased the test score. Capitalization seems to carry some context and the GloVe embeddings capture that.",
      "votes": null
    },
    {
      "id": "433101",
      "postDate": "12/04/2018 17:46:46",
      "content": "<p>I must add here that different network models are affected differently by the tokenization method used. This probably explains the different observations.</p>",
      "rawMarkdown": "I must add here that different network models are affected differently by the tokenization method used. This probably explains the different observations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 425050,
      "author_name": "fengsi",
      "author_url": "",
      "post_date": "11/21/2018 04:01:48",
      "content": "<p>can't access you link~</p>",
      "votes": null,
      "replies": [
        {
          "id": 425393,
          "author_name": "alhalimi",
          "author_url": "",
          "post_date": "11/21/2018 14:51:39",
          "content": "<p>Thanks for letting me know. Not sure what happened there!  It should work now. But just to be sure here is the link: <a href=\"https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility\">https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 425655,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/21/2018 23:01:29",
      "content": "<p>Thanks for sharing.  I tried this but didn't find it increase accuracy. It is same in this kernel: <a href=\"https://www.kaggle.com/theoviel/should-you-clean-your-data\">https://www.kaggle.com/theoviel/should-you-clean-your-data</a></p>\n\n<p>Have you tried it into the model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 425704,
          "author_name": "alhalimi",
          "author_url": "",
          "post_date": "11/22/2018 01:52:43",
          "content": "<p>I haven't seen that post so I'm not sure how it differs from mine. However, I have used my findings from the Tokenization and Word Embedding Compatibility post in my classification model and found it to affect the final test score.  Paragram embeddings, for example, missed many words due to contractions. I found that expanding contractions improved the model's accuracy. And GloVe does not require converting all capitals to lower case. The embedding has both capitals and lower cases in the index. I found that converting all my training text to lower case decreased the test score. Capitalization seems to carry some context and the GloVe embeddings capture that. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433101,
          "author_name": "alhalimi",
          "author_url": "",
          "post_date": "12/04/2018 17:46:46",
          "content": "<p>I must add here that different network models are affected differently by the tokenization method used. This probably explains the different observations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "424908": "The Quora Insincere Question Classification competition allows us to use the four embeddings: glove.840B.300d (GloVe), paragram_300_sl999 (paragram), wiki-news-300d-1M (wiki) and GoogleNews-vectors-negative300 (GoogleNews). In the kernel \"How to: Preprocessing when Using Embeddings\", the author raises the issue of tokenization and its effect on how much of the training vocabulary is covered by words in an embedding. The author uses Google news embeddings to illustrate this point. What about the other embeddings? Which tokenization techniques maximize the number of training vocabulary words covered by the embedding? \n\nWhy is this important? Our classifier will have a hard time learning proper classification features if   our embedding misses a significant portion of the vocabulary.\n\nI explore the remaining three embeddings in this kernel:  [Tokenization and Word Embedding Compatibility][1]\n\n\n  [1]: https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility \"Tokenization and Word Embedding Compatibility\"",
    "425050": "can't access you link~",
    "425393": "Thanks for letting me know. Not sure what happened there!  It should work now. But just to be sure here is the link: https://www.kaggle.com/alhalimi/tokenization-and-word-embedding-compatibility",
    "425655": "Thanks for sharing.  I tried this but didn't find it increase accuracy. It is same in this kernel: https://www.kaggle.com/theoviel/should-you-clean-your-data\n\nHave you tried it into the model?",
    "425704": "I haven't seen that post so I'm not sure how it differs from mine. However, I have used my findings from the Tokenization and Word Embedding Compatibility post in my classification model and found it to affect the final test score.  Paragram embeddings, for example, missed many words due to contractions. I found that expanding contractions improved the model's accuracy. And GloVe does not require converting all capitals to lower case. The embedding has both capitals and lower cases in the index. I found that converting all my training text to lower case decreased the test score. Capitalization seems to carry some context and the GloVe embeddings capture that.",
    "433101": "I must add here that different network models are affected differently by the tokenization method used. This probably explains the different observations."
  },
  "source": "meta"
}