{
  "id": 77075,
  "title": "Is tokenizer eating your information ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77075",
  "author_name": "",
  "post_date": "2019-01-09T09:03:20.207770200Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I can see many public kernels are using </p>\n\n<pre><code>tokenizer = Tokenizer(num_words=max_features)\n</code></pre>\n\n<p>Documentation from keras says that</p>\n\n<pre><code>keras.preprocessing.text.Tokenizer(num_words=None, filters='!\"#$%&amp;amp;()*+,-./:;&amp;lt;=&amp;gt;?@[\\]^_`{|}~ \n', lower=True, split=' ', char_level=False, oov_token=None, document_count=0)\n</code></pre>\n\n<p>It's lowering texts by default,so it can cause loss of information of starting words,or SWEARING WORDS.\nAlso it will replace all characters in \"filters\" string.So again it's removing all punctuations,which have lots of information.</p>\n\n<p>So make sure to use lower=False.</p>",
  "messages": [
    {
      "id": "452849",
      "postDate": "01/09/2019 09:03:20",
      "content": "<p>I can see many public kernels are using </p>\n\n<pre><code>tokenizer = Tokenizer(num_words=max_features)\n</code></pre>\n\n<p>Documentation from keras says that</p>\n\n<pre><code>keras.preprocessing.text.Tokenizer(num_words=None, filters='!\"#$%&amp;amp;()*+,-./:;&amp;lt;=&amp;gt;?@[\\]^_`{|}~ \n', lower=True, split=' ', char_level=False, oov_token=None, document_count=0)\n</code></pre>\n\n<p>It's lowering texts by default,so it can cause loss of information of starting words,or SWEARING WORDS.\nAlso it will replace all characters in \"filters\" string.So again it's removing all punctuations,which have lots of information.</p>\n\n<p>So make sure to use lower=False.</p>",
      "rawMarkdown": "I can see many public kernels are using \n\n    tokenizer = Tokenizer(num_words=max_features)\n\nDocumentation from keras says that\n\n    keras.preprocessing.text.Tokenizer(num_words=None, filters='!\"#$%&amp;()*+,-./:;&lt;=&gt;?@[\\]^_`{|}~ \n    ', lower=True, split=' ', char_level=False, oov_token=None, document_count=0)\n\nIt's lowering texts by default,so it can cause loss of information of starting words,or SWEARING WORDS.\nAlso it will replace all characters in \"filters\" string.So again it's removing all punctuations,which have lots of information.\n\nSo make sure to use lower=False.",
      "votes": null
    },
    {
      "id": "454128",
      "postDate": "01/11/2019 06:57:13",
      "content": "<p>I think it depends. Changing <em>lower</em>, <em>filter</em> and <em>document_count</em> actually gave me lower or roughly the same LB score, even though it improved the coverage of \"max_features\" vocabulary from Glove (possibly more prone to overfitting). It would also slow down loading fasttext embeddings due to extra work on case mapping. </p>",
      "rawMarkdown": "I think it depends. Changing *lower*, *filter* and *document\\_count* actually gave me lower or roughly the same LB score, even though it improved the coverage of \"max\\_features\" vocabulary from Glove (possibly more prone to overfitting). It would also slow down loading fasttext embeddings due to extra work on case mapping.",
      "votes": null
    },
    {
      "id": "454443",
      "postDate": "01/11/2019 16:39:47",
      "content": "<p>I think it is not a rule of thumb but just different approach. \nIf you stay with the default tokenizer use (with lower) than of course you loose some information but you will better match with embedings.\nI would rather suggest to leave the lower but the information which is lost there could be used for engineering some features.</p>",
      "rawMarkdown": "I think it is not a rule of thumb but just different approach. \nIf you stay with the default tokenizer use (with lower) than of course you loose some information but you will better match with embedings.\nI would rather suggest to leave the lower but the information which is lost there could be used for engineering some features.",
      "votes": null
    },
    {
      "id": "457574",
      "postDate": "01/17/2019 18:38:55",
      "content": "<p>I've done some experiments and TweetTokenizer from nltk.tokenize with default settings (preserve_case=True) works the best for me. \n<a href=\"https://www.nltk.org/api/nltk.tokenize.html#nltk.tokenize.casual.TweetTokenizer\">https://www.nltk.org/api/nltk.tokenize.html#nltk.tokenize.casual.TweetTokenizer</a> </p>",
      "rawMarkdown": "I've done some experiments and TweetTokenizer from nltk.tokenize with default settings (preserve_case=True) works the best for me. \nhttps://www.nltk.org/api/nltk.tokenize.html#nltk.tokenize.casual.TweetTokenizer",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454128,
      "author_name": "chengham",
      "author_url": "",
      "post_date": "01/11/2019 06:57:13",
      "content": "<p>I think it depends. Changing <em>lower</em>, <em>filter</em> and <em>document_count</em> actually gave me lower or roughly the same LB score, even though it improved the coverage of \"max_features\" vocabulary from Glove (possibly more prone to overfitting). It would also slow down loading fasttext embeddings due to extra work on case mapping. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454443,
      "author_name": "akuropatwinski",
      "author_url": "",
      "post_date": "01/11/2019 16:39:47",
      "content": "<p>I think it is not a rule of thumb but just different approach. \nIf you stay with the default tokenizer use (with lower) than of course you loose some information but you will better match with embedings.\nI would rather suggest to leave the lower but the information which is lost there could be used for engineering some features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457574,
      "author_name": "nicke1",
      "author_url": "",
      "post_date": "01/17/2019 18:38:55",
      "content": "<p>I've done some experiments and TweetTokenizer from nltk.tokenize with default settings (preserve_case=True) works the best for me. \n<a href=\"https://www.nltk.org/api/nltk.tokenize.html#nltk.tokenize.casual.TweetTokenizer\">https://www.nltk.org/api/nltk.tokenize.html#nltk.tokenize.casual.TweetTokenizer</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "452849": "I can see many public kernels are using \n\n    tokenizer = Tokenizer(num_words=max_features)\n\nDocumentation from keras says that\n\n    keras.preprocessing.text.Tokenizer(num_words=None, filters='!\"#$%&amp;()*+,-./:;&lt;=&gt;?@[\\]^_`{|}~ \n    ', lower=True, split=' ', char_level=False, oov_token=None, document_count=0)\n\nIt's lowering texts by default,so it can cause loss of information of starting words,or SWEARING WORDS.\nAlso it will replace all characters in \"filters\" string.So again it's removing all punctuations,which have lots of information.\n\nSo make sure to use lower=False.",
    "454128": "I think it depends. Changing *lower*, *filter* and *document\\_count* actually gave me lower or roughly the same LB score, even though it improved the coverage of \"max\\_features\" vocabulary from Glove (possibly more prone to overfitting). It would also slow down loading fasttext embeddings due to extra work on case mapping.",
    "454443": "I think it is not a rule of thumb but just different approach. \nIf you stay with the default tokenizer use (with lower) than of course you loose some information but you will better match with embedings.\nI would rather suggest to leave the lower but the information which is lost there could be used for engineering some features.",
    "457574": "I've done some experiments and TweetTokenizer from nltk.tokenize with default settings (preserve_case=True) works the best for me. \nhttps://www.nltk.org/api/nltk.tokenize.html#nltk.tokenize.casual.TweetTokenizer"
  },
  "source": "meta"
}