{
  "id": 71545,
  "title": "Removing rare words from training data",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71545",
  "author_name": "",
  "post_date": "2018-11-14T17:03:24.132125800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I found that there are 197177 words in the training set (after stemming and removing stop words), but most of them appeared only a few times in the entire training set (only 25055 of them had 10 occurrences or more). Is it a good idea to remove the rare words from the training data before implementing algorithms such as word2vec, or is it unnecessary to do so?</p>\n\n<p>If it's a good idea to remove rare words, I can see two ways to do this. One is to automatically remove all words with less than 10 (or another small number) occurrences. Another way is to remove all words if the ratio of \"insincere\" to \"sincere\" questions containing this word is statistically significant. For example, about 6.4% of the questions in the training data are insincere. In a set of 10 questions, there is a 13.1% probability that at least 2 questions are insincere (using the binomial distribution with p=0.064), but a 0.02% probability that at least 5 questions are insincere. So a word that appears in 10 questions, 2 of which are insincere, should be removed, but a word that appears in 10 questions, 5 of which are insincere, should be kept. This would retain offensive words that are infrequent but are nevertheless strong predictors that a question is insincere. Which method would be better to use?</p>",
  "messages": [
    {
      "id": "421170",
      "postDate": "11/14/2018 17:03:24",
      "content": "<p>I found that there are 197177 words in the training set (after stemming and removing stop words), but most of them appeared only a few times in the entire training set (only 25055 of them had 10 occurrences or more). Is it a good idea to remove the rare words from the training data before implementing algorithms such as word2vec, or is it unnecessary to do so?</p>\n\n<p>If it's a good idea to remove rare words, I can see two ways to do this. One is to automatically remove all words with less than 10 (or another small number) occurrences. Another way is to remove all words if the ratio of \"insincere\" to \"sincere\" questions containing this word is statistically significant. For example, about 6.4% of the questions in the training data are insincere. In a set of 10 questions, there is a 13.1% probability that at least 2 questions are insincere (using the binomial distribution with p=0.064), but a 0.02% probability that at least 5 questions are insincere. So a word that appears in 10 questions, 2 of which are insincere, should be removed, but a word that appears in 10 questions, 5 of which are insincere, should be kept. This would retain offensive words that are infrequent but are nevertheless strong predictors that a question is insincere. Which method would be better to use?</p>",
      "rawMarkdown": "I found that there are 197177 words in the training set (after stemming and removing stop words), but most of them appeared only a few times in the entire training set (only 25055 of them had 10 occurrences or more). Is it a good idea to remove the rare words from the training data before implementing algorithms such as word2vec, or is it unnecessary to do so?\n\nIf it's a good idea to remove rare words, I can see two ways to do this. One is to automatically remove all words with less than 10 (or another small number) occurrences. Another way is to remove all words if the ratio of \"insincere\" to \"sincere\" questions containing this word is statistically significant. For example, about 6.4% of the questions in the training data are insincere. In a set of 10 questions, there is a 13.1% probability that at least 2 questions are insincere (using the binomial distribution with p=0.064), but a 0.02% probability that at least 5 questions are insincere. So a word that appears in 10 questions, 2 of which are insincere, should be removed, but a word that appears in 10 questions, 5 of which are insincere, should be kept. This would retain offensive words that are infrequent but are nevertheless strong predictors that a question is insincere. Which method would be better to use?",
      "votes": null
    },
    {
      "id": "421350",
      "postDate": "11/14/2018 23:50:01",
      "content": "<p>If you use a NLP model like bag-of-words or TF-IDF, in almost all implementations of these models (in packages like nltk, keras, etc.) there's options for minimum word frequency, which is simple to do and it quite a common solution to the problem you're describing. In other words, there is precedent for only selecting words that occur in a minimum number of documents. You're second idea, of removing words based on different occurrence frequencies between different classes is implemented in the \"delta TF-IDF\" model, which has a very interesting paper, worth a read. </p>\n\n<p>Ultimately, delta TF-IDF (the idea to remove them based on there being a statistically significant difference in occurrence frequency between classes) is a newer and IMHO, better way of doing things than classic TF-IDF (at least for classification problems). This isn't to say the first idea is bad, the idea of having a minimum number of occurrences for a word to matter is nearly ubiquitous in NLP applications. There's no reason not to use both. First remove all words that occur in less than, say... 1% of the data samples, next look into ratios. I wouldn't go through all that effort using binomial distributions and the like, I'd just find some existing delta TF-IDF algorithm and implement that. </p>",
      "rawMarkdown": "If you use a NLP model like bag-of-words or TF-IDF, in almost all implementations of these models (in packages like nltk, keras, etc.) there's options for minimum word frequency, which is simple to do and it quite a common solution to the problem you're describing. In other words, there is precedent for only selecting words that occur in a minimum number of documents. You're second idea, of removing words based on different occurrence frequencies between different classes is implemented in the \"delta TF-IDF\" model, which has a very interesting paper, worth a read. \n\nUltimately, delta TF-IDF (the idea to remove them based on there being a statistically significant difference in occurrence frequency between classes) is a newer and IMHO, better way of doing things than classic TF-IDF (at least for classification problems). This isn't to say the first idea is bad, the idea of having a minimum number of occurrences for a word to matter is nearly ubiquitous in NLP applications. There's no reason not to use both. First remove all words that occur in less than, say... 1% of the data samples, next look into ratios. I wouldn't go through all that effort using binomial distributions and the like, I'd just find some existing delta TF-IDF algorithm and implement that.",
      "votes": null
    },
    {
      "id": "421810",
      "postDate": "11/15/2018 12:57:20",
      "content": "<p>If you are going to use embeddings, you might want to use Keras' Tokenizer.\nThe parameter num_words lets you control the number of words you want to keep in your vocabulary, keeping the most frequent ones.</p>",
      "rawMarkdown": "If you are going to use embeddings, you might want to use Keras' Tokenizer.\nThe parameter num_words lets you control the number of words you want to keep in your vocabulary, keeping the most frequent ones.",
      "votes": null
    },
    {
      "id": "422187",
      "postDate": "11/15/2018 22:04:57",
      "content": "<p>You could build a kernel to explore such words and ratios. Would be interesting, likely get some upvotes, and give you better insights into your question. I would be very much interested to see one myself..</p>",
      "rawMarkdown": "You could build a kernel to explore such words and ratios. Would be interesting, likely get some upvotes, and give you better insights into your question. I would be very much interested to see one myself..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 421350,
      "author_name": "alecthekulak",
      "author_url": "",
      "post_date": "11/14/2018 23:50:01",
      "content": "<p>If you use a NLP model like bag-of-words or TF-IDF, in almost all implementations of these models (in packages like nltk, keras, etc.) there's options for minimum word frequency, which is simple to do and it quite a common solution to the problem you're describing. In other words, there is precedent for only selecting words that occur in a minimum number of documents. You're second idea, of removing words based on different occurrence frequencies between different classes is implemented in the \"delta TF-IDF\" model, which has a very interesting paper, worth a read. </p>\n\n<p>Ultimately, delta TF-IDF (the idea to remove them based on there being a statistically significant difference in occurrence frequency between classes) is a newer and IMHO, better way of doing things than classic TF-IDF (at least for classification problems). This isn't to say the first idea is bad, the idea of having a minimum number of occurrences for a word to matter is nearly ubiquitous in NLP applications. There's no reason not to use both. First remove all words that occur in less than, say... 1% of the data samples, next look into ratios. I wouldn't go through all that effort using binomial distributions and the like, I'd just find some existing delta TF-IDF algorithm and implement that. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 421810,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/15/2018 12:57:20",
      "content": "<p>If you are going to use embeddings, you might want to use Keras' Tokenizer.\nThe parameter num_words lets you control the number of words you want to keep in your vocabulary, keeping the most frequent ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 422187,
      "author_name": "donkeys",
      "author_url": "",
      "post_date": "11/15/2018 22:04:57",
      "content": "<p>You could build a kernel to explore such words and ratios. Would be interesting, likely get some upvotes, and give you better insights into your question. I would be very much interested to see one myself..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "421170": "I found that there are 197177 words in the training set (after stemming and removing stop words), but most of them appeared only a few times in the entire training set (only 25055 of them had 10 occurrences or more). Is it a good idea to remove the rare words from the training data before implementing algorithms such as word2vec, or is it unnecessary to do so?\n\nIf it's a good idea to remove rare words, I can see two ways to do this. One is to automatically remove all words with less than 10 (or another small number) occurrences. Another way is to remove all words if the ratio of \"insincere\" to \"sincere\" questions containing this word is statistically significant. For example, about 6.4% of the questions in the training data are insincere. In a set of 10 questions, there is a 13.1% probability that at least 2 questions are insincere (using the binomial distribution with p=0.064), but a 0.02% probability that at least 5 questions are insincere. So a word that appears in 10 questions, 2 of which are insincere, should be removed, but a word that appears in 10 questions, 5 of which are insincere, should be kept. This would retain offensive words that are infrequent but are nevertheless strong predictors that a question is insincere. Which method would be better to use?",
    "421350": "If you use a NLP model like bag-of-words or TF-IDF, in almost all implementations of these models (in packages like nltk, keras, etc.) there's options for minimum word frequency, which is simple to do and it quite a common solution to the problem you're describing. In other words, there is precedent for only selecting words that occur in a minimum number of documents. You're second idea, of removing words based on different occurrence frequencies between different classes is implemented in the \"delta TF-IDF\" model, which has a very interesting paper, worth a read. \n\nUltimately, delta TF-IDF (the idea to remove them based on there being a statistically significant difference in occurrence frequency between classes) is a newer and IMHO, better way of doing things than classic TF-IDF (at least for classification problems). This isn't to say the first idea is bad, the idea of having a minimum number of occurrences for a word to matter is nearly ubiquitous in NLP applications. There's no reason not to use both. First remove all words that occur in less than, say... 1% of the data samples, next look into ratios. I wouldn't go through all that effort using binomial distributions and the like, I'd just find some existing delta TF-IDF algorithm and implement that.",
    "421810": "If you are going to use embeddings, you might want to use Keras' Tokenizer.\nThe parameter num_words lets you control the number of words you want to keep in your vocabulary, keeping the most frequent ones.",
    "422187": "You could build a kernel to explore such words and ratios. Would be interesting, likely get some upvotes, and give you better insights into your question. I would be very much interested to see one myself.."
  },
  "source": "meta"
}