{
  "id": 78457,
  "title": "Replacing bad words",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78457",
  "author_name": "",
  "post_date": "2019-01-24T03:19:40.218983100Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Is it a good idea to replace all bad words will a single one?\n- I found a list of bad words from facebook I used it to replace bad words here to improve my model. but didn't see much improvement, instead, it becomes slow.</p>\n\n<pre><code>    def clean_text(x):\n    x = str(x)\n    for word in bad_words:\n        if word in x:\n            x = x.replace(word, \"ZZZZZZ\")\n</code></pre>",
  "messages": [
    {
      "id": "460590",
      "postDate": "01/24/2019 03:19:40",
      "content": "<p>Is it a good idea to replace all bad words will a single one?\n- I found a list of bad words from facebook I used it to replace bad words here to improve my model. but didn't see much improvement, instead, it becomes slow.</p>\n\n<pre><code>    def clean_text(x):\n    x = str(x)\n    for word in bad_words:\n        if word in x:\n            x = x.replace(word, \"ZZZZZZ\")\n</code></pre>",
      "rawMarkdown": "Is it a good idea to replace all bad words will a single one?\n- I found a list of bad words from facebook I used it to replace bad words here to improve my model. but didn't see much improvement, instead, it becomes slow.\n\n        def clean_text(x):\n        x = str(x)\n        for word in bad_words:\n            if word in x:\n                x = x.replace(word, \"ZZZZZZ\")",
      "votes": null
    },
    {
      "id": "460634",
      "postDate": "01/24/2019 05:35:26",
      "content": "<p>I don't think it's a good idea, because you may lose information about these bad words and all you got is noise (\"ZZZZZZ\").</p>",
      "rawMarkdown": "I don't think it's a good idea, because you may lose information about these bad words and all you got is noise (\"ZZZZZZ\").",
      "votes": null
    },
    {
      "id": "460711",
      "postDate": "01/24/2019 09:28:11",
      "content": "<p>Of course, it's slow! I saw the same (wrong) code in kernels. If you have, say, 100 words then you would call x.replace() 100 times for each of 1M strings. You can do replacement only once per string if you use your mind :)</p>",
      "rawMarkdown": "Of course, it's slow! I saw the same (wrong) code in kernels. If you have, say, 100 words then you would call x.replace() 100 times for each of 1M strings. You can do replacement only once per string if you use your mind :)",
      "votes": null
    },
    {
      "id": "460747",
      "postDate": "01/24/2019 10:52:55",
      "content": "<p>@JM100 I am expecting the questions having <code>ZZZZZZ</code> are highly likely to be insincere as they may not suit the guidelines as per this topic <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691</a> . So it might be a good feature for classification. The goal is to have the best performing model. If it improves your score then go ahead and replace but make sure you are not imputing an important feature by mistake as it might not work because of obvious reasons.</p>",
      "rawMarkdown": "JM100 I am expecting the questions having `ZZZZZZ` are highly likely to be insincere as they may not suit the guidelines as per this topic https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691 . So it might be a good feature for classification. The goal is to have the best performing model. If it improves your score then go ahead and replace but make sure you are not imputing an important feature by mistake as it might not work because of obvious reasons.",
      "votes": null
    },
    {
      "id": "460871",
      "postDate": "01/24/2019 15:21:28",
      "content": "<p>Thanks for the reply :-)</p>",
      "rawMarkdown": "Thanks for the reply :-)",
      "votes": null
    },
    {
      "id": "460874",
      "postDate": "01/24/2019 15:25:48",
      "content": "<p>First: Thanks for the reply :-)</p>\n\n<ul>\n<li>The author of this kernel <a href=\"https://www.kaggle.com/syhens/speed-up-your-preprocessing\">https://www.kaggle.com/syhens/speed-up-your-preprocessing</a>  where he Proves that we \"do not create a new string object if you can use <strong>in</strong> operation in python.\"  to reduce calculation time.\nthe problem that I was having is just if this idea of replacing could be reasonable or not</li>\n</ul>",
      "rawMarkdown": "First: Thanks for the reply :-)\n\n- The author of this kernel https://www.kaggle.com/syhens/speed-up-your-preprocessing  where he Proves that we \"do not create a new string object if you can use **in** operation in python.\"  to reduce calculation time.\nthe problem that I was having is just if this idea of replacing could be reasonable or not",
      "votes": null
    },
    {
      "id": "460875",
      "postDate": "01/24/2019 15:26:49",
      "content": "<p>Thanks for the reply :-)</p>",
      "rawMarkdown": "Thanks for the reply :-)",
      "votes": null
    },
    {
      "id": "460938",
      "postDate": "01/24/2019 19:11:42",
      "content": "<p>One thing, I would like to suggest is that you try to impute using some word like 'bad' or 'idiot' only rather than 'ZZZZ'. Why? Since ZZZZ has no embedding in the embedding matrix. \nOtherwise you can also try to get the average embedding of the words in your bad_words list and use that as an embedding for ZZZZ</p>",
      "rawMarkdown": "One thing, I would like to suggest is that you try to impute using some word like 'bad' or 'idiot' only rather than 'ZZZZ'. Why? Since ZZZZ has no embedding in the embedding matrix. \nOtherwise you can also try to get the average embedding of the words in your bad_words list and use that as an embedding for ZZZZ",
      "votes": null
    },
    {
      "id": "460977",
      "postDate": "01/24/2019 22:51:26",
      "content": "<p>Since you didn't reference the source of the \"bad word list\" from Facebook it's hard to know for sure, but the one I found from Facebook contains roughly 1705 words, but they aren't all profanity or swears. The list includes nearly all words, slang and phrases referencing genitals, the anus, narcotics, acts of violence, references to 'hell', references to nazis/Hitler, and various sexual acts. Also phrases like \"how to kill\" are in the list, so anyone asking about exterminating a rat, insect or other infestation will have this tag on their question. So that means that for plenty of words that aren't necessary profanity (like genuine swears) but just reference any of the above, you'll be excluding them from consideration and lumping them in with the rest. </p>\n\n<p>On Quora there are plenty of legitimate questions concerning genitals, drugs, and violent or sexual acts. Since there are lots of people asking sincere questions about drugs, sex, etc. when you lump all these terms together and exclude any hope for your model to pick up on them or their differences. Furthermore, this 'ZZZZZZ' tag will likely be associated with insincere questions from the instances where it actually does replace profanity, and depending on your model, this correlation to insincerity can leak into samples that are actually sincere just because they hold a word/phrase that's being lumped into this group. This could likely impact your accuracy, though to what extent depends on how many replacements is done by your loop above and how many of those replacements occur incorrectly. </p>\n\n<p>If I were you, I wouldn't abandon this idea, I'd continue but with an actual profanity list, not just some expansive list of terms that Facebook (known for being overzealous in its suppression of speech) disagrees with. Use a much smaller list, and maybe consider a model that can account for semantic meanings (like word embeddings or a system like SpaCy) which would already have an idea of if something is profanity or not (and this could in fact be based in-part on context, not just on text). </p>",
      "rawMarkdown": "Since you didn't reference the source of the \"bad word list\" from Facebook it's hard to know for sure, but the one I found from Facebook contains roughly 1705 words, but they aren't all profanity or swears. The list includes nearly all words, slang and phrases referencing genitals, the anus, narcotics, acts of violence, references to 'hell', references to nazis/Hitler, and various sexual acts. Also phrases like \"how to kill\" are in the list, so anyone asking about exterminating a rat, insect or other infestation will have this tag on their question. So that means that for plenty of words that aren't necessary profanity (like genuine swears) but just reference any of the above, you'll be excluding them from consideration and lumping them in with the rest. \n\nOn Quora there are plenty of legitimate questions concerning genitals, drugs, and violent or sexual acts. Since there are lots of people asking sincere questions about drugs, sex, etc. when you lump all these terms together and exclude any hope for your model to pick up on them or their differences. Furthermore, this 'ZZZZZZ' tag will likely be associated with insincere questions from the instances where it actually does replace profanity, and depending on your model, this correlation to insincerity can leak into samples that are actually sincere just because they hold a word/phrase that's being lumped into this group. This could likely impact your accuracy, though to what extent depends on how many replacements is done by your loop above and how many of those replacements occur incorrectly. \n\nIf I were you, I wouldn't abandon this idea, I'd continue but with an actual profanity list, not just some expansive list of terms that Facebook (known for being overzealous in its suppression of speech) disagrees with. Use a much smaller list, and maybe consider a model that can account for semantic meanings (like word embeddings or a system like SpaCy) which would already have an idea of if something is profanity or not (and this could in fact be based in-part on context, not just on text).",
      "votes": null
    },
    {
      "id": "460992",
      "postDate": "01/25/2019 00:21:29",
      "content": "<p>Thanks for the reply :-), I will consider your suggestions</p>",
      "rawMarkdown": "Thanks for the reply :-), I will consider your suggestions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 460634,
      "author_name": "playif1",
      "author_url": "",
      "post_date": "01/24/2019 05:35:26",
      "content": "<p>I don't think it's a good idea, because you may lose information about these bad words and all you got is noise (\"ZZZZZZ\").</p>",
      "votes": null,
      "replies": [
        {
          "id": 460871,
          "author_name": "jmourad100",
          "author_url": "",
          "post_date": "01/24/2019 15:21:28",
          "content": "<p>Thanks for the reply :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 460711,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "01/24/2019 09:28:11",
      "content": "<p>Of course, it's slow! I saw the same (wrong) code in kernels. If you have, say, 100 words then you would call x.replace() 100 times for each of 1M strings. You can do replacement only once per string if you use your mind :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 460874,
          "author_name": "jmourad100",
          "author_url": "",
          "post_date": "01/24/2019 15:25:48",
          "content": "<p>First: Thanks for the reply :-)</p>\n\n<ul>\n<li>The author of this kernel <a href=\"https://www.kaggle.com/syhens/speed-up-your-preprocessing\">https://www.kaggle.com/syhens/speed-up-your-preprocessing</a>  where he Proves that we \"do not create a new string object if you can use <strong>in</strong> operation in python.\"  to reduce calculation time.\nthe problem that I was having is just if this idea of replacing could be reasonable or not</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 460747,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "01/24/2019 10:52:55",
      "content": "<p>@JM100 I am expecting the questions having <code>ZZZZZZ</code> are highly likely to be insincere as they may not suit the guidelines as per this topic <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691</a> . So it might be a good feature for classification. The goal is to have the best performing model. If it improves your score then go ahead and replace but make sure you are not imputing an important feature by mistake as it might not work because of obvious reasons.</p>",
      "votes": null,
      "replies": [
        {
          "id": 460875,
          "author_name": "jmourad100",
          "author_url": "",
          "post_date": "01/24/2019 15:26:49",
          "content": "<p>Thanks for the reply :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 460938,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "01/24/2019 19:11:42",
      "content": "<p>One thing, I would like to suggest is that you try to impute using some word like 'bad' or 'idiot' only rather than 'ZZZZ'. Why? Since ZZZZ has no embedding in the embedding matrix. \nOtherwise you can also try to get the average embedding of the words in your bad_words list and use that as an embedding for ZZZZ</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460977,
      "author_name": "alecthekulak",
      "author_url": "",
      "post_date": "01/24/2019 22:51:26",
      "content": "<p>Since you didn't reference the source of the \"bad word list\" from Facebook it's hard to know for sure, but the one I found from Facebook contains roughly 1705 words, but they aren't all profanity or swears. The list includes nearly all words, slang and phrases referencing genitals, the anus, narcotics, acts of violence, references to 'hell', references to nazis/Hitler, and various sexual acts. Also phrases like \"how to kill\" are in the list, so anyone asking about exterminating a rat, insect or other infestation will have this tag on their question. So that means that for plenty of words that aren't necessary profanity (like genuine swears) but just reference any of the above, you'll be excluding them from consideration and lumping them in with the rest. </p>\n\n<p>On Quora there are plenty of legitimate questions concerning genitals, drugs, and violent or sexual acts. Since there are lots of people asking sincere questions about drugs, sex, etc. when you lump all these terms together and exclude any hope for your model to pick up on them or their differences. Furthermore, this 'ZZZZZZ' tag will likely be associated with insincere questions from the instances where it actually does replace profanity, and depending on your model, this correlation to insincerity can leak into samples that are actually sincere just because they hold a word/phrase that's being lumped into this group. This could likely impact your accuracy, though to what extent depends on how many replacements is done by your loop above and how many of those replacements occur incorrectly. </p>\n\n<p>If I were you, I wouldn't abandon this idea, I'd continue but with an actual profanity list, not just some expansive list of terms that Facebook (known for being overzealous in its suppression of speech) disagrees with. Use a much smaller list, and maybe consider a model that can account for semantic meanings (like word embeddings or a system like SpaCy) which would already have an idea of if something is profanity or not (and this could in fact be based in-part on context, not just on text). </p>",
      "votes": null,
      "replies": [
        {
          "id": 460992,
          "author_name": "jmourad100",
          "author_url": "",
          "post_date": "01/25/2019 00:21:29",
          "content": "<p>Thanks for the reply :-), I will consider your suggestions</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "460590": "Is it a good idea to replace all bad words will a single one?\n- I found a list of bad words from facebook I used it to replace bad words here to improve my model. but didn't see much improvement, instead, it becomes slow.\n\n        def clean_text(x):\n        x = str(x)\n        for word in bad_words:\n            if word in x:\n                x = x.replace(word, \"ZZZZZZ\")",
    "460634": "I don't think it's a good idea, because you may lose information about these bad words and all you got is noise (\"ZZZZZZ\").",
    "460711": "Of course, it's slow! I saw the same (wrong) code in kernels. If you have, say, 100 words then you would call x.replace() 100 times for each of 1M strings. You can do replacement only once per string if you use your mind :)",
    "460747": "JM100 I am expecting the questions having `ZZZZZZ` are highly likely to be insincere as they may not suit the guidelines as per this topic https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691 . So it might be a good feature for classification. The goal is to have the best performing model. If it improves your score then go ahead and replace but make sure you are not imputing an important feature by mistake as it might not work because of obvious reasons.",
    "460871": "Thanks for the reply :-)",
    "460874": "First: Thanks for the reply :-)\n\n- The author of this kernel https://www.kaggle.com/syhens/speed-up-your-preprocessing  where he Proves that we \"do not create a new string object if you can use **in** operation in python.\"  to reduce calculation time.\nthe problem that I was having is just if this idea of replacing could be reasonable or not",
    "460875": "Thanks for the reply :-)",
    "460938": "One thing, I would like to suggest is that you try to impute using some word like 'bad' or 'idiot' only rather than 'ZZZZ'. Why? Since ZZZZ has no embedding in the embedding matrix. \nOtherwise you can also try to get the average embedding of the words in your bad_words list and use that as an embedding for ZZZZ",
    "460977": "Since you didn't reference the source of the \"bad word list\" from Facebook it's hard to know for sure, but the one I found from Facebook contains roughly 1705 words, but they aren't all profanity or swears. The list includes nearly all words, slang and phrases referencing genitals, the anus, narcotics, acts of violence, references to 'hell', references to nazis/Hitler, and various sexual acts. Also phrases like \"how to kill\" are in the list, so anyone asking about exterminating a rat, insect or other infestation will have this tag on their question. So that means that for plenty of words that aren't necessary profanity (like genuine swears) but just reference any of the above, you'll be excluding them from consideration and lumping them in with the rest. \n\nOn Quora there are plenty of legitimate questions concerning genitals, drugs, and violent or sexual acts. Since there are lots of people asking sincere questions about drugs, sex, etc. when you lump all these terms together and exclude any hope for your model to pick up on them or their differences. Furthermore, this 'ZZZZZZ' tag will likely be associated with insincere questions from the instances where it actually does replace profanity, and depending on your model, this correlation to insincerity can leak into samples that are actually sincere just because they hold a word/phrase that's being lumped into this group. This could likely impact your accuracy, though to what extent depends on how many replacements is done by your loop above and how many of those replacements occur incorrectly. \n\nIf I were you, I wouldn't abandon this idea, I'd continue but with an actual profanity list, not just some expansive list of terms that Facebook (known for being overzealous in its suppression of speech) disagrees with. Use a much smaller list, and maybe consider a model that can account for semantic meanings (like word embeddings or a system like SpaCy) which would already have an idea of if something is profanity or not (and this could in fact be based in-part on context, not just on text).",
    "460992": "Thanks for the reply :-), I will consider your suggestions"
  },
  "source": "meta"
}