{
  "id": 74838,
  "title": "List of Toxic Words And Other NLP Resources",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74838",
  "author_name": "",
  "post_date": "2018-12-16T13:20:37.807357200Z",
  "votes": 23,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I found decent amont of insincere questions contain pure toxicity. Recently, Kaggle hosted such competition,  <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/\">Toxic Comment Classification</a> where the toxic list was among winning recipe along with fastText trained on common crawl (<em>I miss it. Kaggle team can you consider uploading fastText trained on common crawl instead of Word2Vec?</em>).</p>\n\n<p>Here are few resources related to making meta-features, list of toxic words...etc. I will be updating them over the period of competition. </p>\n\n<h1>Resources</h1>\n\n<ol>\n<li><a href=\"https://github.com/RobertJGabriel/Google-profanity-words/blob/master/list.txt\">Full List of Bad and Swear Words Banned by Google</a></li>\n<li><a href=\"https://github.com/neptune-ml/open-solution-toxic-comments\">Open Solution of Toxic Comment Challenge</a> (lots of models, functions in there )</li>\n<li><a href=\"https://towardsdatascience.com/understanding-feature-engineering-part-3-traditional-methods-for-text-data-f6f7d70acd41\">Understanding Feature Engineering — Traditional Methods for Text Data</a></li>\n<li><a href=\"https://www.kaggle.com/tunguz/how-toxic-are-quora-questions\">How Toxic Are Quora Questions?</a>  <a href=\"/tunguz\">@tunguz</a> </li>\n</ol>",
  "messages": [
    {
      "id": "439823",
      "postDate": "12/16/2018 13:20:37",
      "content": "<p>I found decent amont of insincere questions contain pure toxicity. Recently, Kaggle hosted such competition,  <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/\">Toxic Comment Classification</a> where the toxic list was among winning recipe along with fastText trained on common crawl (<em>I miss it. Kaggle team can you consider uploading fastText trained on common crawl instead of Word2Vec?</em>).</p>\n\n<p>Here are few resources related to making meta-features, list of toxic words...etc. I will be updating them over the period of competition. </p>\n\n<h1>Resources</h1>\n\n<ol>\n<li><a href=\"https://github.com/RobertJGabriel/Google-profanity-words/blob/master/list.txt\">Full List of Bad and Swear Words Banned by Google</a></li>\n<li><a href=\"https://github.com/neptune-ml/open-solution-toxic-comments\">Open Solution of Toxic Comment Challenge</a> (lots of models, functions in there )</li>\n<li><a href=\"https://towardsdatascience.com/understanding-feature-engineering-part-3-traditional-methods-for-text-data-f6f7d70acd41\">Understanding Feature Engineering — Traditional Methods for Text Data</a></li>\n<li><a href=\"https://www.kaggle.com/tunguz/how-toxic-are-quora-questions\">How Toxic Are Quora Questions?</a>  <a href=\"/tunguz\">@tunguz</a> </li>\n</ol>",
      "rawMarkdown": "I found decent amont of insincere questions contain pure toxicity. Recently, Kaggle hosted such competition,  [Toxic Comment Classification][1] where the toxic list was among winning recipe along with fastText trained on common crawl (_I miss it. Kaggle team can you consider uploading fastText trained on common crawl instead of Word2Vec?_).\n\nHere are few resources related to making meta-features, list of toxic words...etc. I will be updating them over the period of competition. \n\n\n# Resources\n\n1.  [Full List of Bad and Swear Words Banned by Google][4]\n2. [Open Solution of Toxic Comment Challenge][3] (lots of models, functions in there )\n3. [Understanding Feature Engineering — Traditional Methods for Text Data][5]\n4.  [How Toxic Are Quora Questions?][2]  @tunguz \n\n\n  [1]: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/\n  [2]: https://www.kaggle.com/tunguz/how-toxic-are-quora-questions\n  [3]: https://github.com/neptune-ml/open-solution-toxic-comments\n  [4]: https://github.com/RobertJGabriel/Google-profanity-words/blob/master/list.txt\n  [5]: https://towardsdatascience.com/understanding-feature-engineering-part-3-traditional-methods-for-text-data-f6f7d70acd41",
      "votes": null
    },
    {
      "id": "439945",
      "postDate": "12/16/2018 17:51:06",
      "content": "<p>Hi, Shahebaz. Once again a very good job!\nThe feature engineering links and previous solution might be especially helpful for those which are not so much into Deep Learning. But I would say, the \"bad / swear word list\" is not allowed according to the competition rules. Because it is external data. But If one wants to risk using it, I would recommend to take the top x words with the best normalized insincere / sincere occurrence ratio. That gives a pretty good correlation as I have found out - before I remembered that external data isn't allowed.</p>",
      "rawMarkdown": "Hi, Shahebaz. Once again a very good job!\nThe feature engineering links and previous solution might be especially helpful for those which are not so much into Deep Learning. But I would say, the \"bad / swear word list\" is not allowed according to the competition rules. Because it is external data. But If one wants to risk using it, I would recommend to take the top x words with the best normalized insincere / sincere occurrence ratio. That gives a pretty good correlation as I have found out - before I remembered that external data isn't allowed.",
      "votes": null
    },
    {
      "id": "439956",
      "postDate": "12/16/2018 18:03:09",
      "content": "<p>I second you. We are using contractions_mapping and extracted punctuations. And these <strong>are allowed for the competition</strong> as these are kernel generated lists. But, using direct external list as such this is <strong>out of rules and not allowed</strong> (IMO).</p>\n\n<p>&gt; <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/70715#417913\">Yes. You can utilize any data that is generated during the Kernel run.</a></p>\n\n<p>To get a way around this. You can extract similar words that begin with \"bad_word\" extract its embedding and take top n similar words and make a list. Thus, that becomes kernel generated using given data. Using the list linked in this thread, teams can test whether toxic words features works? And if works gives how much F1 score boost? And figure out features that can be generated within dataset and in accordance to rules. Thanks :) </p>",
      "rawMarkdown": "I second you. We are using contractions_mapping and extracted punctuations. And these **are allowed for the competition** as these are kernel generated lists. But, using direct external list as such this is **out of rules and not allowed** (IMO).\n\n&gt; [Yes. You can utilize any data that is generated during the Kernel run.][1]\n\n\n\nTo get a way around this. You can extract similar words that begin with \"bad_word\" extract its embedding and take top n similar words and make a list. Thus, that becomes kernel generated using given data. Using the list linked in this thread, teams can test whether toxic words features works? And if works gives how much F1 score boost? And figure out features that can be generated within dataset and in accordance to rules. Thanks :) \n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/70715#417913",
      "votes": null
    },
    {
      "id": "439974",
      "postDate": "12/16/2018 18:54:34",
      "content": "<p>I agree. Although I read the \"utilize any data [..] generated during Kernel run\" as: \"You are not allowed to save this as result from another Kernel but have to calculate it live.\" So only features which you can calculate within a feasible time frame are a good idea in the first place. I actually asked a question in this direction but it got never answered.\nSo the \"bad word\" approach still wouldn't be allowed. Because you have to have a ground truth somewhere. The only way out of it would therefore probably be a sentiment analysis for words only. SpaCy should have one. But this however will take a long time to compute.</p>",
      "rawMarkdown": "I agree. Although I read the \"utilize any data [..] generated during Kernel run\" as: \"You are not allowed to save this as result from another Kernel but have to calculate it live.\" So only features which you can calculate within a feasible time frame are a good idea in the first place. I actually asked a question in this direction but it got never answered.\nSo the \"bad word\" approach still wouldn't be allowed. Because you have to have a ground truth somewhere. The only way out of it would therefore probably be a sentiment analysis for words only. SpaCy should have one. But this however will take a long time to compute.",
      "votes": null
    },
    {
      "id": "3146987",
      "postDate": "03/11/2025 14:30:30",
      "content": "<p>any update for this dataset</p>",
      "rawMarkdown": "any update for this dataset",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3146987,
      "author_name": "borhaneddineboudiaf",
      "author_url": "",
      "post_date": "03/11/2025 14:30:30",
      "content": "<p>any update for this dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 439945,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "12/16/2018 17:51:06",
      "content": "<p>Hi, Shahebaz. Once again a very good job!\nThe feature engineering links and previous solution might be especially helpful for those which are not so much into Deep Learning. But I would say, the \"bad / swear word list\" is not allowed according to the competition rules. Because it is external data. But If one wants to risk using it, I would recommend to take the top x words with the best normalized insincere / sincere occurrence ratio. That gives a pretty good correlation as I have found out - before I remembered that external data isn't allowed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 439956,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/16/2018 18:03:09",
          "content": "<p>I second you. We are using contractions_mapping and extracted punctuations. And these <strong>are allowed for the competition</strong> as these are kernel generated lists. But, using direct external list as such this is <strong>out of rules and not allowed</strong> (IMO).</p>\n\n<p>&gt; <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/70715#417913\">Yes. You can utilize any data that is generated during the Kernel run.</a></p>\n\n<p>To get a way around this. You can extract similar words that begin with \"bad_word\" extract its embedding and take top n similar words and make a list. Thus, that becomes kernel generated using given data. Using the list linked in this thread, teams can test whether toxic words features works? And if works gives how much F1 score boost? And figure out features that can be generated within dataset and in accordance to rules. Thanks :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439974,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "12/16/2018 18:54:34",
          "content": "<p>I agree. Although I read the \"utilize any data [..] generated during Kernel run\" as: \"You are not allowed to save this as result from another Kernel but have to calculate it live.\" So only features which you can calculate within a feasible time frame are a good idea in the first place. I actually asked a question in this direction but it got never answered.\nSo the \"bad word\" approach still wouldn't be allowed. Because you have to have a ground truth somewhere. The only way out of it would therefore probably be a sentiment analysis for words only. SpaCy should have one. But this however will take a long time to compute.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "439823": "I found decent amont of insincere questions contain pure toxicity. Recently, Kaggle hosted such competition,  [Toxic Comment Classification][1] where the toxic list was among winning recipe along with fastText trained on common crawl (_I miss it. Kaggle team can you consider uploading fastText trained on common crawl instead of Word2Vec?_).\n\nHere are few resources related to making meta-features, list of toxic words...etc. I will be updating them over the period of competition. \n\n\n# Resources\n\n1.  [Full List of Bad and Swear Words Banned by Google][4]\n2. [Open Solution of Toxic Comment Challenge][3] (lots of models, functions in there )\n3. [Understanding Feature Engineering — Traditional Methods for Text Data][5]\n4.  [How Toxic Are Quora Questions?][2]  @tunguz \n\n\n  [1]: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/\n  [2]: https://www.kaggle.com/tunguz/how-toxic-are-quora-questions\n  [3]: https://github.com/neptune-ml/open-solution-toxic-comments\n  [4]: https://github.com/RobertJGabriel/Google-profanity-words/blob/master/list.txt\n  [5]: https://towardsdatascience.com/understanding-feature-engineering-part-3-traditional-methods-for-text-data-f6f7d70acd41",
    "439945": "Hi, Shahebaz. Once again a very good job!\nThe feature engineering links and previous solution might be especially helpful for those which are not so much into Deep Learning. But I would say, the \"bad / swear word list\" is not allowed according to the competition rules. Because it is external data. But If one wants to risk using it, I would recommend to take the top x words with the best normalized insincere / sincere occurrence ratio. That gives a pretty good correlation as I have found out - before I remembered that external data isn't allowed.",
    "439956": "I second you. We are using contractions_mapping and extracted punctuations. And these **are allowed for the competition** as these are kernel generated lists. But, using direct external list as such this is **out of rules and not allowed** (IMO).\n\n&gt; [Yes. You can utilize any data that is generated during the Kernel run.][1]\n\n\n\nTo get a way around this. You can extract similar words that begin with \"bad_word\" extract its embedding and take top n similar words and make a list. Thus, that becomes kernel generated using given data. Using the list linked in this thread, teams can test whether toxic words features works? And if works gives how much F1 score boost? And figure out features that can be generated within dataset and in accordance to rules. Thanks :) \n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/70715#417913",
    "439974": "I agree. Although I read the \"utilize any data [..] generated during Kernel run\" as: \"You are not allowed to save this as result from another Kernel but have to calculate it live.\" So only features which you can calculate within a feasible time frame are a good idea in the first place. I actually asked a question in this direction but it got never answered.\nSo the \"bad word\" approach still wouldn't be allowed. Because you have to have a ground truth somewhere. The only way out of it would therefore probably be a sentiment analysis for words only. SpaCy should have one. But this however will take a long time to compute.",
    "3146987": "any update for this dataset"
  },
  "source": "meta"
}