{
  "id": 77758,
  "title": "Cleaning heavily may hurts the performance!?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77758",
  "author_name": "",
  "post_date": "2019-01-16T12:17:18.195233300Z",
  "votes": 22,
  "comment_count": 22,
  "views": 0,
  "content": "<p>After reading through some positive samples and found them quite disturbing... for example:</p>\n\n<pre><code>prerict_prob: 0.83226292, target: 0, Exactly how does Rejuvalex work?\nprerict_prob: 0.00101633, target: 1, Do you know how Rejuvalex does work?\nprerict_prob: 0.0000388, target: 1, What is freight forwarder management system?\nprerict_prob: 0.00010277, target: 1, How do you unlock your phone from a Gmail account?\nprerict_prob: 0.00015185, target: 1, What is engineering in India?\nprerict_prob: 0.00015750, target: 1, What are the requirements to study and work in Canada for international students?\nprerict_prob: 0.00019671, target: 1, What is the PNP program in Canada?\n......\n</code></pre>\n\n<p>There are so many Confused questions which look sincere but lebeled as insincere. Some of these question contain Proper nouns (eg. <code>Rejuvalex</code>). Repalcing the words to their interpretation may confused the model... I am just wondering cleaning heavily may hurts the performance!?</p>",
  "messages": [
    {
      "id": "456734",
      "postDate": "01/16/2019 12:17:18",
      "content": "<p>After reading through some positive samples and found them quite disturbing... for example:</p>\n\n<pre><code>prerict_prob: 0.83226292, target: 0, Exactly how does Rejuvalex work?\nprerict_prob: 0.00101633, target: 1, Do you know how Rejuvalex does work?\nprerict_prob: 0.0000388, target: 1, What is freight forwarder management system?\nprerict_prob: 0.00010277, target: 1, How do you unlock your phone from a Gmail account?\nprerict_prob: 0.00015185, target: 1, What is engineering in India?\nprerict_prob: 0.00015750, target: 1, What are the requirements to study and work in Canada for international students?\nprerict_prob: 0.00019671, target: 1, What is the PNP program in Canada?\n......\n</code></pre>\n\n<p>There are so many Confused questions which look sincere but lebeled as insincere. Some of these question contain Proper nouns (eg. <code>Rejuvalex</code>). Repalcing the words to their interpretation may confused the model... I am just wondering cleaning heavily may hurts the performance!?</p>",
      "rawMarkdown": "After reading through some positive samples and found them quite disturbing... for example:\n\n    prerict_prob: 0.83226292, target: 0, Exactly how does Rejuvalex work?\n    prerict_prob: 0.00101633, target: 1, Do you know how Rejuvalex does work?\n    prerict_prob: 0.0000388, target: 1, What is freight forwarder management system?\n    prerict_prob: 0.00010277, target: 1, How do you unlock your phone from a Gmail account?\n    prerict_prob: 0.00015185, target: 1, What is engineering in India?\n    prerict_prob: 0.00015750, target: 1, What are the requirements to study and work in Canada for international students?\n    prerict_prob: 0.00019671, target: 1, What is the PNP program in Canada?\n    ......\n\nThere are so many Confused questions which look sincere but lebeled as insincere. Some of these question contain Proper nouns (eg. `Rejuvalex`). Repalcing the words to their interpretation may confused the model... I am just wondering cleaning heavily may hurts the performance!?",
      "votes": null
    },
    {
      "id": "456740",
      "postDate": "01/16/2019 12:29:51",
      "content": "<p>I think it might, one of the criteria Quota use for deciding if a question is valid is if everything is spelt correctly</p>\n\n<p><a href=\"https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora\">https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora</a></p>",
      "rawMarkdown": "I think it might, one of the criteria Quota use for deciding if a question is valid is if everything is spelt correctly\n\nhttps://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora",
      "votes": null
    },
    {
      "id": "456848",
      "postDate": "01/16/2019 16:46:30",
      "content": "<p>@QingLiu Have a look on this topic <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691#456365\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691#456365</a>\nFeatures which follow the guidelines would have a target value 0, 1 otherwise.</p>",
      "rawMarkdown": "QingLiu Have a look on this topic https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691#456365\nFeatures which follow the guidelines would have a target value 0, 1 otherwise.",
      "votes": null
    },
    {
      "id": "456899",
      "postDate": "01/16/2019 18:07:02",
      "content": "<p>Yes, I have found most cleaning to reduce LB score (even if sometimes it will lead to increased f1 score on a local hold-out test set). Not only this, but some of the most popular kernels are doing something completely useless: These kernels use a form of cleaning during which they replace all contractions with the multi-word form. However, if you look at the order in which they do their cleaning steps, it is usually done after they have already inserted spaces around all special characters, including apostrophes. So then when they get around to replacing contractions, there will be no replacements. To illustrate better what I mean, I'll give you an example: Their cleaning method will be searching for instances of \"haven't\" but all instances of \"haven't\" have already been changed to \"haven ' t\". Therefore, nothing will be changed. So they are basically just looping through all of the train and test data searching for something that is certainly not there.</p>\n\n<p>Take a look at this kernel: <a href=\"https://www.kaggle.com/hung96ad/pytorch-starter\">https://www.kaggle.com/hung96ad/pytorch-starter</a>. When they run the code commented with \"Clean the text\", they are inserting spaces around all punctuation. So when they run \"Clean speelings\", it doesn't do anything.</p>\n\n<p>Other things I have tried that didn't work: </p>\n\n<p>replace diacritic characters with the equivalent standard character (é to e), though I don't think I got all of the possibilities so maybe that is worth a try. I will need to take a closer look at the glove and paragram embeddings to see if they include many diacritic characters. I know Glove has at least have a few.</p>\n\n<p>Replacing digits with equivalent number of '#'</p>\n\n<p>Instead of inserting spaces around all special characters, remove them entirely. (Keras tokenizer has default setting to remove most punctuation, but it only removes the following: !\"#$%&amp;()*+,-./:;&lt;=&gt;?@[]^_`{|}~       The third place single model from the google toxic comment competition at least kept !?. (and a few other marks I can't remember) so it might be worth trying to change default tokenizer setting to not remove those marks. Glove has embeddings for some special characters. Maybe I should find out which special single characters glove has embeddings for.</p>\n\n<p>Replacing different apostrophe-like characters with true apostrophes (ie. ` becomes ' ). This helped local hold-out test score, but not LB score. Not sure if I should be doing it.</p>",
      "rawMarkdown": "Yes, I have found most cleaning to reduce LB score (even if sometimes it will lead to increased f1 score on a local hold-out test set). Not only this, but some of the most popular kernels are doing something completely useless: These kernels use a form of cleaning during which they replace all contractions with the multi-word form. However, if you look at the order in which they do their cleaning steps, it is usually done after they have already inserted spaces around all special characters, including apostrophes. So then when they get around to replacing contractions, there will be no replacements. To illustrate better what I mean, I'll give you an example: Their cleaning method will be searching for instances of \"haven't\" but all instances of \"haven't\" have already been changed to \"haven ' t\". Therefore, nothing will be changed. So they are basically just looping through all of the train and test data searching for something that is certainly not there.\n\nTake a look at this kernel: https://www.kaggle.com/hung96ad/pytorch-starter. When they run the code commented with \"Clean the text\", they are inserting spaces around all punctuation. So when they run \"Clean speelings\", it doesn't do anything.\n\nOther things I have tried that didn't work: \n\nreplace diacritic characters with the equivalent standard character (é to e), though I don't think I got all of the possibilities so maybe that is worth a try. I will need to take a closer look at the glove and paragram embeddings to see if they include many diacritic characters. I know Glove has at least have a few.\n\nReplacing digits with equivalent number of '#'\n\nInstead of inserting spaces around all special characters, remove them entirely. (Keras tokenizer has default setting to remove most punctuation, but it only removes the following: !\"#$%&amp;()*+,-./:;&lt;=&gt;?@[\\]^_`{|}~       The third place single model from the google toxic comment competition at least kept !?. (and a few other marks I can't remember) so it might be worth trying to change default tokenizer setting to not remove those marks. Glove has embeddings for some special characters. Maybe I should find out which special single characters glove has embeddings for.\n\nReplacing different apostrophe-like characters with true apostrophes (ie. ` becomes ' ). This helped local hold-out test score, but not LB score. Not sure if I should be doing it.",
      "votes": null
    },
    {
      "id": "456940",
      "postDate": "01/16/2019 19:15:50",
      "content": "<p>This is all fantastic advance, thank you!</p>",
      "rawMarkdown": "This is all fantastic advance, thank you!",
      "votes": null
    },
    {
      "id": "457177",
      "postDate": "01/17/2019 03:44:12",
      "content": "<p>nice finding, but when I run \"Clean speelings\" before \"Clean the text\" , I get worse results. This is quite wired.</p>",
      "rawMarkdown": "nice finding, but when I run \"Clean speelings\" before \"Clean the text\" , I get worse results. This is quite wired.",
      "votes": null
    },
    {
      "id": "457712",
      "postDate": "01/18/2019 01:02:21",
      "content": "<p>i just keep all tokens if you can find a embedding</p>",
      "rawMarkdown": "i just keep all tokens if you can find a embedding",
      "votes": null
    },
    {
      "id": "457772",
      "postDate": "01/18/2019 04:27:07",
      "content": "<p>Yes, I also got worse results when actually changing it out.</p>",
      "rawMarkdown": "Yes, I also got worse results when actually changing it out.",
      "votes": null
    },
    {
      "id": "457961",
      "postDate": "01/18/2019 12:33:55",
      "content": "<p>Very helpful!  I am still struggling for cleaning data and doing some bad case analyze. I just curious about the little badly labeled dataset?</p>",
      "rawMarkdown": "Very helpful!  I am still struggling for cleaning data and doing some bad case analyze. I just curious about the little badly labeled dataset?",
      "votes": null
    },
    {
      "id": "457962",
      "postDate": "01/18/2019 12:34:34",
      "content": "<p>Yeap, I have noticed that. Thanks again.</p>",
      "rawMarkdown": "Yeap, I have noticed that. Thanks again.",
      "votes": null
    },
    {
      "id": "457963",
      "postDate": "01/18/2019 12:35:19",
      "content": "<p>Thanks for sharing! I am trying for that.</p>",
      "rawMarkdown": "Thanks for sharing! I am trying for that.",
      "votes": null
    },
    {
      "id": "457968",
      "postDate": "01/18/2019 12:44:55",
      "content": "<p>Thanks for sharing. Sorry for my late reply. Anyway, did you mean extract some features based on the official guidelines and concat the features with NN extracted features, and then use Dense layer to predict? Or: Extract features valued 0/1 based on the guidelines for every words and concat with the embedding layer?\n I guess the first one, because after cleaning the length of the question has changed.</p>",
      "rawMarkdown": "Thanks for sharing. Sorry for my late reply. Anyway, did you mean extract some features based on the official guidelines and concat the features with NN extracted features, and then use Dense layer to predict? Or: Extract features valued 0/1 based on the guidelines for every words and concat with the embedding layer?\n I guess the first one, because after cleaning the length of the question has changed.",
      "votes": null
    },
    {
      "id": "457973",
      "postDate": "01/18/2019 12:56:36",
      "content": "<p>But if a word is mispelled, there will be no embedding for it. The model will receive a blank vector. </p>",
      "rawMarkdown": "But if a word is mispelled, there will be no embedding for it. The model will receive a blank vector.",
      "votes": null
    },
    {
      "id": "458040",
      "postDate": "01/18/2019 16:25:03",
      "content": "<p>These samples look like they have an incorrect target in the training dataset. Correcting these before training might help the score, or?</p>",
      "rawMarkdown": "These samples look like they have an incorrect target in the training dataset. Correcting these before training might help the score, or?",
      "votes": null
    },
    {
      "id": "458083",
      "postDate": "01/18/2019 18:29:47",
      "content": "<p><a href=\"/isikkuntay\">@isikkuntay</a> that depends a lot on your tokenizer/word to embedding set up.</p>\n\n<p>If a word isn't in your tokenizer it won't get to the embedding part. To get around that you could define a OOV token in your tokenizer to represent unknown/misspelt words. Here's a kernel where I've done that <a href=\"https://www.kaggle.com/hamishdickson/using-keras-oov-tokens\">https://www.kaggle.com/hamishdickson/using-keras-oov-tokens</a></p>\n\n<p>I guess in theory it would be good to replace misspelt words with another token, but I'm not sure how you would decide if a word is spelt correctly vs just rare.</p>\n\n<p>Then for the embedding part, if you have words which aren't in your embedding you can assign them a random vector rather than just zeros (which I think is what you're saying). If you assign every unknown word the same value then you can't really differentiate between them - if they're random at least you can do that</p>",
      "rawMarkdown": "isikkuntay that depends a lot on your tokenizer/word to embedding set up.\n\nIf a word isn't in your tokenizer it won't get to the embedding part. To get around that you could define a OOV token in your tokenizer to represent unknown/misspelt words. Here's a kernel where I've done that https://www.kaggle.com/hamishdickson/using-keras-oov-tokens\n\nI guess in theory it would be good to replace misspelt words with another token, but I'm not sure how you would decide if a word is spelt correctly vs just rare.\n\nThen for the embedding part, if you have words which aren't in your embedding you can assign them a random vector rather than just zeros (which I think is what you're saying). If you assign every unknown word the same value then you can't really differentiate between them - if they're random at least you can do that",
      "votes": null
    },
    {
      "id": "458106",
      "postDate": "01/18/2019 19:30:39",
      "content": "<p>Thanks for that reply Hamish. I was wondering how to do that (I am new). In the public kernels I saw they were assigning zero vectors, which did not feel right. I had reservation towards random vectors too though, since they could end up near other words, but that is better than zeros at least.</p>",
      "rawMarkdown": "Thanks for that reply Hamish. I was wondering how to do that (I am new). In the public kernels I saw they were assigning zero vectors, which did not feel right. I had reservation towards random vectors too though, since they could end up near other words, but that is better than zeros at least.",
      "votes": null
    },
    {
      "id": "458128",
      "postDate": "01/18/2019 20:55:03",
      "content": "<p>yeah, it feels wrong doesn't it!</p>\n\n<p>You can make your embeddings trainable, as you train your model they should be \"corrected\" and end up somewhere better - buuuuut in practice that means you're adding LOTS of new trainable parameters and you can easily end up with a worse result :(</p>",
      "rawMarkdown": "yeah, it feels wrong doesn't it!\n\nYou can make your embeddings trainable, as you train your model they should be \"corrected\" and end up somewhere better - buuuuut in practice that means you're adding LOTS of new trainable parameters and you can easily end up with a worse result :(",
      "votes": null
    },
    {
      "id": "458154",
      "postDate": "01/18/2019 22:57:17",
      "content": "<p>It depends on if testing dataset is correctly labelled or not. If we assume that they are labelled using the same method as training set, then correcting those could end up being detrimental to the score even though it would lead to a model that performed better in reality. I think Quora should have done a better job labelling the questions. Or maybe the problem lies in the fact that we only have access to the title of the question (Quora question also have a section for question details). It could be that questions that look insincere to us are not actually insincere when you have access to the question details. For example, there is one question which has only \"Hello sir?\" as its question text. This question is labelled as a sincere question. It is possible that the asker of this question inserted the rest of the question into the question details. There are a number of other questions that begin the same way, ie. \"Hello sir, among NICMAR and l&amp;T, which one do you think is better for construction management?\"</p>\n\n<p>To sum it up, I and others believe that correcting the labels will not prove to be helpful.</p>",
      "rawMarkdown": "It depends on if testing dataset is correctly labelled or not. If we assume that they are labelled using the same method as training set, then correcting those could end up being detrimental to the score even though it would lead to a model that performed better in reality. I think Quora should have done a better job labelling the questions. Or maybe the problem lies in the fact that we only have access to the title of the question (Quora question also have a section for question details). It could be that questions that look insincere to us are not actually insincere when you have access to the question details. For example, there is one question which has only \"Hello sir?\" as its question text. This question is labelled as a sincere question. It is possible that the asker of this question inserted the rest of the question into the question details. There are a number of other questions that begin the same way, ie. \"Hello sir, among NICMAR and l&amp;T, which one do you think is better for construction management?\"\n\nTo sum it up, I and others believe that correcting the labels will not prove to be helpful.",
      "votes": null
    },
    {
      "id": "458231",
      "postDate": "01/19/2019 06:13:37",
      "content": "<p>thank for your sharing!</p>",
      "rawMarkdown": "thank for your sharing!",
      "votes": null
    },
    {
      "id": "458237",
      "postDate": "01/19/2019 06:26:36",
      "content": "<p>I have similar misgivings. Now I do not understand how do I go about cleaning data.\nDon't understand if I should go ahead and correct some anomaly considering it noise in training data (as described in  problem statement) or is it the same way in the test data also</p>",
      "rawMarkdown": "I have similar misgivings. Now I do not understand how do I go about cleaning data.\nDon't understand if I should go ahead and correct some anomaly considering it noise in training data (as described in  problem statement) or is it the same way in the test data also",
      "votes": null
    },
    {
      "id": "458460",
      "postDate": "01/19/2019 17:42:55",
      "content": "<p>this could be \"noise\" added deliberately ???</p>",
      "rawMarkdown": "this could be \"noise\" added deliberately ???",
      "votes": null
    },
    {
      "id": "458544",
      "postDate": "01/19/2019 22:26:01",
      "content": "<p>For me,</p>\n\n<pre><code>def clean_text(x):\n    x = str(x)\n    for word in bad_words:\n        x = x.replace(word, \"ZZZZZZ\")\n    for punct in puncts:\n        # Add space after and before punct\n        x = x.replace(punct, f' {punct} ')\n    return x\n</code></pre>\n\n<p>I replaced some bad words before Adding space after and before punctuations.</p>",
      "rawMarkdown": "For me,\n\n    def clean_text(x):\n        x = str(x)\n        for word in bad_words:\n            x = x.replace(word, \"ZZZZZZ\")\n        for punct in puncts:\n            # Add space after and before punct\n            x = x.replace(punct, f' {punct} ')\n        return x\n\nI replaced some bad words before Adding space after and before punctuations.",
      "votes": null
    },
    {
      "id": "458659",
      "postDate": "01/20/2019 07:44:08",
      "content": "<p>This is a paper I just found related to performance effects of preprocessing data from the Google  toxic comment competition. Effects of each transformation are tested on multiple models, and oftentimes results vary between models. Still, the conclusions are pretty much exactly what others have already been observing for this competition: some things help slightly, other things are neutral, still others decrease performance.</p>\n\n<p>I think the most important quote is as follows: \"Based on the results we have, our recommendation is not to spend too much time on the transformations rather focus on the selection of the best algorithms.\"</p>\n\n<p><a href=\"https://csce.ucmss.com/cr/books/2018/LFS/CSREA2018/ICA4290.pdf\">https://csce.ucmss.com/cr/books/2018/LFS/CSREA2018/ICA4290.pdf</a></p>",
      "rawMarkdown": "This is a paper I just found related to performance effects of preprocessing data from the Google  toxic comment competition. Effects of each transformation are tested on multiple models, and oftentimes results vary between models. Still, the conclusions are pretty much exactly what others have already been observing for this competition: some things help slightly, other things are neutral, still others decrease performance.\n\nI think the most important quote is as follows: \"Based on the results we have, our recommendation is not to spend too much time on the transformations rather focus on the selection of the best algorithms.\"\n\nhttps://csce.ucmss.com/cr/books/2018/LFS/CSREA2018/ICA4290.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 456740,
      "author_name": "hamishdickson",
      "author_url": "",
      "post_date": "01/16/2019 12:29:51",
      "content": "<p>I think it might, one of the criteria Quota use for deciding if a question is valid is if everything is spelt correctly</p>\n\n<p><a href=\"https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora\">https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 457962,
          "author_name": "sunnymarkliu",
          "author_url": "",
          "post_date": "01/18/2019 12:34:34",
          "content": "<p>Yeap, I have noticed that. Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457973,
          "author_name": "isikkuntay",
          "author_url": "",
          "post_date": "01/18/2019 12:56:36",
          "content": "<p>But if a word is mispelled, there will be no embedding for it. The model will receive a blank vector. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458083,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "01/18/2019 18:29:47",
          "content": "<p><a href=\"/isikkuntay\">@isikkuntay</a> that depends a lot on your tokenizer/word to embedding set up.</p>\n\n<p>If a word isn't in your tokenizer it won't get to the embedding part. To get around that you could define a OOV token in your tokenizer to represent unknown/misspelt words. Here's a kernel where I've done that <a href=\"https://www.kaggle.com/hamishdickson/using-keras-oov-tokens\">https://www.kaggle.com/hamishdickson/using-keras-oov-tokens</a></p>\n\n<p>I guess in theory it would be good to replace misspelt words with another token, but I'm not sure how you would decide if a word is spelt correctly vs just rare.</p>\n\n<p>Then for the embedding part, if you have words which aren't in your embedding you can assign them a random vector rather than just zeros (which I think is what you're saying). If you assign every unknown word the same value then you can't really differentiate between them - if they're random at least you can do that</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458106,
          "author_name": "isikkuntay",
          "author_url": "",
          "post_date": "01/18/2019 19:30:39",
          "content": "<p>Thanks for that reply Hamish. I was wondering how to do that (I am new). In the public kernels I saw they were assigning zero vectors, which did not feel right. I had reservation towards random vectors too though, since they could end up near other words, but that is better than zeros at least.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458128,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "01/18/2019 20:55:03",
          "content": "<p>yeah, it feels wrong doesn't it!</p>\n\n<p>You can make your embeddings trainable, as you train your model they should be \"corrected\" and end up somewhere better - buuuuut in practice that means you're adding LOTS of new trainable parameters and you can easily end up with a worse result :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456848,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "01/16/2019 16:46:30",
      "content": "<p>@QingLiu Have a look on this topic <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691#456365\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691#456365</a>\nFeatures which follow the guidelines would have a target value 0, 1 otherwise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 457968,
          "author_name": "sunnymarkliu",
          "author_url": "",
          "post_date": "01/18/2019 12:44:55",
          "content": "<p>Thanks for sharing. Sorry for my late reply. Anyway, did you mean extract some features based on the official guidelines and concat the features with NN extracted features, and then use Dense layer to predict? Or: Extract features valued 0/1 based on the guidelines for every words and concat with the embedding layer?\n I guess the first one, because after cleaning the length of the question has changed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456899,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "01/16/2019 18:07:02",
      "content": "<p>Yes, I have found most cleaning to reduce LB score (even if sometimes it will lead to increased f1 score on a local hold-out test set). Not only this, but some of the most popular kernels are doing something completely useless: These kernels use a form of cleaning during which they replace all contractions with the multi-word form. However, if you look at the order in which they do their cleaning steps, it is usually done after they have already inserted spaces around all special characters, including apostrophes. So then when they get around to replacing contractions, there will be no replacements. To illustrate better what I mean, I'll give you an example: Their cleaning method will be searching for instances of \"haven't\" but all instances of \"haven't\" have already been changed to \"haven ' t\". Therefore, nothing will be changed. So they are basically just looping through all of the train and test data searching for something that is certainly not there.</p>\n\n<p>Take a look at this kernel: <a href=\"https://www.kaggle.com/hung96ad/pytorch-starter\">https://www.kaggle.com/hung96ad/pytorch-starter</a>. When they run the code commented with \"Clean the text\", they are inserting spaces around all punctuation. So when they run \"Clean speelings\", it doesn't do anything.</p>\n\n<p>Other things I have tried that didn't work: </p>\n\n<p>replace diacritic characters with the equivalent standard character (é to e), though I don't think I got all of the possibilities so maybe that is worth a try. I will need to take a closer look at the glove and paragram embeddings to see if they include many diacritic characters. I know Glove has at least have a few.</p>\n\n<p>Replacing digits with equivalent number of '#'</p>\n\n<p>Instead of inserting spaces around all special characters, remove them entirely. (Keras tokenizer has default setting to remove most punctuation, but it only removes the following: !\"#$%&amp;()*+,-./:;&lt;=&gt;?@[]^_`{|}~       The third place single model from the google toxic comment competition at least kept !?. (and a few other marks I can't remember) so it might be worth trying to change default tokenizer setting to not remove those marks. Glove has embeddings for some special characters. Maybe I should find out which special single characters glove has embeddings for.</p>\n\n<p>Replacing different apostrophe-like characters with true apostrophes (ie. ` becomes ' ). This helped local hold-out test score, but not LB score. Not sure if I should be doing it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 456940,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "01/16/2019 19:15:50",
          "content": "<p>This is all fantastic advance, thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457177,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "01/17/2019 03:44:12",
          "content": "<p>nice finding, but when I run \"Clean speelings\" before \"Clean the text\" , I get worse results. This is quite wired.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457772,
          "author_name": "julius6",
          "author_url": "",
          "post_date": "01/18/2019 04:27:07",
          "content": "<p>Yes, I also got worse results when actually changing it out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457961,
          "author_name": "sunnymarkliu",
          "author_url": "",
          "post_date": "01/18/2019 12:33:55",
          "content": "<p>Very helpful!  I am still struggling for cleaning data and doing some bad case analyze. I just curious about the little badly labeled dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458231,
          "author_name": "karwik",
          "author_url": "",
          "post_date": "01/19/2019 06:13:37",
          "content": "<p>thank for your sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 457712,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "01/18/2019 01:02:21",
      "content": "<p>i just keep all tokens if you can find a embedding</p>",
      "votes": null,
      "replies": [
        {
          "id": 457963,
          "author_name": "sunnymarkliu",
          "author_url": "",
          "post_date": "01/18/2019 12:35:19",
          "content": "<p>Thanks for sharing! I am trying for that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 458040,
      "author_name": "davero",
      "author_url": "",
      "post_date": "01/18/2019 16:25:03",
      "content": "<p>These samples look like they have an incorrect target in the training dataset. Correcting these before training might help the score, or?</p>",
      "votes": null,
      "replies": [
        {
          "id": 458154,
          "author_name": "julius6",
          "author_url": "",
          "post_date": "01/18/2019 22:57:17",
          "content": "<p>It depends on if testing dataset is correctly labelled or not. If we assume that they are labelled using the same method as training set, then correcting those could end up being detrimental to the score even though it would lead to a model that performed better in reality. I think Quora should have done a better job labelling the questions. Or maybe the problem lies in the fact that we only have access to the title of the question (Quora question also have a section for question details). It could be that questions that look insincere to us are not actually insincere when you have access to the question details. For example, there is one question which has only \"Hello sir?\" as its question text. This question is labelled as a sincere question. It is possible that the asker of this question inserted the rest of the question into the question details. There are a number of other questions that begin the same way, ie. \"Hello sir, among NICMAR and l&amp;T, which one do you think is better for construction management?\"</p>\n\n<p>To sum it up, I and others believe that correcting the labels will not prove to be helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458237,
          "author_name": "adityaguru149",
          "author_url": "",
          "post_date": "01/19/2019 06:26:36",
          "content": "<p>I have similar misgivings. Now I do not understand how do I go about cleaning data.\nDon't understand if I should go ahead and correct some anomaly considering it noise in training data (as described in  problem statement) or is it the same way in the test data also</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 458460,
      "author_name": "rahulbakshee",
      "author_url": "",
      "post_date": "01/19/2019 17:42:55",
      "content": "<p>this could be \"noise\" added deliberately ???</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458544,
      "author_name": "jmourad100",
      "author_url": "",
      "post_date": "01/19/2019 22:26:01",
      "content": "<p>For me,</p>\n\n<pre><code>def clean_text(x):\n    x = str(x)\n    for word in bad_words:\n        x = x.replace(word, \"ZZZZZZ\")\n    for punct in puncts:\n        # Add space after and before punct\n        x = x.replace(punct, f' {punct} ')\n    return x\n</code></pre>\n\n<p>I replaced some bad words before Adding space after and before punctuations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458659,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "01/20/2019 07:44:08",
      "content": "<p>This is a paper I just found related to performance effects of preprocessing data from the Google  toxic comment competition. Effects of each transformation are tested on multiple models, and oftentimes results vary between models. Still, the conclusions are pretty much exactly what others have already been observing for this competition: some things help slightly, other things are neutral, still others decrease performance.</p>\n\n<p>I think the most important quote is as follows: \"Based on the results we have, our recommendation is not to spend too much time on the transformations rather focus on the selection of the best algorithms.\"</p>\n\n<p><a href=\"https://csce.ucmss.com/cr/books/2018/LFS/CSREA2018/ICA4290.pdf\">https://csce.ucmss.com/cr/books/2018/LFS/CSREA2018/ICA4290.pdf</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "456734": "After reading through some positive samples and found them quite disturbing... for example:\n\n    prerict_prob: 0.83226292, target: 0, Exactly how does Rejuvalex work?\n    prerict_prob: 0.00101633, target: 1, Do you know how Rejuvalex does work?\n    prerict_prob: 0.0000388, target: 1, What is freight forwarder management system?\n    prerict_prob: 0.00010277, target: 1, How do you unlock your phone from a Gmail account?\n    prerict_prob: 0.00015185, target: 1, What is engineering in India?\n    prerict_prob: 0.00015750, target: 1, What are the requirements to study and work in Canada for international students?\n    prerict_prob: 0.00019671, target: 1, What is the PNP program in Canada?\n    ......\n\nThere are so many Confused questions which look sincere but lebeled as insincere. Some of these question contain Proper nouns (eg. `Rejuvalex`). Repalcing the words to their interpretation may confused the model... I am just wondering cleaning heavily may hurts the performance!?",
    "456740": "I think it might, one of the criteria Quota use for deciding if a question is valid is if everything is spelt correctly\n\nhttps://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora",
    "456848": "QingLiu Have a look on this topic https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/77691#456365\nFeatures which follow the guidelines would have a target value 0, 1 otherwise.",
    "456899": "Yes, I have found most cleaning to reduce LB score (even if sometimes it will lead to increased f1 score on a local hold-out test set). Not only this, but some of the most popular kernels are doing something completely useless: These kernels use a form of cleaning during which they replace all contractions with the multi-word form. However, if you look at the order in which they do their cleaning steps, it is usually done after they have already inserted spaces around all special characters, including apostrophes. So then when they get around to replacing contractions, there will be no replacements. To illustrate better what I mean, I'll give you an example: Their cleaning method will be searching for instances of \"haven't\" but all instances of \"haven't\" have already been changed to \"haven ' t\". Therefore, nothing will be changed. So they are basically just looping through all of the train and test data searching for something that is certainly not there.\n\nTake a look at this kernel: https://www.kaggle.com/hung96ad/pytorch-starter. When they run the code commented with \"Clean the text\", they are inserting spaces around all punctuation. So when they run \"Clean speelings\", it doesn't do anything.\n\nOther things I have tried that didn't work: \n\nreplace diacritic characters with the equivalent standard character (é to e), though I don't think I got all of the possibilities so maybe that is worth a try. I will need to take a closer look at the glove and paragram embeddings to see if they include many diacritic characters. I know Glove has at least have a few.\n\nReplacing digits with equivalent number of '#'\n\nInstead of inserting spaces around all special characters, remove them entirely. (Keras tokenizer has default setting to remove most punctuation, but it only removes the following: !\"#$%&amp;()*+,-./:;&lt;=&gt;?@[\\]^_`{|}~       The third place single model from the google toxic comment competition at least kept !?. (and a few other marks I can't remember) so it might be worth trying to change default tokenizer setting to not remove those marks. Glove has embeddings for some special characters. Maybe I should find out which special single characters glove has embeddings for.\n\nReplacing different apostrophe-like characters with true apostrophes (ie. ` becomes ' ). This helped local hold-out test score, but not LB score. Not sure if I should be doing it.",
    "456940": "This is all fantastic advance, thank you!",
    "457177": "nice finding, but when I run \"Clean speelings\" before \"Clean the text\" , I get worse results. This is quite wired.",
    "457712": "i just keep all tokens if you can find a embedding",
    "457772": "Yes, I also got worse results when actually changing it out.",
    "457961": "Very helpful!  I am still struggling for cleaning data and doing some bad case analyze. I just curious about the little badly labeled dataset?",
    "457962": "Yeap, I have noticed that. Thanks again.",
    "457963": "Thanks for sharing! I am trying for that.",
    "457968": "Thanks for sharing. Sorry for my late reply. Anyway, did you mean extract some features based on the official guidelines and concat the features with NN extracted features, and then use Dense layer to predict? Or: Extract features valued 0/1 based on the guidelines for every words and concat with the embedding layer?\n I guess the first one, because after cleaning the length of the question has changed.",
    "457973": "But if a word is mispelled, there will be no embedding for it. The model will receive a blank vector.",
    "458040": "These samples look like they have an incorrect target in the training dataset. Correcting these before training might help the score, or?",
    "458083": "isikkuntay that depends a lot on your tokenizer/word to embedding set up.\n\nIf a word isn't in your tokenizer it won't get to the embedding part. To get around that you could define a OOV token in your tokenizer to represent unknown/misspelt words. Here's a kernel where I've done that https://www.kaggle.com/hamishdickson/using-keras-oov-tokens\n\nI guess in theory it would be good to replace misspelt words with another token, but I'm not sure how you would decide if a word is spelt correctly vs just rare.\n\nThen for the embedding part, if you have words which aren't in your embedding you can assign them a random vector rather than just zeros (which I think is what you're saying). If you assign every unknown word the same value then you can't really differentiate between them - if they're random at least you can do that",
    "458106": "Thanks for that reply Hamish. I was wondering how to do that (I am new). In the public kernels I saw they were assigning zero vectors, which did not feel right. I had reservation towards random vectors too though, since they could end up near other words, but that is better than zeros at least.",
    "458128": "yeah, it feels wrong doesn't it!\n\nYou can make your embeddings trainable, as you train your model they should be \"corrected\" and end up somewhere better - buuuuut in practice that means you're adding LOTS of new trainable parameters and you can easily end up with a worse result :(",
    "458154": "It depends on if testing dataset is correctly labelled or not. If we assume that they are labelled using the same method as training set, then correcting those could end up being detrimental to the score even though it would lead to a model that performed better in reality. I think Quora should have done a better job labelling the questions. Or maybe the problem lies in the fact that we only have access to the title of the question (Quora question also have a section for question details). It could be that questions that look insincere to us are not actually insincere when you have access to the question details. For example, there is one question which has only \"Hello sir?\" as its question text. This question is labelled as a sincere question. It is possible that the asker of this question inserted the rest of the question into the question details. There are a number of other questions that begin the same way, ie. \"Hello sir, among NICMAR and l&amp;T, which one do you think is better for construction management?\"\n\nTo sum it up, I and others believe that correcting the labels will not prove to be helpful.",
    "458231": "thank for your sharing!",
    "458237": "I have similar misgivings. Now I do not understand how do I go about cleaning data.\nDon't understand if I should go ahead and correct some anomaly considering it noise in training data (as described in  problem statement) or is it the same way in the test data also",
    "458460": "this could be \"noise\" added deliberately ???",
    "458544": "For me,\n\n    def clean_text(x):\n        x = str(x)\n        for word in bad_words:\n            x = x.replace(word, \"ZZZZZZ\")\n        for punct in puncts:\n            # Add space after and before punct\n            x = x.replace(punct, f' {punct} ')\n        return x\n\nI replaced some bad words before Adding space after and before punctuations.",
    "458659": "This is a paper I just found related to performance effects of preprocessing data from the Google  toxic comment competition. Effects of each transformation are tested on multiple models, and oftentimes results vary between models. Still, the conclusions are pretty much exactly what others have already been observing for this competition: some things help slightly, other things are neutral, still others decrease performance.\n\nI think the most important quote is as follows: \"Based on the results we have, our recommendation is not to spend too much time on the transformations rather focus on the selection of the best algorithms.\"\n\nhttps://csce.ucmss.com/cr/books/2018/LFS/CSREA2018/ICA4290.pdf"
  },
  "source": "meta"
}