{
  "id": 72120,
  "title": "Frequent Spelling Mistakes : Qoura, Mastrubation and Demonitisation",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72120",
  "author_name": "Theo Viel",
  "post_date": "2018-11-20T14:50:19.946000",
  "votes": 49,
  "comment_count": 24,
  "views": 0,
  "content": "<p>So I was checking the most frequent words in the database, for which no embeddings were available (after some text treatment). Here are the ones that were unknown because of spelling mistakes :</p>\n\n<ul>\n<li>Qoura : 82 times</li>\n<li>mastrubation : 33 times</li>\n<li>demonitisation : 29 times</li>\n<li>Whst : 27 times</li>\n<li>watsapp : 24 times</li>\n<li>mastrubate : 20 times</li>\n<li>qouta : 16 times</li>\n<li>demonitization : 14 times</li>\n<li>narcissit : 13 times</li>\n<li>mastrubating : 13 times</li>\n<li>narcisist : 12 times</li>\n<li>...</li>\n</ul>\n\n<p>I found it quite funny, but it can be useful if you want to add a spelling corrector in your preprocessing.</p>\n\n<p>EDIT : Made my work  on how I found those public, check : \n<a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2\">https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2</a></p>",
  "messages": [
    {
      "id": 424702,
      "postDate": "2018-11-20T14:50:19.947Z",
      "content": "<p>So I was checking the most frequent words in the database, for which no embeddings were available (after some text treatment). Here are the ones that were unknown because of spelling mistakes :</p>\n\n<ul>\n<li>Qoura : 82 times</li>\n<li>mastrubation : 33 times</li>\n<li>demonitisation : 29 times</li>\n<li>Whst : 27 times</li>\n<li>watsapp : 24 times</li>\n<li>mastrubate : 20 times</li>\n<li>qouta : 16 times</li>\n<li>demonitization : 14 times</li>\n<li>narcissit : 13 times</li>\n<li>mastrubating : 13 times</li>\n<li>narcisist : 12 times</li>\n<li>...</li>\n</ul>\n\n<p>I found it quite funny, but it can be useful if you want to add a spelling corrector in your preprocessing.</p>\n\n<p>EDIT : Made my work  on how I found those public, check : \n<a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2\">https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2</a></p>",
      "rawMarkdown": "So I was checking the most frequent words in the database, for which no embeddings were available (after some text treatment). Here are the ones that were unknown because of spelling mistakes :\n\n- Qoura : 82 times\n- mastrubation : 33 times\n- demonitisation : 29 times\n- Whst : 27 times\n- watsapp : 24 times\n- mastrubate : 20 times\n- qouta : 16 times\n- demonitization : 14 times\n- narcissit : 13 times\n- mastrubating : 13 times\n- narcisist : 12 times\n- ...\n\nI found it quite funny, but it can be useful if you want to add a spelling corrector in your preprocessing.\n\nEDIT : Made my work  on how I found those public, check : \nhttps://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2",
      "votes": 48
    },
    {
      "id": 424717,
      "postDate": "2018-11-20T15:10:25.650Z",
      "content": "<p>Intersting find. I cant believ some poeple are so bad at speling!</p>",
      "rawMarkdown": "Intersting find. I cant believ some poeple are so bad at speling!",
      "votes": 9
    },
    {
      "id": 425609,
      "postDate": "2018-11-21T20:59:51.743Z",
      "content": "<p>Haha. Is that possible that people just write in this way to pass Quora's checking system?</p>",
      "rawMarkdown": "Haha. Is that possible that people just write in this way to pass Quora's checking system?",
      "votes": 5,
      "replies": [
        {
          "id": 425621,
          "postDate": "2018-11-21T21:17:02.967Z",
          "content": "<p>Could have been that, but there's a lot of sentences with the correctly spelled word in it. So that has to be something else.</p>",
          "rawMarkdown": "Could have been that, but there's a lot of sentences with the correctly spelled word in it. So that has to be something else.",
          "votes": 1
        },
        {
          "id": 425622,
          "postDate": "2018-11-21T21:18:24.517Z",
          "content": "<p>Plus I don't think there's a checking system on Quora, but I'm not sure though. :)</p>",
          "rawMarkdown": "Plus I don't think there's a checking system on Quora, but I'm not sure though. :)",
          "votes": 1
        },
        {
          "id": 425629,
          "postDate": "2018-11-21T21:32:21.810Z",
          "content": "<p>Is that case, this can be a feature :)</p>",
          "rawMarkdown": "Is that case, this can be a feature :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 424888,
      "postDate": "2018-11-20T20:40:00.760Z",
      "content": "<p>Im gonna have to mark your topic as insincere :P</p>",
      "rawMarkdown": "Im gonna have to mark your topic as insincere :P",
      "votes": 6
    },
    {
      "id": 429955,
      "postDate": "2018-11-29T15:37:32.213Z",
      "content": "<p>some indian words is creating problem hmmm , so we can find a list of bad words generally indian spoken up then we can replace them with similar english words , might solve the problem , even what i feel we have to negative samples augmentation i have made on my kernels and product research with Codenscious technologies , for financial news classifier i build those techniques on basis of english grammar </p>",
      "rawMarkdown": "some indian words is creating problem hmmm , so we can find a list of bad words generally indian spoken up then we can replace them with similar english words , might solve the problem , even what i feel we have to negative samples augmentation i have made on my kernels and product research with Codenscious technologies , for financial news classifier i build those techniques on basis of english grammar ",
      "votes": 4,
      "replies": [
        {
          "id": 429966,
          "postDate": "2018-11-29T15:56:30.223Z",
          "content": "<p>I don't speak Indian so it's going to be tough for me to do it manually :)\nBut it's a good idea.</p>",
          "rawMarkdown": "I don't speak Indian so it's going to be tough for me to do it manually :)\nBut it's a good idea.",
          "votes": 1
        },
        {
          "id": 429988,
          "postDate": "2018-11-29T16:18:44.620Z",
          "rawMarkdown": "",
          "votes": 2
        }
      ]
    },
    {
      "id": 430026,
      "postDate": "2018-11-29T17:11:56.883Z",
      "content": "<p>It made me laugh, then it made me think. Great work! (Y)</p>",
      "rawMarkdown": "It made me laugh, then it made me think. Great work! (Y)",
      "votes": 1
    },
    {
      "id": 426601,
      "postDate": "2018-11-23T14:19:24.697Z",
      "content": "<p>Very interesting insight. Thanks for sharing!!</p>",
      "rawMarkdown": "Very interesting insight. Thanks for sharing!!",
      "votes": 1
    },
    {
      "id": 426529,
      "postDate": "2018-11-23T11:57:50.820Z",
      "content": "<p>That is actually really helpfull. Thanks! How did you find these? Do you mind sharing the code as well?</p>",
      "rawMarkdown": "That is actually really helpfull. Thanks! How did you find these? Do you mind sharing the code as well?",
      "votes": 1,
      "replies": [
        {
          "id": 426550,
          "postDate": "2018-11-23T12:35:12.830Z",
          "content": "<p>I found those manually checking words that were in the corpus and not for which embeddings were not available. I'll make a kernel about it soon, I'll keep you updated.</p>",
          "rawMarkdown": "I found those manually checking words that were in the corpus and not for which embeddings were not available. I'll make a kernel about it soon, I'll keep you updated.",
          "votes": 1
        },
        {
          "id": 426605,
          "postDate": "2018-11-23T14:29:04.803Z",
          "content": "<p>Here's the code : \n<a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\">https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing</a></p>",
          "rawMarkdown": "Here's the code : \nhttps://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing",
          "votes": 2
        }
      ]
    },
    {
      "id": 425797,
      "postDate": "2018-11-22T05:51:42.260Z",
      "content": "<p>Nice finding LOL</p>",
      "rawMarkdown": "Nice finding LOL",
      "votes": 1
    },
    {
      "id": 426100,
      "postDate": "2018-11-22T16:02:52.033Z",
      "content": "<p>interesting , maybe we could do char by char match on all words in glove dictionary and replace with most correct match.</p>",
      "rawMarkdown": "interesting , maybe we could do char by char match on all words in glove dictionary and replace with most correct match.",
      "votes": 2
    },
    {
      "id": 425782,
      "postDate": "2018-11-22T05:12:46.073Z",
      "content": "<p>haha it is funny. It hints us that people who asked insincere question may be not that well educated.</p>",
      "rawMarkdown": "haha it is funny. It hints us that people who asked insincere question may be not that well educated.",
      "votes": 2,
      "replies": [
        {
          "id": 425858,
          "postDate": "2018-11-22T08:11:53.273Z",
          "content": "<p>I did not check if the questions where the spelling mistakes are were insincere, but I'll give it a look!</p>",
          "rawMarkdown": "I did not check if the questions where the spelling mistakes are were insincere, but I'll give it a look!",
          "votes": 1
        },
        {
          "id": 429627,
          "postDate": "2018-11-29T05:02:09.237Z",
          "content": "<p>Is there a correlation between spelling mistakes and the target variability?</p>",
          "rawMarkdown": "Is there a correlation between spelling mistakes and the target variability?",
          "votes": 1
        },
        {
          "id": 429996,
          "postDate": "2018-11-29T16:26:18.230Z",
          "content": "<p>There could be, I'll definitely check that.\nSpelling mistakes can be used as word level features, not sure if it is useful though.</p>",
          "rawMarkdown": "There could be, I'll definitely check that.\nSpelling mistakes can be used as word level features, not sure if it is useful though.",
          "votes": 1
        }
      ]
    },
    {
      "id": 425466,
      "postDate": "2018-11-21T16:42:05.083Z",
      "content": "<p>this cracked me up man!</p>",
      "rawMarkdown": "this cracked me up man!",
      "votes": 2
    },
    {
      "id": 429602,
      "postDate": "2018-11-29T03:58:34.593Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 428234,
      "postDate": "2018-11-27T00:55:31.620Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 425694,
      "postDate": "2018-11-22T01:30:26.350Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 424717,
      "author_name": "Allen",
      "author_url": "",
      "post_date": "2018-11-20T15:10:25.650000",
      "content": "<p>Intersting find. I cant believ some poeple are so bad at speling!</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 425609,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-11-21T20:59:51.743000",
      "content": "<p>Haha. Is that possible that people just write in this way to pass Quora's checking system?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 425621,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-21T21:17:02.967000",
          "content": "<p>Could have been that, but there's a lot of sentences with the correctly spelled word in it. So that has to be something else.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425622,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-21T21:18:24.517000",
          "content": "<p>Plus I don't think there's a checking system on Quora, but I'm not sure though. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 425629,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-21T21:32:21.810000",
          "content": "<p>Is that case, this can be a feature :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 424888,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2018-11-20T20:40:00.760000",
      "content": "<p>Im gonna have to mark your topic as insincere :P</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 429955,
      "author_name": "rohan",
      "author_url": "",
      "post_date": "2018-11-29T15:37:32.213000",
      "content": "<p>some indian words is creating problem hmmm , so we can find a list of bad words generally indian spoken up then we can replace them with similar english words , might solve the problem , even what i feel we have to negative samples augmentation i have made on my kernels and product research with Codenscious technologies , for financial news classifier i build those techniques on basis of english grammar </p>",
      "votes": 4,
      "replies": [
        {
          "id": 429966,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-29T15:56:30.223000",
          "content": "<p>I don't speak Indian so it's going to be tough for me to do it manually :)\nBut it's a good idea.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429988,
          "author_name": "rohan",
          "author_url": "",
          "post_date": "2018-11-29T16:18:44.620000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 430026,
      "author_name": "Nephtali Garrido",
      "author_url": "",
      "post_date": "2018-11-29T17:11:56.883000",
      "content": "<p>It made me laugh, then it made me think. Great work! (Y)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426601,
      "author_name": "Abhishekmamidi",
      "author_url": "",
      "post_date": "2018-11-23T14:19:24.697000",
      "content": "<p>Very interesting insight. Thanks for sharing!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426529,
      "author_name": "Robin Denz",
      "author_url": "",
      "post_date": "2018-11-23T11:57:50.820000",
      "content": "<p>That is actually really helpfull. Thanks! How did you find these? Do you mind sharing the code as well?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 426550,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-23T12:35:12.830000",
          "content": "<p>I found those manually checking words that were in the corpus and not for which embeddings were not available. I'll make a kernel about it soon, I'll keep you updated.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 426605,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-23T14:29:04.803000",
          "content": "<p>Here's the code : \n<a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\">https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 425797,
      "author_name": "Sunil Vikram",
      "author_url": "",
      "post_date": "2018-11-22T05:51:42.260000",
      "content": "<p>Nice finding LOL</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426100,
      "author_name": "chans",
      "author_url": "",
      "post_date": "2018-11-22T16:02:52.033000",
      "content": "<p>interesting , maybe we could do char by char match on all words in glove dictionary and replace with most correct match.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 425782,
      "author_name": "jhajhria",
      "author_url": "",
      "post_date": "2018-11-22T05:12:46.073000",
      "content": "<p>haha it is funny. It hints us that people who asked insincere question may be not that well educated.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 425858,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-22T08:11:53.273000",
          "content": "<p>I did not check if the questions where the spelling mistakes are were insincere, but I'll give it a look!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429627,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-11-29T05:02:09.237000",
          "content": "<p>Is there a correlation between spelling mistakes and the target variability?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429996,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-29T16:26:18.230000",
          "content": "<p>There could be, I'll definitely check that.\nSpelling mistakes can be used as word level features, not sure if it is useful though.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 425466,
      "author_name": "Firat Gonen",
      "author_url": "",
      "post_date": "2018-11-21T16:42:05.083000",
      "content": "<p>this cracked me up man!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 429602,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-29T03:58:34.593000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 428234,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-27T00:55:31.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425694,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-22T01:30:26.350000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "424702": "So I was checking the most frequent words in the database, for which no embeddings were available (after some text treatment). Here are the ones that were unknown because of spelling mistakes :\n\n- Qoura : 82 times\n- mastrubation : 33 times\n- demonitisation : 29 times\n- Whst : 27 times\n- watsapp : 24 times\n- mastrubate : 20 times\n- qouta : 16 times\n- demonitization : 14 times\n- narcissit : 13 times\n- mastrubating : 13 times\n- narcisist : 12 times\n- ...\n\nI found it quite funny, but it can be useful if you want to add a spelling corrector in your preprocessing.\n\nEDIT : Made my work  on how I found those public, check : \nhttps://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2",
    "424717": "Intersting find. I cant believ some poeple are so bad at speling!",
    "425609": "Haha. Is that possible that people just write in this way to pass Quora's checking system?",
    "424888": "Im gonna have to mark your topic as insincere :P",
    "429955": "some indian words is creating problem hmmm , so we can find a list of bad words generally indian spoken up then we can replace them with similar english words , might solve the problem , even what i feel we have to negative samples augmentation i have made on my kernels and product research with Codenscious technologies , for financial news classifier i build those techniques on basis of english grammar ",
    "430026": "It made me laugh, then it made me think. Great work! (Y)",
    "426601": "Very interesting insight. Thanks for sharing!!",
    "426529": "That is actually really helpfull. Thanks! How did you find these? Do you mind sharing the code as well?",
    "425797": "Nice finding LOL",
    "426100": "interesting , maybe we could do char by char match on all words in glove dictionary and replace with most correct match.",
    "425782": "haha it is funny. It hints us that people who asked insincere question may be not that well educated.",
    "425466": "this cracked me up man!",
    "429602": "",
    "428234": "",
    "425694": ""
  }
}