{
  "id": 71372,
  "title": "Why no cleaning.?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71372",
  "author_name": "",
  "post_date": "2018-11-13T05:05:45.823212400Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was going though the public kernels and observed that most of the top kernels don't do anything to clean the text like removing punctuation, converting to same case, stemming etc.</p>\n\n<p>Is there any specific reason for not doing cleaning in this competition.? </p>",
  "messages": [
    {
      "id": "420122",
      "postDate": "11/13/2018 05:05:45",
      "content": "<p>I was going though the public kernels and observed that most of the top kernels don't do anything to clean the text like removing punctuation, converting to same case, stemming etc.</p>\n\n<p>Is there any specific reason for not doing cleaning in this competition.? </p>",
      "rawMarkdown": "I was going though the public kernels and observed that most of the top kernels don't do anything to clean the text like removing punctuation, converting to same case, stemming etc.\n\nIs there any specific reason for not doing cleaning in this competition.?",
      "votes": null
    },
    {
      "id": "420124",
      "postDate": "11/13/2018 05:09:28",
      "content": "<p>Ideally, we should let the model learn these things instead manually clean the text ... but in my case, I am just lazy...</p>",
      "rawMarkdown": "Ideally, we should let the model learn these things instead manually clean the text ... but in my case, I am just lazy...",
      "votes": null
    },
    {
      "id": "420169",
      "postDate": "11/13/2018 07:35:40",
      "content": "<p>Me too. We are in an early phase of the competition and cleaning (at least for me) takes longer than to spit out a Keras model. Maybe I will add an educational kernel on cleaning.</p>",
      "rawMarkdown": "Me too. We are in an early phase of the competition and cleaning (at least for me) takes longer than to spit out a Keras model. Maybe I will add an educational kernel on cleaning.",
      "votes": null
    },
    {
      "id": "420190",
      "postDate": "11/13/2018 08:34:59",
      "content": "<p>That's great. Looking forward for that. :)</p>",
      "rawMarkdown": "That's great. Looking forward for that. :)",
      "votes": null
    },
    {
      "id": "420242",
      "postDate": "11/13/2018 10:32:11",
      "content": "<p>you can find it here <a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings#\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings#</a></p>",
      "rawMarkdown": "you can find it here https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings#",
      "votes": null
    },
    {
      "id": "420304",
      "postDate": "11/13/2018 12:41:40",
      "content": "<p>Thanks, The kernel is very informative</p>",
      "rawMarkdown": "Thanks, The kernel is very informative",
      "votes": null
    },
    {
      "id": "420369",
      "postDate": "11/13/2018 14:04:13",
      "content": "<p>Actually, in several top kernels like this one: <a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\">https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings</a>, there's cleaning happening under the hood via keras' Tokenizer - <a href=\"https://keras.io/preprocessing/text/\">https://keras.io/preprocessing/text/</a>. It turns out that keras even makes some simple cleaning easy: the tokenizer does lowercasing and punctuation removal. </p>\n\n<p>As for other more aggressive cleaning methods like stemming, you may want to avoid that when using pretrained word vectors, as they can already do a nice recognizing similarities (and key differences) between different word forms. I.e. stemming involves information loss that sometimes removes more noise than signal, but that's often not the case. </p>",
      "rawMarkdown": "Actually, in several top kernels like this one: https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings, there's cleaning happening under the hood via keras' Tokenizer - https://keras.io/preprocessing/text/. It turns out that keras even makes some simple cleaning easy: the tokenizer does lowercasing and punctuation removal. \n\nAs for other more aggressive cleaning methods like stemming, you may want to avoid that when using pretrained word vectors, as they can already do a nice recognizing similarities (and key differences) between different word forms. I.e. stemming involves information loss that sometimes removes more noise than signal, but that's often not the case.",
      "votes": null
    },
    {
      "id": "420410",
      "postDate": "11/13/2018 15:35:45",
      "content": "<p>I was wondering the same, and it appears that usual text cleaning steps do not improve results. Check my work on this question for more infos: <a href=\"https://www.kaggle.com/theoviel/should-you-clean-your-data\">https://www.kaggle.com/theoviel/should-you-clean-your-data</a></p>",
      "rawMarkdown": "I was wondering the same, and it appears that usual text cleaning steps do not improve results. Check my work on this question for more infos: https://www.kaggle.com/theoviel/should-you-clean-your-data",
      "votes": null
    },
    {
      "id": "420423",
      "postDate": "11/13/2018 15:52:42",
      "content": "<p>Thank you for this! I was missing a preprocessing kernel. </p>",
      "rawMarkdown": "Thank you for this! I was missing a preprocessing kernel.",
      "votes": null
    },
    {
      "id": "420614",
      "postDate": "11/13/2018 22:31:44",
      "content": "<p>Thanks for the kernel. It is really thorough. Now we need to check for the other models as well ;)</p>",
      "rawMarkdown": "Thanks for the kernel. It is really thorough. Now we need to check for the other models as well ;)",
      "votes": null
    },
    {
      "id": "420618",
      "postDate": "11/13/2018 22:34:24",
      "content": "<p>The tokeniser cleans the data. It automatically removes punct and lower all the letters and splits using space. Taken from the class API: filters='!\"#$%&amp;()*+,-./:;&lt;=&gt;?@[]^_`{|}~ ', lower=True, split=' ', char_level=False</p>",
      "rawMarkdown": "The tokeniser cleans the data. It automatically removes punct and lower all the letters and splits using space. Taken from the class API: filters='!\"#$%&amp;()*+,-./:;&lt;=&gt;?@[\\]^_`{|}~ ', lower=True, split=' ', char_level=False",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 420124,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/13/2018 05:09:28",
      "content": "<p>Ideally, we should let the model learn these things instead manually clean the text ... but in my case, I am just lazy...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420169,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "11/13/2018 07:35:40",
      "content": "<p>Me too. We are in an early phase of the competition and cleaning (at least for me) takes longer than to spit out a Keras model. Maybe I will add an educational kernel on cleaning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 420190,
          "author_name": "sreeram1234",
          "author_url": "",
          "post_date": "11/13/2018 08:34:59",
          "content": "<p>That's great. Looking forward for that. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420242,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "11/13/2018 10:32:11",
          "content": "<p>you can find it here <a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings#\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings#</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420304,
          "author_name": "sreeram1234",
          "author_url": "",
          "post_date": "11/13/2018 12:41:40",
          "content": "<p>Thanks, The kernel is very informative</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420423,
          "author_name": "anebzt",
          "author_url": "",
          "post_date": "11/13/2018 15:52:42",
          "content": "<p>Thank you for this! I was missing a preprocessing kernel. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420614,
          "author_name": "mrbeer",
          "author_url": "",
          "post_date": "11/13/2018 22:31:44",
          "content": "<p>Thanks for the kernel. It is really thorough. Now we need to check for the other models as well ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 420369,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "11/13/2018 14:04:13",
      "content": "<p>Actually, in several top kernels like this one: <a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\">https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings</a>, there's cleaning happening under the hood via keras' Tokenizer - <a href=\"https://keras.io/preprocessing/text/\">https://keras.io/preprocessing/text/</a>. It turns out that keras even makes some simple cleaning easy: the tokenizer does lowercasing and punctuation removal. </p>\n\n<p>As for other more aggressive cleaning methods like stemming, you may want to avoid that when using pretrained word vectors, as they can already do a nice recognizing similarities (and key differences) between different word forms. I.e. stemming involves information loss that sometimes removes more noise than signal, but that's often not the case. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420410,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/13/2018 15:35:45",
      "content": "<p>I was wondering the same, and it appears that usual text cleaning steps do not improve results. Check my work on this question for more infos: <a href=\"https://www.kaggle.com/theoviel/should-you-clean-your-data\">https://www.kaggle.com/theoviel/should-you-clean-your-data</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 420618,
          "author_name": "mrbeer",
          "author_url": "",
          "post_date": "11/13/2018 22:34:24",
          "content": "<p>The tokeniser cleans the data. It automatically removes punct and lower all the letters and splits using space. Taken from the class API: filters='!\"#$%&amp;()*+,-./:;&lt;=&gt;?@[]^_`{|}~ ', lower=True, split=' ', char_level=False</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "420122": "I was going though the public kernels and observed that most of the top kernels don't do anything to clean the text like removing punctuation, converting to same case, stemming etc.\n\nIs there any specific reason for not doing cleaning in this competition.?",
    "420124": "Ideally, we should let the model learn these things instead manually clean the text ... but in my case, I am just lazy...",
    "420169": "Me too. We are in an early phase of the competition and cleaning (at least for me) takes longer than to spit out a Keras model. Maybe I will add an educational kernel on cleaning.",
    "420190": "That's great. Looking forward for that. :)",
    "420242": "you can find it here https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings#",
    "420304": "Thanks, The kernel is very informative",
    "420369": "Actually, in several top kernels like this one: https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings, there's cleaning happening under the hood via keras' Tokenizer - https://keras.io/preprocessing/text/. It turns out that keras even makes some simple cleaning easy: the tokenizer does lowercasing and punctuation removal. \n\nAs for other more aggressive cleaning methods like stemming, you may want to avoid that when using pretrained word vectors, as they can already do a nice recognizing similarities (and key differences) between different word forms. I.e. stemming involves information loss that sometimes removes more noise than signal, but that's often not the case.",
    "420410": "I was wondering the same, and it appears that usual text cleaning steps do not improve results. Check my work on this question for more infos: https://www.kaggle.com/theoviel/should-you-clean-your-data",
    "420423": "Thank you for this! I was missing a preprocessing kernel.",
    "420614": "Thanks for the kernel. It is really thorough. Now we need to check for the other models as well ;)",
    "420618": "The tokeniser cleans the data. It automatically removes punct and lower all the letters and splits using space. Taken from the class API: filters='!\"#$%&amp;()*+,-./:;&lt;=&gt;?@[\\]^_`{|}~ ', lower=True, split=' ', char_level=False"
  },
  "source": "meta"
}