{
  "id": 75416,
  "title": "Text preprocessing misunderstandings",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75416",
  "author_name": "",
  "post_date": "2018-12-21T12:03:35.129613900Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I'm a newbie in ML and I tried to do some text preprocessing but I don't understand some things:</p>\n\n<p>1) When we replacing numbers by ## or  ###, depending on number length, a lot of words like: '##th', '##k', '##s', '####s' .....\nBut when I tried to add an additional space before and after the replaced number, my score decreased on 0.001, by the time % of recognized vocabulary words in embedings increased. \nWhat we should do with these words?</p>\n\n<p>2) When I tied to remove double spaces and spaces at the end of a sentence my score decreased on 0.002 and I don't understand why?</p>\n\n<p>3) The data contains a lot of additional symbols ('谷', '精', '╔', 'រ', '☹',  .....). I had used code that adds an additional space before and after this symbol and this trick improved my score on 0.002, but when I changed logic to replace all the additional symbols by space it decreased my score on  0.003 and I don't understand why. Are these symbols really affect a toxicity of the sentence?</p>\n\n<p>I will be very pleased with your explanation.</p>",
  "messages": [
    {
      "id": "443311",
      "postDate": "12/21/2018 12:03:35",
      "content": "<p>I'm a newbie in ML and I tried to do some text preprocessing but I don't understand some things:</p>\n\n<p>1) When we replacing numbers by ## or  ###, depending on number length, a lot of words like: '##th', '##k', '##s', '####s' .....\nBut when I tried to add an additional space before and after the replaced number, my score decreased on 0.001, by the time % of recognized vocabulary words in embedings increased. \nWhat we should do with these words?</p>\n\n<p>2) When I tied to remove double spaces and spaces at the end of a sentence my score decreased on 0.002 and I don't understand why?</p>\n\n<p>3) The data contains a lot of additional symbols ('谷', '精', '╔', 'រ', '☹',  .....). I had used code that adds an additional space before and after this symbol and this trick improved my score on 0.002, but when I changed logic to replace all the additional symbols by space it decreased my score on  0.003 and I don't understand why. Are these symbols really affect a toxicity of the sentence?</p>\n\n<p>I will be very pleased with your explanation.</p>",
      "rawMarkdown": "I'm a newbie in ML and I tried to do some text preprocessing but I don't understand some things:\n\n1) When we replacing numbers by ## or  ###, depending on number length, a lot of words like: '##th', '##k', '##s', '####s' .....\nBut when I tried to add an additional space before and after the replaced number, my score decreased on 0.001, by the time % of recognized vocabulary words in embedings increased. \nWhat we should do with these words?\n\n2) When I tied to remove double spaces and spaces at the end of a sentence my score decreased on 0.002 and I don't understand why?\n\n3) The data contains a lot of additional symbols ('谷', '精', '╔', 'រ', '☹',  .....). I had used code that adds an additional space before and after this symbol and this trick improved my score on 0.002, but when I changed logic to replace all the additional symbols by space it decreased my score on  0.003 and I don't understand why. Are these symbols really affect a toxicity of the sentence?\n\nI will be very pleased with your explanation.",
      "votes": null
    },
    {
      "id": "443315",
      "postDate": "12/21/2018 12:12:19",
      "content": "<p>1) A function related to the first question</p>\n\n<pre><code>def clean_numbers(x):\n   x = re.sub('[0-9]{5,}', ' ##### ', x)\n   x = re.sub('[0-9]{4}', ' #### ', x)\n   x = re.sub('[0-9]{3}', ' ### ', x)\n   x = re.sub('[0-9]{2}', ' ## ', x)\n   return x\n</code></pre>\n\n<p>2) A function related to the second question</p>\n\n<pre><code>def remove_extra_spaces(text):    \n  text = re.sub(r'\\s+', r' ', text)\n  # Remove ending space if any\n  text = re.sub(r'\\s+$', r'', text)\n\n  return text\n</code></pre>\n\n<p>3) Functions related to the third question</p>\n\n<pre><code>def clean_text(x):\n   x = str(x)\n   for punct in puncts:\n       x = x.replace(punct, f' {punct} ')\n   return x  \n\ndef clean_text(x):\n   x = str(x)\n   for punct in puncts:\n    x = x.replace(punct,  '  ')\n   return x\n</code></pre>",
      "rawMarkdown": "1) A function related to the first question\n\n    def clean_numbers(x):\n       x = re.sub('[0-9]{5,}', ' ##### ', x)\n       x = re.sub('[0-9]{4}', ' #### ', x)\n       x = re.sub('[0-9]{3}', ' ### ', x)\n       x = re.sub('[0-9]{2}', ' ## ', x)\n       return x\n\n2) A function related to the second question\n\n    def remove_extra_spaces(text):    \n      text = re.sub(r'\\s+', r' ', text)\n      # Remove ending space if any\n      text = re.sub(r'\\s+$', r'', text)\n   \n      return text\n\n3) Functions related to the third question\n\n    def clean_text(x):\n       x = str(x)\n       for punct in puncts:\n           x = x.replace(punct, f' {punct} ')\n       return x  \n\n    def clean_text(x):\n       x = str(x)\n       for punct in puncts:\n        x = x.replace(punct,  '  ')\n       return x",
      "votes": null
    },
    {
      "id": "443316",
      "postDate": "12/21/2018 12:12:57",
      "content": "<p>If you're using Keras' CuDNN layers results have a high variance so such a small variation means nothing.\nAre your results reproductible ?</p>",
      "rawMarkdown": "If you're using Keras' CuDNN layers results have a high variance so such a small variation means nothing.\nAre your results reproductible ?",
      "votes": null
    },
    {
      "id": "443325",
      "postDate": "12/21/2018 12:45:39",
      "content": "<p>My results are reproducible so if I run my kernel 2 times I would have the same results. Also, I tried to run my kernels with these changes not once.</p>",
      "rawMarkdown": "My results are reproducible so if I run my kernel 2 times I would have the same results. Also, I tried to run my kernels with these changes not once.",
      "votes": null
    },
    {
      "id": "443527",
      "postDate": "12/21/2018 20:05:00",
      "content": "<p>2) Hmm. Actually you shouldn't have to do anything like this, since you will be splitting the strings into words anyway, so double-spaces will be ignored. Also, a simple <code>.strip()</code> should remove spaces at the end and at the beginning.</p>\n\n<p>The only reason how I can see this changes anything, is that <code>\\s</code> in regex also matches <code>\\t</code> (tab). But how that could affect the score, I'm really not sure....</p>",
      "rawMarkdown": "2) Hmm. Actually you shouldn't have to do anything like this, since you will be splitting the strings into words anyway, so double-spaces will be ignored. Also, a simple `.strip()` should remove spaces at the end and at the beginning.\n\nThe only reason how I can see this changes anything, is that `\\s` in regex also matches `\\t` (tab). But how that could affect the score, I'm really not sure....",
      "votes": null
    },
    {
      "id": "443548",
      "postDate": "12/21/2018 20:54:47",
      "content": "<p>I'm glad you brought up this topic. </p>\n\n<p>The different embeddings we've been provided all had different pre-processing transforms applied to their source corpuses prior to training. To get the max out of a pre-trained embedding, it stands to reason that the pre-processing we apply matches the one used to generate the associated embedding. Glove's preprocessing code is pretty well known; does anyone have a snippet or know where to get the processing code for the other source vectors?</p>",
      "rawMarkdown": "I'm glad you brought up this topic. \n\nThe different embeddings we've been provided all had different pre-processing transforms applied to their source corpuses prior to training. To get the max out of a pre-trained embedding, it stands to reason that the pre-processing we apply matches the one used to generate the associated embedding. Glove's preprocessing code is pretty well known; does anyone have a snippet or know where to get the processing code for the other source vectors?",
      "votes": null
    },
    {
      "id": "443720",
      "postDate": "12/22/2018 08:21:21",
      "content": "<p>A real basic question re. preprocessing. I am new to Kaggle.</p>\n\n<p>When we finally make a submission, we need to apply the model to the test file.\nAre we allowed to pre-process the test data?\nIf not, then does it make sense to preprocess training data?</p>",
      "rawMarkdown": "A real basic question re. preprocessing. I am new to Kaggle.\n\nWhen we finally make a submission, we need to apply the model to the test file.\nAre we allowed to pre-process the test data?\nIf not, then does it make sense to preprocess training data?",
      "votes": null
    },
    {
      "id": "443745",
      "postDate": "12/22/2018 10:13:52",
      "content": "<p>In the end, we should submit a submission file with tags is that toxic comment or not. So, if preprocessing of the text improve your result, why do not to use it?</p>",
      "rawMarkdown": "In the end, we should submit a submission file with tags is that toxic comment or not. So, if preprocessing of the text improve your result, why do not to use it?",
      "votes": null
    },
    {
      "id": "443816",
      "postDate": "12/22/2018 13:44:59",
      "content": "<p>Preprocessing is part of the data scientist's toolbox and absolutely important to get a good result. There's also no reason to forbid it. Go for it! :)</p>",
      "rawMarkdown": "Preprocessing is part of the data scientist's toolbox and absolutely important to get a good result. There's also no reason to forbid it. Go for it! :)",
      "votes": null
    },
    {
      "id": "444381",
      "postDate": "12/23/2018 23:08:44",
      "content": "<p>Thanks Max and Vladyslav.</p>",
      "rawMarkdown": "Thanks Max and Vladyslav.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 443315,
      "author_name": "psyfaker",
      "author_url": "",
      "post_date": "12/21/2018 12:12:19",
      "content": "<p>1) A function related to the first question</p>\n\n<pre><code>def clean_numbers(x):\n   x = re.sub('[0-9]{5,}', ' ##### ', x)\n   x = re.sub('[0-9]{4}', ' #### ', x)\n   x = re.sub('[0-9]{3}', ' ### ', x)\n   x = re.sub('[0-9]{2}', ' ## ', x)\n   return x\n</code></pre>\n\n<p>2) A function related to the second question</p>\n\n<pre><code>def remove_extra_spaces(text):    \n  text = re.sub(r'\\s+', r' ', text)\n  # Remove ending space if any\n  text = re.sub(r'\\s+$', r'', text)\n\n  return text\n</code></pre>\n\n<p>3) Functions related to the third question</p>\n\n<pre><code>def clean_text(x):\n   x = str(x)\n   for punct in puncts:\n       x = x.replace(punct, f' {punct} ')\n   return x  \n\ndef clean_text(x):\n   x = str(x)\n   for punct in puncts:\n    x = x.replace(punct,  '  ')\n   return x\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 443527,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/21/2018 20:05:00",
          "content": "<p>2) Hmm. Actually you shouldn't have to do anything like this, since you will be splitting the strings into words anyway, so double-spaces will be ignored. Also, a simple <code>.strip()</code> should remove spaces at the end and at the beginning.</p>\n\n<p>The only reason how I can see this changes anything, is that <code>\\s</code> in regex also matches <code>\\t</code> (tab). But how that could affect the score, I'm really not sure....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443316,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "12/21/2018 12:12:57",
      "content": "<p>If you're using Keras' CuDNN layers results have a high variance so such a small variation means nothing.\nAre your results reproductible ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 443325,
          "author_name": "psyfaker",
          "author_url": "",
          "post_date": "12/21/2018 12:45:39",
          "content": "<p>My results are reproducible so if I run my kernel 2 times I would have the same results. Also, I tried to run my kernels with these changes not once.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443548,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/21/2018 20:54:47",
      "content": "<p>I'm glad you brought up this topic. </p>\n\n<p>The different embeddings we've been provided all had different pre-processing transforms applied to their source corpuses prior to training. To get the max out of a pre-trained embedding, it stands to reason that the pre-processing we apply matches the one used to generate the associated embedding. Glove's preprocessing code is pretty well known; does anyone have a snippet or know where to get the processing code for the other source vectors?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443720,
      "author_name": "takeseven",
      "author_url": "",
      "post_date": "12/22/2018 08:21:21",
      "content": "<p>A real basic question re. preprocessing. I am new to Kaggle.</p>\n\n<p>When we finally make a submission, we need to apply the model to the test file.\nAre we allowed to pre-process the test data?\nIf not, then does it make sense to preprocess training data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 443745,
          "author_name": "psyfaker",
          "author_url": "",
          "post_date": "12/22/2018 10:13:52",
          "content": "<p>In the end, we should submit a submission file with tags is that toxic comment or not. So, if preprocessing of the text improve your result, why do not to use it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443816,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/22/2018 13:44:59",
          "content": "<p>Preprocessing is part of the data scientist's toolbox and absolutely important to get a good result. There's also no reason to forbid it. Go for it! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444381,
          "author_name": "takeseven",
          "author_url": "",
          "post_date": "12/23/2018 23:08:44",
          "content": "<p>Thanks Max and Vladyslav.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "443311": "I'm a newbie in ML and I tried to do some text preprocessing but I don't understand some things:\n\n1) When we replacing numbers by ## or  ###, depending on number length, a lot of words like: '##th', '##k', '##s', '####s' .....\nBut when I tried to add an additional space before and after the replaced number, my score decreased on 0.001, by the time % of recognized vocabulary words in embedings increased. \nWhat we should do with these words?\n\n2) When I tied to remove double spaces and spaces at the end of a sentence my score decreased on 0.002 and I don't understand why?\n\n3) The data contains a lot of additional symbols ('谷', '精', '╔', 'រ', '☹',  .....). I had used code that adds an additional space before and after this symbol and this trick improved my score on 0.002, but when I changed logic to replace all the additional symbols by space it decreased my score on  0.003 and I don't understand why. Are these symbols really affect a toxicity of the sentence?\n\nI will be very pleased with your explanation.",
    "443315": "1) A function related to the first question\n\n    def clean_numbers(x):\n       x = re.sub('[0-9]{5,}', ' ##### ', x)\n       x = re.sub('[0-9]{4}', ' #### ', x)\n       x = re.sub('[0-9]{3}', ' ### ', x)\n       x = re.sub('[0-9]{2}', ' ## ', x)\n       return x\n\n2) A function related to the second question\n\n    def remove_extra_spaces(text):    \n      text = re.sub(r'\\s+', r' ', text)\n      # Remove ending space if any\n      text = re.sub(r'\\s+$', r'', text)\n   \n      return text\n\n3) Functions related to the third question\n\n    def clean_text(x):\n       x = str(x)\n       for punct in puncts:\n           x = x.replace(punct, f' {punct} ')\n       return x  \n\n    def clean_text(x):\n       x = str(x)\n       for punct in puncts:\n        x = x.replace(punct,  '  ')\n       return x",
    "443316": "If you're using Keras' CuDNN layers results have a high variance so such a small variation means nothing.\nAre your results reproductible ?",
    "443325": "My results are reproducible so if I run my kernel 2 times I would have the same results. Also, I tried to run my kernels with these changes not once.",
    "443527": "2) Hmm. Actually you shouldn't have to do anything like this, since you will be splitting the strings into words anyway, so double-spaces will be ignored. Also, a simple `.strip()` should remove spaces at the end and at the beginning.\n\nThe only reason how I can see this changes anything, is that `\\s` in regex also matches `\\t` (tab). But how that could affect the score, I'm really not sure....",
    "443548": "I'm glad you brought up this topic. \n\nThe different embeddings we've been provided all had different pre-processing transforms applied to their source corpuses prior to training. To get the max out of a pre-trained embedding, it stands to reason that the pre-processing we apply matches the one used to generate the associated embedding. Glove's preprocessing code is pretty well known; does anyone have a snippet or know where to get the processing code for the other source vectors?",
    "443720": "A real basic question re. preprocessing. I am new to Kaggle.\n\nWhen we finally make a submission, we need to apply the model to the test file.\nAre we allowed to pre-process the test data?\nIf not, then does it make sense to preprocess training data?",
    "443745": "In the end, we should submit a submission file with tags is that toxic comment or not. So, if preprocessing of the text improve your result, why do not to use it?",
    "443816": "Preprocessing is part of the data scientist's toolbox and absolutely important to get a good result. There's also no reason to forbid it. Go for it! :)",
    "444381": "Thanks Max and Vladyslav."
  },
  "source": "meta"
}