{
  "id": 76627,
  "title": "how much is cleaning helping?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76627",
  "author_name": "",
  "post_date": "2019-01-05T00:15:39.212722400Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>How much lift are you getting with all the different types of cleaning discussed here?</p>\n\n<p>I'm currently doing quite minimal cleaning and getting a CV of 0.690. LB is lower :)</p>",
  "messages": [
    {
      "id": "450437",
      "postDate": "01/05/2019 00:15:39",
      "content": "<p>How much lift are you getting with all the different types of cleaning discussed here?</p>\n\n<p>I'm currently doing quite minimal cleaning and getting a CV of 0.690. LB is lower :)</p>",
      "rawMarkdown": "How much lift are you getting with all the different types of cleaning discussed here?\n\nI'm currently doing quite minimal cleaning and getting a CV of 0.690. LB is lower :)",
      "votes": null
    },
    {
      "id": "450487",
      "postDate": "01/05/2019 04:42:10",
      "content": "<p>As much as +0.01 CV, the LB is too unstable but I'd say about +0.003.</p>",
      "rawMarkdown": "As much as +0.01 CV, the LB is too unstable but I'd say about +0.003.",
      "votes": null
    },
    {
      "id": "450720",
      "postDate": "01/05/2019 15:20:04",
      "content": "<p>interesting, my cv now is 0.695  with LB of 0.684 with minimal cleaning. </p>",
      "rawMarkdown": "interesting, my cv now is 0.695  with LB of 0.684 with minimal cleaning.",
      "votes": null
    },
    {
      "id": "451853",
      "postDate": "01/07/2019 19:15:10",
      "content": "<p>I find that leaning numbers actually reduces my score. I think that most people are cleaning numbers because the top models from the google toxic comment competition used number cleaning. But the google comment competition allowed external data so you could use any embedding files. I didn't participate in that competition, but maybe the embeddings that were used for that competition replaced most numbers with '#'. For this competition, pretty much everyone is using a combination of the Glove and Paragram embeddings. If you look at those files, you will find that both of these embeddings haven't replaced all numbers with 2 digits or more the equivalent number of '#'. So cleaning the digits is unnecessarily removing data. </p>",
      "rawMarkdown": "I find that leaning numbers actually reduces my score. I think that most people are cleaning numbers because the top models from the google toxic comment competition used number cleaning. But the google comment competition allowed external data so you could use any embedding files. I didn't participate in that competition, but maybe the embeddings that were used for that competition replaced most numbers with '#'. For this competition, pretty much everyone is using a combination of the Glove and Paragram embeddings. If you look at those files, you will find that both of these embeddings haven't replaced all numbers with 2 digits or more the equivalent number of '#'. So cleaning the digits is unnecessarily removing data.",
      "votes": null
    },
    {
      "id": "452002",
      "postDate": "01/08/2019 02:47:56",
      "content": "<p>That is a pretty big jump on your local since your original post. Did you add more cleaning to your previous cleaning scheme?</p>",
      "rawMarkdown": "That is a pretty big jump on your local since your original post. Did you add more cleaning to your previous cleaning scheme?",
      "votes": null
    },
    {
      "id": "452718",
      "postDate": "01/09/2019 04:32:14",
      "content": "<p>Cleaning too much may causes over-fitting.</p>",
      "rawMarkdown": "Cleaning too much may causes over-fitting.",
      "votes": null
    },
    {
      "id": "452721",
      "postDate": "01/09/2019 04:42:22",
      "content": "<p>Actually \"cleaning\" is not the correct word, more like <em>tokenization</em>, different tokenization styles caused large jumps in CV score for me.</p>",
      "rawMarkdown": "Actually \"cleaning\" is not the correct word, more like *tokenization*, different tokenization styles caused large jumps in CV score for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 450487,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "01/05/2019 04:42:10",
      "content": "<p>As much as +0.01 CV, the LB is too unstable but I'd say about +0.003.</p>",
      "votes": null,
      "replies": [
        {
          "id": 450720,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "01/05/2019 15:20:04",
          "content": "<p>interesting, my cv now is 0.695  with LB of 0.684 with minimal cleaning. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452002,
          "author_name": "learnmower",
          "author_url": "",
          "post_date": "01/08/2019 02:47:56",
          "content": "<p>That is a pretty big jump on your local since your original post. Did you add more cleaning to your previous cleaning scheme?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452721,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "01/09/2019 04:42:22",
          "content": "<p>Actually \"cleaning\" is not the correct word, more like <em>tokenization</em>, different tokenization styles caused large jumps in CV score for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 451853,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "01/07/2019 19:15:10",
      "content": "<p>I find that leaning numbers actually reduces my score. I think that most people are cleaning numbers because the top models from the google toxic comment competition used number cleaning. But the google comment competition allowed external data so you could use any embedding files. I didn't participate in that competition, but maybe the embeddings that were used for that competition replaced most numbers with '#'. For this competition, pretty much everyone is using a combination of the Glove and Paragram embeddings. If you look at those files, you will find that both of these embeddings haven't replaced all numbers with 2 digits or more the equivalent number of '#'. So cleaning the digits is unnecessarily removing data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452718,
      "author_name": "sunnymarkliu",
      "author_url": "",
      "post_date": "01/09/2019 04:32:14",
      "content": "<p>Cleaning too much may causes over-fitting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "450437": "How much lift are you getting with all the different types of cleaning discussed here?\n\nI'm currently doing quite minimal cleaning and getting a CV of 0.690. LB is lower :)",
    "450487": "As much as +0.01 CV, the LB is too unstable but I'd say about +0.003.",
    "450720": "interesting, my cv now is 0.695  with LB of 0.684 with minimal cleaning.",
    "451853": "I find that leaning numbers actually reduces my score. I think that most people are cleaning numbers because the top models from the google toxic comment competition used number cleaning. But the google comment competition allowed external data so you could use any embedding files. I didn't participate in that competition, but maybe the embeddings that were used for that competition replaced most numbers with '#'. For this competition, pretty much everyone is using a combination of the Glove and Paragram embeddings. If you look at those files, you will find that both of these embeddings haven't replaced all numbers with 2 digits or more the equivalent number of '#'. So cleaning the digits is unnecessarily removing data.",
    "452002": "That is a pretty big jump on your local since your original post. Did you add more cleaning to your previous cleaning scheme?",
    "452718": "Cleaning too much may causes over-fitting.",
    "452721": "Actually \"cleaning\" is not the correct word, more like *tokenization*, different tokenization styles caused large jumps in CV score for me."
  },
  "source": "meta"
}