{
  "id": 143547,
  "title": "data cleaning?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143547",
  "author_name": "Giuseppe",
  "post_date": "2020-04-15T13:25:22.743000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>in jigsaw-toxic-comment-train.csv, some \"comment_text\" fields contain something like </p>\n\n<p>== A barnstar for you! == </p>\n\n<p>at the beginning, which looks like metadata not belonging to the text itself. Should one delete them?</p>",
  "messages": [
    {
      "id": 842444,
      "postDate": "2020-05-11T11:56:59.537Z",
      "content": "<p>The question of data cleaning with BERT is a tough one. From my own experiments, I tried cleaning the text but didn't have any meaningful jump in scores on LB (the score got worse).\nMaybe you can remove the \\n in one of the training file but BERT has its own inner mechanisms to properly tokenize a text and handle dirty data. By cleaning it you risk to remove valuable information.</p>",
      "rawMarkdown": "The question of data cleaning with BERT is a tough one. From my own experiments, I tried cleaning the text but didn't have any meaningful jump in scores on LB (the score got worse).\nMaybe you can remove the \\n in one of the training file but BERT has its own inner mechanisms to properly tokenize a text and handle dirty data. By cleaning it you risk to remove valuable information."
    },
    {
      "id": 808545,
      "postDate": "2020-04-15T13:25:22.743Z",
      "content": "<p>in jigsaw-toxic-comment-train.csv, some \"comment_text\" fields contain something like </p>\n\n<p>== A barnstar for you! == </p>\n\n<p>at the beginning, which looks like metadata not belonging to the text itself. Should one delete them?</p>",
      "rawMarkdown": "in jigsaw-toxic-comment-train.csv, some \"comment_text\" fields contain something like \n\n== A barnstar for you! == \n\nat the beginning, which looks like metadata not belonging to the text itself. Should one delete them?"
    },
    {
      "id": 840847,
      "postDate": "2020-05-10T12:36:21.930Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 842444,
      "author_name": "PAB97",
      "author_url": "",
      "post_date": "2020-05-11T11:56:59.537000",
      "content": "<p>The question of data cleaning with BERT is a tough one. From my own experiments, I tried cleaning the text but didn't have any meaningful jump in scores on LB (the score got worse).\nMaybe you can remove the \\n in one of the training file but BERT has its own inner mechanisms to properly tokenize a text and handle dirty data. By cleaning it you risk to remove valuable information.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 840847,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-10T12:36:21.930000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "842444": "The question of data cleaning with BERT is a tough one. From my own experiments, I tried cleaning the text but didn't have any meaningful jump in scores on LB (the score got worse).\nMaybe you can remove the \\n in one of the training file but BERT has its own inner mechanisms to properly tokenize a text and handle dirty data. By cleaning it you risk to remove valuable information.",
    "808545": "in jigsaw-toxic-comment-train.csv, some \"comment_text\" fields contain something like \n\n== A barnstar for you! == \n\nat the beginning, which looks like metadata not belonging to the text itself. Should one delete them?",
    "840847": ""
  }
}