{
  "id": 157025,
  "title": "(Non-)Fun stuff found in data :)",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/157025",
  "author_name": "Mukharbek Organokov",
  "post_date": "2020-06-09T01:38:11.617000",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2Fa59bd4aa160cff5fe7bba1e358a0bca4%2FScreenshot%20from%202020-06-09%20033716.png?generation=1591666681691381&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 878891,
      "postDate": "2020-06-09T01:38:11.617Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2Fa59bd4aa160cff5fe7bba1e358a0bca4%2FScreenshot%20from%202020-06-09%20033716.png?generation=1591666681691381&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2Fa59bd4aa160cff5fe7bba1e358a0bca4%2FScreenshot%20from%202020-06-09%20033716.png?generation=1591666681691381&amp;alt=media)\n",
      "votes": 15
    },
    {
      "id": 879929,
      "postDate": "2020-06-09T20:56:40.687Z",
      "content": "<p>Toxicity score: 1.1 😂 </p>",
      "rawMarkdown": "Toxicity score: 1.1 😂 ",
      "votes": 4,
      "replies": [
        {
          "id": 886296,
          "postDate": "2020-06-14T21:53:37.833Z",
          "content": "<p>😂 😂 😂 </p>",
          "rawMarkdown": "😂 😂 😂 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 880912,
      "postDate": "2020-06-10T16:32:34.203Z",
      "content": "<p>Haha we use that to do pseudo labelling !</p>",
      "rawMarkdown": "Haha we use that to do pseudo labelling !",
      "votes": 1
    },
    {
      "id": 887648,
      "postDate": "2020-06-15T19:33:29.500Z",
      "content": "<p>Some texts contain just one word. And after tokenizing, the array would be with similar values. </p>\n\n<p>I think such bullshit could be easy to identify and classify. Due to length, or repeats.\nAny used it particularly?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2F2f7354d790f268b2e7649d3f37c3a91a%2Foie_5z3KFawMi2wy.png?generation=1592249560839834&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Some texts contain just one word. And after tokenizing, the array would be with similar values. \n\nI think such bullshit could be easy to identify and classify. Due to length, or repeats.\nAny used it particularly?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2F2f7354d790f268b2e7649d3f37c3a91a%2Foie_5z3KFawMi2wy.png?generation=1592249560839834&amp;alt=media)\n",
      "votes": 2
    },
    {
      "id": 887640,
      "postDate": "2020-06-15T19:30:03.997Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 880395,
      "postDate": "2020-06-10T08:56:48.683Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 879929,
      "author_name": "BryanB",
      "author_url": "",
      "post_date": "2020-06-09T20:56:40.687000",
      "content": "<p>Toxicity score: 1.1 😂 </p>",
      "votes": 4,
      "replies": [
        {
          "id": 886296,
          "author_name": "Joao Pedro Medeiros ",
          "author_url": "",
          "post_date": "2020-06-14T21:53:37.833000",
          "content": "<p>😂 😂 😂 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 880912,
      "author_name": "Haythem Tellili",
      "author_url": "",
      "post_date": "2020-06-10T16:32:34.203000",
      "content": "<p>Haha we use that to do pseudo labelling !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 887648,
      "author_name": "Mukharbek Organokov",
      "author_url": "",
      "post_date": "2020-06-15T19:33:29.500000",
      "content": "<p>Some texts contain just one word. And after tokenizing, the array would be with similar values. </p>\n\n<p>I think such bullshit could be easy to identify and classify. Due to length, or repeats.\nAny used it particularly?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2F2f7354d790f268b2e7649d3f37c3a91a%2Foie_5z3KFawMi2wy.png?generation=1592249560839834&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 887640,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-15T19:30:03.997000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 880395,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-10T08:56:48.683000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "878891": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2Fa59bd4aa160cff5fe7bba1e358a0bca4%2FScreenshot%20from%202020-06-09%20033716.png?generation=1591666681691381&amp;alt=media)\n",
    "879929": "Toxicity score: 1.1 😂 ",
    "880912": "Haha we use that to do pseudo labelling !",
    "887648": "Some texts contain just one word. And after tokenizing, the array would be with similar values. \n\nI think such bullshit could be easy to identify and classify. Due to length, or repeats.\nAny used it particularly?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1788308%2F2f7354d790f268b2e7649d3f37c3a91a%2Foie_5z3KFawMi2wy.png?generation=1592249560839834&amp;alt=media)\n",
    "887640": "",
    "880395": ""
  }
}