{
  "id": 148530,
  "title": "Pseudo-Labeled Open-Subtitles dataset",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/148530",
  "author_name": "",
  "post_date": "2020-05-04T17:57:47.556324600Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>Recently I have published <a href=\"https://www.kaggle.com/shonenkov/hack-with-parallel-corpus\">kernel</a> about parallel corpora,  pseudo-labeling and label smoothing.</p>\n\n<p>I would like to share with you <a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">my dataset</a>. </p>\n\n<p>It contains samples with token length &lt; 80</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/mini-eda\">mini-eda</a></p>\n\n<p>You are welcome!</p>",
  "messages": [
    {
      "id": "833292",
      "postDate": "05/04/2020 17:57:47",
      "content": "<p>Hi everyone!</p>\n\n<p>Recently I have published <a href=\"https://www.kaggle.com/shonenkov/hack-with-parallel-corpus\">kernel</a> about parallel corpora,  pseudo-labeling and label smoothing.</p>\n\n<p>I would like to share with you <a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">my dataset</a>. </p>\n\n<p>It contains samples with token length &lt; 80</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/mini-eda\">mini-eda</a></p>\n\n<p>You are welcome!</p>",
      "rawMarkdown": "Hi everyone!\n\nRecently I have published [kernel](https://www.kaggle.com/shonenkov/hack-with-parallel-corpus) about parallel corpora,  pseudo-labeling and label smoothing.\n\nI would like to share with you [my dataset](https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling). \n\nIt contains samples with token length &lt; 80\n\n[mini-eda](https://www.kaggle.com/shonenkov/mini-eda)\n\nYou are welcome!",
      "votes": null
    },
    {
      "id": "833733",
      "postDate": "05/05/2020 03:14:39",
      "content": "<p>wow, thanks <a href=\"/shonenkov\">@shonenkov</a> for sharing :) </p>",
      "rawMarkdown": "wow, thanks @shonenkov for sharing :)",
      "votes": null
    },
    {
      "id": "841977",
      "postDate": "05/11/2020 05:41:23",
      "content": "<p>Thanks for sharing <a href=\"/shonenkov\">@shonenkov</a> . 💯 </p>",
      "rawMarkdown": "Thanks for sharing @shonenkov . 💯",
      "votes": null
    },
    {
      "id": "842241",
      "postDate": "05/11/2020 09:18:02",
      "content": "<p>Thanks for sharing <a href=\"/shonenkov\">@shonenkov</a>!</p>",
      "rawMarkdown": "Thanks for sharing @shonenkov!",
      "votes": null
    },
    {
      "id": "842552",
      "postDate": "05/11/2020 13:11:22",
      "content": "<p>That's great. </p>",
      "rawMarkdown": "That's great.",
      "votes": null
    },
    {
      "id": "848178",
      "postDate": "05/14/2020 19:33:38",
      "content": "<p>fyi the token length distribution of this dataset is very different from the training set and the validation set. </p>\n\n<p>in general, the open subtitles datasets contain a large proportion of extremely short phrases, which don't appear nearly as often in the training set and validation set. </p>",
      "rawMarkdown": "fyi the token length distribution of this dataset is very different from the training set and the validation set. \n\nin general, the open subtitles datasets contain a large proportion of extremely short phrases, which don't appear nearly as often in the training set and validation set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 833733,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "05/05/2020 03:14:39",
      "content": "<p>wow, thanks <a href=\"/shonenkov\">@shonenkov</a> for sharing :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 841977,
      "author_name": "",
      "author_url": "",
      "post_date": "05/11/2020 05:41:23",
      "content": "<p>Thanks for sharing <a href=\"/shonenkov\">@shonenkov</a> . 💯 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 842241,
      "author_name": "divyansh22",
      "author_url": "",
      "post_date": "05/11/2020 09:18:02",
      "content": "<p>Thanks for sharing <a href=\"/shonenkov\">@shonenkov</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 842552,
      "author_name": "amithasanshuvo",
      "author_url": "",
      "post_date": "05/11/2020 13:11:22",
      "content": "<p>That's great. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 848178,
      "author_name": "hooong",
      "author_url": "",
      "post_date": "05/14/2020 19:33:38",
      "content": "<p>fyi the token length distribution of this dataset is very different from the training set and the validation set. </p>\n\n<p>in general, the open subtitles datasets contain a large proportion of extremely short phrases, which don't appear nearly as often in the training set and validation set. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "833292": "Hi everyone!\n\nRecently I have published [kernel](https://www.kaggle.com/shonenkov/hack-with-parallel-corpus) about parallel corpora,  pseudo-labeling and label smoothing.\n\nI would like to share with you [my dataset](https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling). \n\nIt contains samples with token length &lt; 80\n\n[mini-eda](https://www.kaggle.com/shonenkov/mini-eda)\n\nYou are welcome!",
    "833733": "wow, thanks @shonenkov for sharing :)",
    "841977": "Thanks for sharing @shonenkov . 💯",
    "842241": "Thanks for sharing @shonenkov!",
    "842552": "That's great.",
    "848178": "fyi the token length distribution of this dataset is very different from the training set and the validation set. \n\nin general, the open subtitles datasets contain a large proportion of extremely short phrases, which don't appear nearly as often in the training set and validation set."
  },
  "source": "meta"
}