{
  "id": 159589,
  "title": "Read BERT preprocessed dataset",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159589",
  "author_name": "",
  "post_date": "2020-06-17T22:41:19.540127600Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In the dataset some files are preprocessed for BERT, by tokenizing the 'comment_text'. My question is how do you read the files such that the token columns is correctly read as arrays? When I do it it is only read as objects of string. I have been reading people's shared notebooks and all of them tokenize the raw dataset again in their code.</p>",
  "messages": [
    {
      "id": "891104",
      "postDate": "06/17/2020 22:41:19",
      "content": "<p>In the dataset some files are preprocessed for BERT, by tokenizing the 'comment_text'. My question is how do you read the files such that the token columns is correctly read as arrays? When I do it it is only read as objects of string. I have been reading people's shared notebooks and all of them tokenize the raw dataset again in their code.</p>",
      "rawMarkdown": "In the dataset some files are preprocessed for BERT, by tokenizing the 'comment_text'. My question is how do you read the files such that the token columns is correctly read as arrays? When I do it it is only read as objects of string. I have been reading people's shared notebooks and all of them tokenize the raw dataset again in their code.",
      "votes": null
    },
    {
      "id": "891954",
      "postDate": "06/18/2020 14:56:08",
      "content": "<p>You can write a function to take the string, remove the parentheses, split it by commas, and then turn it into an array of ints. Hope that helps!</p>",
      "rawMarkdown": "You can write a function to take the string, remove the parentheses, split it by commas, and then turn it into an array of ints. Hope that helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 891954,
      "author_name": "anasofiauzsoy",
      "author_url": "",
      "post_date": "06/18/2020 14:56:08",
      "content": "<p>You can write a function to take the string, remove the parentheses, split it by commas, and then turn it into an array of ints. Hope that helps!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "891104": "In the dataset some files are preprocessed for BERT, by tokenizing the 'comment_text'. My question is how do you read the files such that the token columns is correctly read as arrays? When I do it it is only read as objects of string. I have been reading people's shared notebooks and all of them tokenize the raw dataset again in their code.",
    "891954": "You can write a function to take the string, remove the parentheses, split it by commas, and then turn it into an array of ints. Hope that helps!"
  },
  "source": "meta"
}