{
  "id": 143469,
  "title": "Better tokenizers",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143469",
  "author_name": "arseny-n",
  "post_date": "2020-04-15T09:00:38.941000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello, currently I am using a huggingface XLM-Roberta model, the tokenization takes from 7 to 10 minutes. The default tokenizer uses under the hood a sentencepice tokenizer. I am currently trying to use a fast hugging face tokenizer. </p>\n\n<p>I can't find an easy way to convert a <code>sentencepiece.bpe.model</code> file which is created by the default tokenizer to  <code>vocab.txt</code> and <code>merges.txt</code> files required by the hugging face tokenizer. </p>\n\n<p>So I have two questions:\n- Has anyone encountered similar issues ? \n- And will it be useful to do so ? I couldn't find a benchmark of tokenizers but both sentencepice and huggingface tokenizers seem to be implemented with speed in mind, so maybe the current tokenizer is the optimal choice.</p>",
  "messages": [
    {
      "id": 808265,
      "postDate": "2020-04-15T09:05:14.043Z",
      "content": "<p>You are not using // processing; <a href=\"https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly?scriptVersionId=31854117\">self-plug</a></p>",
      "rawMarkdown": "You are not using // processing; [self-plug](https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly?scriptVersionId=31854117)",
      "votes": -1,
      "replies": [
        {
          "id": 809048,
          "postDate": "2020-04-15T19:47:16.247Z",
          "content": "<p>Thanks for the hint! I arrived to a similar conclusion - now I use the joblib library to parallelize the tokenization the time went form 10 mins to 3, but it still would be interesting to know weather there is a way to use the faster tokenizers. </p>",
          "rawMarkdown": "Thanks for the hint! I arrived to a similar conclusion - now I use the joblib library to parallelize the tokenization the time went form 10 mins to 3, but it still would be interesting to know weather there is a way to use the faster tokenizers. "
        }
      ]
    },
    {
      "id": 808259,
      "postDate": "2020-04-15T09:00:38.943Z",
      "content": "<p>Hello, currently I am using a huggingface XLM-Roberta model, the tokenization takes from 7 to 10 minutes. The default tokenizer uses under the hood a sentencepice tokenizer. I am currently trying to use a fast hugging face tokenizer. </p>\n\n<p>I can't find an easy way to convert a <code>sentencepiece.bpe.model</code> file which is created by the default tokenizer to  <code>vocab.txt</code> and <code>merges.txt</code> files required by the hugging face tokenizer. </p>\n\n<p>So I have two questions:\n- Has anyone encountered similar issues ? \n- And will it be useful to do so ? I couldn't find a benchmark of tokenizers but both sentencepice and huggingface tokenizers seem to be implemented with speed in mind, so maybe the current tokenizer is the optimal choice.</p>",
      "rawMarkdown": "Hello, currently I am using a huggingface XLM-Roberta model, the tokenization takes from 7 to 10 minutes. The default tokenizer uses under the hood a sentencepice tokenizer. I am currently trying to use a fast hugging face tokenizer. \n\nI can't find an easy way to convert a `sentencepiece.bpe.model` file which is created by the default tokenizer to  `vocab.txt` and `merges.txt` files required by the hugging face tokenizer. \n\nSo I have two questions:\n- Has anyone encountered similar issues ? \n- And will it be useful to do so ? I couldn't find a benchmark of tokenizers but both sentencepice and huggingface tokenizers seem to be implemented with speed in mind, so maybe the current tokenizer is the optimal choice."
    },
    {
      "id": 829982,
      "postDate": "2020-05-02T08:19:09.057Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 808265,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-04-15T09:05:14.043000",
      "content": "<p>You are not using // processing; <a href=\"https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly?scriptVersionId=31854117\">self-plug</a></p>",
      "votes": -1,
      "replies": [
        {
          "id": 809048,
          "author_name": "arseny-n",
          "author_url": "",
          "post_date": "2020-04-15T19:47:16.247000",
          "content": "<p>Thanks for the hint! I arrived to a similar conclusion - now I use the joblib library to parallelize the tokenization the time went form 10 mins to 3, but it still would be interesting to know weather there is a way to use the faster tokenizers. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 829982,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-02T08:19:09.057000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "808265": "You are not using // processing; [self-plug](https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly?scriptVersionId=31854117)",
    "808259": "Hello, currently I am using a huggingface XLM-Roberta model, the tokenization takes from 7 to 10 minutes. The default tokenizer uses under the hood a sentencepice tokenizer. I am currently trying to use a fast hugging face tokenizer. \n\nI can't find an easy way to convert a `sentencepiece.bpe.model` file which is created by the default tokenizer to  `vocab.txt` and `merges.txt` files required by the hugging face tokenizer. \n\nSo I have two questions:\n- Has anyone encountered similar issues ? \n- And will it be useful to do so ? I couldn't find a benchmark of tokenizers but both sentencepice and huggingface tokenizers seem to be implemented with speed in mind, so maybe the current tokenizer is the optimal choice.",
    "829982": ""
  }
}