{
  "id": 431175,
  "title": "Get the vocabulary list from tokenizer for Bengali",
  "url": "/competitions/bengaliai-speech/discussion/431175",
  "author_name": "Pritam Sinha",
  "post_date": "2023-08-12T10:52:16.152000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi! I was playing with the tokenizer of Whisper. I see the length of the tokenizer vocabulary is 50364. Among these, how do I get only the vocab for the Bengali language? I want to see all the tokens used in the tokenizer for the Bengali language. How do I get those tokens and their respective IDs from the tokenizer?<br>\nSecondly, What type of tokenization does Whisper tokenizer use? Is it char level, word level, or BPE?<br>\nThirdly, I see 2 functions <code>convert_ids_to_tokens</code> and <code>batch_decode</code> in the tokenizer, how are these 2 different?</p>\n<p>Can someone clarify</p>",
  "messages": [
    {
      "id": 2387025,
      "postDate": "2023-08-12T10:52:16.153Z",
      "content": "<p>Hi! I was playing with the tokenizer of Whisper. I see the length of the tokenizer vocabulary is 50364. Among these, how do I get only the vocab for the Bengali language? I want to see all the tokens used in the tokenizer for the Bengali language. How do I get those tokens and their respective IDs from the tokenizer?<br>\nSecondly, What type of tokenization does Whisper tokenizer use? Is it char level, word level, or BPE?<br>\nThirdly, I see 2 functions <code>convert_ids_to_tokens</code> and <code>batch_decode</code> in the tokenizer, how are these 2 different?</p>\n<p>Can someone clarify</p>",
      "rawMarkdown": "Hi! I was playing with the tokenizer of Whisper. I see the length of the tokenizer vocabulary is 50364. Among these, how do I get only the vocab for the Bengali language? I want to see all the tokens used in the tokenizer for the Bengali language. How do I get those tokens and their respective IDs from the tokenizer?\nSecondly, What type of tokenization does Whisper tokenizer use? Is it char level, word level, or BPE?\nThirdly, I see 2 functions `convert_ids_to_tokens` and `batch_decode` in the tokenizer, how are these 2 different?\n\nCan someone clarify",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2387025": "Hi! I was playing with the tokenizer of Whisper. I see the length of the tokenizer vocabulary is 50364. Among these, how do I get only the vocab for the Bengali language? I want to see all the tokens used in the tokenizer for the Bengali language. How do I get those tokens and their respective IDs from the tokenizer?\nSecondly, What type of tokenization does Whisper tokenizer use? Is it char level, word level, or BPE?\nThirdly, I see 2 functions `convert_ids_to_tokens` and `batch_decode` in the tokenizer, how are these 2 different?\n\nCan someone clarify"
  }
}