{
  "id": 451204,
  "title": "75th Place Solution - Bronze - My First Competition Medal ",
  "url": "/competitions/bengaliai-speech/discussion/451204",
  "author_name": "",
  "post_date": "2023-10-27T14:37:11.148161600Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I wanted to take a moment to extend my heartfelt thanks to you and your team for organizing the Competition - Bengali AI Speech Recognition. It has been an incredible journey and a memorable experience.<br>\nI am grateful for the opportunity to participate in this competition. <br>\nI also want to acknowledge the camaraderie and the sense of community that was fostered throughout the competition. The chance to interact with fellow participants, learn from one another, and share our experiences was truly invaluable.</p>\n<p>Here is my approach,</p>\n<p><strong>Solution Overview</strong><br>\n<em>1. Importing Required Libraries</em><br>\nThe solution began by importing necessary Python packages and dependencies, including packages for text normalization, decoding, and audio processing.</p>\n<p><em>2. Model and Preprocessing</em><br>\n<strong><em>Load Model and Decoder</em></strong><br>\nThe solution used the Wav2Vec2 model for CTC (Connectionist Temporal Classification) from Hugging Face.<br>\nA language model (LM) was loaded to assist in decoding.<br>\nData Preprocessing<br>\nThe audio data was loaded and preprocessed using the librosa library.<br>\nThe audio was transformed into features suitable for the model.<br>\nVocabulary and decoders were initialized for later use.</p>\n<p><em>3. Data Preparation</em><br>\nThe solution prepared the test dataset, loading audio files, and creating a DataLoader for batch processing.</p>\n<p><em>4. Inference</em><br>\nThe model was used for inference, where it predicted the text transcription for the test audio data.<br>\nThe predictions were post-processed to improve the readability of the output.</p>\n<p><em>5. Make Submission</em><br>\nThe final step was to format the predictions into a submission file for the Kaggle competition. The output sentences were further processed to ensure proper formatting with Bengali punctuation.</p>\n<p><em>Results</em><br>\nThe solution achieved a significant milestone in converting Bengali speech into text. It is important to note that the performance of this solution depends on the quality and size of the training data and the specific design of the language model and decoder.</p>\n<h3>Solution Overview</h3>\n<p>The solution is detailed below, along with visualizations to help understand the process.</p>\n<h4>1. Importing Required Libraries</h4>\n<p>The process begins with importing essential libraries and dependencies for the solution:</p>\n<pre><code> pandas  pd\n pyctcdecode\n kenlm\n torch\n librosa\n numpy  np\n</code></pre>\n<h4>2. Model and Preprocessing</h4>\n<h5>Load Model and Decoder</h5>\n<p>The Wav2Vec2 model and language model decoder are loaded. The language model assists in decoding the transcriptions.</p>\n<h4>3. Data Preparation</h4>\n<p>The solution prepares the test dataset by loading audio files and creating a DataLoader for batch processing. The test dataset is stored in a DataFrame, as seen below:</p>\n<pre><code>test = pd.read_csv(DATA / , dtype={: })\n(test.head())\n</code></pre>\n<h4>4. Inference</h4>\n<p>Inference is a crucial step in converting audio to text. The code leverages the Wav2Vec2 model to predict transcriptions for the test audio data. Here is an example of inference:</p>\n<pre><code> torch.no_grad():\n     batch  tqdm(test_loader):\n        x = batch[]\n        x = x.to(device, non-blocking=)\n         torch.cuda.amp.autocast():\n            y = model(x).logits\n        y = y.detach().cpu().numpy()\n\n        \n         l  y:\n            beam = decoder.decode_beams(l, beam_width=)\n            s = beam[][]\n            pred_sentence_list.append(s)\n</code></pre>\n<h4>5. Make Submission</h4>\n<p>The final step involves formatting the predictions into a submission file for the Kaggle competition. The output sentences are further processed to ensure proper formatting with Bengali punctuation:</p>\n<pre><code> ():\n    \n    period_set = ([, , , ])\n    _words = [bnorm(word)[]   word  sentence.split()]\n    sentence = .join([word  word  _words  word   ])\n    :\n         sentence[-]   period_set:\n            sentence += \n    :\n        sentence = \n     sentence\n</code></pre>\n<h4>Visualizations</h4>\n<p>To visualize the process, here's an illustration of the data flow:</p>\n<p>Start<br>\n|<br>\n|--- Import Libraries<br>\n|   |<br>\n|   |--- Pandas, pyctcdecode, kenlm, torch, librosa, numpy, etc.<br>\n|<br>\n|--- Load Test Dataset<br>\n|   |<br>\n|   |--- Read sample_submission.csv<br>\n|<br>\n|--- Inference<br>\n|   |<br>\n|   |--- For each batch of test audio data:<br>\n|   |   |<br>\n|   |   |--- Preprocess Audio<br>\n|   |   |   |<br>\n|   |   |   |--- Extract features<br>\n|   |   |<br>\n|   |   |--- Wav2Vec2 Model Inference<br>\n|   |   |   |<br>\n|   |   |   |--- Convert audio to text transcriptions<br>\n|   |<br>\n|   |--- Decode Transcriptions<br>\n|   |   |<br>\n|   |   |--- Language Model (KenLM) Decoding<br>\n|<br>\n|--- Postprocess Transcriptions<br>\n|   |<br>\n|   |--- Normalize and format transcriptions<br>\n|<br>\n|--- Make Submission<br>\n|   |<br>\n|   |--- Format predictions for submission<br>\n|<br>\nEnd</p>\n<ul>\n<li>Data flow starts with importing libraries and dependencies.</li>\n<li>The test dataset is loaded from CSV.</li>\n<li>Audio files are processed through the Wav2Vec2 model for inference.</li>\n<li>Predicted transcriptions are post-processed for better readability.</li>\n</ul>\n<h3>Results</h3>\n<p>The solution achieved a significant milestone in converting Bengali speech into text. </p>\n<h3>Conclusion</h3>\n<p>The Bengali Speech Recognition Kaggle Competition is a challenging task with numerous real-world applications. The provided solution leveraged advanced deep learning models, pre-processing techniques, and post-processing steps to provide accurate transcriptions of Bengali speech.</p>\n<p>Thanks a lot Everyone </p>\n<p>Happy Learning !!!!</p>",
  "messages": [
    {
      "id": "2501587",
      "postDate": "10/27/2023 14:37:11",
      "content": "<p>I wanted to take a moment to extend my heartfelt thanks to you and your team for organizing the Competition - Bengali AI Speech Recognition. It has been an incredible journey and a memorable experience.<br>\nI am grateful for the opportunity to participate in this competition. <br>\nI also want to acknowledge the camaraderie and the sense of community that was fostered throughout the competition. The chance to interact with fellow participants, learn from one another, and share our experiences was truly invaluable.</p>\n<p>Here is my approach,</p>\n<p><strong>Solution Overview</strong><br>\n<em>1. Importing Required Libraries</em><br>\nThe solution began by importing necessary Python packages and dependencies, including packages for text normalization, decoding, and audio processing.</p>\n<p><em>2. Model and Preprocessing</em><br>\n<strong><em>Load Model and Decoder</em></strong><br>\nThe solution used the Wav2Vec2 model for CTC (Connectionist Temporal Classification) from Hugging Face.<br>\nA language model (LM) was loaded to assist in decoding.<br>\nData Preprocessing<br>\nThe audio data was loaded and preprocessed using the librosa library.<br>\nThe audio was transformed into features suitable for the model.<br>\nVocabulary and decoders were initialized for later use.</p>\n<p><em>3. Data Preparation</em><br>\nThe solution prepared the test dataset, loading audio files, and creating a DataLoader for batch processing.</p>\n<p><em>4. Inference</em><br>\nThe model was used for inference, where it predicted the text transcription for the test audio data.<br>\nThe predictions were post-processed to improve the readability of the output.</p>\n<p><em>5. Make Submission</em><br>\nThe final step was to format the predictions into a submission file for the Kaggle competition. The output sentences were further processed to ensure proper formatting with Bengali punctuation.</p>\n<p><em>Results</em><br>\nThe solution achieved a significant milestone in converting Bengali speech into text. It is important to note that the performance of this solution depends on the quality and size of the training data and the specific design of the language model and decoder.</p>\n<h3>Solution Overview</h3>\n<p>The solution is detailed below, along with visualizations to help understand the process.</p>\n<h4>1. Importing Required Libraries</h4>\n<p>The process begins with importing essential libraries and dependencies for the solution:</p>\n<pre><code> pandas  pd\n pyctcdecode\n kenlm\n torch\n librosa\n numpy  np\n</code></pre>\n<h4>2. Model and Preprocessing</h4>\n<h5>Load Model and Decoder</h5>\n<p>The Wav2Vec2 model and language model decoder are loaded. The language model assists in decoding the transcriptions.</p>\n<h4>3. Data Preparation</h4>\n<p>The solution prepares the test dataset by loading audio files and creating a DataLoader for batch processing. The test dataset is stored in a DataFrame, as seen below:</p>\n<pre><code>test = pd.read_csv(DATA / , dtype={: })\n(test.head())\n</code></pre>\n<h4>4. Inference</h4>\n<p>Inference is a crucial step in converting audio to text. The code leverages the Wav2Vec2 model to predict transcriptions for the test audio data. Here is an example of inference:</p>\n<pre><code> torch.no_grad():\n     batch  tqdm(test_loader):\n        x = batch[]\n        x = x.to(device, non-blocking=)\n         torch.cuda.amp.autocast():\n            y = model(x).logits\n        y = y.detach().cpu().numpy()\n\n        \n         l  y:\n            beam = decoder.decode_beams(l, beam_width=)\n            s = beam[][]\n            pred_sentence_list.append(s)\n</code></pre>\n<h4>5. Make Submission</h4>\n<p>The final step involves formatting the predictions into a submission file for the Kaggle competition. The output sentences are further processed to ensure proper formatting with Bengali punctuation:</p>\n<pre><code> ():\n    \n    period_set = ([, , , ])\n    _words = [bnorm(word)[]   word  sentence.split()]\n    sentence = .join([word  word  _words  word   ])\n    :\n         sentence[-]   period_set:\n            sentence += \n    :\n        sentence = \n     sentence\n</code></pre>\n<h4>Visualizations</h4>\n<p>To visualize the process, here's an illustration of the data flow:</p>\n<p>Start<br>\n|<br>\n|--- Import Libraries<br>\n|   |<br>\n|   |--- Pandas, pyctcdecode, kenlm, torch, librosa, numpy, etc.<br>\n|<br>\n|--- Load Test Dataset<br>\n|   |<br>\n|   |--- Read sample_submission.csv<br>\n|<br>\n|--- Inference<br>\n|   |<br>\n|   |--- For each batch of test audio data:<br>\n|   |   |<br>\n|   |   |--- Preprocess Audio<br>\n|   |   |   |<br>\n|   |   |   |--- Extract features<br>\n|   |   |<br>\n|   |   |--- Wav2Vec2 Model Inference<br>\n|   |   |   |<br>\n|   |   |   |--- Convert audio to text transcriptions<br>\n|   |<br>\n|   |--- Decode Transcriptions<br>\n|   |   |<br>\n|   |   |--- Language Model (KenLM) Decoding<br>\n|<br>\n|--- Postprocess Transcriptions<br>\n|   |<br>\n|   |--- Normalize and format transcriptions<br>\n|<br>\n|--- Make Submission<br>\n|   |<br>\n|   |--- Format predictions for submission<br>\n|<br>\nEnd</p>\n<ul>\n<li>Data flow starts with importing libraries and dependencies.</li>\n<li>The test dataset is loaded from CSV.</li>\n<li>Audio files are processed through the Wav2Vec2 model for inference.</li>\n<li>Predicted transcriptions are post-processed for better readability.</li>\n</ul>\n<h3>Results</h3>\n<p>The solution achieved a significant milestone in converting Bengali speech into text. </p>\n<h3>Conclusion</h3>\n<p>The Bengali Speech Recognition Kaggle Competition is a challenging task with numerous real-world applications. The provided solution leveraged advanced deep learning models, pre-processing techniques, and post-processing steps to provide accurate transcriptions of Bengali speech.</p>\n<p>Thanks a lot Everyone </p>\n<p>Happy Learning !!!!</p>",
      "rawMarkdown": "I wanted to take a moment to extend my heartfelt thanks to you and your team for organizing the Competition - Bengali AI Speech Recognition. It has been an incredible journey and a memorable experience.\nI am grateful for the opportunity to participate in this competition. \nI also want to acknowledge the camaraderie and the sense of community that was fostered throughout the competition. The chance to interact with fellow participants, learn from one another, and share our experiences was truly invaluable.\n\nHere is my approach,\n\n**Solution Overview**\n*1. Importing Required Libraries*\nThe solution began by importing necessary Python packages and dependencies, including packages for text normalization, decoding, and audio processing.\n\n*2. Model and Preprocessing*\n***Load Model and Decoder***\nThe solution used the Wav2Vec2 model for CTC (Connectionist Temporal Classification) from Hugging Face.\nA language model (LM) was loaded to assist in decoding.\nData Preprocessing\nThe audio data was loaded and preprocessed using the librosa library.\nThe audio was transformed into features suitable for the model.\nVocabulary and decoders were initialized for later use.\n\n*3. Data Preparation*\nThe solution prepared the test dataset, loading audio files, and creating a DataLoader for batch processing.\n\n*4. Inference*\nThe model was used for inference, where it predicted the text transcription for the test audio data.\nThe predictions were post-processed to improve the readability of the output.\n\n*5. Make Submission*\nThe final step was to format the predictions into a submission file for the Kaggle competition. The output sentences were further processed to ensure proper formatting with Bengali punctuation.\n\n*Results*\nThe solution achieved a significant milestone in converting Bengali speech into text. It is important to note that the performance of this solution depends on the quality and size of the training data and the specific design of the language model and decoder.\n\n### Solution Overview\n\nThe solution is detailed below, along with visualizations to help understand the process.\n\n#### 1. Importing Required Libraries\n\nThe process begins with importing essential libraries and dependencies for the solution:\n\n```python\nimport pandas as pd\nimport pyctcdecode\nimport kenlm\nimport torch\nimport librosa\nimport numpy as np\n```\n\n#### 2. Model and Preprocessing\n\n##### Load Model and Decoder\n\nThe Wav2Vec2 model and language model decoder are loaded. The language model assists in decoding the transcriptions.\n\n#### 3. Data Preparation\n\nThe solution prepares the test dataset by loading audio files and creating a DataLoader for batch processing. The test dataset is stored in a DataFrame, as seen below:\n\n```python\ntest = pd.read_csv(DATA / \"sample_submission.csv\", dtype={\"id\": str})\nprint(test.head())\n```\n\n#### 4. Inference\n\nInference is a crucial step in converting audio to text. The code leverages the Wav2Vec2 model to predict transcriptions for the test audio data. Here is an example of inference:\n\n```python\nwith torch.no_grad():\n    for batch in tqdm(test_loader):\n        x = batch[\"input_values\"]\n        x = x.to(device, non-blocking=True)\n        with torch.cuda.amp.autocast(True):\n            y = model(x).logits\n        y = y.detach().cpu().numpy()\n        \n        # Decoding the transcriptions\n        for l in y:\n            beam = decoder.decode_beams(l, beam_width=512)\n            s = beam[0][0]\n            pred_sentence_list.append(s)\n```\n\n#### 5. Make Submission\n\nThe final step involves formatting the predictions into a submission file for the Kaggle competition. The output sentences are further processed to ensure proper formatting with Bengali punctuation:\n\n```python\ndef postprocess(sentence):\n    # Post-processing for better readability\n    period_set = set([\".\", \"?\", \"!\", \"।\"])\n    _words = [bnorm(word)['normalized']  for word in sentence.split()]\n    sentence = \" \".join([word for word in _words if word is not None])\n    try:\n        if sentence[-1] not in period_set:\n            sentence += \"।\"\n    except:\n        sentence = \"।\"\n    return sentence\n```\n\n#### Visualizations\n\nTo visualize the process, here's an illustration of the data flow:\n\nStart\n|\n|--- Import Libraries\n|   |\n|   |--- Pandas, pyctcdecode, kenlm, torch, librosa, numpy, etc.\n|\n|--- Load Test Dataset\n|   |\n|   |--- Read sample_submission.csv\n|\n|--- Inference\n|   |\n|   |--- For each batch of test audio data:\n|   |   |\n|   |   |--- Preprocess Audio\n|   |   |   |\n|   |   |   |--- Extract features\n|   |   |\n|   |   |--- Wav2Vec2 Model Inference\n|   |   |   |\n|   |   |   |--- Convert audio to text transcriptions\n|   |\n|   |--- Decode Transcriptions\n|   |   |\n|   |   |--- Language Model (KenLM) Decoding\n|\n|--- Postprocess Transcriptions\n|   |\n|   |--- Normalize and format transcriptions\n|\n|--- Make Submission\n|   |\n|   |--- Format predictions for submission\n|\nEnd\n\n\n- Data flow starts with importing libraries and dependencies.\n- The test dataset is loaded from CSV.\n- Audio files are processed through the Wav2Vec2 model for inference.\n- Predicted transcriptions are post-processed for better readability.\n\n### Results\n\nThe solution achieved a significant milestone in converting Bengali speech into text. \n\n### Conclusion\n\nThe Bengali Speech Recognition Kaggle Competition is a challenging task with numerous real-world applications. The provided solution leveraged advanced deep learning models, pre-processing techniques, and post-processing steps to provide accurate transcriptions of Bengali speech.\n\nThanks a lot Everyone \n\nHappy Learning !!!!",
      "votes": null
    },
    {
      "id": "2501656",
      "postDate": "10/27/2023 15:34:27",
      "content": "<p>Wow, that is awesome!!!!</p>",
      "rawMarkdown": "Wow, that is awesome!!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2501656,
      "author_name": "asif00",
      "author_url": "",
      "post_date": "10/27/2023 15:34:27",
      "content": "<p>Wow, that is awesome!!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2501587": "I wanted to take a moment to extend my heartfelt thanks to you and your team for organizing the Competition - Bengali AI Speech Recognition. It has been an incredible journey and a memorable experience.\nI am grateful for the opportunity to participate in this competition. \nI also want to acknowledge the camaraderie and the sense of community that was fostered throughout the competition. The chance to interact with fellow participants, learn from one another, and share our experiences was truly invaluable.\n\nHere is my approach,\n\n**Solution Overview**\n*1. Importing Required Libraries*\nThe solution began by importing necessary Python packages and dependencies, including packages for text normalization, decoding, and audio processing.\n\n*2. Model and Preprocessing*\n***Load Model and Decoder***\nThe solution used the Wav2Vec2 model for CTC (Connectionist Temporal Classification) from Hugging Face.\nA language model (LM) was loaded to assist in decoding.\nData Preprocessing\nThe audio data was loaded and preprocessed using the librosa library.\nThe audio was transformed into features suitable for the model.\nVocabulary and decoders were initialized for later use.\n\n*3. Data Preparation*\nThe solution prepared the test dataset, loading audio files, and creating a DataLoader for batch processing.\n\n*4. Inference*\nThe model was used for inference, where it predicted the text transcription for the test audio data.\nThe predictions were post-processed to improve the readability of the output.\n\n*5. Make Submission*\nThe final step was to format the predictions into a submission file for the Kaggle competition. The output sentences were further processed to ensure proper formatting with Bengali punctuation.\n\n*Results*\nThe solution achieved a significant milestone in converting Bengali speech into text. It is important to note that the performance of this solution depends on the quality and size of the training data and the specific design of the language model and decoder.\n\n### Solution Overview\n\nThe solution is detailed below, along with visualizations to help understand the process.\n\n#### 1. Importing Required Libraries\n\nThe process begins with importing essential libraries and dependencies for the solution:\n\n```python\nimport pandas as pd\nimport pyctcdecode\nimport kenlm\nimport torch\nimport librosa\nimport numpy as np\n```\n\n#### 2. Model and Preprocessing\n\n##### Load Model and Decoder\n\nThe Wav2Vec2 model and language model decoder are loaded. The language model assists in decoding the transcriptions.\n\n#### 3. Data Preparation\n\nThe solution prepares the test dataset by loading audio files and creating a DataLoader for batch processing. The test dataset is stored in a DataFrame, as seen below:\n\n```python\ntest = pd.read_csv(DATA / \"sample_submission.csv\", dtype={\"id\": str})\nprint(test.head())\n```\n\n#### 4. Inference\n\nInference is a crucial step in converting audio to text. The code leverages the Wav2Vec2 model to predict transcriptions for the test audio data. Here is an example of inference:\n\n```python\nwith torch.no_grad():\n    for batch in tqdm(test_loader):\n        x = batch[\"input_values\"]\n        x = x.to(device, non-blocking=True)\n        with torch.cuda.amp.autocast(True):\n            y = model(x).logits\n        y = y.detach().cpu().numpy()\n        \n        # Decoding the transcriptions\n        for l in y:\n            beam = decoder.decode_beams(l, beam_width=512)\n            s = beam[0][0]\n            pred_sentence_list.append(s)\n```\n\n#### 5. Make Submission\n\nThe final step involves formatting the predictions into a submission file for the Kaggle competition. The output sentences are further processed to ensure proper formatting with Bengali punctuation:\n\n```python\ndef postprocess(sentence):\n    # Post-processing for better readability\n    period_set = set([\".\", \"?\", \"!\", \"।\"])\n    _words = [bnorm(word)['normalized']  for word in sentence.split()]\n    sentence = \" \".join([word for word in _words if word is not None])\n    try:\n        if sentence[-1] not in period_set:\n            sentence += \"।\"\n    except:\n        sentence = \"।\"\n    return sentence\n```\n\n#### Visualizations\n\nTo visualize the process, here's an illustration of the data flow:\n\nStart\n|\n|--- Import Libraries\n|   |\n|   |--- Pandas, pyctcdecode, kenlm, torch, librosa, numpy, etc.\n|\n|--- Load Test Dataset\n|   |\n|   |--- Read sample_submission.csv\n|\n|--- Inference\n|   |\n|   |--- For each batch of test audio data:\n|   |   |\n|   |   |--- Preprocess Audio\n|   |   |   |\n|   |   |   |--- Extract features\n|   |   |\n|   |   |--- Wav2Vec2 Model Inference\n|   |   |   |\n|   |   |   |--- Convert audio to text transcriptions\n|   |\n|   |--- Decode Transcriptions\n|   |   |\n|   |   |--- Language Model (KenLM) Decoding\n|\n|--- Postprocess Transcriptions\n|   |\n|   |--- Normalize and format transcriptions\n|\n|--- Make Submission\n|   |\n|   |--- Format predictions for submission\n|\nEnd\n\n\n- Data flow starts with importing libraries and dependencies.\n- The test dataset is loaded from CSV.\n- Audio files are processed through the Wav2Vec2 model for inference.\n- Predicted transcriptions are post-processed for better readability.\n\n### Results\n\nThe solution achieved a significant milestone in converting Bengali speech into text. \n\n### Conclusion\n\nThe Bengali Speech Recognition Kaggle Competition is a challenging task with numerous real-world applications. The provided solution leveraged advanced deep learning models, pre-processing techniques, and post-processing steps to provide accurate transcriptions of Bengali speech.\n\nThanks a lot Everyone \n\nHappy Learning !!!!",
    "2501656": "Wow, that is awesome!!!!"
  },
  "source": "meta"
}