{
  "id": 451281,
  "title": "699th Rank for Bengali.AI Speech Recognition Challenge!",
  "url": "/competitions/bengaliai-speech/discussion/451281",
  "author_name": "Saeideh Mousavi",
  "post_date": "2023-10-27T23:26:19.756000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<ul>\n<li>Business context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech</a></li>\n<li>Data context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a></li>\n</ul>\n<p>I'm relatively new to competitions and submissions, and also this is my first time working with Audio data. I had zero knowledge about ASR and Audio models. During this competition, I gained a lot of experience with ASR and Transformers and implementing and training models.</p>\n<p><strong>Model</strong> : OpenAI Whisper-small pre-trained model <br>\n<strong>Training step</strong> : trained up to step 200,000 <br>\n<strong>Learning rate</strong> :1e-5<br>\n<strong>Batch_size</strong> =12<br>\n<strong>Trained on</strong> : 12GB RTX 3080 Ti<br>\n<strong>Trainer</strong>: Hugging face trainer<br>\n<strong>External Dataset</strong>: MUSAN<br>\n<strong>Normalization</strong>: Normalize Training Data with librosa.util.normalize<br>\n<strong>Audio Augmentation</strong>:Add 4 types of Audio Augmentation randomly to the data during the preparation process for training</p>\n<ul>\n<li><strong>Pitch</strong>: Change the pitch of the audio clip with librosa.effects.pitch_shift </li>\n<li><strong>Speed</strong>: Change the speed of the audio clip with librosa.effects.time_stretch </li>\n<li><strong>Noise</strong>: Add randomly chosen background noise audio clips from the MUSAN dataset to the Training Data</li>\n<li><strong>Music</strong>: Add randomly chosen background music instruments audio clips from the MUSAN dataset to the Training Data</li>\n</ul>\n<p>The reason why I chose the Whisper model was because, in the paper of the dataset, they mentioned that Whisper had better results than any other model. <br>\nI have done some audio augmentation on training data but it did not help a lot in getting a better score, which means these audio augmentation steps did not close the gap between test data and train data. <br>\nAfter the competitions and reading the leaderboard solution, I noticed that I should have spent more time on cleaning data from noisy audio clips, poorly annotated data, and poor-quality audio clips and making the training data more generic towards the test dataset and its diverse domain.</p>\n<p>Thank you to everyone who takes the time to read this, and I really appreciate any suggestions you may have to help me enhance my machine learning skills.</p>\n<p>Inference notebook: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/saeidehmousavi/bengali-competition-notebook</a><br>\nModel Training notebook:<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/saeidehmousavi/audio-augmentation-notebook</a></p>",
  "messages": [
    {
      "id": 2502074,
      "postDate": "2023-10-27T23:26:19.757Z",
      "content": "<ul>\n<li>Business context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech</a></li>\n<li>Data context: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a></li>\n</ul>\n<p>I'm relatively new to competitions and submissions, and also this is my first time working with Audio data. I had zero knowledge about ASR and Audio models. During this competition, I gained a lot of experience with ASR and Transformers and implementing and training models.</p>\n<p><strong>Model</strong> : OpenAI Whisper-small pre-trained model <br>\n<strong>Training step</strong> : trained up to step 200,000 <br>\n<strong>Learning rate</strong> :1e-5<br>\n<strong>Batch_size</strong> =12<br>\n<strong>Trained on</strong> : 12GB RTX 3080 Ti<br>\n<strong>Trainer</strong>: Hugging face trainer<br>\n<strong>External Dataset</strong>: MUSAN<br>\n<strong>Normalization</strong>: Normalize Training Data with librosa.util.normalize<br>\n<strong>Audio Augmentation</strong>:Add 4 types of Audio Augmentation randomly to the data during the preparation process for training</p>\n<ul>\n<li><strong>Pitch</strong>: Change the pitch of the audio clip with librosa.effects.pitch_shift </li>\n<li><strong>Speed</strong>: Change the speed of the audio clip with librosa.effects.time_stretch </li>\n<li><strong>Noise</strong>: Add randomly chosen background noise audio clips from the MUSAN dataset to the Training Data</li>\n<li><strong>Music</strong>: Add randomly chosen background music instruments audio clips from the MUSAN dataset to the Training Data</li>\n</ul>\n<p>The reason why I chose the Whisper model was because, in the paper of the dataset, they mentioned that Whisper had better results than any other model. <br>\nI have done some audio augmentation on training data but it did not help a lot in getting a better score, which means these audio augmentation steps did not close the gap between test data and train data. <br>\nAfter the competitions and reading the leaderboard solution, I noticed that I should have spent more time on cleaning data from noisy audio clips, poorly annotated data, and poor-quality audio clips and making the training data more generic towards the test dataset and its diverse domain.</p>\n<p>Thank you to everyone who takes the time to read this, and I really appreciate any suggestions you may have to help me enhance my machine learning skills.</p>\n<p>Inference notebook: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/saeidehmousavi/bengali-competition-notebook</a><br>\nModel Training notebook:<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/saeidehmousavi/audio-augmentation-notebook</a></p>",
      "rawMarkdown": "- Business context: [https://www.kaggle.com/competitions/bengaliai-speech](url)\n- Data context: [https://www.kaggle.com/competitions/bengaliai-speech/data](url)\n\nI'm relatively new to competitions and submissions, and also this is my first time working with Audio data. I had zero knowledge about ASR and Audio models. During this competition, I gained a lot of experience with ASR and Transformers and implementing and training models.\n\n**Model** : OpenAI Whisper-small pre-trained model \n**Training step** : trained up to step 200,000 \n**Learning rate** :1e-5\n**Batch_size** =12\n**Trained on** : 12GB RTX 3080 Ti\n**Trainer**: Hugging face trainer\n**External Dataset**: MUSAN\n**Normalization**: Normalize Training Data with librosa.util.normalize\n**Audio Augmentation**:Add 4 types of Audio Augmentation randomly to the data during the preparation process for training\n-  **Pitch**: Change the pitch of the audio clip with librosa.effects.pitch_shift \n-  **Speed**: Change the speed of the audio clip with librosa.effects.time_stretch \n-  **Noise**: Add randomly chosen background noise audio clips from the MUSAN dataset to the Training Data\n-  **Music**: Add randomly chosen background music instruments audio clips from the MUSAN dataset to the Training Data\n\n\n\nThe reason why I chose the Whisper model was because, in the paper of the dataset, they mentioned that Whisper had better results than any other model. \nI have done some audio augmentation on training data but it did not help a lot in getting a better score, which means these audio augmentation steps did not close the gap between test data and train data. \nAfter the competitions and reading the leaderboard solution, I noticed that I should have spent more time on cleaning data from noisy audio clips, poorly annotated data, and poor-quality audio clips and making the training data more generic towards the test dataset and its diverse domain.\n \nThank you to everyone who takes the time to read this, and I really appreciate any suggestions you may have to help me enhance my machine learning skills.\n\nInference notebook: [https://www.kaggle.com/code/saeidehmousavi/bengali-competition-notebook](url)\nModel Training notebook:[https://www.kaggle.com/code/saeidehmousavi/audio-augmentation-notebook](url)\n\n\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2502074": "- Business context: [https://www.kaggle.com/competitions/bengaliai-speech](url)\n- Data context: [https://www.kaggle.com/competitions/bengaliai-speech/data](url)\n\nI'm relatively new to competitions and submissions, and also this is my first time working with Audio data. I had zero knowledge about ASR and Audio models. During this competition, I gained a lot of experience with ASR and Transformers and implementing and training models.\n\n**Model** : OpenAI Whisper-small pre-trained model \n**Training step** : trained up to step 200,000 \n**Learning rate** :1e-5\n**Batch_size** =12\n**Trained on** : 12GB RTX 3080 Ti\n**Trainer**: Hugging face trainer\n**External Dataset**: MUSAN\n**Normalization**: Normalize Training Data with librosa.util.normalize\n**Audio Augmentation**:Add 4 types of Audio Augmentation randomly to the data during the preparation process for training\n-  **Pitch**: Change the pitch of the audio clip with librosa.effects.pitch_shift \n-  **Speed**: Change the speed of the audio clip with librosa.effects.time_stretch \n-  **Noise**: Add randomly chosen background noise audio clips from the MUSAN dataset to the Training Data\n-  **Music**: Add randomly chosen background music instruments audio clips from the MUSAN dataset to the Training Data\n\n\n\nThe reason why I chose the Whisper model was because, in the paper of the dataset, they mentioned that Whisper had better results than any other model. \nI have done some audio augmentation on training data but it did not help a lot in getting a better score, which means these audio augmentation steps did not close the gap between test data and train data. \nAfter the competitions and reading the leaderboard solution, I noticed that I should have spent more time on cleaning data from noisy audio clips, poorly annotated data, and poor-quality audio clips and making the training data more generic towards the test dataset and its diverse domain.\n \nThank you to everyone who takes the time to read this, and I really appreciate any suggestions you may have to help me enhance my machine learning skills.\n\nInference notebook: [https://www.kaggle.com/code/saeidehmousavi/bengali-competition-notebook](url)\nModel Training notebook:[https://www.kaggle.com/code/saeidehmousavi/audio-augmentation-notebook](url)\n\n\n"
  }
}