{
  "id": 450606,
  "title": "🥈24th in two weeks and `No space left on device`!",
  "url": "/competitions/bengaliai-speech/writeups/bayartsogt-yadamsuren-24th-in-two-weeks-and-no-spa",
  "author_name": "",
  "post_date": "2023-10-25T01:56:29.600493800Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>First of all, THANK YOU ALL!</h2>\n<p>As a late joiner, It was so helpful to read insightful discussions of <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> and list goes on!</p>\n<h2>Approach</h2>\n<ul>\n<li>Code: <a href=\"https://github.com/bayartsogt-ya/bengali-speech-2023\" target=\"_blank\">https://github.com/bayartsogt-ya/bengali-speech-2023</a></li>\n<li>Inference: <a href=\"https://www.kaggle.com/code/bayartsogtya/submit-to-restore-punctuation/notebook\" target=\"_blank\">https://www.kaggle.com/code/bayartsogtya/submit-to-restore-punctuation/notebook</a></li>\n<li>Backbone Model: <a href=\"https://huggingface.co/facebook/wav2vec2-xls-r-300m\" target=\"_blank\"><code>facebook/wav2vec2-xls-r-300m</code></a></li>\n<li>LM: KenLM 5-gram (16G) trained on <a href=\"https://github.com/AI4Bharat/IndicBERT#indiccorp-v2\" target=\"_blank\">IndicCorpV2 corpus</a> and <a href=\"https://github.com/rezacsedu/Bengali-Hate-Speech-Dataset/tree/main\" target=\"_blank\">Bengali Hate Speech Dataset</a></li>\n<li>More Data: Competition data + MadASR2023 + OpenSLR53</li>\n<li>Data Augmentation: <code>audiomentations.AddBackgroundNoise</code> using subset of \"Bollywood Music\", \"Applause\" and \"Theme Music\" from <a href=\"https://research.google.com/audioset/dataset/index.html\" target=\"_blank\">AudioSet dataset</a></li>\n<li>Restore Punctuation <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a></li>\n</ul>\n<h2>Important lesson for future me!</h2>\n<ul>\n<li><strong><code>[No space left on device]</code></strong> Just write your own custom dataset class!!!<ul>\n<li>Look at <a href=\"https://github.com/bayartsogt-ya/bengali-speech-2023/blob/main/train2.py\" target=\"_blank\">https://github.com/bayartsogt-ya/bengali-speech-2023/blob/main/train2.py</a>.</li>\n<li>Just increase <code>dataloader_num_workers</code> if you have enough cores. Preparing input and use <code>datasets.Dataset.set_transform</code> is complicated and <strong>not</strong> efficient.</li>\n<li>Be simple! read it from a file system in <code>__getitem__</code> and apply whatever you want on the fly!</li></ul></li>\n<li><strong><code>[Quality vs Quantity]</code></strong> 0.475 on only validation VS 0.421 on train (filtered) validation madasr openslr53 😂<ul>\n<li>It is obvious that filtering on big datasets helps!</li></ul></li>\n<li><strong><code>[Manually Check Output]</code></strong> See where your model is making mistake on your validation data.<ul>\n<li>This helped me to see that punctuations (dari, comma, question mark, etc…) were counted as substitutions and deleted.</li></ul></li>\n<li><strong><code>[Stop procrastinating on small things]</code></strong> You could have checked different chunk_length_s way before deadline. But you did not here! -&gt; This is not calling out you did not try to train whisper!</li>\n</ul>\n<h2>Guilt of Overfitting to LB!</h2>\n<p>Because test data (Out of Distribution) data was so different from train datasets, it was really about overfitting to public leaderboard.</p>\n<pre><code>!kaggle competitions submissions -v bengaliai-speech &gt;&gt; ./bengali-speech-submissions.csv\ndf = pd.read_csv()\nnp.corrcoef(df.publicScore, df.privateScore)[, ]\n\n...\nsns.lineplot(df, y=, x=, hue=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2Ff73eba4301b1d7ec468eb29fba3ddb02%2Fpublic_vs_private.png?generation=1698196214469291&amp;alt=media\" alt=\"\"></p>\n<h2>In the End</h2>\n<p>It is all about learning!<br>\nEven though it is always so frustrating to feel you were so close or so much could have done or should have done, I appreciate this learning path and that's why I joined to Kaggle in the first place! 🫡</p>",
  "messages": [
    {
      "id": "2497907",
      "postDate": "10/25/2023 01:56:29",
      "content": "<h2>First of all, THANK YOU ALL!</h2>\n<p>As a late joiner, It was so helpful to read insightful discussions of <a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> and list goes on!</p>\n<h2>Approach</h2>\n<ul>\n<li>Code: <a href=\"https://github.com/bayartsogt-ya/bengali-speech-2023\" target=\"_blank\">https://github.com/bayartsogt-ya/bengali-speech-2023</a></li>\n<li>Inference: <a href=\"https://www.kaggle.com/code/bayartsogtya/submit-to-restore-punctuation/notebook\" target=\"_blank\">https://www.kaggle.com/code/bayartsogtya/submit-to-restore-punctuation/notebook</a></li>\n<li>Backbone Model: <a href=\"https://huggingface.co/facebook/wav2vec2-xls-r-300m\" target=\"_blank\"><code>facebook/wav2vec2-xls-r-300m</code></a></li>\n<li>LM: KenLM 5-gram (16G) trained on <a href=\"https://github.com/AI4Bharat/IndicBERT#indiccorp-v2\" target=\"_blank\">IndicCorpV2 corpus</a> and <a href=\"https://github.com/rezacsedu/Bengali-Hate-Speech-Dataset/tree/main\" target=\"_blank\">Bengali Hate Speech Dataset</a></li>\n<li>More Data: Competition data + MadASR2023 + OpenSLR53</li>\n<li>Data Augmentation: <code>audiomentations.AddBackgroundNoise</code> using subset of \"Bollywood Music\", \"Applause\" and \"Theme Music\" from <a href=\"https://research.google.com/audioset/dataset/index.html\" target=\"_blank\">AudioSet dataset</a></li>\n<li>Restore Punctuation <a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a></li>\n</ul>\n<h2>Important lesson for future me!</h2>\n<ul>\n<li><strong><code>[No space left on device]</code></strong> Just write your own custom dataset class!!!<ul>\n<li>Look at <a href=\"https://github.com/bayartsogt-ya/bengali-speech-2023/blob/main/train2.py\" target=\"_blank\">https://github.com/bayartsogt-ya/bengali-speech-2023/blob/main/train2.py</a>.</li>\n<li>Just increase <code>dataloader_num_workers</code> if you have enough cores. Preparing input and use <code>datasets.Dataset.set_transform</code> is complicated and <strong>not</strong> efficient.</li>\n<li>Be simple! read it from a file system in <code>__getitem__</code> and apply whatever you want on the fly!</li></ul></li>\n<li><strong><code>[Quality vs Quantity]</code></strong> 0.475 on only validation VS 0.421 on train (filtered) validation madasr openslr53 😂<ul>\n<li>It is obvious that filtering on big datasets helps!</li></ul></li>\n<li><strong><code>[Manually Check Output]</code></strong> See where your model is making mistake on your validation data.<ul>\n<li>This helped me to see that punctuations (dari, comma, question mark, etc…) were counted as substitutions and deleted.</li></ul></li>\n<li><strong><code>[Stop procrastinating on small things]</code></strong> You could have checked different chunk_length_s way before deadline. But you did not here! -&gt; This is not calling out you did not try to train whisper!</li>\n</ul>\n<h2>Guilt of Overfitting to LB!</h2>\n<p>Because test data (Out of Distribution) data was so different from train datasets, it was really about overfitting to public leaderboard.</p>\n<pre><code>!kaggle competitions submissions -v bengaliai-speech &gt;&gt; ./bengali-speech-submissions.csv\ndf = pd.read_csv()\nnp.corrcoef(df.publicScore, df.privateScore)[, ]\n\n...\nsns.lineplot(df, y=, x=, hue=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2Ff73eba4301b1d7ec468eb29fba3ddb02%2Fpublic_vs_private.png?generation=1698196214469291&amp;alt=media\" alt=\"\"></p>\n<h2>In the End</h2>\n<p>It is all about learning!<br>\nEven though it is always so frustrating to feel you were so close or so much could have done or should have done, I appreciate this learning path and that's why I joined to Kaggle in the first place! 🫡</p>",
      "rawMarkdown": "## First of all, THANK YOU ALL!\nAs a late joiner, It was so helpful to read insightful discussions of @imtiazprio @reasat @tugstugi @hengck23 @mbmmurad and list goes on!\n\n## Approach\n* Code: https://github.com/bayartsogt-ya/bengali-speech-2023\n* Inference: https://www.kaggle.com/code/bayartsogtya/submit-to-restore-punctuation/notebook\n* Backbone Model: [`facebook/wav2vec2-xls-r-300m`](https://huggingface.co/facebook/wav2vec2-xls-r-300m)\n* LM: KenLM 5-gram (16G) trained on [IndicCorpV2 corpus](https://github.com/AI4Bharat/IndicBERT#indiccorp-v2) and [Bengali Hate Speech Dataset](https://github.com/rezacsedu/Bengali-Hate-Speech-Dataset/tree/main)\n* More Data: Competition data + MadASR2023 + OpenSLR53\n* Data Augmentation: `audiomentations.AddBackgroundNoise` using subset of \"Bollywood Music\", \"Applause\" and \"Theme Music\" from [AudioSet dataset](https://research.google.com/audioset/dataset/index.html)\n* Restore Punctuation https://github.com/xashru/punctuation-restoration\n\n## Important lesson for future me!\n* **`[No space left on device]`** Just write your own custom dataset class!!!\n    * Look at https://github.com/bayartsogt-ya/bengali-speech-2023/blob/main/train2.py.\n    * Just increase `dataloader_num_workers` if you have enough cores. Preparing input and use `datasets.Dataset.set_transform` is complicated and **not** efficient.\n    * Be simple! read it from a file system in `__getitem__` and apply whatever you want on the fly!\n* **`[Quality vs Quantity]`** 0.475 on only validation VS 0.421 on train (filtered) validation madasr openslr53 😂\n    * It is obvious that filtering on big datasets helps!\n* **`[Manually Check Output]`** See where your model is making mistake on your validation data.\n    * This helped me to see that punctuations (dari, comma, question mark, etc...) were counted as substitutions and deleted.\n* **`[Stop procrastinating on small things]`** You could have checked different chunk_length_s way before deadline. But you did not here! -> This is not calling out you did not try to train whisper!\n\n## Guilt of Overfitting to LB!\n\nBecause test data (Out of Distribution) data was so different from train datasets, it was really about overfitting to public leaderboard.\n```python\n!kaggle competitions submissions -v bengaliai-speech >> ./bengali-speech-submissions.csv\n>>> df = pd.read_csv(\"./data/bengali-speech-submissions.csv\")\n>>> np.corrcoef(df.publicScore, df.privateScore)[0, 1]\n0.990815988867224\n...\n>>> sns.lineplot(df, y=\"score\", x=\"date\", hue=\"split\")\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2Ff73eba4301b1d7ec468eb29fba3ddb02%2Fpublic_vs_private.png?generation=1698196214469291&alt=media)\n\n## In the End\n\nIt is all about learning!\nEven though it is always so frustrating to feel you were so close or so much could have done or should have done, I appreciate this learning path and that's why I joined to Kaggle in the first place! 🫡",
      "votes": null
    },
    {
      "id": "2498065",
      "postDate": "10/25/2023 05:44:03",
      "content": "<p>Considering the short time investment,  I think the result is fantastic!! <br>\nBest regards and all the best for future competitions. The write up is clean and well structured <a href=\"https://www.kaggle.com/bayartsogtya\" target=\"_blank\">@bayartsogtya</a> <br>\nCongratulations!! </p>",
      "rawMarkdown": "Considering the short time investment,  I think the result is fantastic!! \nBest regards and all the best for future competitions. The write up is clean and well structured @bayartsogtya \nCongratulations!!",
      "votes": null
    },
    {
      "id": "2498766",
      "postDate": "10/25/2023 14:13:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Thanks @ravi20076",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2498065,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/25/2023 05:44:03",
      "content": "<p>Considering the short time investment,  I think the result is fantastic!! <br>\nBest regards and all the best for future competitions. The write up is clean and well structured <a href=\"https://www.kaggle.com/bayartsogtya\" target=\"_blank\">@bayartsogtya</a> <br>\nCongratulations!! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2498766,
          "author_name": "bayartsogtya",
          "author_url": "",
          "post_date": "10/25/2023 14:13:22",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2497907": "## First of all, THANK YOU ALL!\nAs a late joiner, It was so helpful to read insightful discussions of @imtiazprio @reasat @tugstugi @hengck23 @mbmmurad and list goes on!\n\n## Approach\n* Code: https://github.com/bayartsogt-ya/bengali-speech-2023\n* Inference: https://www.kaggle.com/code/bayartsogtya/submit-to-restore-punctuation/notebook\n* Backbone Model: [`facebook/wav2vec2-xls-r-300m`](https://huggingface.co/facebook/wav2vec2-xls-r-300m)\n* LM: KenLM 5-gram (16G) trained on [IndicCorpV2 corpus](https://github.com/AI4Bharat/IndicBERT#indiccorp-v2) and [Bengali Hate Speech Dataset](https://github.com/rezacsedu/Bengali-Hate-Speech-Dataset/tree/main)\n* More Data: Competition data + MadASR2023 + OpenSLR53\n* Data Augmentation: `audiomentations.AddBackgroundNoise` using subset of \"Bollywood Music\", \"Applause\" and \"Theme Music\" from [AudioSet dataset](https://research.google.com/audioset/dataset/index.html)\n* Restore Punctuation https://github.com/xashru/punctuation-restoration\n\n## Important lesson for future me!\n* **`[No space left on device]`** Just write your own custom dataset class!!!\n    * Look at https://github.com/bayartsogt-ya/bengali-speech-2023/blob/main/train2.py.\n    * Just increase `dataloader_num_workers` if you have enough cores. Preparing input and use `datasets.Dataset.set_transform` is complicated and **not** efficient.\n    * Be simple! read it from a file system in `__getitem__` and apply whatever you want on the fly!\n* **`[Quality vs Quantity]`** 0.475 on only validation VS 0.421 on train (filtered) validation madasr openslr53 😂\n    * It is obvious that filtering on big datasets helps!\n* **`[Manually Check Output]`** See where your model is making mistake on your validation data.\n    * This helped me to see that punctuations (dari, comma, question mark, etc...) were counted as substitutions and deleted.\n* **`[Stop procrastinating on small things]`** You could have checked different chunk_length_s way before deadline. But you did not here! -> This is not calling out you did not try to train whisper!\n\n## Guilt of Overfitting to LB!\n\nBecause test data (Out of Distribution) data was so different from train datasets, it was really about overfitting to public leaderboard.\n```python\n!kaggle competitions submissions -v bengaliai-speech >> ./bengali-speech-submissions.csv\n>>> df = pd.read_csv(\"./data/bengali-speech-submissions.csv\")\n>>> np.corrcoef(df.publicScore, df.privateScore)[0, 1]\n0.990815988867224\n...\n>>> sns.lineplot(df, y=\"score\", x=\"date\", hue=\"split\")\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2Ff73eba4301b1d7ec468eb29fba3ddb02%2Fpublic_vs_private.png?generation=1698196214469291&alt=media)\n\n## In the End\n\nIt is all about learning!\nEven though it is always so frustrating to feel you were so close or so much could have done or should have done, I appreciate this learning path and that's why I joined to Kaggle in the first place! 🫡",
    "2498065": "Considering the short time investment,  I think the result is fantastic!! \nBest regards and all the best for future competitions. The write up is clean and well structured @bayartsogtya \nCongratulations!!",
    "2498766": "Thanks @ravi20076"
  },
  "source": "meta"
}