{
  "id": 426810,
  "title": "LB improved by 0.002 to 0.506 : Should we zero pad audios or not?",
  "url": "/competitions/bengaliai-speech/discussion/426810",
  "author_name": "",
  "post_date": "2023-07-25T07:48:12.546067100Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947\" target=\"_blank\">Notebook with LB score 0.506</a></p>\n</blockquote>\n<p>In <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> bhai's best scoring <a href=\"https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference\" target=\"_blank\">Notebook</a>, the audios in a batch were padded up to the maximum audio length of the batch during inference. </p>\n<pre><code> ():\n    \n    \n    lengths = torch.tensor([ t.shape[]  t  batch ])\n    \n    batch = [ torch.Tensor(t)  t  batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    \n    mask = (batch != )\n     batch, lengths, mask\n</code></pre>\n<p>I wasn't sure whether this was the right idea since the audio length might vary a lot in batches (consider the case when a 1s audio is in the batch with a 10s audio).  So I ran two experiments. </p>\n<ol>\n<li>In <a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947\" target=\"_blank\">my notebook</a> I passed each of the audios separately into the pipeline, thus not requiring to pad the audio, and the LB score improved to <strong>0.506!</strong></li>\n<li>I sorted the audio paths by the duration of the audio files. </li>\n</ol>\n<pre><code>paths = (os(TEST_DIRECTORY,))\n\ntmp_df = pd({:paths.:paths})\ntmp_df = tmp_df(lambda x: (AudioSegment(x)))\npaths = tmp_df()()\n</code></pre>\n<p>Thus audios having same duration were batched together, so less padding was used comparatively than before. This increased the performance by the same amount, achieving <strong>LB 0.506</strong></p>\n<p>I wasn't able to check the CV score since the inference time is higher in both notebooks. </p>\n<p>Now my question is, does zero padding actually decrease the inference performance of Wav2Vec2? Any ideas/comments?</p>",
  "messages": [
    {
      "id": "2357910",
      "postDate": "07/25/2023 07:48:12",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947\" target=\"_blank\">Notebook with LB score 0.506</a></p>\n</blockquote>\n<p>In <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> bhai's best scoring <a href=\"https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference\" target=\"_blank\">Notebook</a>, the audios in a batch were padded up to the maximum audio length of the batch during inference. </p>\n<pre><code> ():\n    \n    \n    lengths = torch.tensor([ t.shape[]  t  batch ])\n    \n    batch = [ torch.Tensor(t)  t  batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    \n    mask = (batch != )\n     batch, lengths, mask\n</code></pre>\n<p>I wasn't sure whether this was the right idea since the audio length might vary a lot in batches (consider the case when a 1s audio is in the batch with a 10s audio).  So I ran two experiments. </p>\n<ol>\n<li>In <a href=\"https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947\" target=\"_blank\">my notebook</a> I passed each of the audios separately into the pipeline, thus not requiring to pad the audio, and the LB score improved to <strong>0.506!</strong></li>\n<li>I sorted the audio paths by the duration of the audio files. </li>\n</ol>\n<pre><code>paths = (os(TEST_DIRECTORY,))\n\ntmp_df = pd({:paths.:paths})\ntmp_df = tmp_df(lambda x: (AudioSegment(x)))\npaths = tmp_df()()\n</code></pre>\n<p>Thus audios having same duration were batched together, so less padding was used comparatively than before. This increased the performance by the same amount, achieving <strong>LB 0.506</strong></p>\n<p>I wasn't able to check the CV score since the inference time is higher in both notebooks. </p>\n<p>Now my question is, does zero padding actually decrease the inference performance of Wav2Vec2? Any ideas/comments?</p>",
      "rawMarkdown": ">[Notebook with LB score 0.506](https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947)\n\n\nIn @reasat bhai's best scoring [Notebook](https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference), the audios in a batch were padded up to the maximum audio length of the batch during inference. \n\n```\ndef collate_fn_padd(batch):\n    '''\n    Padds batch of variable length\n\n    note: it converts things ToTensor manually here since the ToTensor transform\n    assume it takes in images rather than arbitrary tensors.\n    '''\n    ## get sequence lengths\n    lengths = torch.tensor([ t.shape[0] for t in batch ])\n    ## padd\n    batch = [ torch.Tensor(t) for t in batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    ## compute mask\n    mask = (batch != 0)\n    return batch, lengths, mask\n```\n\nI wasn't sure whether this was the right idea since the audio length might vary a lot in batches (consider the case when a 1s audio is in the batch with a 10s audio).  So I ran two experiments. \n\n1. In [my notebook](https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947) I passed each of the audios separately into the pipeline, thus not requiring to pad the audio, and the LB score improved to **0.506!**\n2. I sorted the audio paths by the duration of the audio files. \n```\npaths = glob(os.path.join(TEST_DIRECTORY,'*.mp3'))\n\ntmp_df = pd.DataFrame({\"id\":paths.\"duration\":paths})\ntmp_df[\"duration\"] = tmp_df[\"duration\"].apply(lambda x: len(AudioSegment.from_file(x)))\npaths = tmp_df.sort_values('duration')['id'].tolist()\n```\n\nThus audios having same duration were batched together, so less padding was used comparatively than before. This increased the performance by the same amount, achieving **LB 0.506**\n\nI wasn't able to check the CV score since the inference time is higher in both notebooks. \n\nNow my question is, does zero padding actually decrease the inference performance of Wav2Vec2? Any ideas/comments?",
      "votes": null
    },
    {
      "id": "2358129",
      "postDate": "07/25/2023 10:00:20",
      "content": "<p>batching results is almost always worse than single inference.<br>\nthis is also proven in local validation.</p>\n<p>the reason is prediction of noise token at the end of wave (i.e. padding region)</p>\n<p>for example:</p>\n<pre><code>with out pad:\npredict : \n\n\nwith pad\npredict :  \n__\n____x___\n\nx   \n</code></pre>\n<p>i observe this when i debug my code. to verify if my padding code is correct or not, i compare results with and without batching.</p>",
      "rawMarkdown": "batching results is almost always worse than single inference.\nthis is also proven in local validation.\n\nthe reason is prediction of noise token at the end of wave (i.e. padding region)\n\nfor example:\n```\nwith out pad:\npredict : \n<s> abd </s>\n\nwith pad\npredict :  \n<s> abd  ____x</s>___\n<s> abd  </s>____x____\n\nx is some \"noise\"\n```\n\ni observe this when i debug my code. to verify if my padding code is correct or not, i compare results with and without batching.",
      "votes": null
    },
    {
      "id": "2358234",
      "postDate": "07/25/2023 11:27:31",
      "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> <br>\nremove this to get 0.505:</p>\n<p>submit_df.loc[:,'sentence'] = submit_df.sentence.apply(lambda x:dari(x))</p>",
      "rawMarkdown": "mbmmurad \nremove this to get 0.505:\n\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(lambda x:dari(x))",
      "votes": null
    },
    {
      "id": "2358817",
      "postDate": "07/25/2023 18:41:37",
      "content": "<p>Thanks for the insight. I think It'll be a dilemma between sacrifing a bit LB score vs Fast inference. </p>",
      "rawMarkdown": "Thanks for the insight. I think It'll be a dilemma between sacrifing a bit LB score vs Fast inference.",
      "votes": null
    },
    {
      "id": "2358868",
      "postDate": "07/25/2023 19:56:04",
      "content": "<p>i manage to find my old results and i will show it here</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f4aa57d966bcfdcfd5ae476e0de552d%2FSelection_999(2807).png?generation=1690314890709034&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F76106326d9984e670608619c544173b9%2FSelection_999(2806).png?generation=1690314900191892&amp;alt=media\" alt=\"\"></p>\n<p>two noise predictions from batch inference. these are absent for single input inference.</p>",
      "rawMarkdown": "i manage to find my old results and i will show it here\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f4aa57d966bcfdcfd5ae476e0de552d%2FSelection_999(2807).png?generation=1690314890709034&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F76106326d9984e670608619c544173b9%2FSelection_999(2806).png?generation=1690314900191892&alt=media)\n\ntwo noise predictions from batch inference. these are absent for single input inference.",
      "votes": null
    },
    {
      "id": "2359889",
      "postDate": "07/26/2023 12:18:50",
      "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> </p>\n<p>I checked in local which results are:</p>\n<ul>\n<li>Single infer is worse than batch infer</li>\n</ul>\n<p>I reviewed your notebooks and find out that you also process .wav files in test folder. Other public notebooks only process .mp3 files in test folder. So, I think the reason that your notebook has score of 0.506 is that you process .wav files in test folder.</p>",
      "rawMarkdown": "mbmmurad \n\nI checked in local which results are:\n+ Single infer is worse than batch infer\n\nI reviewed your notebooks and find out that you also process .wav files in test folder. Other public notebooks only process .mp3 files in test folder. So, I think the reason that your notebook has score of 0.506 is that you process .wav files in test folder.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2358129,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2023 10:00:20",
      "content": "<p>batching results is almost always worse than single inference.<br>\nthis is also proven in local validation.</p>\n<p>the reason is prediction of noise token at the end of wave (i.e. padding region)</p>\n<p>for example:</p>\n<pre><code>with out pad:\npredict : \n\n\nwith pad\npredict :  \n__\n____x___\n\nx   \n</code></pre>\n<p>i observe this when i debug my code. to verify if my padding code is correct or not, i compare results with and without batching.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2358817,
          "author_name": "mbmmurad",
          "author_url": "",
          "post_date": "07/25/2023 18:41:37",
          "content": "<p>Thanks for the insight. I think It'll be a dilemma between sacrifing a bit LB score vs Fast inference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2358868,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/25/2023 19:56:04",
          "content": "<p>i manage to find my old results and i will show it here</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f4aa57d966bcfdcfd5ae476e0de552d%2FSelection_999(2807).png?generation=1690314890709034&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F76106326d9984e670608619c544173b9%2FSelection_999(2806).png?generation=1690314900191892&amp;alt=media\" alt=\"\"></p>\n<p>two noise predictions from batch inference. these are absent for single input inference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2358234,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/25/2023 11:27:31",
      "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> <br>\nremove this to get 0.505:</p>\n<p>submit_df.loc[:,'sentence'] = submit_df.sentence.apply(lambda x:dari(x))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2359889,
      "author_name": "texopher",
      "author_url": "",
      "post_date": "07/26/2023 12:18:50",
      "content": "<p><a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> </p>\n<p>I checked in local which results are:</p>\n<ul>\n<li>Single infer is worse than batch infer</li>\n</ul>\n<p>I reviewed your notebooks and find out that you also process .wav files in test folder. Other public notebooks only process .mp3 files in test folder. So, I think the reason that your notebook has score of 0.506 is that you process .wav files in test folder.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2357910": ">[Notebook with LB score 0.506](https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947)\n\n\nIn @reasat bhai's best scoring [Notebook](https://www.kaggle.com/code/reasat/yellowking-dlsprint-inference), the audios in a batch were padded up to the maximum audio length of the batch during inference. \n\n```\ndef collate_fn_padd(batch):\n    '''\n    Padds batch of variable length\n\n    note: it converts things ToTensor manually here since the ToTensor transform\n    assume it takes in images rather than arbitrary tensors.\n    '''\n    ## get sequence lengths\n    lengths = torch.tensor([ t.shape[0] for t in batch ])\n    ## padd\n    batch = [ torch.Tensor(t) for t in batch ]\n    batch = torch.nn.utils.rnn.pad_sequence(batch)\n    ## compute mask\n    mask = (batch != 0)\n    return batch, lengths, mask\n```\n\nI wasn't sure whether this was the right idea since the audio length might vary a lot in batches (consider the case when a 1s audio is in the batch with a 10s audio).  So I ran two experiments. \n\n1. In [my notebook](https://www.kaggle.com/code/mbmmurad/lb-0-506-inference-w-previous-comp-winner-s-model?scriptVersionId=137803947) I passed each of the audios separately into the pipeline, thus not requiring to pad the audio, and the LB score improved to **0.506!**\n2. I sorted the audio paths by the duration of the audio files. \n```\npaths = glob(os.path.join(TEST_DIRECTORY,'*.mp3'))\n\ntmp_df = pd.DataFrame({\"id\":paths.\"duration\":paths})\ntmp_df[\"duration\"] = tmp_df[\"duration\"].apply(lambda x: len(AudioSegment.from_file(x)))\npaths = tmp_df.sort_values('duration')['id'].tolist()\n```\n\nThus audios having same duration were batched together, so less padding was used comparatively than before. This increased the performance by the same amount, achieving **LB 0.506**\n\nI wasn't able to check the CV score since the inference time is higher in both notebooks. \n\nNow my question is, does zero padding actually decrease the inference performance of Wav2Vec2? Any ideas/comments?",
    "2358129": "batching results is almost always worse than single inference.\nthis is also proven in local validation.\n\nthe reason is prediction of noise token at the end of wave (i.e. padding region)\n\nfor example:\n```\nwith out pad:\npredict : \n<s> abd </s>\n\nwith pad\npredict :  \n<s> abd  ____x</s>___\n<s> abd  </s>____x____\n\nx is some \"noise\"\n```\n\ni observe this when i debug my code. to verify if my padding code is correct or not, i compare results with and without batching.",
    "2358234": "mbmmurad \nremove this to get 0.505:\n\nsubmit_df.loc[:,'sentence'] = submit_df.sentence.apply(lambda x:dari(x))",
    "2358817": "Thanks for the insight. I think It'll be a dilemma between sacrifing a bit LB score vs Fast inference.",
    "2358868": "i manage to find my old results and i will show it here\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5f4aa57d966bcfdcfd5ae476e0de552d%2FSelection_999(2807).png?generation=1690314890709034&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F76106326d9984e670608619c544173b9%2FSelection_999(2806).png?generation=1690314900191892&alt=media)\n\ntwo noise predictions from batch inference. these are absent for single input inference.",
    "2359889": "mbmmurad \n\nI checked in local which results are:\n+ Single infer is worse than batch infer\n\nI reviewed your notebooks and find out that you also process .wav files in test folder. Other public notebooks only process .mp3 files in test folder. So, I think the reason that your notebook has score of 0.506 is that you process .wav files in test folder."
  },
  "source": "meta"
}