{
  "id": 448126,
  "title": "11th place solution",
  "url": "/competitions/bengaliai-speech/discussion/448126",
  "author_name": "DarioAr",
  "post_date": "2023-10-18T15:05:50.227000",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<h1>11th Solution</h1>\n<p>The solution consisted of a single whisper medium model using a beam decoder with a size of 4.</p>\n<h6>Training</h6>\n<ul>\n<li>Like other solutions, the most important step is to use cleaned data. I use following rules based on metadata shared by host:</li>\n</ul>\n<pre><code>cond0 = train_df[\n    (\n        (train_df.ykg_wer &lt; ) | (train_df.ggl_wer &lt; )\n    ) &amp; (\n        (train_df.total_wer_by_client_ykg &lt; ) |\n        (train_df.total_wer_by_client_ggl &lt; )\n    ) &amp; (train_df.mos_pred &gt; )\n]\n</code></pre>\n<ul>\n<li><p>To speed up training, about 80% of the dataset is presented to the model in combinations of two audios, showing only single audios to prevent hallucinations. This also helps to reduce the impact of possible bad annotations.</p></li>\n<li><p>SpecAugment, SpecAugment++, CutOut were used.</p></li>\n</ul>\n<h6>Inference</h6>\n<ul>\n<li><p>For inference, the most important thing is to handle correctly audios longer than 30 secs and sentences longer than 448 tokens. This made an improvement on LB from 0.43 to 0.38.</p></li>\n<li><p>Inference code: <a href=\"https://www.kaggle.com/code/themadrambito/11th-place-whisper-inference\" target=\"_blank\">https://www.kaggle.com/code/themadrambito/11th-place-whisper-inference</a></p></li>\n</ul>",
  "messages": [
    {
      "id": 2487404,
      "postDate": "2023-10-18T15:05:50.227Z",
      "content": "<h1>11th Solution</h1>\n<p>The solution consisted of a single whisper medium model using a beam decoder with a size of 4.</p>\n<h6>Training</h6>\n<ul>\n<li>Like other solutions, the most important step is to use cleaned data. I use following rules based on metadata shared by host:</li>\n</ul>\n<pre><code>cond0 = train_df[\n    (\n        (train_df.ykg_wer &lt; ) | (train_df.ggl_wer &lt; )\n    ) &amp; (\n        (train_df.total_wer_by_client_ykg &lt; ) |\n        (train_df.total_wer_by_client_ggl &lt; )\n    ) &amp; (train_df.mos_pred &gt; )\n]\n</code></pre>\n<ul>\n<li><p>To speed up training, about 80% of the dataset is presented to the model in combinations of two audios, showing only single audios to prevent hallucinations. This also helps to reduce the impact of possible bad annotations.</p></li>\n<li><p>SpecAugment, SpecAugment++, CutOut were used.</p></li>\n</ul>\n<h6>Inference</h6>\n<ul>\n<li><p>For inference, the most important thing is to handle correctly audios longer than 30 secs and sentences longer than 448 tokens. This made an improvement on LB from 0.43 to 0.38.</p></li>\n<li><p>Inference code: <a href=\"https://www.kaggle.com/code/themadrambito/11th-place-whisper-inference\" target=\"_blank\">https://www.kaggle.com/code/themadrambito/11th-place-whisper-inference</a></p></li>\n</ul>",
      "rawMarkdown": "# 11th Solution\n\nThe solution consisted of a single whisper medium model using a beam decoder with a size of 4.\n\n###### Training\n- Like other solutions, the most important step is to use cleaned data. I use following rules based on metadata shared by host:\n\n```python\ncond0 = train_df[\n    (\n        (train_df.ykg_wer < 0.6) | (train_df.ggl_wer < 0.6)\n    ) & (\n        (train_df.total_wer_by_client_ykg < 0.7) |\n        (train_df.total_wer_by_client_ggl < 0.7)\n    ) & (train_df.mos_pred > 1.5)\n]\n```\n\n- To speed up training, about 80% of the dataset is presented to the model in combinations of two audios, showing only single audios to prevent hallucinations. This also helps to reduce the impact of possible bad annotations.\n\n- SpecAugment, SpecAugment++, CutOut were used.\n\n###### Inference\n\n- For inference, the most important thing is to handle correctly audios longer than 30 secs and sentences longer than 448 tokens. This made an improvement on LB from 0.43 to 0.38.\n\n- Inference code: https://www.kaggle.com/code/themadrambito/11th-place-whisper-inference\n\n",
      "votes": 10
    },
    {
      "id": 2491336,
      "postDate": "2023-10-21T14:29:37.053Z",
      "content": "<p>Can you descibe how you tackle with audios longer than 30 secs?</p>",
      "rawMarkdown": "Can you descibe how you tackle with audios longer than 30 secs?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2491457,
          "postDate": "2023-10-21T16:24:59.160Z",
          "content": "<p>Yes, let's say we have an audio with a length of 35 seconds:</p>\n<p>On the first forward pass of the model, we will have:</p>\n<ul>\n<li>The first 30 seconds of audio</li>\n<li>START 0.0 will be our tokens</li>\n</ul>\n<p>Then we have three possible cases based on the response of the model:</p>\n<ul>\n<li><p>First case: START 0.0 sentence…<br>\nIn this case, the model has used all of the 448 tokens it can output and has not found an end to the sentence. For the second iteration, I will advance the audio 15 seconds, and the next initial sequence will be: START 0.0 half sentence.</p></li>\n<li><p>Second case: START 0.0 sentence 18.0<br>\nIn this case, the model has found an end to the sentence, and I will just advance the audio based on that end timestamp. The next tokens will be PREV half sentence START 0.0.</p></li>\n<li><p>Third case: START 0.0 sentence 5.0 sentence…<br>\nIn this case, we have found a timestamp, but the sequence is not finished, so I just advance the audio to the timestamp in the middle, 5 seconds, and continue from there. The next tokens will be PREV some part of the previous sentence START 0.0 half sentence (max context is 224 tokens).</p></li>\n</ul>\n<p>I run this in a loop until we have visited all 35 seconds of audio. I recommend running the inference code.</p>",
          "rawMarkdown": "Yes, let's say we have an audio with a length of 35 seconds:\n\nOn the first forward pass of the model, we will have:\n- The first 30 seconds of audio\n- START 0.0 will be our tokens\n\nThen we have three possible cases based on the response of the model:\n\n- First case: START 0.0 sentence...\nIn this case, the model has used all of the 448 tokens it can output and has not found an end to the sentence. For the second iteration, I will advance the audio 15 seconds, and the next initial sequence will be: START 0.0 half sentence.\n\n- Second case: START 0.0 sentence 18.0\nIn this case, the model has found an end to the sentence, and I will just advance the audio based on that end timestamp. The next tokens will be PREV half sentence START 0.0.\n\n- Third case: START 0.0 sentence 5.0 sentence...\nIn this case, we have found a timestamp, but the sequence is not finished, so I just advance the audio to the timestamp in the middle, 5 seconds, and continue from there. The next tokens will be PREV some part of the previous sentence START 0.0 half sentence (max context is 224 tokens).\n\nI run this in a loop until we have visited all 35 seconds of audio. I recommend running the inference code.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2487948,
      "postDate": "2023-10-18T22:34:26.903Z",
      "content": "<p>Congrats on your solo gold and thanks for sharing your inference code! Whisper cracked our minds for a while :)</p>",
      "rawMarkdown": "Congrats on your solo gold and thanks for sharing your inference code! Whisper cracked our minds for a while :)",
      "votes": 1
    },
    {
      "id": 2487433,
      "postDate": "2023-10-18T15:34:02.537Z",
      "content": "<p>Yours is only the second whisper model in this competition (the first one at #1). Congratulations !</p>",
      "rawMarkdown": "Yours is only the second whisper model in this competition (the first one at #1). Congratulations !",
      "votes": 1
    },
    {
      "id": 2487852,
      "postDate": "2023-10-18T19:51:04.393Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2491336,
      "author_name": "HuBERT",
      "author_url": "",
      "post_date": "2023-10-21T14:29:37.053000",
      "content": "<p>Can you descibe how you tackle with audios longer than 30 secs?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2491457,
          "author_name": "DarioAr",
          "author_url": "",
          "post_date": "2023-10-21T16:24:59.160000",
          "content": "<p>Yes, let's say we have an audio with a length of 35 seconds:</p>\n<p>On the first forward pass of the model, we will have:</p>\n<ul>\n<li>The first 30 seconds of audio</li>\n<li>START 0.0 will be our tokens</li>\n</ul>\n<p>Then we have three possible cases based on the response of the model:</p>\n<ul>\n<li><p>First case: START 0.0 sentence…<br>\nIn this case, the model has used all of the 448 tokens it can output and has not found an end to the sentence. For the second iteration, I will advance the audio 15 seconds, and the next initial sequence will be: START 0.0 half sentence.</p></li>\n<li><p>Second case: START 0.0 sentence 18.0<br>\nIn this case, the model has found an end to the sentence, and I will just advance the audio based on that end timestamp. The next tokens will be PREV half sentence START 0.0.</p></li>\n<li><p>Third case: START 0.0 sentence 5.0 sentence…<br>\nIn this case, we have found a timestamp, but the sequence is not finished, so I just advance the audio to the timestamp in the middle, 5 seconds, and continue from there. The next tokens will be PREV some part of the previous sentence START 0.0 half sentence (max context is 224 tokens).</p></li>\n</ul>\n<p>I run this in a loop until we have visited all 35 seconds of audio. I recommend running the inference code.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2487948,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2023-10-18T22:34:26.903000",
      "content": "<p>Congrats on your solo gold and thanks for sharing your inference code! Whisper cracked our minds for a while :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2487433,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-10-18T15:34:02.537000",
      "content": "<p>Yours is only the second whisper model in this competition (the first one at #1). Congratulations !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2487852,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T19:51:04.393000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2487404": "# 11th Solution\n\nThe solution consisted of a single whisper medium model using a beam decoder with a size of 4.\n\n###### Training\n- Like other solutions, the most important step is to use cleaned data. I use following rules based on metadata shared by host:\n\n```python\ncond0 = train_df[\n    (\n        (train_df.ykg_wer < 0.6) | (train_df.ggl_wer < 0.6)\n    ) & (\n        (train_df.total_wer_by_client_ykg < 0.7) |\n        (train_df.total_wer_by_client_ggl < 0.7)\n    ) & (train_df.mos_pred > 1.5)\n]\n```\n\n- To speed up training, about 80% of the dataset is presented to the model in combinations of two audios, showing only single audios to prevent hallucinations. This also helps to reduce the impact of possible bad annotations.\n\n- SpecAugment, SpecAugment++, CutOut were used.\n\n###### Inference\n\n- For inference, the most important thing is to handle correctly audios longer than 30 secs and sentences longer than 448 tokens. This made an improvement on LB from 0.43 to 0.38.\n\n- Inference code: https://www.kaggle.com/code/themadrambito/11th-place-whisper-inference\n\n",
    "2491336": "Can you descibe how you tackle with audios longer than 30 secs?\n",
    "2487948": "Congrats on your solo gold and thanks for sharing your inference code! Whisper cracked our minds for a while :)",
    "2487433": "Yours is only the second whisper model in this competition (the first one at #1). Congratulations !",
    "2487852": ""
  }
}