{
  "id": 437259,
  "title": "[LB 0.383] WER on example audios",
  "url": "/competitions/bengaliai-speech/discussion/437259",
  "author_name": "",
  "post_date": "2023-09-06T04:22:07.199299Z",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I wanted to see how the model behaves on the example audios. Annotations are from here: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932</a></p>\n<p>WER (without punctuation) on those audios seems to correllate with LB.</p>\n<pre><code> WER=. CER=.\n TV Drama WER=. CER=.\n Advertisement WER=. CER=.\n WER=. CER=.\n WER=. CER=.\n TV Drama WER=. CER=.\n WER=. CER=.\n Presentation WER=. CER=.\n Class WER=. CER=.\n Session WER=. CER=.\n Recital WER=. CER=.\n Literature WER=. CER=.\n Profanity WER=. CER=.\n Drama Jatra WER=. CER=.\n Show Interview WER=. CER=.\n WER=. CER=.\n Islamic Sermon WER=. CER=.\n\n=. CER=. (ignoring punctuation)\n</code></pre>\n<p>Some remarks:</p>\n<ul>\n<li>Telemedicine is the most difficult one, because of 8khz?</li>\n<li>why Bangladeshi TV Drama is more difficult than Indian TV Drama?</li>\n<li>why Stage Drama Jatra is easier than Poem Recital and Puthi Literature?</li>\n</ul>",
  "messages": [
    {
      "id": "2425619",
      "postDate": "09/06/2023 04:22:07",
      "content": "<p>I wanted to see how the model behaves on the example audios. Annotations are from here: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932</a></p>\n<p>WER (without punctuation) on those audios seems to correllate with LB.</p>\n<pre><code> WER=. CER=.\n TV Drama WER=. CER=.\n Advertisement WER=. CER=.\n WER=. CER=.\n WER=. CER=.\n TV Drama WER=. CER=.\n WER=. CER=.\n Presentation WER=. CER=.\n Class WER=. CER=.\n Session WER=. CER=.\n Recital WER=. CER=.\n Literature WER=. CER=.\n Profanity WER=. CER=.\n Drama Jatra WER=. CER=.\n Show Interview WER=. CER=.\n WER=. CER=.\n Islamic Sermon WER=. CER=.\n\n=. CER=. (ignoring punctuation)\n</code></pre>\n<p>Some remarks:</p>\n<ul>\n<li>Telemedicine is the most difficult one, because of 8khz?</li>\n<li>why Bangladeshi TV Drama is more difficult than Indian TV Drama?</li>\n<li>why Stage Drama Jatra is easier than Poem Recital and Puthi Literature?</li>\n</ul>",
      "rawMarkdown": "I wanted to see how the model behaves on the example audios. Annotations are from here: https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\n\nWER (without punctuation) on those audios seems to correllate with LB.\n\n```\nAudiobook WER=16.67 CER=4.46\nBangladeshi TV Drama WER=45.04 CER=15.56\nBengali Advertisement WER=57.14 CER=33.44\nCartoon WER=32.53 CER=12.71\nDebate WER=44.86 CER=23.65\nIndian TV Drama WER=32.08 CER=14.79\nMovie WER=26.00 CER=6.55\nNews Presentation WER=26.39 CER=12.78\nOnline Class WER=30.95 CER=17.66\nParliament Session WER=37.18 CER=17.02\nPoem Recital WER=60.00 CER=18.01\nPuthi Literature WER=60.47 CER=37.72\nSlang Profanity WER=66.05 CER=36.06\nStage Drama Jatra WER=19.48 CER=5.98\nTalk Show Interview WER=30.59 CER=13.15\nTelemedicine WER=68.83 CER=48.55\nWaz Islamic Sermon WER=36.99 CER=16.11\n\nWER=0.406609 CER=0.196573 (ignoring punctuation)\n```\n\nSome remarks:\n* Telemedicine is the most difficult one, because of 8khz?\n* why Bangladeshi TV Drama is more difficult than Indian TV Drama?\n* why Stage Drama Jatra is easier than Poem Recital and Puthi Literature?",
      "votes": null
    },
    {
      "id": "2425645",
      "postDate": "09/06/2023 04:57:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> <br>\nThanks for sharing the results and your insights!<br>\nSome comments on your remarks</p>\n<ol>\n<li>Yes, that might be the reason!</li>\n<li>Maybe Bangladeshi drama have people speaking with diverse dialect compared to Indian TV drama (Can't say for sure without a proper analysis)?</li>\n<li>Stage drama should not be easier, maybe we are looking at one single data point that is not representative of the distribution.</li>\n</ol>",
      "rawMarkdown": "Hi @tugstugi \nThanks for sharing the results and your insights!\nSome comments on your remarks\n1. Yes, that might be the reason!\n2. Maybe Bangladeshi drama have people speaking with diverse dialect compared to Indian TV drama (Can't say for sure without a proper analysis)?\n3. Stage drama should not be easier, maybe we are looking at one single data point that is not representative of the distribution.",
      "votes": null
    },
    {
      "id": "2425889",
      "postDate": "09/06/2023 09:00:53",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> , thanks for sharing your results! I annotated those audios, and here are some insights about them.</p>\n<ol>\n<li>The telemedicine audio that is provided has deficient quality, there are some parts that aren't even understandable by human ears! It contains some dialects too.</li>\n<li>For Indian TV drama and Bangla TV drama, there could be several reasons:</li>\n</ol>\n<blockquote>\n  <p>The Bangla TV drama audio provided here contains several transliterated words. These words are actually challenging for the model to understand<br>\n  The Bangla TV Drama audio contains a fast dialogue between two people, where speaker overlapping occurs sometimes. The Indian TV drama, actually includes a very slow speech of a person. </p>\n</blockquote>\n<ol>\n<li>Poem recital and Puthi Literature are actually two very difficult tasks for the models to understand. The poem recital contains audio where the words are pronounced in a lot more different manner than the usual cases. The same goes for the puthi recital, in this case t the words are pronounced in a musical tone. <br>\nCompared to them, the stage drama audio is a bit simpler!</li>\n</ol>\n<p>One important thing is these are just examples of these domains. They might contain so many more variations! So we can't actually be sure about which models might actually be the toughest to predict. We'd need a robust system that can perform well in any situation!</p>",
      "rawMarkdown": "tugstugi , thanks for sharing your results! I annotated those audios, and here are some insights about them.\n1. The telemedicine audio that is provided has deficient quality, there are some parts that aren't even understandable by human ears! It contains some dialects too.\n2. For Indian TV drama and Bangla TV drama, there could be several reasons:\n>The Bangla TV drama audio provided here contains several transliterated words. These words are actually challenging for the model to understand\n>The Bangla TV Drama audio contains a fast dialogue between two people, where speaker overlapping occurs sometimes. The Indian TV drama, actually includes a very slow speech of a person. \n\n3. Poem recital and Puthi Literature are actually two very difficult tasks for the models to understand. The poem recital contains audio where the words are pronounced in a lot more different manner than the usual cases. The same goes for the puthi recital, in this case t the words are pronounced in a musical tone. \nCompared to them, the stage drama audio is a bit simpler!\n\nOne important thing is these are just examples of these domains. They might contain so many more variations! So we can't actually be sure about which models might actually be the toughest to predict. We'd need a robust system that can perform well in any situation!",
      "votes": null
    },
    {
      "id": "2426065",
      "postDate": "09/06/2023 11:33:19",
      "content": "<p>Another important point is that these audios are each 40s long and their transcriptions are quite long. The test set might not contain such longer audios. It'll be interesting to see how we can correlate these WER with the LB WER</p>",
      "rawMarkdown": "Another important point is that these audios are each 40s long and their transcriptions are quite long. The test set might not contain such longer audios. It'll be interesting to see how we can correlate these WER with the LB WER",
      "votes": null
    },
    {
      "id": "2430156",
      "postDate": "09/09/2023 05:59:44",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> Can you please share the code which you used to preprocess the ground truth labels, and then compute your score ? It will help us better to be on the same page while evaluating on OOD. I am getting a 0.55ish WER compare to your 0.4066, so I think maybe I am missing some step ?</p>",
      "rawMarkdown": "tugstugi Can you please share the code which you used to preprocess the ground truth labels, and then compute your score ? It will help us better to be on the same page while evaluating on OOD. I am getting a 0.55ish WER compare to your 0.4066, so I think maybe I am missing some step ?",
      "votes": null
    },
    {
      "id": "2445616",
      "postDate": "09/19/2023 01:42:26",
      "content": "<p>So those examples audios don't reflect LB. LB0.364 models has WER=0.406970 CER=0.204985.</p>",
      "rawMarkdown": "So those examples audios don't reflect LB. LB0.364 models has WER=0.406970 CER=0.204985.",
      "votes": null
    },
    {
      "id": "2445770",
      "postDate": "09/19/2023 04:34:23",
      "content": "<p>These example audios are not correlated with LB? Do you get any CV &amp; LB correlation &amp; How?  </p>",
      "rawMarkdown": "These example audios are not correlated with LB? Do you get any CV & LB correlation & How?",
      "votes": null
    },
    {
      "id": "2449243",
      "postDate": "09/21/2023 04:55:58",
      "content": "<p>There is no correlation, atleast if you do inference on those long audios directly and compute WER.</p>",
      "rawMarkdown": "There is no correlation, atleast if you do inference on those long audios directly and compute WER.",
      "votes": null
    },
    {
      "id": "2454144",
      "postDate": "09/24/2023 15:47:28",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> </p>\n<p>You are doing an amazing job. Congrats and All the best.</p>\n<p>Can you share some tips to enter the 0.3 club ?</p>",
      "rawMarkdown": "tugstugi \n\nYou are doing an amazing job. Congrats and All the best.\n\nCan you share some tips to enter the 0.3 club ?",
      "votes": null
    },
    {
      "id": "2456361",
      "postDate": "09/26/2023 06:23:32",
      "content": "<p>nothing fancy, normalize and remove punctuations and then call jiwer.wer. As I said above, this WER doesn't reflect LB.</p>",
      "rawMarkdown": "nothing fancy, normalize and remove punctuations and then call jiwer.wer. As I said above, this WER doesn't reflect LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2425645,
      "author_name": "reasat",
      "author_url": "",
      "post_date": "09/06/2023 04:57:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> <br>\nThanks for sharing the results and your insights!<br>\nSome comments on your remarks</p>\n<ol>\n<li>Yes, that might be the reason!</li>\n<li>Maybe Bangladeshi drama have people speaking with diverse dialect compared to Indian TV drama (Can't say for sure without a proper analysis)?</li>\n<li>Stage drama should not be easier, maybe we are looking at one single data point that is not representative of the distribution.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2425889,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "09/06/2023 09:00:53",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> , thanks for sharing your results! I annotated those audios, and here are some insights about them.</p>\n<ol>\n<li>The telemedicine audio that is provided has deficient quality, there are some parts that aren't even understandable by human ears! It contains some dialects too.</li>\n<li>For Indian TV drama and Bangla TV drama, there could be several reasons:</li>\n</ol>\n<blockquote>\n  <p>The Bangla TV drama audio provided here contains several transliterated words. These words are actually challenging for the model to understand<br>\n  The Bangla TV Drama audio contains a fast dialogue between two people, where speaker overlapping occurs sometimes. The Indian TV drama, actually includes a very slow speech of a person. </p>\n</blockquote>\n<ol>\n<li>Poem recital and Puthi Literature are actually two very difficult tasks for the models to understand. The poem recital contains audio where the words are pronounced in a lot more different manner than the usual cases. The same goes for the puthi recital, in this case t the words are pronounced in a musical tone. <br>\nCompared to them, the stage drama audio is a bit simpler!</li>\n</ol>\n<p>One important thing is these are just examples of these domains. They might contain so many more variations! So we can't actually be sure about which models might actually be the toughest to predict. We'd need a robust system that can perform well in any situation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2426065,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "09/06/2023 11:33:19",
      "content": "<p>Another important point is that these audios are each 40s long and their transcriptions are quite long. The test set might not contain such longer audios. It'll be interesting to see how we can correlate these WER with the LB WER</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2430156,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "09/09/2023 05:59:44",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> Can you please share the code which you used to preprocess the ground truth labels, and then compute your score ? It will help us better to be on the same page while evaluating on OOD. I am getting a 0.55ish WER compare to your 0.4066, so I think maybe I am missing some step ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2456361,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "09/26/2023 06:23:32",
          "content": "<p>nothing fancy, normalize and remove punctuations and then call jiwer.wer. As I said above, this WER doesn't reflect LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2445616,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "09/19/2023 01:42:26",
      "content": "<p>So those examples audios don't reflect LB. LB0.364 models has WER=0.406970 CER=0.204985.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2445770,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "09/19/2023 04:34:23",
          "content": "<p>These example audios are not correlated with LB? Do you get any CV &amp; LB correlation &amp; How?  </p>",
          "votes": null,
          "replies": [
            {
              "id": 2449243,
              "author_name": "tugstugi",
              "author_url": "",
              "post_date": "09/21/2023 04:55:58",
              "content": "<p>There is no correlation, atleast if you do inference on those long audios directly and compute WER.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2454144,
      "author_name": "dhakshiin1601",
      "author_url": "",
      "post_date": "09/24/2023 15:47:28",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> </p>\n<p>You are doing an amazing job. Congrats and All the best.</p>\n<p>Can you share some tips to enter the 0.3 club ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2425619": "I wanted to see how the model behaves on the example audios. Annotations are from here: https://www.kaggle.com/competitions/bengaliai-speech/discussion/425932\n\nWER (without punctuation) on those audios seems to correllate with LB.\n\n```\nAudiobook WER=16.67 CER=4.46\nBangladeshi TV Drama WER=45.04 CER=15.56\nBengali Advertisement WER=57.14 CER=33.44\nCartoon WER=32.53 CER=12.71\nDebate WER=44.86 CER=23.65\nIndian TV Drama WER=32.08 CER=14.79\nMovie WER=26.00 CER=6.55\nNews Presentation WER=26.39 CER=12.78\nOnline Class WER=30.95 CER=17.66\nParliament Session WER=37.18 CER=17.02\nPoem Recital WER=60.00 CER=18.01\nPuthi Literature WER=60.47 CER=37.72\nSlang Profanity WER=66.05 CER=36.06\nStage Drama Jatra WER=19.48 CER=5.98\nTalk Show Interview WER=30.59 CER=13.15\nTelemedicine WER=68.83 CER=48.55\nWaz Islamic Sermon WER=36.99 CER=16.11\n\nWER=0.406609 CER=0.196573 (ignoring punctuation)\n```\n\nSome remarks:\n* Telemedicine is the most difficult one, because of 8khz?\n* why Bangladeshi TV Drama is more difficult than Indian TV Drama?\n* why Stage Drama Jatra is easier than Poem Recital and Puthi Literature?",
    "2425645": "Hi @tugstugi \nThanks for sharing the results and your insights!\nSome comments on your remarks\n1. Yes, that might be the reason!\n2. Maybe Bangladeshi drama have people speaking with diverse dialect compared to Indian TV drama (Can't say for sure without a proper analysis)?\n3. Stage drama should not be easier, maybe we are looking at one single data point that is not representative of the distribution.",
    "2425889": "tugstugi , thanks for sharing your results! I annotated those audios, and here are some insights about them.\n1. The telemedicine audio that is provided has deficient quality, there are some parts that aren't even understandable by human ears! It contains some dialects too.\n2. For Indian TV drama and Bangla TV drama, there could be several reasons:\n>The Bangla TV drama audio provided here contains several transliterated words. These words are actually challenging for the model to understand\n>The Bangla TV Drama audio contains a fast dialogue between two people, where speaker overlapping occurs sometimes. The Indian TV drama, actually includes a very slow speech of a person. \n\n3. Poem recital and Puthi Literature are actually two very difficult tasks for the models to understand. The poem recital contains audio where the words are pronounced in a lot more different manner than the usual cases. The same goes for the puthi recital, in this case t the words are pronounced in a musical tone. \nCompared to them, the stage drama audio is a bit simpler!\n\nOne important thing is these are just examples of these domains. They might contain so many more variations! So we can't actually be sure about which models might actually be the toughest to predict. We'd need a robust system that can perform well in any situation!",
    "2426065": "Another important point is that these audios are each 40s long and their transcriptions are quite long. The test set might not contain such longer audios. It'll be interesting to see how we can correlate these WER with the LB WER",
    "2430156": "tugstugi Can you please share the code which you used to preprocess the ground truth labels, and then compute your score ? It will help us better to be on the same page while evaluating on OOD. I am getting a 0.55ish WER compare to your 0.4066, so I think maybe I am missing some step ?",
    "2445616": "So those examples audios don't reflect LB. LB0.364 models has WER=0.406970 CER=0.204985.",
    "2445770": "These example audios are not correlated with LB? Do you get any CV & LB correlation & How?",
    "2449243": "There is no correlation, atleast if you do inference on those long audios directly and compute WER.",
    "2454144": "tugstugi \n\nYou are doing an amazing job. Congrats and All the best.\n\nCan you share some tips to enter the 0.3 club ?",
    "2456361": "nothing fancy, normalize and remove punctuations and then call jiwer.wer. As I said above, this WER doesn't reflect LB."
  },
  "source": "meta"
}