{
  "id": 425932,
  "title": "Hand Annotation of the Example OOD audios",
  "url": "/competitions/bengaliai-speech/discussion/425932",
  "author_name": "",
  "post_date": "2023-07-21T03:34:20.792831400Z",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Since the main challenge of this competition is to <strong>Recognize Out of domain audios</strong> well, it'd be better if we evaluate our models on some OOD data. The organizers have provided some example audios of the 17 OOD domains. I have hand annotated these audios so that we could evaluate our models on these. </p>\n<p>Please refer to the notebook to see the dataset and some insights.<br>\nLink to the Notebook: <a href=\"https://www.kaggle.com/code/mbmmurad/exploring-ood-domains-hand-annotation-of-examples/notebook\" target=\"_blank\">https://www.kaggle.com/code/mbmmurad/exploring-ood-domains-hand-annotation-of-examples/notebook</a></p>\n<p>There still might be some inconsistencies in the annotation, I'll try to improve it gradually. </p>",
  "messages": [
    {
      "id": "2352499",
      "postDate": "07/21/2023 03:34:20",
      "content": "<p>Since the main challenge of this competition is to <strong>Recognize Out of domain audios</strong> well, it'd be better if we evaluate our models on some OOD data. The organizers have provided some example audios of the 17 OOD domains. I have hand annotated these audios so that we could evaluate our models on these. </p>\n<p>Please refer to the notebook to see the dataset and some insights.<br>\nLink to the Notebook: <a href=\"https://www.kaggle.com/code/mbmmurad/exploring-ood-domains-hand-annotation-of-examples/notebook\" target=\"_blank\">https://www.kaggle.com/code/mbmmurad/exploring-ood-domains-hand-annotation-of-examples/notebook</a></p>\n<p>There still might be some inconsistencies in the annotation, I'll try to improve it gradually. </p>",
      "rawMarkdown": "Since the main challenge of this competition is to **Recognize Out of domain audios** well, it'd be better if we evaluate our models on some OOD data. The organizers have provided some example audios of the 17 OOD domains. I have hand annotated these audios so that we could evaluate our models on these. \n\nPlease refer to the notebook to see the dataset and some insights.\nLink to the Notebook: https://www.kaggle.com/code/mbmmurad/exploring-ood-domains-hand-annotation-of-examples/notebook\n\nThere still might be some inconsistencies in the annotation, I'll try to improve it gradually.",
      "votes": null
    },
    {
      "id": "2356583",
      "postDate": "07/24/2023 09:27:39",
      "content": "<p>evaluation using your ground truth.<br>\nmodel is baseline nemo conformer-CTC (which has WER of 0.88 on OOD dataset in the dataset paper)</p>\n<pre><code>                         file       score\n               Audiobook.wav    \n    Bangladeshi TV Drama.wav    \n   Bengali Advertisement.wav    \n                 Cartoon.wav    \n                  Debate.wav    \n         Indian TV Drama.wav    \n                   Movie.wav    \n       News Presentation.wav    \n            Online Class.wav    \n      Parliament Session.wav    \n           Poem Recital.wav    \n       Puthi Literature.wav    \n        Slang Profanity.mp3    \n      Stage Drama Jatra.wav    \n    Talk Show Interview.wav    \n           Telemedicine.mp3    \n     Waz Islamic Sermon.wav    \n\n  x  columns\n0.8416567937027707\n\nProcess finished with exit code 0\n</code></pre>",
      "rawMarkdown": "evaluation using your ground truth.\nmodel is baseline nemo conformer-CTC (which has WER of 0.88 on OOD dataset in the dataset paper)\n\n```\n\n                         file  ...     score\n0               Audiobook.wav  ...  0.555556\n1    Bangladeshi TV Drama.wav  ...  0.808000\n2   Bengali Advertisement.wav  ...  0.990196\n3                 Cartoon.wav  ...  0.949367\n4                  Debate.wav  ...  0.711538\n5         Indian TV Drama.wav  ...  0.897959\n6                   Movie.wav  ...  0.933333\n7       News Presentation.wav  ...  0.735294\n8            Online Class.wav  ...  0.721311\n9      Parliament Session.wav  ...  0.649351\n10           Poem Recital.wav  ...  0.930233\n11       Puthi Literature.wav  ...  0.950000\n12        Slang Profanity.mp3  ...  0.975000\n13      Stage Drama Jatra.wav  ...  0.881579\n14    Talk Show Interview.wav  ...  0.759036\n15           Telemedicine.mp3  ...  0.947368\n16     Waz Islamic Sermon.wav  ...  0.913043\n\n[17 rows x 4 columns]\n0.8416567937027707\n\nProcess finished with exit code 0\n\n\n```",
      "votes": null
    },
    {
      "id": "2357083",
      "postDate": "07/24/2023 14:36:19",
      "content": "<p>Good job! Btw This WER will not actually reflect the performance on the test set. These 17 audios are each 40s-1min long. So it contains around 6-7 sentences. So a better way could be to split these audios into several clips and then infer on the smaller ones.  WER is a very misleading metric. \"How are you\" and \"How are you?\" have a WER of 33.33%. So punctuations also play a major role here.  For longer sentences, models fails to generate correct punctuation.</p>",
      "rawMarkdown": "Good job! Btw This WER will not actually reflect the performance on the test set. These 17 audios are each 40s-1min long. So it contains around 6-7 sentences. So a better way could be to split these audios into several clips and then infer on the smaller ones.  WER is a very misleading metric. \"How are you\" and \"How are you?\" have a WER of 33.33%. So punctuations also play a major role here.  For longer sentences, models fails to generate correct punctuation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2356583,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/24/2023 09:27:39",
      "content": "<p>evaluation using your ground truth.<br>\nmodel is baseline nemo conformer-CTC (which has WER of 0.88 on OOD dataset in the dataset paper)</p>\n<pre><code>                         file       score\n               Audiobook.wav    \n    Bangladeshi TV Drama.wav    \n   Bengali Advertisement.wav    \n                 Cartoon.wav    \n                  Debate.wav    \n         Indian TV Drama.wav    \n                   Movie.wav    \n       News Presentation.wav    \n            Online Class.wav    \n      Parliament Session.wav    \n           Poem Recital.wav    \n       Puthi Literature.wav    \n        Slang Profanity.mp3    \n      Stage Drama Jatra.wav    \n    Talk Show Interview.wav    \n           Telemedicine.mp3    \n     Waz Islamic Sermon.wav    \n\n  x  columns\n0.8416567937027707\n\nProcess finished with exit code 0\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2357083,
          "author_name": "mbmmurad",
          "author_url": "",
          "post_date": "07/24/2023 14:36:19",
          "content": "<p>Good job! Btw This WER will not actually reflect the performance on the test set. These 17 audios are each 40s-1min long. So it contains around 6-7 sentences. So a better way could be to split these audios into several clips and then infer on the smaller ones.  WER is a very misleading metric. \"How are you\" and \"How are you?\" have a WER of 33.33%. So punctuations also play a major role here.  For longer sentences, models fails to generate correct punctuation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2352499": "Since the main challenge of this competition is to **Recognize Out of domain audios** well, it'd be better if we evaluate our models on some OOD data. The organizers have provided some example audios of the 17 OOD domains. I have hand annotated these audios so that we could evaluate our models on these. \n\nPlease refer to the notebook to see the dataset and some insights.\nLink to the Notebook: https://www.kaggle.com/code/mbmmurad/exploring-ood-domains-hand-annotation-of-examples/notebook\n\nThere still might be some inconsistencies in the annotation, I'll try to improve it gradually.",
    "2356583": "evaluation using your ground truth.\nmodel is baseline nemo conformer-CTC (which has WER of 0.88 on OOD dataset in the dataset paper)\n\n```\n\n                         file  ...     score\n0               Audiobook.wav  ...  0.555556\n1    Bangladeshi TV Drama.wav  ...  0.808000\n2   Bengali Advertisement.wav  ...  0.990196\n3                 Cartoon.wav  ...  0.949367\n4                  Debate.wav  ...  0.711538\n5         Indian TV Drama.wav  ...  0.897959\n6                   Movie.wav  ...  0.933333\n7       News Presentation.wav  ...  0.735294\n8            Online Class.wav  ...  0.721311\n9      Parliament Session.wav  ...  0.649351\n10           Poem Recital.wav  ...  0.930233\n11       Puthi Literature.wav  ...  0.950000\n12        Slang Profanity.mp3  ...  0.975000\n13      Stage Drama Jatra.wav  ...  0.881579\n14    Talk Show Interview.wav  ...  0.759036\n15           Telemedicine.mp3  ...  0.947368\n16     Waz Islamic Sermon.wav  ...  0.913043\n\n[17 rows x 4 columns]\n0.8416567937027707\n\nProcess finished with exit code 0\n\n\n```",
    "2357083": "Good job! Btw This WER will not actually reflect the performance on the test set. These 17 audios are each 40s-1min long. So it contains around 6-7 sentences. So a better way could be to split these audios into several clips and then infer on the smaller ones.  WER is a very misleading metric. \"How are you\" and \"How are you?\" have a WER of 33.33%. So punctuations also play a major role here.  For longer sentences, models fails to generate correct punctuation."
  },
  "source": "meta"
}