{
  "id": 429246,
  "title": "Query to organizers : Are the ground-truth sentences normalized?",
  "url": "/competitions/bengaliai-speech/discussion/429246",
  "author_name": "Md Boktiar Mahbub Murad",
  "post_date": "2023-08-04T16:27:39.004000",
  "votes": 17,
  "comment_count": 17,
  "views": 0,
  "content": "<p>They are two popular normalizers available for Bangla text sentences , </p>\n<blockquote>\n  <p><a href=\"https://pypi.org/project/bnunicodenormalizer\" target=\"_blank\">bnunicodenormalizer</a><br>\n  <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">csebuetnlp/normalizer</a>. </p>\n</blockquote>\n<p>It seems like <strong>bnunicodenomlaizer gives better LB result</strong> than unnormalized prediction and the other normalizer.<br>\nCan the organizers kindly confirm <strong>whether the ground truths for the hidden test set are normalized?</strong> If yes, then <strong>by which normalizer?</strong></p>",
  "messages": [
    {
      "id": 2374068,
      "postDate": "2023-08-04T16:27:39.003Z",
      "content": "<p>They are two popular normalizers available for Bangla text sentences , </p>\n<blockquote>\n  <p><a href=\"https://pypi.org/project/bnunicodenormalizer\" target=\"_blank\">bnunicodenormalizer</a><br>\n  <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">csebuetnlp/normalizer</a>. </p>\n</blockquote>\n<p>It seems like <strong>bnunicodenomlaizer gives better LB result</strong> than unnormalized prediction and the other normalizer.<br>\nCan the organizers kindly confirm <strong>whether the ground truths for the hidden test set are normalized?</strong> If yes, then <strong>by which normalizer?</strong></p>",
      "rawMarkdown": "They are two popular normalizers available for Bangla text sentences , \n>[bnunicodenormalizer](https://pypi.org/project/bnunicodenormalizer)\n>[csebuetnlp/normalizer](https://github.com/csebuetnlp/normalizer). \n\nIt seems like **bnunicodenomlaizer gives better LB result** than unnormalized prediction and the other normalizer.\nCan the organizers kindly confirm **whether the ground truths for the hidden test set are normalized?** If yes, then **by which normalizer?**",
      "votes": 17
    },
    {
      "id": 2374380,
      "postDate": "2023-08-04T23:02:01.830Z",
      "content": "<p>Hi Murad,  we used bnunicodenormalizer.</p>",
      "rawMarkdown": "Hi Murad,  we used bnunicodenormalizer.",
      "votes": 5,
      "replies": [
        {
          "id": 2374492,
          "postDate": "2023-08-05T03:20:27.643Z",
          "content": "<p>Thanks for the reply.<br>\nThis is very important infomation, it affects how the model should be trained and train data/prediction should be pre/post-processed.</p>\n<p>This post should be pinned in the discussion or the evaluation page sould be upated.</p>",
          "rawMarkdown": "Thanks for the reply.\nThis is very important infomation, it affects how the model should be trained and train data/prediction should be pre/post-processed.\n\nThis post should be pinned in the discussion or the evaluation page sould be upated.",
          "votes": 3
        },
        {
          "id": 2374496,
          "postDate": "2023-08-05T03:27:13.967Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2374497,
          "postDate": "2023-08-05T03:28:02.243Z",
          "content": "<p><a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> bhai, thanks for confirming. </p>",
          "rawMarkdown": "@reasat bhai, thanks for confirming. "
        },
        {
          "id": 2374651,
          "postDate": "2023-08-05T06:18:35.897Z",
          "content": "<p>So what about labels in the train set, are they normalized?</p>",
          "rawMarkdown": "So what about labels in the train set, are they normalized?",
          "replies": [
            {
              "id": 2374656,
              "postDate": "2023-08-05T06:21:21.723Z",
              "content": "<p>No,  neither of the train/validation split are normalized.</p>",
              "rawMarkdown": "No,  neither of the train/validation split are normalized.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2375383,
          "postDate": "2023-08-05T15:31:16.677Z",
          "content": "<p>Hi, thanks for this information!<br>\nAnother question is that are the punctuation marks such as '।' (Bengali full stop DARI) kept in the ground-truth test-set sentences while calculating WER? Because I notice that adding full stop after each sentence will cause a lot difference in terms of WER.</p>",
          "rawMarkdown": "Hi, thanks for this information!\nAnother question is that are the punctuation marks such as '।' (Bengali full stop DARI) kept in the ground-truth test-set sentences while calculating WER? Because I notice that adding full stop after each sentence will cause a lot difference in terms of WER.",
          "replies": [
            {
              "id": 2375929,
              "postDate": "2023-08-06T03:06:05.473Z",
              "content": "<p>Yeah, that's a critical issue. The punctuation should be kept in the GT sentences and then WER should be calculated. You can check whether they've done it by submitting a prediction with punctuation and another without punctuation. If the one with punctuations give better LB result, that means they've kept the punctuation in the GT's (which I think is the actual case here)</p>",
              "rawMarkdown": "Yeah, that's a critical issue. The punctuation should be kept in the GT sentences and then WER should be calculated. You can check whether they've done it by submitting a prediction with punctuation and another without punctuation. If the one with punctuations give better LB result, that means they've kept the punctuation in the GT's (which I think is the actual case here)"
            },
            {
              "id": 2376600,
              "postDate": "2023-08-06T13:49:37.223Z",
              "content": "<p>there are two choice:<br>\n1) the organizer can change the ground truth to only normalised  text and remove any punctuation. This will be a pure asr problem (instead of transcription). here we just solve the task of convert wave to word.</p>\n<p>2) you have to to predict punctuation, either by:<br>\na. post processing : wave --&gt; character --&gt; word + punctuation<br>\nb. treat punctuation as a character : wave --&gt; character + punctuation --&gt; word + punctuation</p>\n<p>which methods work better depends on the data …. wether you have sufficient data and if the source/target data have domain shift, etc</p>",
              "rawMarkdown": "there are two choice:\n1) the organizer can change the ground truth to only normalised  text and remove any punctuation. This will be a pure asr problem (instead of transcription). here we just solve the task of convert wave to word.\n\n2) you have to to predict punctuation, either by:\na. post processing : wave --> character --> word + punctuation\nb. treat punctuation as a character : wave --> character + punctuation --> word + punctuation\n\nwhich methods work better depends on the data .... wether you have sufficient data and if the source/target data have domain shift, etc",
              "votes": 1
            },
            {
              "id": 2377504,
              "postDate": "2023-08-07T07:16:16.150Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/revospeech\" target=\"_blank\">@revospeech</a> <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!</p>\n<p>Punctuations were left in because it is not completely disconnected from the ASR task. E.g., intonation changes based on the presence/absence of full stop. Since the sentences are human annotated and validated, there shouldn't be label noise wrt punctuations. Best of luck!</p>",
              "rawMarkdown": "Hi @revospeech @mbmmurad @hengck23!\n\nPunctuations were left in because it is not completely disconnected from the ASR task. E.g., intonation changes based on the presence/absence of full stop. Since the sentences are human annotated and validated, there shouldn't be label noise wrt punctuations. Best of luck!"
            },
            {
              "id": 2377522,
              "postDate": "2023-08-07T07:28:07.010Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2377523,
              "postDate": "2023-08-07T07:29:48.457Z",
              "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> bhai, thanks for the clarification! <br>\nWe expected that punctuations would be left in. This now makes the competition much more interesting and challenging too.<br>\nAlso kudos for arranging such a nice competition on Bangla ASR! </p>",
              "rawMarkdown": "@imtiazprio bhai, thanks for the clarification! \nWe expected that punctuations would be left in. This now makes the competition much more interesting and challenging too.\nAlso kudos for arranging such a nice competition on Bangla ASR! "
            },
            {
              "id": 2377593,
              "postDate": "2023-08-07T08:13:24.977Z",
              "content": "<p>thanks! also you guys are doing great congrats!</p>",
              "rawMarkdown": "thanks! also you guys are doing great congrats!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2381288,
      "postDate": "2023-08-09T05:45:24.600Z",
      "content": "<p>some information at discord discussion</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F292868f6b063331aee01c773f277dffd%2FSelection_999(2853).png?generation=1691559896222461&amp;alt=media\" alt=\"\"></p>\n<p>there is a paper on the normaliser used in evalution:</p>\n<p><a href=\"https://arxiv.org/pdf/2306.01743.pdf\" target=\"_blank\">https://arxiv.org/pdf/2306.01743.pdf</a></p>",
      "rawMarkdown": "some information at discord discussion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F292868f6b063331aee01c773f277dffd%2FSelection_999(2853).png?generation=1691559896222461&alt=media)\n\nthere is a paper on the normaliser used in evalution:\n\nhttps://arxiv.org/pdf/2306.01743.pdf",
      "votes": 2
    },
    {
      "id": 2379220,
      "postDate": "2023-08-08T04:46:28.763Z",
      "content": "<p>Sorry I'm new to this, but do these normalizers delete punctuation?</p>",
      "rawMarkdown": "Sorry I'm new to this, but do these normalizers delete punctuation?",
      "replies": [
        {
          "id": 2379240,
          "postDate": "2023-08-08T04:58:37.337Z",
          "content": "<p>Nope they don't delete punctuations, but may remove extra spaces around punctuations if inappropriate!</p>",
          "rawMarkdown": "Nope they don't delete punctuations, but may remove extra spaces around punctuations if inappropriate!",
          "votes": 2,
          "replies": [
            {
              "id": 2379273,
              "postDate": "2023-08-08T05:09:13.803Z",
              "content": "<p>Thank you so much!</p>",
              "rawMarkdown": "Thank you so much!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2374380,
      "author_name": "Tahsin",
      "author_url": "",
      "post_date": "2023-08-04T23:02:01.830000",
      "content": "<p>Hi Murad,  we used bnunicodenormalizer.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2374492,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-08-05T03:20:27.643000",
          "content": "<p>Thanks for the reply.<br>\nThis is very important infomation, it affects how the model should be trained and train data/prediction should be pre/post-processed.</p>\n<p>This post should be pinned in the discussion or the evaluation page sould be upated.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2374496,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-05T03:27:13.967000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2374497,
          "author_name": "Md Boktiar Mahbub Murad",
          "author_url": "",
          "post_date": "2023-08-05T03:28:02.243000",
          "content": "<p><a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> bhai, thanks for confirming. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2374651,
          "author_name": "Aphysict",
          "author_url": "",
          "post_date": "2023-08-05T06:18:35.897000",
          "content": "<p>So what about labels in the train set, are they normalized?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2374656,
              "author_name": "Md Boktiar Mahbub Murad",
              "author_url": "",
              "post_date": "2023-08-05T06:21:21.723000",
              "content": "<p>No,  neither of the train/validation split are normalized.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2375383,
          "author_name": "Modifier",
          "author_url": "",
          "post_date": "2023-08-05T15:31:16.677000",
          "content": "<p>Hi, thanks for this information!<br>\nAnother question is that are the punctuation marks such as '।' (Bengali full stop DARI) kept in the ground-truth test-set sentences while calculating WER? Because I notice that adding full stop after each sentence will cause a lot difference in terms of WER.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2375929,
              "author_name": "Md Boktiar Mahbub Murad",
              "author_url": "",
              "post_date": "2023-08-06T03:06:05.473000",
              "content": "<p>Yeah, that's a critical issue. The punctuation should be kept in the GT sentences and then WER should be calculated. You can check whether they've done it by submitting a prediction with punctuation and another without punctuation. If the one with punctuations give better LB result, that means they've kept the punctuation in the GT's (which I think is the actual case here)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2376600,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-06T13:49:37.223000",
              "content": "<p>there are two choice:<br>\n1) the organizer can change the ground truth to only normalised  text and remove any punctuation. This will be a pure asr problem (instead of transcription). here we just solve the task of convert wave to word.</p>\n<p>2) you have to to predict punctuation, either by:<br>\na. post processing : wave --&gt; character --&gt; word + punctuation<br>\nb. treat punctuation as a character : wave --&gt; character + punctuation --&gt; word + punctuation</p>\n<p>which methods work better depends on the data …. wether you have sufficient data and if the source/target data have domain shift, etc</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2377504,
              "author_name": "Ahmed Imtiaz Humayun",
              "author_url": "",
              "post_date": "2023-08-07T07:16:16.150000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/revospeech\" target=\"_blank\">@revospeech</a> <a href=\"https://www.kaggle.com/mbmmurad\" target=\"_blank\">@mbmmurad</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!</p>\n<p>Punctuations were left in because it is not completely disconnected from the ASR task. E.g., intonation changes based on the presence/absence of full stop. Since the sentences are human annotated and validated, there shouldn't be label noise wrt punctuations. Best of luck!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2377522,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-07T07:28:07.010000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2377523,
              "author_name": "Md Boktiar Mahbub Murad",
              "author_url": "",
              "post_date": "2023-08-07T07:29:48.457000",
              "content": "<p><a href=\"https://www.kaggle.com/imtiazprio\" target=\"_blank\">@imtiazprio</a> bhai, thanks for the clarification! <br>\nWe expected that punctuations would be left in. This now makes the competition much more interesting and challenging too.<br>\nAlso kudos for arranging such a nice competition on Bangla ASR! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2377593,
              "author_name": "Ahmed Imtiaz Humayun",
              "author_url": "",
              "post_date": "2023-08-07T08:13:24.977000",
              "content": "<p>thanks! also you guys are doing great congrats!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2381288,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-09T05:45:24.600000",
      "content": "<p>some information at discord discussion</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F292868f6b063331aee01c773f277dffd%2FSelection_999(2853).png?generation=1691559896222461&amp;alt=media\" alt=\"\"></p>\n<p>there is a paper on the normaliser used in evalution:</p>\n<p><a href=\"https://arxiv.org/pdf/2306.01743.pdf\" target=\"_blank\">https://arxiv.org/pdf/2306.01743.pdf</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2379220,
      "author_name": "Larry Gan",
      "author_url": "",
      "post_date": "2023-08-08T04:46:28.763000",
      "content": "<p>Sorry I'm new to this, but do these normalizers delete punctuation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2379240,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2023-08-08T04:58:37.337000",
          "content": "<p>Nope they don't delete punctuations, but may remove extra spaces around punctuations if inappropriate!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2379273,
              "author_name": "Larry Gan",
              "author_url": "",
              "post_date": "2023-08-08T05:09:13.803000",
              "content": "<p>Thank you so much!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2374068": "They are two popular normalizers available for Bangla text sentences , \n>[bnunicodenormalizer](https://pypi.org/project/bnunicodenormalizer)\n>[csebuetnlp/normalizer](https://github.com/csebuetnlp/normalizer). \n\nIt seems like **bnunicodenomlaizer gives better LB result** than unnormalized prediction and the other normalizer.\nCan the organizers kindly confirm **whether the ground truths for the hidden test set are normalized?** If yes, then **by which normalizer?**",
    "2374380": "Hi Murad,  we used bnunicodenormalizer.",
    "2381288": "some information at discord discussion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F292868f6b063331aee01c773f277dffd%2FSelection_999(2853).png?generation=1691559896222461&alt=media)\n\nthere is a paper on the normaliser used in evalution:\n\nhttps://arxiv.org/pdf/2306.01743.pdf",
    "2379220": "Sorry I'm new to this, but do these normalizers delete punctuation?"
  }
}