{
  "id": 491463,
  "title": "Query to the organisers : What does the symbol '<>' in the training transcriptions refer to?",
  "url": "/competitions/ben10/discussion/491463",
  "author_name": "",
  "post_date": "2024-04-06T00:09:08.283339700Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<blockquote>\n  <p>Many training samples contain the <strong>'&lt;&gt;' symbol</strong> in the transcripts. Some of the transcripts have it in the middle, some have it in the beginning/end. <br>\n  My guess is it refers to new/different speaker segments.<br>\n  Can the organisers clarify this and any other notations used in the training data? Do the ground truths in the test set contain them?</p>\n</blockquote>\n<p>Two other questions. Please answer if possible.</p>\n<blockquote>\n  <ol>\n  <li>Were the ground truths normalised using bnunicodenormalizer?</li>\n  <li>Do the ground truths contain punctuations? Adding/removing punctuations seemed to have little effect on the public LB</li>\n  </ol>\n</blockquote>\n<p>Thanks. Also kudos for arranging another nice competition on Bangla ASR!</p>",
  "messages": [
    {
      "id": "2737791",
      "postDate": "04/06/2024 00:09:08",
      "content": "<blockquote>\n  <p>Many training samples contain the <strong>'&lt;&gt;' symbol</strong> in the transcripts. Some of the transcripts have it in the middle, some have it in the beginning/end. <br>\n  My guess is it refers to new/different speaker segments.<br>\n  Can the organisers clarify this and any other notations used in the training data? Do the ground truths in the test set contain them?</p>\n</blockquote>\n<p>Two other questions. Please answer if possible.</p>\n<blockquote>\n  <ol>\n  <li>Were the ground truths normalised using bnunicodenormalizer?</li>\n  <li>Do the ground truths contain punctuations? Adding/removing punctuations seemed to have little effect on the public LB</li>\n  </ol>\n</blockquote>\n<p>Thanks. Also kudos for arranging another nice competition on Bangla ASR!</p>",
      "rawMarkdown": "> Many training samples contain the **'<>' symbol** in the transcripts. Some of the transcripts have it in the middle, some have it in the beginning/end. \nMy guess is it refers to new/different speaker segments.\nCan the organisers clarify this and any other notations used in the training data? Do the ground truths in the test set contain them?\n\nTwo other questions. Please answer if possible.\n\n>1. Were the ground truths normalised using bnunicodenormalizer?\n> 2. Do the ground truths contain punctuations? Adding/removing punctuations seemed to have little effect on the public LB\n\nThanks. Also kudos for arranging another nice competition on Bangla ASR!",
      "votes": null
    },
    {
      "id": "2737818",
      "postDate": "04/06/2024 00:45:12",
      "content": "<p>I think, deffinitely ground truth contains punctuations. My submission with punctuations improved LB by a good margin.</p>",
      "rawMarkdown": "I think, deffinitely ground truth contains punctuations. My submission with punctuations improved LB by a good margin.",
      "votes": null
    },
    {
      "id": "2738006",
      "postDate": "04/06/2024 04:01:30",
      "content": "<p>It was difficult to properly hear the spoken word in some cases. The annotators used the &lt;&gt; symbol to signify the existence of incomprehensible speech. The test set contains them. There shouldn’t be any more unexpected symbols. </p>\n<p>The ground truth contains punctuation.</p>\n<p>I'll get back to you about the bnunicodenormalizer question.</p>",
      "rawMarkdown": "It was difficult to properly hear the spoken word in some cases. The annotators used the <> symbol to signify the existence of incomprehensible speech. The test set contains them. There shouldn’t be any more unexpected symbols. \n\nThe ground truth contains punctuation.\n\nI'll get back to you about the bnunicodenormalizer question.",
      "votes": null
    },
    {
      "id": "2738305",
      "postDate": "04/06/2024 08:57:59",
      "content": "<p>Only the valid/test partition is normalized using bnunicodenormalizer</p>",
      "rawMarkdown": "Only the valid/test partition is normalized using bnunicodenormalizer",
      "votes": null
    },
    {
      "id": "2738537",
      "postDate": "04/06/2024 11:58:26",
      "content": "<p>Thanks a lot, bhai. The inclusion of the &lt;&gt; symbol will make it more challenging. Have you considered it not to include them in the test set? Because most of the sota Bengali ASR models don’t contain this information. Now it will be more difficult to finetune them with this symbol representing incomprehensive speech. I think it would be better not to include a challenge to detect an incomprehensive segment in a dialect-focused competition.</p>",
      "rawMarkdown": "Thanks a lot, bhai. The inclusion of the <> symbol will make it more challenging. Have you considered it not to include them in the test set? Because most of the sota Bengali ASR models don’t contain this information. Now it will be more difficult to finetune them with this symbol representing incomprehensive speech. I think it would be better not to include a challenge to detect an incomprehensive segment in a dialect-focused competition.",
      "votes": null
    },
    {
      "id": "2738905",
      "postDate": "04/06/2024 17:35:35",
      "content": "<p>It’s a complicated decision. Let me check with the group. </p>",
      "rawMarkdown": "It’s a complicated decision. Let me check with the group.",
      "votes": null
    },
    {
      "id": "2739052",
      "postDate": "04/06/2024 19:36:32",
      "content": "<p>Sure bhai. Thanks a lot for considering. This would help the competitors being more focused on handling the dialects!</p>",
      "rawMarkdown": "Sure bhai. Thanks a lot for considering. This would help the competitors being more focused on handling the dialects!",
      "votes": null
    },
    {
      "id": "2756636",
      "postDate": "04/17/2024 05:49:18",
      "content": "<p>Hello bhaiya, is there any update about the '&lt;&gt;' notation?</p>",
      "rawMarkdown": "Hello bhaiya, is there any update about the '<>' notation?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2737818,
      "author_name": "mahfuzulkabirsourav",
      "author_url": "",
      "post_date": "04/06/2024 00:45:12",
      "content": "<p>I think, deffinitely ground truth contains punctuations. My submission with punctuations improved LB by a good margin.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2738006,
      "author_name": "reasat",
      "author_url": "",
      "post_date": "04/06/2024 04:01:30",
      "content": "<p>It was difficult to properly hear the spoken word in some cases. The annotators used the &lt;&gt; symbol to signify the existence of incomprehensible speech. The test set contains them. There shouldn’t be any more unexpected symbols. </p>\n<p>The ground truth contains punctuation.</p>\n<p>I'll get back to you about the bnunicodenormalizer question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2738305,
          "author_name": "reasat",
          "author_url": "",
          "post_date": "04/06/2024 08:57:59",
          "content": "<p>Only the valid/test partition is normalized using bnunicodenormalizer</p>",
          "votes": null,
          "replies": [
            {
              "id": 2738537,
              "author_name": "mbmmurad",
              "author_url": "",
              "post_date": "04/06/2024 11:58:26",
              "content": "<p>Thanks a lot, bhai. The inclusion of the &lt;&gt; symbol will make it more challenging. Have you considered it not to include them in the test set? Because most of the sota Bengali ASR models don’t contain this information. Now it will be more difficult to finetune them with this symbol representing incomprehensive speech. I think it would be better not to include a challenge to detect an incomprehensive segment in a dialect-focused competition.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2738905,
                  "author_name": "reasat",
                  "author_url": "",
                  "post_date": "04/06/2024 17:35:35",
                  "content": "<p>It’s a complicated decision. Let me check with the group. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2739052,
                      "author_name": "mbmmurad",
                      "author_url": "",
                      "post_date": "04/06/2024 19:36:32",
                      "content": "<p>Sure bhai. Thanks a lot for considering. This would help the competitors being more focused on handling the dialects!</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2756636,
                      "author_name": "mahfuzulkabirsourav",
                      "author_url": "",
                      "post_date": "04/17/2024 05:49:18",
                      "content": "<p>Hello bhaiya, is there any update about the '&lt;&gt;' notation?</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2737791": "> Many training samples contain the **'<>' symbol** in the transcripts. Some of the transcripts have it in the middle, some have it in the beginning/end. \nMy guess is it refers to new/different speaker segments.\nCan the organisers clarify this and any other notations used in the training data? Do the ground truths in the test set contain them?\n\nTwo other questions. Please answer if possible.\n\n>1. Were the ground truths normalised using bnunicodenormalizer?\n> 2. Do the ground truths contain punctuations? Adding/removing punctuations seemed to have little effect on the public LB\n\nThanks. Also kudos for arranging another nice competition on Bangla ASR!",
    "2737818": "I think, deffinitely ground truth contains punctuations. My submission with punctuations improved LB by a good margin.",
    "2738006": "It was difficult to properly hear the spoken word in some cases. The annotators used the <> symbol to signify the existence of incomprehensible speech. The test set contains them. There shouldn’t be any more unexpected symbols. \n\nThe ground truth contains punctuation.\n\nI'll get back to you about the bnunicodenormalizer question.",
    "2738305": "Only the valid/test partition is normalized using bnunicodenormalizer",
    "2738537": "Thanks a lot, bhai. The inclusion of the <> symbol will make it more challenging. Have you considered it not to include them in the test set? Because most of the sota Bengali ASR models don’t contain this information. Now it will be more difficult to finetune them with this symbol representing incomprehensive speech. I think it would be better not to include a challenge to detect an incomprehensive segment in a dialect-focused competition.",
    "2738905": "It’s a complicated decision. Let me check with the group.",
    "2739052": "Sure bhai. Thanks a lot for considering. This would help the competitors being more focused on handling the dialects!",
    "2756636": "Hello bhaiya, is there any update about the '<>' notation?"
  },
  "source": "meta"
}