{
  "id": 673165,
  "title": "Unicode forms affecting scores",
  "url": "/competitions/dl-sprint-4-0-bengali-long-form-speech-recognition/discussion/673165",
  "author_name": "",
  "post_date": "2026-02-12T20:00:57.890359100Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I came across something that might be affecting all of our scores, and I wanted to share in case it’s helpful.</p>\n<p>I noticed that some of the ground truth transcripts use decomposed Unicode forms for characters like ড়, ঢ়, and য় (base letter + separate nukta), while others use the standard precomposed forms (U+09DC, etc.). They look exactly the same on screen, but because the byte sequences differ, correct predictions are sometimes being counted as errors.</p>\n<p>What’s more, the training labels themselves aren’t consistent — different files use different forms, and some even mix both in the same line. So the same prediction might be scored as correct or incorrect depending on which encoding happened to be used.</p>\n<p>When I tested on one sample, just applying NFC normalization to both my output and the reference lowered the WER from about 28% to 17% — with no other changes. That made me wonder if others might be seeing similar issues without realizing it.</p>\n<p>I’m not sure what the best fix is, but perhaps applying Unicode normalization (like NFC) in the evaluation script could help make scoring more consistent. Or if it’s easier, re-exporting the labels with a single encoding might also do the trick.</p>\n<p>I hope this is useful — thank you for considering it, and I really appreciate all the work everyone has put into this.</p>",
  "messages": [
    {
      "id": "3405348",
      "postDate": "02/12/2026 20:00:57",
      "content": "<p>Hi everyone,</p>\n<p>I came across something that might be affecting all of our scores, and I wanted to share in case it’s helpful.</p>\n<p>I noticed that some of the ground truth transcripts use decomposed Unicode forms for characters like ড়, ঢ়, and য় (base letter + separate nukta), while others use the standard precomposed forms (U+09DC, etc.). They look exactly the same on screen, but because the byte sequences differ, correct predictions are sometimes being counted as errors.</p>\n<p>What’s more, the training labels themselves aren’t consistent — different files use different forms, and some even mix both in the same line. So the same prediction might be scored as correct or incorrect depending on which encoding happened to be used.</p>\n<p>When I tested on one sample, just applying NFC normalization to both my output and the reference lowered the WER from about 28% to 17% — with no other changes. That made me wonder if others might be seeing similar issues without realizing it.</p>\n<p>I’m not sure what the best fix is, but perhaps applying Unicode normalization (like NFC) in the evaluation script could help make scoring more consistent. Or if it’s easier, re-exporting the labels with a single encoding might also do the trick.</p>\n<p>I hope this is useful — thank you for considering it, and I really appreciate all the work everyone has put into this.</p>",
      "rawMarkdown": "Hi everyone,\n\nI came across something that might be affecting all of our scores, and I wanted to share in case it’s helpful.\n\nI noticed that some of the ground truth transcripts use decomposed Unicode forms for characters like ড়, ঢ়, and য় (base letter + separate nukta), while others use the standard precomposed forms (U+09DC, etc.). They look exactly the same on screen, but because the byte sequences differ, correct predictions are sometimes being counted as errors.\n\nWhat’s more, the training labels themselves aren’t consistent — different files use different forms, and some even mix both in the same line. So the same prediction might be scored as correct or incorrect depending on which encoding happened to be used.\n\nWhen I tested on one sample, just applying NFC normalization to both my output and the reference lowered the WER from about 28% to 17% — with no other changes. That made me wonder if others might be seeing similar issues without realizing it.\n\nI’m not sure what the best fix is, but perhaps applying Unicode normalization (like NFC) in the evaluation script could help make scoring more consistent. Or if it’s easier, re-exporting the labels with a single encoding might also do the trick.\n\nI hope this is useful — thank you for considering it, and I really appreciate all the work everyone has put into this.",
      "votes": null
    },
    {
      "id": "3406019",
      "postDate": "02/14/2026 10:46:49",
      "content": "<p>Hi, please go through the pinned discussion section regarding this issue. Thanks for your concern</p>",
      "rawMarkdown": "Hi, please go through the pinned discussion section regarding this issue. Thanks for your concern",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3406019,
      "author_name": "shadmantabib",
      "author_url": "",
      "post_date": "02/14/2026 10:46:49",
      "content": "<p>Hi, please go through the pinned discussion section regarding this issue. Thanks for your concern</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3405348": "Hi everyone,\n\nI came across something that might be affecting all of our scores, and I wanted to share in case it’s helpful.\n\nI noticed that some of the ground truth transcripts use decomposed Unicode forms for characters like ড়, ঢ়, and য় (base letter + separate nukta), while others use the standard precomposed forms (U+09DC, etc.). They look exactly the same on screen, but because the byte sequences differ, correct predictions are sometimes being counted as errors.\n\nWhat’s more, the training labels themselves aren’t consistent — different files use different forms, and some even mix both in the same line. So the same prediction might be scored as correct or incorrect depending on which encoding happened to be used.\n\nWhen I tested on one sample, just applying NFC normalization to both my output and the reference lowered the WER from about 28% to 17% — with no other changes. That made me wonder if others might be seeing similar issues without realizing it.\n\nI’m not sure what the best fix is, but perhaps applying Unicode normalization (like NFC) in the evaluation script could help make scoring more consistent. Or if it’s easier, re-exporting the labels with a single encoding might also do the trick.\n\nI hope this is useful — thank you for considering it, and I really appreciate all the work everyone has put into this.",
    "3406019": "Hi, please go through the pinned discussion section regarding this issue. Thanks for your concern"
  },
  "source": "meta"
}