{
  "id": 673438,
  "title": "IMPORTANT EVALUATION UPDATE: Leaderboard Rescoring & Unicode Normalization",
  "url": "/competitions/dl-sprint-4-0-bengali-long-form-speech-recognition/discussion/673438",
  "author_name": "Shadman Tabib",
  "post_date": "2026-02-14T10:44:31.276000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Dear Participants,</strong></p>\n<p>We have rescored the entire leaderboard using a consistent ground truth evaluation pipeline. The leaderboard may remain temporarily frozen. Please check the updated final scores.</p>\n<p>To ensure fairness and remove encoding-related inconsistencies, we applied Unicode normalization (NFC) and cleaned invisible characters in the CSV ground truth.</p>\n<p><strong>Reason:</strong><br>\nVisually identical Bengali text may have different Unicode representations (composite vs decomposed), which can introduce artificial scoring penalties.</p>\n<p><strong>Post-processing applied to ground truth:</strong></p>\n<pre><code>ZW = r\"[\\u200B-\\u200D\\uFEFF]\"  # zero-width space/joiners + BOM\n\ndef normalize_bn_text(s: str) -&gt; str:\n    if s is None:\n        return \"\"\n    s = str(s)\n    s = unicodedata.normalize(\"NFC\", s)\n    s = re.sub(ZW, \"\", s)\n    s = s.replace(\"\\u00A0\", \" \")\n    s = \" \".join(s.split())\n    return s\n</code></pre>\n<p><strong>Important Notice for Future Submissions:</strong><br>\nParticipants are requested to submit predictions using <strong>NFC-normalized CSV files</strong> to avoid Unicode-related inconsistencies during evaluation.</p>\n<p>Only ground truth text representation was standardized. Model outputs were not modified.</p>\n<p>Thank you for your understanding. Best of luck for the remaining days of the competition.</p>",
  "messages": [
    {
      "id": 3406018,
      "postDate": "2026-02-14T10:44:31.277Z",
      "content": "<p><strong>Dear Participants,</strong></p>\n<p>We have rescored the entire leaderboard using a consistent ground truth evaluation pipeline. The leaderboard may remain temporarily frozen. Please check the updated final scores.</p>\n<p>To ensure fairness and remove encoding-related inconsistencies, we applied Unicode normalization (NFC) and cleaned invisible characters in the CSV ground truth.</p>\n<p><strong>Reason:</strong><br>\nVisually identical Bengali text may have different Unicode representations (composite vs decomposed), which can introduce artificial scoring penalties.</p>\n<p><strong>Post-processing applied to ground truth:</strong></p>\n<pre><code>ZW = r\"[\\u200B-\\u200D\\uFEFF]\"  # zero-width space/joiners + BOM\n\ndef normalize_bn_text(s: str) -&gt; str:\n    if s is None:\n        return \"\"\n    s = str(s)\n    s = unicodedata.normalize(\"NFC\", s)\n    s = re.sub(ZW, \"\", s)\n    s = s.replace(\"\\u00A0\", \" \")\n    s = \" \".join(s.split())\n    return s\n</code></pre>\n<p><strong>Important Notice for Future Submissions:</strong><br>\nParticipants are requested to submit predictions using <strong>NFC-normalized CSV files</strong> to avoid Unicode-related inconsistencies during evaluation.</p>\n<p>Only ground truth text representation was standardized. Model outputs were not modified.</p>\n<p>Thank you for your understanding. Best of luck for the remaining days of the competition.</p>",
      "rawMarkdown": "**Dear Participants,**\n\nWe have rescored the entire leaderboard using a consistent ground truth evaluation pipeline. The leaderboard may remain temporarily frozen. Please check the updated final scores.\n\nTo ensure fairness and remove encoding-related inconsistencies, we applied Unicode normalization (NFC) and cleaned invisible characters in the CSV ground truth.\n\n**Reason:**  \nVisually identical Bengali text may have different Unicode representations (composite vs decomposed), which can introduce artificial scoring penalties.\n\n**Post-processing applied to ground truth:**\n\n```python\nZW = r\"[\\u200B-\\u200D\\uFEFF]\"  # zero-width space/joiners + BOM\n\ndef normalize_bn_text(s: str) -> str:\n    if s is None:\n        return \"\"\n    s = str(s)\n    s = unicodedata.normalize(\"NFC\", s)\n    s = re.sub(ZW, \"\", s)\n    s = s.replace(\"\\u00A0\", \" \")\n    s = \" \".join(s.split())\n    return s\n```\n\n**Important Notice for Future Submissions:**  \nParticipants are requested to submit predictions using **NFC-normalized CSV files** to avoid Unicode-related inconsistencies during evaluation.\n\nOnly ground truth text representation was standardized. Model outputs were not modified.\n\nThank you for your understanding. Best of luck for the remaining days of the competition.\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3406018": "**Dear Participants,**\n\nWe have rescored the entire leaderboard using a consistent ground truth evaluation pipeline. The leaderboard may remain temporarily frozen. Please check the updated final scores.\n\nTo ensure fairness and remove encoding-related inconsistencies, we applied Unicode normalization (NFC) and cleaned invisible characters in the CSV ground truth.\n\n**Reason:**  \nVisually identical Bengali text may have different Unicode representations (composite vs decomposed), which can introduce artificial scoring penalties.\n\n**Post-processing applied to ground truth:**\n\n```python\nZW = r\"[\\u200B-\\u200D\\uFEFF]\"  # zero-width space/joiners + BOM\n\ndef normalize_bn_text(s: str) -> str:\n    if s is None:\n        return \"\"\n    s = str(s)\n    s = unicodedata.normalize(\"NFC\", s)\n    s = re.sub(ZW, \"\", s)\n    s = s.replace(\"\\u00A0\", \" \")\n    s = \" \".join(s.split())\n    return s\n```\n\n**Important Notice for Future Submissions:**  \nParticipants are requested to submit predictions using **NFC-normalized CSV files** to avoid Unicode-related inconsistencies during evaluation.\n\nOnly ground truth text representation was standardized. Model outputs were not modified.\n\nThank you for your understanding. Best of luck for the remaining days of the competition.\n"
  }
}