{
  "id": 445104,
  "title": "Help needed from bengali language speakers. Can you compare 2 sentences for me?",
  "url": "/competitions/bengaliai-speech/discussion/445104",
  "author_name": "",
  "post_date": "2023-10-05T08:25:36.319240Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<pre><code>from datasets import load_metric\nwer_metric = ()\n#wer=\nlabel_str =  \n\npred_str =   \n\nwer = wer_metric(predictions=, references=)\n\n</code></pre>\n<p>I am getting wer:  0.2727272727272727 <br>\nAre the strings above equal? if they are why am I getting 0.27 for WER?</p>",
  "messages": [
    {
      "id": "2468149",
      "postDate": "10/05/2023 08:25:36",
      "content": "<pre><code>from datasets import load_metric\nwer_metric = ()\n#wer=\nlabel_str =  \n\npred_str =   \n\nwer = wer_metric(predictions=, references=)\n\n</code></pre>\n<p>I am getting wer:  0.2727272727272727 <br>\nAre the strings above equal? if they are why am I getting 0.27 for WER?</p>",
      "rawMarkdown": "```\nfrom datasets import load_metric\nwer_metric = load_metric(\"wer\")\n#wer=0.27\nlabel_str =  \"ব্যাংকার ওয়েবসাইট গিয়ে হামরা হামাদের ভার্চুয়াল বা ডিজিটাল কার্ড দ্যাখতে পারতেছি\"\n\npred_str =   \"ব্যাংকার ওয়েবসাইট গিয়ে হামরা হামাদের ভার্চুয়াল বা ডিজিটাল কার্ড দ্যাখতে পারতেছি\"\n\nwer = wer_metric.compute(predictions=[pred_str], references=[label_str])\nprint(\"wer: \", wer)\n\n```\nI am getting wer:  0.2727272727272727 \nAre the strings above equal? if they are why am I getting 0.27 for WER?",
      "votes": null
    },
    {
      "id": "2468389",
      "postDate": "10/05/2023 13:01:02",
      "content": "<p>ওয়েবসাইট != ওয়েবসাইট</p>",
      "rawMarkdown": "ওয়েবসাইট != ওয়েবসাইট",
      "votes": null
    },
    {
      "id": "2468470",
      "postDate": "10/05/2023 14:19:03",
      "content": "<p>different way to spell \"website\" ?  I can't tell the difference :( </p>",
      "rawMarkdown": "different way to spell \"website\" ?  I can't tell the difference :(",
      "votes": null
    },
    {
      "id": "2468526",
      "postDate": "10/05/2023 15:21:21",
      "content": "<p>Hi, I just tried checking your case. So apparently word-by-word comparisons yield that \"য়ে\" != \"য়ে\". However, if you try a normalizer such as the bnunicodenormalizer, you can immediately see that the two strings are exactly the same after normalization. </p>",
      "rawMarkdown": "Hi, I just tried checking your case. So apparently word-by-word comparisons yield that \"য়ে\" != \"য়ে\". However, if you try a normalizer such as the bnunicodenormalizer, you can immediately see that the two strings are exactly the same after normalization.",
      "votes": null
    },
    {
      "id": "2468561",
      "postDate": "10/05/2023 15:49:38",
      "content": "<p>Check in this way: </p>\n<pre><code> difflib\n\n ():\n    differ = difflib.Differ()\n    diff = (differ.compare(ground_truth, prediction))\n\n    highlighted_diff = []\n\n     item  diff:\n         item.startswith():  \n            highlighted_diff.append( + item[-] + )  \n         item.startswith():  \n            highlighted_diff.append( + item[-] + )  \n        :\n            highlighted_diff.append(item[-])\n\n     .join(highlighted_diff)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2125668%2Fa738cd4984034fe688de1b0673c12e17%2Fobraz_2023-10-05_175354744.png?generation=1696521235534481&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Check in this way: \n\n```python\nimport difflib\n\ndef highlight_differences(ground_truth, prediction):\n    differ = difflib.Differ()\n    diff = list(differ.compare(ground_truth, prediction))\n\n    highlighted_diff = []\n\n    for item in diff:\n        if item.startswith('-'):  # For deleted fragments\n            highlighted_diff.append('\\033[91m\\033[43m' + item[-1] + '\\033[0m')  # Red text on yellow backgound\n        elif item.startswith('+'):  # For added framgents\n            highlighted_diff.append('\\033[94m\\033[47m' + item[-1] + '\\033[0m')  # Blue text on light gray background\n        else:\n            highlighted_diff.append(item[-1])\n\n    return ''.join(highlighted_diff)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2125668%2Fa738cd4984034fe688de1b0673c12e17%2Fobraz_2023-10-05_175354744.png?generation=1696521235534481&alt=media)",
      "votes": null
    },
    {
      "id": "2468884",
      "postDate": "10/06/2023 00:09:54",
      "content": "<p>Hi. The problem is mainly with the character 'য়' . It can be written in two ways. </p>\n<blockquote>\n  <p>'য়' -&gt; 'য', '়' ( The 2nd one is called Nukta)<br>\n  'য়' -&gt; 'য়'</p>\n</blockquote>\n<p>Of the two sentences you mentioned, one of them contains the first format and the other one has the 2nd one. That's why it's showing a high WER although there's no difference visibly.</p>\n<p>To solve this exact issue, normalizer is used. The normalizer transforms all the sentences in the same canonical form.</p>",
      "rawMarkdown": "Hi. The problem is mainly with the character 'য়' . It can be written in two ways. \n> 'য়' -> 'য', '়' ( The 2nd one is called Nukta)\n> 'য়' -> 'য়'\n\nOf the two sentences you mentioned, one of them contains the first format and the other one has the 2nd one. That's why it's showing a high WER although there's no difference visibly.\n\nTo solve this exact issue, normalizer is used. The normalizer transforms all the sentences in the same canonical form.",
      "votes": null
    },
    {
      "id": "2469612",
      "postDate": "10/06/2023 13:26:17",
      "content": "<p>Hi, in your <a href=\"https://www.kaggle.com/code/mbmmurad/detailed-eda-normalizer-and-wer\" target=\"_blank\">EDA notebook</a>, you used <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">https://github.com/csebuetnlp/normalizer</a> to normalize the sentence. May I ask what is the difference between this normalizer and the bnunicodenormalizer </p>",
      "rawMarkdown": "Hi, in your [EDA notebook](https://www.kaggle.com/code/mbmmurad/detailed-eda-normalizer-and-wer), you used https://github.com/csebuetnlp/normalizer to normalize the sentence. May I ask what is the difference between this normalizer and the bnunicodenormalizer",
      "votes": null
    },
    {
      "id": "2471896",
      "postDate": "10/06/2023 18:06:26",
      "content": "<p>I also noticed that the normalizer has Non Comercial Use license [Contents of this repository are restricted to non-commercial research purposes only under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).] If you are working towards prize money I don't think we can use it.</p>",
      "rawMarkdown": "I also noticed that the normalizer has Non Comercial Use license [Contents of this repository are restricted to non-commercial research purposes only under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).] If you are working towards prize money I don't think we can use it.",
      "votes": null
    },
    {
      "id": "2474720",
      "postDate": "10/09/2023 12:10:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ly8388\" target=\"_blank\">@ly8388</a>, does this indicate using bnunicodenormalizer while training can lead to performance dip? or is this only limited to naked eye confusion?</p>",
      "rawMarkdown": "Hi @ly8388, does this indicate using bnunicodenormalizer while training can lead to performance dip? or is this only limited to naked eye confusion?",
      "votes": null
    },
    {
      "id": "2474735",
      "postDate": "10/09/2023 12:23:25",
      "content": "<p>I think a normalizer affect the performance.<br>\nYou also should be careful of what normalizer is used in the preprocess of the pretraining model.</p>",
      "rawMarkdown": "I think a normalizer affect the performance.\nYou also should be careful of what normalizer is used in the preprocess of the pretraining model.",
      "votes": null
    },
    {
      "id": "2482325",
      "postDate": "10/14/2023 22:06:32",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2F4da3f336c2bd69798b2c164e85795b8f%2FScreenshot%202023-10-14%20at%205.05.56PM.png?generation=1697321181363157&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2F4da3f336c2bd69798b2c164e85795b8f%2FScreenshot%202023-10-14%20at%205.05.56PM.png?generation=1697321181363157&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2468389,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "10/05/2023 13:01:02",
      "content": "<p>ওয়েবসাইট != ওয়েবসাইট</p>",
      "votes": null,
      "replies": [
        {
          "id": 2468470,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "10/05/2023 14:19:03",
          "content": "<p>different way to spell \"website\" ?  I can't tell the difference :( </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2482325,
          "author_name": "bayartsogtya",
          "author_url": "",
          "post_date": "10/14/2023 22:06:32",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2F4da3f336c2bd69798b2c164e85795b8f%2FScreenshot%202023-10-14%20at%205.05.56PM.png?generation=1697321181363157&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2468526,
      "author_name": "ly8388",
      "author_url": "",
      "post_date": "10/05/2023 15:21:21",
      "content": "<p>Hi, I just tried checking your case. So apparently word-by-word comparisons yield that \"য়ে\" != \"য়ে\". However, if you try a normalizer such as the bnunicodenormalizer, you can immediately see that the two strings are exactly the same after normalization. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2474720,
          "author_name": "anilreddyvv",
          "author_url": "",
          "post_date": "10/09/2023 12:10:29",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ly8388\" target=\"_blank\">@ly8388</a>, does this indicate using bnunicodenormalizer while training can lead to performance dip? or is this only limited to naked eye confusion?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2474735,
              "author_name": "yamashitamotokazu",
              "author_url": "",
              "post_date": "10/09/2023 12:23:25",
              "content": "<p>I think a normalizer affect the performance.<br>\nYou also should be careful of what normalizer is used in the preprocess of the pretraining model.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2468561,
      "author_name": "hubert101",
      "author_url": "",
      "post_date": "10/05/2023 15:49:38",
      "content": "<p>Check in this way: </p>\n<pre><code> difflib\n\n ():\n    differ = difflib.Differ()\n    diff = (differ.compare(ground_truth, prediction))\n\n    highlighted_diff = []\n\n     item  diff:\n         item.startswith():  \n            highlighted_diff.append( + item[-] + )  \n         item.startswith():  \n            highlighted_diff.append( + item[-] + )  \n        :\n            highlighted_diff.append(item[-])\n\n     .join(highlighted_diff)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2125668%2Fa738cd4984034fe688de1b0673c12e17%2Fobraz_2023-10-05_175354744.png?generation=1696521235534481&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2468884,
      "author_name": "mbmmurad",
      "author_url": "",
      "post_date": "10/06/2023 00:09:54",
      "content": "<p>Hi. The problem is mainly with the character 'য়' . It can be written in two ways. </p>\n<blockquote>\n  <p>'য়' -&gt; 'য', '়' ( The 2nd one is called Nukta)<br>\n  'য়' -&gt; 'য়'</p>\n</blockquote>\n<p>Of the two sentences you mentioned, one of them contains the first format and the other one has the 2nd one. That's why it's showing a high WER although there's no difference visibly.</p>\n<p>To solve this exact issue, normalizer is used. The normalizer transforms all the sentences in the same canonical form.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2469612,
          "author_name": "baohaoliao",
          "author_url": "",
          "post_date": "10/06/2023 13:26:17",
          "content": "<p>Hi, in your <a href=\"https://www.kaggle.com/code/mbmmurad/detailed-eda-normalizer-and-wer\" target=\"_blank\">EDA notebook</a>, you used <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">https://github.com/csebuetnlp/normalizer</a> to normalize the sentence. May I ask what is the difference between this normalizer and the bnunicodenormalizer </p>",
          "votes": null,
          "replies": [
            {
              "id": 2471896,
              "author_name": "robot2020",
              "author_url": "",
              "post_date": "10/06/2023 18:06:26",
              "content": "<p>I also noticed that the normalizer has Non Comercial Use license [Contents of this repository are restricted to non-commercial research purposes only under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).] If you are working towards prize money I don't think we can use it.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2468149": "```\nfrom datasets import load_metric\nwer_metric = load_metric(\"wer\")\n#wer=0.27\nlabel_str =  \"ব্যাংকার ওয়েবসাইট গিয়ে হামরা হামাদের ভার্চুয়াল বা ডিজিটাল কার্ড দ্যাখতে পারতেছি\"\n\npred_str =   \"ব্যাংকার ওয়েবসাইট গিয়ে হামরা হামাদের ভার্চুয়াল বা ডিজিটাল কার্ড দ্যাখতে পারতেছি\"\n\nwer = wer_metric.compute(predictions=[pred_str], references=[label_str])\nprint(\"wer: \", wer)\n\n```\nI am getting wer:  0.2727272727272727 \nAre the strings above equal? if they are why am I getting 0.27 for WER?",
    "2468389": "ওয়েবসাইট != ওয়েবসাইট",
    "2468470": "different way to spell \"website\" ?  I can't tell the difference :(",
    "2468526": "Hi, I just tried checking your case. So apparently word-by-word comparisons yield that \"য়ে\" != \"য়ে\". However, if you try a normalizer such as the bnunicodenormalizer, you can immediately see that the two strings are exactly the same after normalization.",
    "2468561": "Check in this way: \n\n```python\nimport difflib\n\ndef highlight_differences(ground_truth, prediction):\n    differ = difflib.Differ()\n    diff = list(differ.compare(ground_truth, prediction))\n\n    highlighted_diff = []\n\n    for item in diff:\n        if item.startswith('-'):  # For deleted fragments\n            highlighted_diff.append('\\033[91m\\033[43m' + item[-1] + '\\033[0m')  # Red text on yellow backgound\n        elif item.startswith('+'):  # For added framgents\n            highlighted_diff.append('\\033[94m\\033[47m' + item[-1] + '\\033[0m')  # Blue text on light gray background\n        else:\n            highlighted_diff.append(item[-1])\n\n    return ''.join(highlighted_diff)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2125668%2Fa738cd4984034fe688de1b0673c12e17%2Fobraz_2023-10-05_175354744.png?generation=1696521235534481&alt=media)",
    "2468884": "Hi. The problem is mainly with the character 'য়' . It can be written in two ways. \n> 'য়' -> 'য', '়' ( The 2nd one is called Nukta)\n> 'য়' -> 'য়'\n\nOf the two sentences you mentioned, one of them contains the first format and the other one has the 2nd one. That's why it's showing a high WER although there's no difference visibly.\n\nTo solve this exact issue, normalizer is used. The normalizer transforms all the sentences in the same canonical form.",
    "2469612": "Hi, in your [EDA notebook](https://www.kaggle.com/code/mbmmurad/detailed-eda-normalizer-and-wer), you used https://github.com/csebuetnlp/normalizer to normalize the sentence. May I ask what is the difference between this normalizer and the bnunicodenormalizer",
    "2471896": "I also noticed that the normalizer has Non Comercial Use license [Contents of this repository are restricted to non-commercial research purposes only under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).] If you are working towards prize money I don't think we can use it.",
    "2474720": "Hi @ly8388, does this indicate using bnunicodenormalizer while training can lead to performance dip? or is this only limited to naked eye confusion?",
    "2474735": "I think a normalizer affect the performance.\nYou also should be careful of what normalizer is used in the preprocess of the pretraining model.",
    "2482325": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055010%2F4da3f336c2bd69798b2c164e85795b8f%2FScreenshot%202023-10-14%20at%205.05.56PM.png?generation=1697321181363157&alt=media)"
  },
  "source": "meta"
}