{
  "id": 432305,
  "title": "How to deal with punctuation characters?",
  "url": "/competitions/bengaliai-speech/discussion/432305",
  "author_name": "",
  "post_date": "2023-08-16T21:10:44.623466200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I've made some statistics of characters and noticed that some punctuation symbols are significantly common. </p>\n<table>\n<thead>\n<tr>\n<th>symbol</th>\n<th>count</th>\n<th>comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>,</td>\n<td>282795</td>\n<td>this is very common, language model should help with this, just wonder if use of comma can be sometimes arbitrary in Bengali</td>\n</tr>\n<tr>\n<td>-</td>\n<td>75769</td>\n<td>connects two words,  e.g. চট্টগ্রাম-কক্সবাজার . To address this named entity recognition is required</td>\n</tr>\n<tr>\n<td>?</td>\n<td>63923</td>\n<td>quite common, usually at the end</td>\n</tr>\n<tr>\n<td>\"</td>\n<td>16329</td>\n<td>titles, etc. তাঁর প্রথম উপন্যাস \"উত্তম পুরুষ\". To address this named entity recognition required. Some cases are hard and would require language model, eg. 'তাঁর নামের অর্থ \"পূজা-উৎসর্গের দেবী\" বা \"সন্তুষ্ট নারী\"।' (Her name means \"goddess of worship\" or \"satisfied woman\")</td>\n</tr>\n<tr>\n<td>'</td>\n<td>4502</td>\n<td>same as above \" symbol, but is there any rule of choosing between \" and ' ?</td>\n</tr>\n<tr>\n<td>’</td>\n<td>2703</td>\n<td>this one and ‘ below are very often (but as statistics suggests, not always) together and encapsulate short words or characters</td>\n</tr>\n<tr>\n<td>–</td>\n<td>2609</td>\n<td>just an evil twin of a hyphen above…</td>\n</tr>\n<tr>\n<td>—</td>\n<td>2104</td>\n<td>dash, this seems like a hard task for prediction</td>\n</tr>\n<tr>\n<td>‘</td>\n<td>1598</td>\n<td></td>\n</tr>\n<tr>\n<td>:</td>\n<td>1137</td>\n<td>sometimes found in named entity, sometimes much harder to predict in a sentence</td>\n</tr>\n<tr>\n<td>;</td>\n<td>837</td>\n<td>semicolons are tough, luckily they are scarce</td>\n</tr>\n<tr>\n<td>.</td>\n<td>582</td>\n<td>occurs only with initials and acronyms, like J.R.R. Tolkien, M.A. (Master of Arts).</td>\n</tr>\n<tr>\n<td>/</td>\n<td>482</td>\n<td>not possible to distinguish between this and comma</td>\n</tr>\n<tr>\n<td>“</td>\n<td>104</td>\n<td>anything about 100 frequency we can safely ignore I think</td>\n</tr>\n<tr>\n<td>”</td>\n<td>90</td>\n<td></td>\n</tr>\n<tr>\n<td>৷</td>\n<td>12</td>\n<td></td>\n</tr>\n<tr>\n<td>…</td>\n<td>9</td>\n<td></td>\n</tr>\n<tr>\n<td>‚</td>\n<td>4</td>\n<td></td>\n</tr>\n<tr>\n<td>॥</td>\n<td>4</td>\n<td></td>\n</tr>\n<tr>\n<td>table format helperr</td>\n<td>table format helperrrrrrrrrrrrr</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>Are they important when it comes to evaluation? Or we can just remove them? I think task is hard enough without them, but that's just my opinion :)</p>",
  "messages": [
    {
      "id": "2394352",
      "postDate": "08/16/2023 21:10:44",
      "content": "<p>I've made some statistics of characters and noticed that some punctuation symbols are significantly common. </p>\n<table>\n<thead>\n<tr>\n<th>symbol</th>\n<th>count</th>\n<th>comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>,</td>\n<td>282795</td>\n<td>this is very common, language model should help with this, just wonder if use of comma can be sometimes arbitrary in Bengali</td>\n</tr>\n<tr>\n<td>-</td>\n<td>75769</td>\n<td>connects two words,  e.g. চট্টগ্রাম-কক্সবাজার . To address this named entity recognition is required</td>\n</tr>\n<tr>\n<td>?</td>\n<td>63923</td>\n<td>quite common, usually at the end</td>\n</tr>\n<tr>\n<td>\"</td>\n<td>16329</td>\n<td>titles, etc. তাঁর প্রথম উপন্যাস \"উত্তম পুরুষ\". To address this named entity recognition required. Some cases are hard and would require language model, eg. 'তাঁর নামের অর্থ \"পূজা-উৎসর্গের দেবী\" বা \"সন্তুষ্ট নারী\"।' (Her name means \"goddess of worship\" or \"satisfied woman\")</td>\n</tr>\n<tr>\n<td>'</td>\n<td>4502</td>\n<td>same as above \" symbol, but is there any rule of choosing between \" and ' ?</td>\n</tr>\n<tr>\n<td>’</td>\n<td>2703</td>\n<td>this one and ‘ below are very often (but as statistics suggests, not always) together and encapsulate short words or characters</td>\n</tr>\n<tr>\n<td>–</td>\n<td>2609</td>\n<td>just an evil twin of a hyphen above…</td>\n</tr>\n<tr>\n<td>—</td>\n<td>2104</td>\n<td>dash, this seems like a hard task for prediction</td>\n</tr>\n<tr>\n<td>‘</td>\n<td>1598</td>\n<td></td>\n</tr>\n<tr>\n<td>:</td>\n<td>1137</td>\n<td>sometimes found in named entity, sometimes much harder to predict in a sentence</td>\n</tr>\n<tr>\n<td>;</td>\n<td>837</td>\n<td>semicolons are tough, luckily they are scarce</td>\n</tr>\n<tr>\n<td>.</td>\n<td>582</td>\n<td>occurs only with initials and acronyms, like J.R.R. Tolkien, M.A. (Master of Arts).</td>\n</tr>\n<tr>\n<td>/</td>\n<td>482</td>\n<td>not possible to distinguish between this and comma</td>\n</tr>\n<tr>\n<td>“</td>\n<td>104</td>\n<td>anything about 100 frequency we can safely ignore I think</td>\n</tr>\n<tr>\n<td>”</td>\n<td>90</td>\n<td></td>\n</tr>\n<tr>\n<td>৷</td>\n<td>12</td>\n<td></td>\n</tr>\n<tr>\n<td>…</td>\n<td>9</td>\n<td></td>\n</tr>\n<tr>\n<td>‚</td>\n<td>4</td>\n<td></td>\n</tr>\n<tr>\n<td>॥</td>\n<td>4</td>\n<td></td>\n</tr>\n<tr>\n<td>table format helperr</td>\n<td>table format helperrrrrrrrrrrrr</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>Are they important when it comes to evaluation? Or we can just remove them? I think task is hard enough without them, but that's just my opinion :)</p>",
      "rawMarkdown": "I've made some statistics of characters and noticed that some punctuation symbols are significantly common. \n\n| symbol |   count   | comment |\n| :-------------: | :-------------: | ----- |\n| , | 282795 | this is very common, language model should help with this, just wonder if use of comma can be sometimes arbitrary in Bengali |\n| - | 75769 | connects two words,  e.g. চট্টগ্রাম-কক্সবাজার . To address this named entity recognition is required |\n| ? | 63923| quite common, usually at the end |\n| \" | 16329 | titles, etc. তাঁর প্রথম উপন্যাস \"উত্তম পুরুষ\". To address this named entity recognition required. Some cases are hard and would require language model, eg. 'তাঁর নামের অর্থ \"পূজা-উৎসর্গের দেবী\" বা \"সন্তুষ্ট নারী\"।' (Her name means \"goddess of worship\" or \"satisfied woman\") |\n|  ' | 4502 | same as above \" symbol, but is there any rule of choosing between \" and ' ? |\n| ’ | 2703 | this one and ‘ below are very often (but as statistics suggests, not always) together and encapsulate short words or characters |\n| – | 2609 | just an evil twin of a hyphen above... |\n| — | 2104 | dash, this seems like a hard task for prediction |\n| ‘ | 1598 |  |\n| : | 1137 | sometimes found in named entity, sometimes much harder to predict in a sentence |\n| ; | 837 | semicolons are tough, luckily they are scarce |\n| . | 582 | occurs only with initials and acronyms, like J.R.R. Tolkien, M.A. (Master of Arts). |\n| / | 482 | not possible to distinguish between this and comma |\n| “ | 104 | anything about 100 frequency we can safely ignore I think |\n| ” | 90 | |\n| ৷ | 12 | |\n| … | 9 | |\n| ‚ |  4 | |\n| ॥ |  4 | |\n| table format helperr | table format helperrrrrrrrrrrrr | |\n\n\nAre they important when it comes to evaluation? Or we can just remove them? I think task is hard enough without them, but that's just my opinion :)",
      "votes": null
    },
    {
      "id": "2400110",
      "postDate": "08/20/2023 19:21:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/Daniel\" target=\"_blank\">@Daniel</a>,</p>\n<p>Just use bnunicodenormalizer <a href=\"https://pypi.org/project/bnunicodenormalizer/\" target=\"_blank\">https://pypi.org/project/bnunicodenormalizer/</a>. This will get rid of conflicts and reduce the full list of allowable punctuations.</p>\n<p>Punctuations are important and cant be completely discarded as it will be equivalent to introducing label noise :( </p>",
      "rawMarkdown": "Hi @Daniel,\n\nJust use bnunicodenormalizer https://pypi.org/project/bnunicodenormalizer/. This will get rid of conflicts and reduce the full list of allowable punctuations.\n\nPunctuations are important and cant be completely discarded as it will be equivalent to introducing label noise :(",
      "votes": null
    },
    {
      "id": "2403226",
      "postDate": "08/22/2023 15:09:21",
      "content": "<p><a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a>    This one may help. </p>",
      "rawMarkdown": "https://github.com/xashru/punctuation-restoration    This one may help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2400110,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "08/20/2023 19:21:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/Daniel\" target=\"_blank\">@Daniel</a>,</p>\n<p>Just use bnunicodenormalizer <a href=\"https://pypi.org/project/bnunicodenormalizer/\" target=\"_blank\">https://pypi.org/project/bnunicodenormalizer/</a>. This will get rid of conflicts and reduce the full list of allowable punctuations.</p>\n<p>Punctuations are important and cant be completely discarded as it will be equivalent to introducing label noise :( </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2403226,
      "author_name": "berserker408",
      "author_url": "",
      "post_date": "08/22/2023 15:09:21",
      "content": "<p><a href=\"https://github.com/xashru/punctuation-restoration\" target=\"_blank\">https://github.com/xashru/punctuation-restoration</a>    This one may help. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2394352": "I've made some statistics of characters and noticed that some punctuation symbols are significantly common. \n\n| symbol |   count   | comment |\n| :-------------: | :-------------: | ----- |\n| , | 282795 | this is very common, language model should help with this, just wonder if use of comma can be sometimes arbitrary in Bengali |\n| - | 75769 | connects two words,  e.g. চট্টগ্রাম-কক্সবাজার . To address this named entity recognition is required |\n| ? | 63923| quite common, usually at the end |\n| \" | 16329 | titles, etc. তাঁর প্রথম উপন্যাস \"উত্তম পুরুষ\". To address this named entity recognition required. Some cases are hard and would require language model, eg. 'তাঁর নামের অর্থ \"পূজা-উৎসর্গের দেবী\" বা \"সন্তুষ্ট নারী\"।' (Her name means \"goddess of worship\" or \"satisfied woman\") |\n|  ' | 4502 | same as above \" symbol, but is there any rule of choosing between \" and ' ? |\n| ’ | 2703 | this one and ‘ below are very often (but as statistics suggests, not always) together and encapsulate short words or characters |\n| – | 2609 | just an evil twin of a hyphen above... |\n| — | 2104 | dash, this seems like a hard task for prediction |\n| ‘ | 1598 |  |\n| : | 1137 | sometimes found in named entity, sometimes much harder to predict in a sentence |\n| ; | 837 | semicolons are tough, luckily they are scarce |\n| . | 582 | occurs only with initials and acronyms, like J.R.R. Tolkien, M.A. (Master of Arts). |\n| / | 482 | not possible to distinguish between this and comma |\n| “ | 104 | anything about 100 frequency we can safely ignore I think |\n| ” | 90 | |\n| ৷ | 12 | |\n| … | 9 | |\n| ‚ |  4 | |\n| ॥ |  4 | |\n| table format helperr | table format helperrrrrrrrrrrrr | |\n\n\nAre they important when it comes to evaluation? Or we can just remove them? I think task is hard enough without them, but that's just my opinion :)",
    "2400110": "Hi @Daniel,\n\nJust use bnunicodenormalizer https://pypi.org/project/bnunicodenormalizer/. This will get rid of conflicts and reduce the full list of allowable punctuations.\n\nPunctuations are important and cant be completely discarded as it will be equivalent to introducing label noise :(",
    "2403226": "https://github.com/xashru/punctuation-restoration    This one may help."
  },
  "source": "meta"
}