{
  "id": 434310,
  "title": "Normalized Texts and 5-gram LM for Wav2Vec2",
  "url": "/competitions/bengaliai-speech/discussion/434310",
  "author_name": "",
  "post_date": "2023-08-24T18:29:15.714686400Z",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>I have recently normalized and fixed some problematic texts using <a href=\"https://github.com/mnansary/bnUnicodeNormalizer\" target=\"_blank\">bnunicodenormalizer</a> as well as the normalizer developed by <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">CSE BUET NLP</a>. Additionally, I have created a 5-gram language model using KenLM after removing punctuation from the normalized sentences. The following links lead to public notebooks that you are welcome to explore:</p>\n<ul>\n<li><strong>Normalized MaCro</strong>: <a href=\"https://www.kaggle.com/umongsain/macro-normalization\" target=\"_blank\">MaCro | Normalization</a></li>\n<li><strong>n-gram with KenLM</strong>: <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">Build an n-gram with KenLM | MaCro</a></li>\n</ul>\n<p>I am open to any feedback you might have. Cheers! 🍻</p>",
  "messages": [
    {
      "id": "2407000",
      "postDate": "08/24/2023 18:29:15",
      "content": "<p>Hello everyone!</p>\n<p>I have recently normalized and fixed some problematic texts using <a href=\"https://github.com/mnansary/bnUnicodeNormalizer\" target=\"_blank\">bnunicodenormalizer</a> as well as the normalizer developed by <a href=\"https://github.com/csebuetnlp/normalizer\" target=\"_blank\">CSE BUET NLP</a>. Additionally, I have created a 5-gram language model using KenLM after removing punctuation from the normalized sentences. The following links lead to public notebooks that you are welcome to explore:</p>\n<ul>\n<li><strong>Normalized MaCro</strong>: <a href=\"https://www.kaggle.com/umongsain/macro-normalization\" target=\"_blank\">MaCro | Normalization</a></li>\n<li><strong>n-gram with KenLM</strong>: <a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">Build an n-gram with KenLM | MaCro</a></li>\n</ul>\n<p>I am open to any feedback you might have. Cheers! 🍻</p>",
      "rawMarkdown": "Hello everyone!\n\nI have recently normalized and fixed some problematic texts using [bnunicodenormalizer](https://github.com/mnansary/bnUnicodeNormalizer) as well as the normalizer developed by [CSE BUET NLP](https://github.com/csebuetnlp/normalizer). Additionally, I have created a 5-gram language model using KenLM after removing punctuation from the normalized sentences. The following links lead to public notebooks that you are welcome to explore:\n\n- **Normalized MaCro**: [MaCro | Normalization](https://www.kaggle.com/umongsain/macro-normalization)\n- **n-gram with KenLM**: [Build an n-gram with KenLM | MaCro](https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro)\n\nI am open to any feedback you might have. Cheers! 🍻",
      "votes": null
    },
    {
      "id": "2407482",
      "postDate": "08/25/2023 05:32:06",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a>, thanks for sharing your work.</p>",
      "rawMarkdown": "Great work @umongsain, thanks for sharing your work.",
      "votes": null
    },
    {
      "id": "2422798",
      "postDate": "09/04/2023 08:54:28",
      "content": "<p>You could turn/limit the output of the cell they are too long </p>",
      "rawMarkdown": "You could turn/limit the output of the cell they are too long",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2407482,
      "author_name": "faysalmiah1721758",
      "author_url": "",
      "post_date": "08/25/2023 05:32:06",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a>, thanks for sharing your work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2422798,
      "author_name": "andrewsmith1993",
      "author_url": "",
      "post_date": "09/04/2023 08:54:28",
      "content": "<p>You could turn/limit the output of the cell they are too long </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407000": "Hello everyone!\n\nI have recently normalized and fixed some problematic texts using [bnunicodenormalizer](https://github.com/mnansary/bnUnicodeNormalizer) as well as the normalizer developed by [CSE BUET NLP](https://github.com/csebuetnlp/normalizer). Additionally, I have created a 5-gram language model using KenLM after removing punctuation from the normalized sentences. The following links lead to public notebooks that you are welcome to explore:\n\n- **Normalized MaCro**: [MaCro | Normalization](https://www.kaggle.com/umongsain/macro-normalization)\n- **n-gram with KenLM**: [Build an n-gram with KenLM | MaCro](https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro)\n\nI am open to any feedback you might have. Cheers! 🍻",
    "2407482": "Great work @umongsain, thanks for sharing your work.",
    "2422798": "You could turn/limit the output of the cell they are too long"
  },
  "source": "meta"
}