{
  "id": 442421,
  "title": "Question about baseline 5gram",
  "url": "/competitions/bengaliai-speech/discussion/442421",
  "author_name": "Roy Wei",
  "post_date": "2023-09-22T14:17:39.179000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I wonder how effective is 5gram that is used in the baseline. Thus, I use the following code to test it:</p>\n<p><code>import kenlm</code></p>\n<p><code>LM_PATH = INPUT / \"bengali-sr-download-public-trained-models/wav2vec2-xls-r-300m-bengali/language_model/\"</code></p>\n<p><code>language_model = kenlm.Model(str(LM_PATH / '5gram.bin'))</code><br>\n<code>sentene = \"\"</code><br>\n<code>score = language_model.score(sentence)</code><br>\n<code>print(score)</code></p>\n<p>Yet, whatever I send in the score is -30ish, which is presumably pretty low. I'm not sure if it's an issue with my implementation or that 5gram requires too much of data to be effective. </p>\n<p>In addition, do anyone know how to train ngram?</p>",
  "messages": [
    {
      "id": 2451521,
      "postDate": "2023-09-22T15:01:52.937Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro</a></p>\n<p>I have not tried it though (I don't know what corpus to use)</p>",
      "rawMarkdown": "https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\n\nI have not tried it though (I don't know what corpus to use)",
      "replies": [
        {
          "id": 2451967,
          "postDate": "2023-09-23T01:17:43.633Z",
          "content": "<p>I have seen that the ngram was based on 30M Bengali words or sentences. So I guess it's the problem with my implementation?</p>",
          "rawMarkdown": "I have seen that the ngram was based on 30M Bengali words or sentences. So I guess it's the problem with my implementation?"
        }
      ]
    },
    {
      "id": 2451445,
      "postDate": "2023-09-22T14:17:39.180Z",
      "content": "<p>I wonder how effective is 5gram that is used in the baseline. Thus, I use the following code to test it:</p>\n<p><code>import kenlm</code></p>\n<p><code>LM_PATH = INPUT / \"bengali-sr-download-public-trained-models/wav2vec2-xls-r-300m-bengali/language_model/\"</code></p>\n<p><code>language_model = kenlm.Model(str(LM_PATH / '5gram.bin'))</code><br>\n<code>sentene = \"\"</code><br>\n<code>score = language_model.score(sentence)</code><br>\n<code>print(score)</code></p>\n<p>Yet, whatever I send in the score is -30ish, which is presumably pretty low. I'm not sure if it's an issue with my implementation or that 5gram requires too much of data to be effective. </p>\n<p>In addition, do anyone know how to train ngram?</p>",
      "rawMarkdown": "I wonder how effective is 5gram that is used in the baseline. Thus, I use the following code to test it:\n\n`import kenlm`\n\n`LM_PATH = INPUT / \"bengali-sr-download-public-trained-models/wav2vec2-xls-r-300m-bengali/language_model/\"`\n\n`language_model = kenlm.Model(str(LM_PATH / '5gram.bin'))`\n`sentene = \"\"`\n`score = language_model.score(sentence)`\n`print(score)`\n\nYet, whatever I send in the score is -30ish, which is presumably pretty low. I'm not sure if it's an issue with my implementation or that 5gram requires too much of data to be effective. \n\nIn addition, do anyone know how to train ngram?"
    },
    {
      "id": 2451503,
      "postDate": "2023-09-22T14:57:24.533Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2451521,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-09-22T15:01:52.937000",
      "content": "<p><a href=\"https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\" target=\"_blank\">https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro</a></p>\n<p>I have not tried it though (I don't know what corpus to use)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2451967,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2023-09-23T01:17:43.633000",
          "content": "<p>I have seen that the ngram was based on 30M Bengali words or sentences. So I guess it's the problem with my implementation?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2451503,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-22T14:57:24.533000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2451521": "https://www.kaggle.com/code/umongsain/build-an-n-gram-with-kenlm-macro\n\nI have not tried it though (I don't know what corpus to use)",
    "2451445": "I wonder how effective is 5gram that is used in the baseline. Thus, I use the following code to test it:\n\n`import kenlm`\n\n`LM_PATH = INPUT / \"bengali-sr-download-public-trained-models/wav2vec2-xls-r-300m-bengali/language_model/\"`\n\n`language_model = kenlm.Model(str(LM_PATH / '5gram.bin'))`\n`sentene = \"\"`\n`score = language_model.score(sentence)`\n`print(score)`\n\nYet, whatever I send in the score is -30ish, which is presumably pretty low. I'm not sure if it's an issue with my implementation or that 5gram requires too much of data to be effective. \n\nIn addition, do anyone know how to train ngram?",
    "2451503": ""
  }
}