{
  "id": 614125,
  "title": "Question about the language model, especially the nGram module.",
  "url": "/competitions/brain-to-text-25/discussion/614125",
  "author_name": "",
  "post_date": "2025-11-01T09:44:14.367770Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Can you briefly explain how the language model was trained? Was the n-gram model trained on ground truth phoneme labels or on the RNN outputs?</p>",
  "messages": [
    {
      "id": "3309743",
      "postDate": "11/01/2025 09:44:14",
      "content": "<p>Can you briefly explain how the language model was trained? Was the n-gram model trained on ground truth phoneme labels or on the RNN outputs?</p>",
      "rawMarkdown": "Can you briefly explain how the language model was trained? Was the n-gram model trained on ground truth phoneme labels or on the RNN outputs?",
      "votes": null
    },
    {
      "id": "3309911",
      "postDate": "11/01/2025 17:43:48",
      "content": "<p>I think the ngram model is pretrained on a very large english-sentence dataset ( this is why the baseline 5-gram model takes up ~300 GB of memory), to get a accurate representation of word-combination probabilities. Not sure the english-sentence dataset is specifically designed to only include conversational telephone script ( switchboard) , or just for general english sentences. I think you can fine-tune the ngram model by including more switchboard sentences to the ngram model training, as we can see that the train/val/test data is dominated by switchboard sentences. </p>\n<p>fyi: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14683498%2Fe7c51ba9470e4c19e584b5af7731a2dc%2FScreenshot%202025-11-02%20014321.png?generation=1762019012760695&amp;alt=media\" alt=\"distribution of sentence corpus\"></p>",
      "rawMarkdown": "I think the ngram model is pretrained on a very large english-sentence dataset ( this is why the baseline 5-gram model takes up ~300 GB of memory), to get a accurate representation of word-combination probabilities. Not sure the english-sentence dataset is specifically designed to only include conversational telephone script ( switchboard) , or just for general english sentences. I think you can fine-tune the ngram model by including more switchboard sentences to the ngram model training, as we can see that the train/val/test data is dominated by switchboard sentences. \n\nfyi: ![distribution of sentence corpus](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14683498%2Fe7c51ba9470e4c19e584b5af7731a2dc%2FScreenshot%202025-11-02%20014321.png?generation=1762019012760695&alt=media)",
      "votes": null
    },
    {
      "id": "3310613",
      "postDate": "11/03/2025 09:39:50",
      "content": "<p>Ok, i thought the nGram model is responsible for the Phoneme-to-Word decoding but its just a model for rescoring the word predictions, right? What is responsible for the Phoneme-to-Word decoding? I am asking because we got better Phoneme prediction results but the Word prediction got worse than Baseline, which doesnt makes much sense.</p>",
      "rawMarkdown": "Ok, i thought the nGram model is responsible for the Phoneme-to-Word decoding but its just a model for rescoring the word predictions, right? What is responsible for the Phoneme-to-Word decoding? I am asking because we got better Phoneme prediction results but the Word prediction got worse than Baseline, which doesnt makes much sense.",
      "votes": null
    },
    {
      "id": "3310847",
      "postDate": "11/03/2025 18:42:40",
      "content": "<p><strong>brain-data to Phoneme prediction</strong> and <strong>Phoneme to word prediction</strong> are separated in the baseline architecture,\nit is definitely good news if you're getting netter results in brain-data to Phoneme prediction, but Phoneme to word prediction are done by rescoring the possible phoneme combinations (which forms words --&gt; sentence). </p>\n<p>The rescoring mechanism for each sentence are controlled by 3 parts, </p>\n<ol>\n<li>accoustic score ( how likely does your phoneme prediction model predicts the current phonemes in the sentence)</li>\n<li>ngram model score( how likely does this set of words appears in this order, \ne.g. \"I love you\", score = P(\"I\" | ,\"start token\") + P ( \"love\" |  \" I\" ) + P(\"you\" | \"  I love\") + \nP(\" | \"  I love you\")</li>\n<li>LLM rescoring (how likely does this sentence appears in the LLM)</li>\n</ol>",
      "rawMarkdown": "**brain-data to Phoneme prediction** and **Phoneme to word prediction** are separated in the baseline architecture,\nit is definitely good news if you're getting netter results in brain-data to Phoneme prediction, but Phoneme to word prediction are done by rescoring the possible phoneme combinations (which forms words --> sentence). \n\nThe rescoring mechanism for each sentence are controlled by 3 parts, \n1. accoustic score ( how likely does your phoneme prediction model predicts the current phonemes in the sentence)\n2.  ngram model score( how likely does this set of words appears in this order, \ne.g. \"I love you\", score = P(\"I\" | ,\"start token\") + P ( \"love\" |  \"<start token> I\" ) + P(\"you\" | \" <start token> I love\") + \nP(\"<end token> | \" <start token> I love you\")\n3. LLM rescoring (how likely does this sentence appears in the LLM)",
      "votes": null
    },
    {
      "id": "3311114",
      "postDate": "11/04/2025 09:43:46",
      "content": "<p>If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?. Otherwise I don’t see why a stronger phoneme predictor would yield worse performance once the LM is applied.</p>",
      "rawMarkdown": "If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?. Otherwise I don’t see why a stronger phoneme predictor would yield worse performance once the LM is applied.",
      "votes": null
    },
    {
      "id": "3311619",
      "postDate": "11/05/2025 12:18:52",
      "content": "<p>did you set the seed so that your performance is reproducible? because your observation may be due to random initialization of the model and training process</p>",
      "rawMarkdown": "did you set the seed so that your performance is reproducible? because your observation may be due to random initialization of the model and training process",
      "votes": null
    },
    {
      "id": "3311663",
      "postDate": "11/05/2025 14:14:42",
      "content": "<p><a href=\"url\" target=\"_blank\">https://arxiv.org/abs/1507.08240</a>  This paper has a clear explanation.</p>",
      "rawMarkdown": "[https://arxiv.org/abs/1507.08240](url)  This paper has a clear explanation.",
      "votes": null
    },
    {
      "id": "3311673",
      "postDate": "11/05/2025 14:44:41",
      "content": "<p>I think youre getting it wrong here. We improved the Phoneme model into producing Phonemes with lower PER. After that we pluged our new models output into the language Model and got worse WER results like i mentioned. This cant be the fault of the seed.</p>",
      "rawMarkdown": "I think youre getting it wrong here. We improved the Phoneme model into producing Phonemes with lower PER. After that we pluged our new models output into the language Model and got worse WER results like i mentioned. This cant be the fault of the seed.",
      "votes": null
    },
    {
      "id": "3311675",
      "postDate": "11/05/2025 14:49:38",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "3311771",
      "postDate": "11/05/2025 18:29:42",
      "content": "<p>I think your earlier intuition is probably correct: </p>\n<blockquote>\n  <p>If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?</p>\n</blockquote>\n<p>I'd suggest you try (1) visualizing your phoneme logits (especially for incorrect predictions) and (2) playing with the language model parameters described above by <a href=\"https://www.kaggle.com/heyyousum\" target=\"_blank\">@heyyousum</a> to see if you can get a better WER with different settings.</p>",
      "rawMarkdown": "I think your earlier intuition is probably correct: \n>If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?\n\nI'd suggest you try (1) visualizing your phoneme logits (especially for incorrect predictions) and (2) playing with the language model parameters described above by @heyyousum to see if you can get a better WER with different settings.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3309911,
      "author_name": "heyyousum",
      "author_url": "",
      "post_date": "11/01/2025 17:43:48",
      "content": "<p>I think the ngram model is pretrained on a very large english-sentence dataset ( this is why the baseline 5-gram model takes up ~300 GB of memory), to get a accurate representation of word-combination probabilities. Not sure the english-sentence dataset is specifically designed to only include conversational telephone script ( switchboard) , or just for general english sentences. I think you can fine-tune the ngram model by including more switchboard sentences to the ngram model training, as we can see that the train/val/test data is dominated by switchboard sentences. </p>\n<p>fyi: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14683498%2Fe7c51ba9470e4c19e584b5af7731a2dc%2FScreenshot%202025-11-02%20014321.png?generation=1762019012760695&amp;alt=media\" alt=\"distribution of sentence corpus\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3310613,
          "author_name": "apollos1301",
          "author_url": "",
          "post_date": "11/03/2025 09:39:50",
          "content": "<p>Ok, i thought the nGram model is responsible for the Phoneme-to-Word decoding but its just a model for rescoring the word predictions, right? What is responsible for the Phoneme-to-Word decoding? I am asking because we got better Phoneme prediction results but the Word prediction got worse than Baseline, which doesnt makes much sense.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3310847,
              "author_name": "heyyousum",
              "author_url": "",
              "post_date": "11/03/2025 18:42:40",
              "content": "<p><strong>brain-data to Phoneme prediction</strong> and <strong>Phoneme to word prediction</strong> are separated in the baseline architecture,\nit is definitely good news if you're getting netter results in brain-data to Phoneme prediction, but Phoneme to word prediction are done by rescoring the possible phoneme combinations (which forms words --&gt; sentence). </p>\n<p>The rescoring mechanism for each sentence are controlled by 3 parts, </p>\n<ol>\n<li>accoustic score ( how likely does your phoneme prediction model predicts the current phonemes in the sentence)</li>\n<li>ngram model score( how likely does this set of words appears in this order, \ne.g. \"I love you\", score = P(\"I\" | ,\"start token\") + P ( \"love\" |  \" I\" ) + P(\"you\" | \"  I love\") + \nP(\" | \"  I love you\")</li>\n<li>LLM rescoring (how likely does this sentence appears in the LLM)</li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 3311114,
                  "author_name": "apollos1301",
                  "author_url": "",
                  "post_date": "11/04/2025 09:43:46",
                  "content": "<p>If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?. Otherwise I don’t see why a stronger phoneme predictor would yield worse performance once the LM is applied.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3311619,
                      "author_name": "heyyousum",
                      "author_url": "",
                      "post_date": "11/05/2025 12:18:52",
                      "content": "<p>did you set the seed so that your performance is reproducible? because your observation may be due to random initialization of the model and training process</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3311673,
                          "author_name": "apollos1301",
                          "author_url": "",
                          "post_date": "11/05/2025 14:44:41",
                          "content": "<p>I think youre getting it wrong here. We improved the Phoneme model into producing Phonemes with lower PER. After that we pluged our new models output into the language Model and got worse WER results like i mentioned. This cant be the fault of the seed.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3311771,
                              "author_name": "notnickc",
                              "author_url": "",
                              "post_date": "11/05/2025 18:29:42",
                              "content": "<p>I think your earlier intuition is probably correct: </p>\n<blockquote>\n  <p>If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?</p>\n</blockquote>\n<p>I'd suggest you try (1) visualizing your phoneme logits (especially for incorrect predictions) and (2) playing with the language model parameters described above by <a href=\"https://www.kaggle.com/heyyousum\" target=\"_blank\">@heyyousum</a> to see if you can get a better WER with different settings.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3311663,
      "author_name": "pandadapan",
      "author_url": "",
      "post_date": "11/05/2025 14:14:42",
      "content": "<p><a href=\"url\" target=\"_blank\">https://arxiv.org/abs/1507.08240</a>  This paper has a clear explanation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3311675,
          "author_name": "apollos1301",
          "author_url": "",
          "post_date": "11/05/2025 14:49:38",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3309743": "Can you briefly explain how the language model was trained? Was the n-gram model trained on ground truth phoneme labels or on the RNN outputs?",
    "3309911": "I think the ngram model is pretrained on a very large english-sentence dataset ( this is why the baseline 5-gram model takes up ~300 GB of memory), to get a accurate representation of word-combination probabilities. Not sure the english-sentence dataset is specifically designed to only include conversational telephone script ( switchboard) , or just for general english sentences. I think you can fine-tune the ngram model by including more switchboard sentences to the ngram model training, as we can see that the train/val/test data is dominated by switchboard sentences. \n\nfyi: ![distribution of sentence corpus](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14683498%2Fe7c51ba9470e4c19e584b5af7731a2dc%2FScreenshot%202025-11-02%20014321.png?generation=1762019012760695&alt=media)",
    "3310613": "Ok, i thought the nGram model is responsible for the Phoneme-to-Word decoding but its just a model for rescoring the word predictions, right? What is responsible for the Phoneme-to-Word decoding? I am asking because we got better Phoneme prediction results but the Word prediction got worse than Baseline, which doesnt makes much sense.",
    "3310847": "**brain-data to Phoneme prediction** and **Phoneme to word prediction** are separated in the baseline architecture,\nit is definitely good news if you're getting netter results in brain-data to Phoneme prediction, but Phoneme to word prediction are done by rescoring the possible phoneme combinations (which forms words --> sentence). \n\nThe rescoring mechanism for each sentence are controlled by 3 parts, \n1. accoustic score ( how likely does your phoneme prediction model predicts the current phonemes in the sentence)\n2.  ngram model score( how likely does this set of words appears in this order, \ne.g. \"I love you\", score = P(\"I\" | ,\"start token\") + P ( \"love\" |  \"<start token> I\" ) + P(\"you\" | \" <start token> I love\") + \nP(\"<end token> | \" <start token> I love you\")\n3. LLM rescoring (how likely does this sentence appears in the LLM)",
    "3311114": "If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?. Otherwise I don’t see why a stronger phoneme predictor would yield worse performance once the LM is applied.",
    "3311619": "did you set the seed so that your performance is reproducible? because your observation may be due to random initialization of the model and training process",
    "3311663": "[https://arxiv.org/abs/1507.08240](url)  This paper has a clear explanation.",
    "3311673": "I think youre getting it wrong here. We improved the Phoneme model into producing Phonemes with lower PER. After that we pluged our new models output into the language Model and got worse WER results like i mentioned. This cant be the fault of the seed.",
    "3311675": "Thank you!",
    "3311771": "I think your earlier intuition is probably correct: \n>If the phoneme model’s output distribution is too peaked, the language model won’t be able to correct errors, it has too little probability mass on alternative phonemes to work with, right?\n\nI'd suggest you try (1) visualizing your phoneme logits (especially for incorrect predictions) and (2) playing with the language model parameters described above by @heyyousum to see if you can get a better WER with different settings."
  },
  "source": "meta"
}