{
  "id": 612159,
  "title": "where do you guys run the n-gram model and LLM?",
  "url": "/competitions/brain-to-text-25/discussion/612159",
  "author_name": "",
  "post_date": "2025-10-17T08:01:31.354127900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>the baseline model uses two NVIDIA RTX 4090 GPUs, with 512GB of RAM, the LLM requires 12GB of VRAM, but my home computer / kaggle / google colab does not have that much RAM/ VRAM storage, to run the models. Wondering where do y'all run te models and LLMs on? </p>",
  "messages": [
    {
      "id": "3303134",
      "postDate": "10/17/2025 08:01:31",
      "content": "<p>the baseline model uses two NVIDIA RTX 4090 GPUs, with 512GB of RAM, the LLM requires 12GB of VRAM, but my home computer / kaggle / google colab does not have that much RAM/ VRAM storage, to run the models. Wondering where do y'all run te models and LLMs on? </p>",
      "rawMarkdown": "the baseline model uses two NVIDIA RTX 4090 GPUs, with 512GB of RAM, the LLM requires 12GB of VRAM, but my home computer / kaggle / google colab does not have that much RAM/ VRAM storage, to run the models. Wondering where do y'all run te models and LLMs on?",
      "votes": null
    },
    {
      "id": "3303352",
      "postDate": "10/17/2025 18:34:12",
      "content": "<p>The resource requirements are admittedly a little crazy, yeah. We have a workstation computer that we run it on locally.</p>\n<p>Unfortunately, this WFST+LLM approach is still the best that we've found for converting our CTC phoneme logits to words. Other more light-weight solutions exist (like kenlm or torchaudio's ctc decoder), but AFAIK they have not matched the WFST accuracy. One could also use an end-to-end decoder or a transducer + LLM decoder to avoid having to use the WFST. Hopefully through this competition, someone comes up with a solution!</p>",
      "rawMarkdown": "The resource requirements are admittedly a little crazy, yeah. We have a workstation computer that we run it on locally.\n\nUnfortunately, this WFST+LLM approach is still the best that we've found for converting our CTC phoneme logits to words. Other more light-weight solutions exist (like kenlm or torchaudio's ctc decoder), but AFAIK they have not matched the WFST accuracy. One could also use an end-to-end decoder or a transducer + LLM decoder to avoid having to use the WFST. Hopefully through this competition, someone comes up with a solution!",
      "votes": null
    },
    {
      "id": "3303508",
      "postDate": "10/18/2025 07:42:42",
      "content": "<p>thanks for your reply Nick, I agree that reducing the hardware requirement would open this technology to more clinical / daily settings, but yea right now I just want to replicate the experiment results, i am sure some of the participants are able to replicate the results in a more hardware-friendly way, hope to hear from others~</p>",
      "rawMarkdown": "thanks for your reply Nick, I agree that reducing the hardware requirement would open this technology to more clinical / daily settings, but yea right now I just want to replicate the experiment results, i am sure some of the participants are able to replicate the results in a more hardware-friendly way, hope to hear from others~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3303352,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "10/17/2025 18:34:12",
      "content": "<p>The resource requirements are admittedly a little crazy, yeah. We have a workstation computer that we run it on locally.</p>\n<p>Unfortunately, this WFST+LLM approach is still the best that we've found for converting our CTC phoneme logits to words. Other more light-weight solutions exist (like kenlm or torchaudio's ctc decoder), but AFAIK they have not matched the WFST accuracy. One could also use an end-to-end decoder or a transducer + LLM decoder to avoid having to use the WFST. Hopefully through this competition, someone comes up with a solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3303508,
          "author_name": "attorneyevil",
          "author_url": "",
          "post_date": "10/18/2025 07:42:42",
          "content": "<p>thanks for your reply Nick, I agree that reducing the hardware requirement would open this technology to more clinical / daily settings, but yea right now I just want to replicate the experiment results, i am sure some of the participants are able to replicate the results in a more hardware-friendly way, hope to hear from others~</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3303134": "the baseline model uses two NVIDIA RTX 4090 GPUs, with 512GB of RAM, the LLM requires 12GB of VRAM, but my home computer / kaggle / google colab does not have that much RAM/ VRAM storage, to run the models. Wondering where do y'all run te models and LLMs on?",
    "3303352": "The resource requirements are admittedly a little crazy, yeah. We have a workstation computer that we run it on locally.\n\nUnfortunately, this WFST+LLM approach is still the best that we've found for converting our CTC phoneme logits to words. Other more light-weight solutions exist (like kenlm or torchaudio's ctc decoder), but AFAIK they have not matched the WFST accuracy. One could also use an end-to-end decoder or a transducer + LLM decoder to avoid having to use the WFST. Hopefully through this competition, someone comes up with a solution!",
    "3303508": "thanks for your reply Nick, I agree that reducing the hardware requirement would open this technology to more clinical / daily settings, but yea right now I just want to replicate the experiment results, i am sure some of the participants are able to replicate the results in a more hardware-friendly way, hope to hear from others~"
  },
  "source": "meta"
}