{
  "id": 638854,
  "title": "WFST training details",
  "url": "/competitions/brain-to-text-25/discussion/638854",
  "author_name": "Suhas Dara",
  "post_date": "2025-11-24T06:57:26.684000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello, I was wondering if there are more detailed descriptions of how the WFSTs were trained beyond what is in the paper, and if the 3-gram and 5-gram models were trained similarly. I am currently trying to reverse engineer from the T.fst and L.fst provided in the 3-gram zip, but I find it to be an inefficient way of going about it. The details about disambiguation tokens, pronunciation probabilities of alternate pronunciations, and sil vs eps tokens, etc., are still a little unclear to me after looking at the tools provided, such as make_lexicon_fst.pl</p>",
  "messages": [
    {
      "id": 3346187,
      "postDate": "2025-11-24T06:57:26.683Z",
      "content": "<p>Hello, I was wondering if there are more detailed descriptions of how the WFSTs were trained beyond what is in the paper, and if the 3-gram and 5-gram models were trained similarly. I am currently trying to reverse engineer from the T.fst and L.fst provided in the 3-gram zip, but I find it to be an inefficient way of going about it. The details about disambiguation tokens, pronunciation probabilities of alternate pronunciations, and sil vs eps tokens, etc., are still a little unclear to me after looking at the tools provided, such as make_lexicon_fst.pl</p>",
      "rawMarkdown": "Hello, I was wondering if there are more detailed descriptions of how the WFSTs were trained beyond what is in the paper, and if the 3-gram and 5-gram models were trained similarly. I am currently trying to reverse engineer from the T.fst and L.fst provided in the 3-gram zip, but I find it to be an inefficient way of going about it. The details about disambiguation tokens, pronunciation probabilities of alternate pronunciations, and sil vs eps tokens, etc., are still a little unclear to me after looking at the tools provided, such as make_lexicon_fst.pl"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3346187": "Hello, I was wondering if there are more detailed descriptions of how the WFSTs were trained beyond what is in the paper, and if the 3-gram and 5-gram models were trained similarly. I am currently trying to reverse engineer from the T.fst and L.fst provided in the 3-gram zip, but I find it to be an inefficient way of going about it. The details about disambiguation tokens, pronunciation probabilities of alternate pronunciations, and sil vs eps tokens, etc., are still a little unclear to me after looking at the tools provided, such as make_lexicon_fst.pl"
  }
}