{
  "id": 665412,
  "title": "[7th place solution] Mamba + GRU + KenLM (with code)",
  "url": "/competitions/brain-to-text-25/discussion/665412",
  "author_name": "Summer",
  "post_date": "2026-01-01T08:57:07.322000",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>[7th Place Solution] Mamba + GRU + KenLM (with code)</h1>\n<p>First off, thanks to Nicholas Card, the UC Davis Neuroprosthetics Lab, and the organizers for this fascinating competition. It has always been my dream to explore how ML can apply to clinical settings. Huge thanks to my teammate <a href=\"https://www.kaggle.com/Kyle\" target=\"_blank\">@Kyle</a> Hui.</p>\n<h3>🔗 Resources</h3>\n<ul>\n<li><strong>📖 Detailed Writeup (Medium):</strong> <a href=\"https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d\" target=\"_blank\">Click Here</a></li>\n<li><strong>💻 Full Code (GitHub):</strong> <a href=\"https://github.com/greentree327/brain-to-text-mamba-decoder\" target=\"_blank\">Click Here</a></li>\n</ul>\n<p><em>If you find this writeup or the code useful, please support us by <strong>starring the GitHub repo</strong> and leaving a <strong>clap/comment on Medium</strong>!</em></p>\n<h2>TL;DR</h2>\n<p>Our solution is a hybrid ensemble of <strong>SoftWindow Bi-Mamba</strong> and <strong>GRU</strong> models. We focused heavily on memory efficiency (fitting a custom KenLM into 19GB RAM vs 300GB baseline) and a <strong>Context-Aware Inference Pipeline (LISA)</strong> that gates LLM rescoring based on signal confidence.</p>\n<h2>1. The Challenge</h2>\n<p>We faced severe overfitting (0.0 training loss vs high validation loss) and strict compute limits (Single A100 on Colab). We needed a model light on RAM but heavy on regularization.</p>\n<h2>2. The Solution: Hybrid Architecture</h2>\n<p>We found <strong>Mamba</strong> captured long-range semantic context while <strong>GRU</strong> stabilized short-range acoustic modeling. Their error patterns were highly uncorrelated.</p>\n<ul>\n<li><strong>Bi-directional Mamba:</strong> 3-layer input compression with heavy dropout.</li>\n<li><strong>Soft Sliding Window:</strong> Biased SSM parameters to prioritize short-term memory.</li>\n<li><strong>Temporal Consistency Loss:</strong> A custom loss minimizing the difference between adjacent day matrices to handle neural drift.</li>\n</ul>\n<h2>3. Memory Optimization: Custom KenLM</h2>\n<p>Standard 5-gram models (~300GB) were impossible. We compiled a custom <strong>flashlight decoder</strong> with <code>kTrieMaxLabel=14</code>, fitting a robust 4-gram model into just <strong>19GB RAM</strong>.</p>\n<h2>4. Inference Pipeline (LISA)</h2>\n<p>We didn't rely on a single model. Our <strong>LISA (LLM-Integrated Scoring Aggregation)</strong> pipeline routes samples dynamically:</p>\n<ol>\n<li><strong>Logit Averaging:</strong> We averaged logits within model families to create constructive interference (signal amplification).</li>\n<li><strong>Gating:</strong> We calculated a normalized n-gram score for every sentence.<ul>\n<li><strong>High Confidence (≥ -3.76):</strong> Routed to <strong>Mistral-7B</strong> for semantic rescoring.</li>\n<li><strong>Low Confidence (&lt; -3.76):</strong> Bypassed the LLM to avoid hallucinations on random sequences.</li></ul></li>\n<li><strong>TTA:</strong> Batched Test-Time Adaptation on high-confidence predictions.</li>\n</ol>\n<h2>5. Results</h2>\n<p>This approach allowed us to decode both coherent speech and random patterns with high fidelity, achieving 7th place despite hardware constraints.</p>\n<hr>\n<p><em>Again, if you enjoyed this summary, please check out the full detailed breakdown on <a href=\"https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d\" target=\"_blank\">Medium</a> and star the <a href=\"https://github.com/greentree327/brain-to-text-mamba-decoder\" target=\"_blank\">GitHub repo</a>.</em></p>",
  "messages": [
    {
      "id": 3384442,
      "postDate": "2026-01-01T08:57:07.323Z",
      "content": "<h1>[7th Place Solution] Mamba + GRU + KenLM (with code)</h1>\n<p>First off, thanks to Nicholas Card, the UC Davis Neuroprosthetics Lab, and the organizers for this fascinating competition. It has always been my dream to explore how ML can apply to clinical settings. Huge thanks to my teammate <a href=\"https://www.kaggle.com/Kyle\" target=\"_blank\">@Kyle</a> Hui.</p>\n<h3>🔗 Resources</h3>\n<ul>\n<li><strong>📖 Detailed Writeup (Medium):</strong> <a href=\"https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d\" target=\"_blank\">Click Here</a></li>\n<li><strong>💻 Full Code (GitHub):</strong> <a href=\"https://github.com/greentree327/brain-to-text-mamba-decoder\" target=\"_blank\">Click Here</a></li>\n</ul>\n<p><em>If you find this writeup or the code useful, please support us by <strong>starring the GitHub repo</strong> and leaving a <strong>clap/comment on Medium</strong>!</em></p>\n<h2>TL;DR</h2>\n<p>Our solution is a hybrid ensemble of <strong>SoftWindow Bi-Mamba</strong> and <strong>GRU</strong> models. We focused heavily on memory efficiency (fitting a custom KenLM into 19GB RAM vs 300GB baseline) and a <strong>Context-Aware Inference Pipeline (LISA)</strong> that gates LLM rescoring based on signal confidence.</p>\n<h2>1. The Challenge</h2>\n<p>We faced severe overfitting (0.0 training loss vs high validation loss) and strict compute limits (Single A100 on Colab). We needed a model light on RAM but heavy on regularization.</p>\n<h2>2. The Solution: Hybrid Architecture</h2>\n<p>We found <strong>Mamba</strong> captured long-range semantic context while <strong>GRU</strong> stabilized short-range acoustic modeling. Their error patterns were highly uncorrelated.</p>\n<ul>\n<li><strong>Bi-directional Mamba:</strong> 3-layer input compression with heavy dropout.</li>\n<li><strong>Soft Sliding Window:</strong> Biased SSM parameters to prioritize short-term memory.</li>\n<li><strong>Temporal Consistency Loss:</strong> A custom loss minimizing the difference between adjacent day matrices to handle neural drift.</li>\n</ul>\n<h2>3. Memory Optimization: Custom KenLM</h2>\n<p>Standard 5-gram models (~300GB) were impossible. We compiled a custom <strong>flashlight decoder</strong> with <code>kTrieMaxLabel=14</code>, fitting a robust 4-gram model into just <strong>19GB RAM</strong>.</p>\n<h2>4. Inference Pipeline (LISA)</h2>\n<p>We didn't rely on a single model. Our <strong>LISA (LLM-Integrated Scoring Aggregation)</strong> pipeline routes samples dynamically:</p>\n<ol>\n<li><strong>Logit Averaging:</strong> We averaged logits within model families to create constructive interference (signal amplification).</li>\n<li><strong>Gating:</strong> We calculated a normalized n-gram score for every sentence.<ul>\n<li><strong>High Confidence (≥ -3.76):</strong> Routed to <strong>Mistral-7B</strong> for semantic rescoring.</li>\n<li><strong>Low Confidence (&lt; -3.76):</strong> Bypassed the LLM to avoid hallucinations on random sequences.</li></ul></li>\n<li><strong>TTA:</strong> Batched Test-Time Adaptation on high-confidence predictions.</li>\n</ol>\n<h2>5. Results</h2>\n<p>This approach allowed us to decode both coherent speech and random patterns with high fidelity, achieving 7th place despite hardware constraints.</p>\n<hr>\n<p><em>Again, if you enjoyed this summary, please check out the full detailed breakdown on <a href=\"https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d\" target=\"_blank\">Medium</a> and star the <a href=\"https://github.com/greentree327/brain-to-text-mamba-decoder\" target=\"_blank\">GitHub repo</a>.</em></p>",
      "rawMarkdown": "# [7th Place Solution] Mamba + GRU + KenLM (with code)\n\nFirst off, thanks to Nicholas Card, the UC Davis Neuroprosthetics Lab, and the organizers for this fascinating competition. It has always been my dream to explore how ML can apply to clinical settings. Huge thanks to my teammate @Kyle Hui.\n\n### 🔗 Resources\n* **📖 Detailed Writeup (Medium):** [Click Here](https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d)\n* **💻 Full Code (GitHub):** [Click Here](https://github.com/greentree327/brain-to-text-mamba-decoder)\n\n*If you find this writeup or the code useful, please support us by **starring the GitHub repo** and leaving a **clap/comment on Medium**!*\n\n## TL;DR\nOur solution is a hybrid ensemble of **SoftWindow Bi-Mamba** and **GRU** models. We focused heavily on memory efficiency (fitting a custom KenLM into 19GB RAM vs 300GB baseline) and a **Context-Aware Inference Pipeline (LISA)** that gates LLM rescoring based on signal confidence.\n\n## 1. The Challenge\nWe faced severe overfitting (0.0 training loss vs high validation loss) and strict compute limits (Single A100 on Colab). We needed a model light on RAM but heavy on regularization.\n\n## 2. The Solution: Hybrid Architecture\nWe found **Mamba** captured long-range semantic context while **GRU** stabilized short-range acoustic modeling. Their error patterns were highly uncorrelated.\n* **Bi-directional Mamba:** 3-layer input compression with heavy dropout.\n* **Soft Sliding Window:** Biased SSM parameters to prioritize short-term memory.\n* **Temporal Consistency Loss:** A custom loss minimizing the difference between adjacent day matrices to handle neural drift.\n\n## 3. Memory Optimization: Custom KenLM\nStandard 5-gram models (~300GB) were impossible. We compiled a custom **flashlight decoder** with `kTrieMaxLabel=14`, fitting a robust 4-gram model into just **19GB RAM**.\n\n## 4. Inference Pipeline (LISA)\nWe didn't rely on a single model. Our **LISA (LLM-Integrated Scoring Aggregation)** pipeline routes samples dynamically:\n1. **Logit Averaging:** We averaged logits within model families to create constructive interference (signal amplification).\n2. **Gating:** We calculated a normalized n-gram score for every sentence.\n    * **High Confidence (≥ -3.76):** Routed to **Mistral-7B** for semantic rescoring.\n    * **Low Confidence (< -3.76):** Bypassed the LLM to avoid hallucinations on random sequences.\n3. **TTA:** Batched Test-Time Adaptation on high-confidence predictions.\n\n## 5. Results\nThis approach allowed us to decode both coherent speech and random patterns with high fidelity, achieving 7th place despite hardware constraints.\n\n---\n\n*Again, if you enjoyed this summary, please check out the full detailed breakdown on [Medium](https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d) and star the [GitHub repo](https://github.com/greentree327/brain-to-text-mamba-decoder).*\n",
      "votes": 7
    },
    {
      "id": 3384634,
      "postDate": "2026-01-01T16:23:59.273Z",
      "content": "<p>Thanks for participating and for your detailed write-up, I will check it out!</p>",
      "rawMarkdown": "Thanks for participating and for your detailed write-up, I will check it out!"
    },
    {
      "id": 3384484,
      "postDate": "2026-01-01T10:36:12.687Z",
      "content": "<p>Hi, Thanks! but the notebook is not available. </p>",
      "rawMarkdown": "Hi, Thanks! but the notebook is not available. ",
      "replies": [
        {
          "id": 3384512,
          "postDate": "2026-01-01T12:21:06.113Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jamalsaeedi\" target=\"_blank\">@jamalsaeedi</a> , thanks for pointing out, i have just changed the notebook link, please check if it is accessible now  , appreciate it🙏</p>",
          "rawMarkdown": "Hi @jamalsaeedi , thanks for pointing out, i have just changed the notebook link, please check if it is accessible now  , appreciate it🙏",
          "replies": [
            {
              "id": 3384633,
              "postDate": "2026-01-01T16:23:13.873Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3384634,
      "author_name": "Nick Card",
      "author_url": "",
      "post_date": "2026-01-01T16:23:59.273000",
      "content": "<p>Thanks for participating and for your detailed write-up, I will check it out!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3384484,
      "author_name": "Cyrus",
      "author_url": "",
      "post_date": "2026-01-01T10:36:12.687000",
      "content": "<p>Hi, Thanks! but the notebook is not available. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3384512,
          "author_name": "Summer",
          "author_url": "",
          "post_date": "2026-01-01T12:21:06.113000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jamalsaeedi\" target=\"_blank\">@jamalsaeedi</a> , thanks for pointing out, i have just changed the notebook link, please check if it is accessible now  , appreciate it🙏</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3384633,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-01-01T16:23:13.873000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3384442": "# [7th Place Solution] Mamba + GRU + KenLM (with code)\n\nFirst off, thanks to Nicholas Card, the UC Davis Neuroprosthetics Lab, and the organizers for this fascinating competition. It has always been my dream to explore how ML can apply to clinical settings. Huge thanks to my teammate @Kyle Hui.\n\n### 🔗 Resources\n* **📖 Detailed Writeup (Medium):** [Click Here](https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d)\n* **💻 Full Code (GitHub):** [Click Here](https://github.com/greentree327/brain-to-text-mamba-decoder)\n\n*If you find this writeup or the code useful, please support us by **starring the GitHub repo** and leaving a **clap/comment on Medium**!*\n\n## TL;DR\nOur solution is a hybrid ensemble of **SoftWindow Bi-Mamba** and **GRU** models. We focused heavily on memory efficiency (fitting a custom KenLM into 19GB RAM vs 300GB baseline) and a **Context-Aware Inference Pipeline (LISA)** that gates LLM rescoring based on signal confidence.\n\n## 1. The Challenge\nWe faced severe overfitting (0.0 training loss vs high validation loss) and strict compute limits (Single A100 on Colab). We needed a model light on RAM but heavy on regularization.\n\n## 2. The Solution: Hybrid Architecture\nWe found **Mamba** captured long-range semantic context while **GRU** stabilized short-range acoustic modeling. Their error patterns were highly uncorrelated.\n* **Bi-directional Mamba:** 3-layer input compression with heavy dropout.\n* **Soft Sliding Window:** Biased SSM parameters to prioritize short-term memory.\n* **Temporal Consistency Loss:** A custom loss minimizing the difference between adjacent day matrices to handle neural drift.\n\n## 3. Memory Optimization: Custom KenLM\nStandard 5-gram models (~300GB) were impossible. We compiled a custom **flashlight decoder** with `kTrieMaxLabel=14`, fitting a robust 4-gram model into just **19GB RAM**.\n\n## 4. Inference Pipeline (LISA)\nWe didn't rely on a single model. Our **LISA (LLM-Integrated Scoring Aggregation)** pipeline routes samples dynamically:\n1. **Logit Averaging:** We averaged logits within model families to create constructive interference (signal amplification).\n2. **Gating:** We calculated a normalized n-gram score for every sentence.\n    * **High Confidence (≥ -3.76):** Routed to **Mistral-7B** for semantic rescoring.\n    * **Low Confidence (< -3.76):** Bypassed the LLM to avoid hallucinations on random sequences.\n3. **TTA:** Batched Test-Time Adaptation on high-confidence predictions.\n\n## 5. Results\nThis approach allowed us to decode both coherent speech and random patterns with high fidelity, achieving 7th place despite hardware constraints.\n\n---\n\n*Again, if you enjoyed this summary, please check out the full detailed breakdown on [Medium](https://medium.com/@jackson3b04/7th-place-solution-mamba-gru-kenlm-with-code-brain-to-text-25-00f1c69dcd0d) and star the [GitHub repo](https://github.com/greentree327/brain-to-text-mamba-decoder).*\n",
    "3384634": "Thanks for participating and for your detailed write-up, I will check it out!",
    "3384484": "Hi, Thanks! but the notebook is not available. "
  }
}