{
  "id": 603760,
  "title": "Brain-to-text '25 ",
  "url": "/competitions/brain-to-text-25/discussion/603760",
  "author_name": "Prathamesh Mistry",
  "post_date": "2025-09-04T07:55:47.292000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8486392%2F99cf78f4cb7d1dbfe677f97d44d7e544%2FScreenshot%202025-09-04%20112956.png?generation=1756972575998983&amp;alt=media\" alt=\"\">WER↔CER gap (~0.58 vs 0.385): Any tricks you used to close this (error types suggest word-level LM or better tokenization)?</p>\n<p>CTC blank bias / logit scaling: Have you tuned blank prior or applied temperature on logits before decoding? Worth it here?</p>\n<p>Language model rescoring: Best lightweight option you’ve tried—KenLM 3-gram/5-gram via pyctcdecode, or shallow fusion with a small seq2seq (e.g., TinyT5)?</p>\n<p>Feature-space regularization: Has channel dropout / spec-augment-like masks on time×feature helped for iEEG features?</p>\n<p>Alignment: I used linear interpolation. Anyone tried DTW or learned alignment that improved robustness?</p>\n<p>Optimizer schedule: Cosine vs linear  did cosine + warm restarts noticeably help beyond ~80 epochs?</p>\n<p>EMA weights: Did Exponential Moving Average of weights reduce generalization gap for you?</p>\n<p>Mixup/CutMix (feature wise) across trials  yay or nay on neural features?</p>\n<p>RNN vs Conformer/Temporal CNN: If you switched encoders, what was the best trade-off on T4×2?</p>\n<p>Post-processing: Any simple spell/wordpiece correction step that gave &gt;0.01 WER improvement?</p>",
  "messages": [
    {
      "id": 3281392,
      "postDate": "2025-09-04T07:55:47.293Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8486392%2F99cf78f4cb7d1dbfe677f97d44d7e544%2FScreenshot%202025-09-04%20112956.png?generation=1756972575998983&amp;alt=media\" alt=\"\">WER↔CER gap (~0.58 vs 0.385): Any tricks you used to close this (error types suggest word-level LM or better tokenization)?</p>\n<p>CTC blank bias / logit scaling: Have you tuned blank prior or applied temperature on logits before decoding? Worth it here?</p>\n<p>Language model rescoring: Best lightweight option you’ve tried—KenLM 3-gram/5-gram via pyctcdecode, or shallow fusion with a small seq2seq (e.g., TinyT5)?</p>\n<p>Feature-space regularization: Has channel dropout / spec-augment-like masks on time×feature helped for iEEG features?</p>\n<p>Alignment: I used linear interpolation. Anyone tried DTW or learned alignment that improved robustness?</p>\n<p>Optimizer schedule: Cosine vs linear  did cosine + warm restarts noticeably help beyond ~80 epochs?</p>\n<p>EMA weights: Did Exponential Moving Average of weights reduce generalization gap for you?</p>\n<p>Mixup/CutMix (feature wise) across trials  yay or nay on neural features?</p>\n<p>RNN vs Conformer/Temporal CNN: If you switched encoders, what was the best trade-off on T4×2?</p>\n<p>Post-processing: Any simple spell/wordpiece correction step that gave &gt;0.01 WER improvement?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8486392%2F99cf78f4cb7d1dbfe677f97d44d7e544%2FScreenshot%202025-09-04%20112956.png?generation=1756972575998983&alt=media)WER↔CER gap (~0.58 vs 0.385): Any tricks you used to close this (error types suggest word-level LM or better tokenization)?\n\nCTC blank bias / logit scaling: Have you tuned blank prior or applied temperature on logits before decoding? Worth it here?\n\nLanguage model rescoring: Best lightweight option you’ve tried—KenLM 3-gram/5-gram via pyctcdecode, or shallow fusion with a small seq2seq (e.g., TinyT5)?\n\nFeature-space regularization: Has channel dropout / spec-augment-like masks on time×feature helped for iEEG features?\n\nAlignment: I used linear interpolation. Anyone tried DTW or learned alignment that improved robustness?\n\nOptimizer schedule: Cosine vs linear  did cosine + warm restarts noticeably help beyond ~80 epochs?\n\nEMA weights: Did Exponential Moving Average of weights reduce generalization gap for you?\n\nMixup/CutMix (feature wise) across trials  yay or nay on neural features?\n\nRNN vs Conformer/Temporal CNN: If you switched encoders, what was the best trade-off on T4×2?\n\nPost-processing: Any simple spell/wordpiece correction step that gave >0.01 WER improvement?",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3281392": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8486392%2F99cf78f4cb7d1dbfe677f97d44d7e544%2FScreenshot%202025-09-04%20112956.png?generation=1756972575998983&alt=media)WER↔CER gap (~0.58 vs 0.385): Any tricks you used to close this (error types suggest word-level LM or better tokenization)?\n\nCTC blank bias / logit scaling: Have you tuned blank prior or applied temperature on logits before decoding? Worth it here?\n\nLanguage model rescoring: Best lightweight option you’ve tried—KenLM 3-gram/5-gram via pyctcdecode, or shallow fusion with a small seq2seq (e.g., TinyT5)?\n\nFeature-space regularization: Has channel dropout / spec-augment-like masks on time×feature helped for iEEG features?\n\nAlignment: I used linear interpolation. Anyone tried DTW or learned alignment that improved robustness?\n\nOptimizer schedule: Cosine vs linear  did cosine + warm restarts noticeably help beyond ~80 epochs?\n\nEMA weights: Did Exponential Moving Average of weights reduce generalization gap for you?\n\nMixup/CutMix (feature wise) across trials  yay or nay on neural features?\n\nRNN vs Conformer/Temporal CNN: If you switched encoders, what was the best trade-off on T4×2?\n\nPost-processing: Any simple spell/wordpiece correction step that gave >0.01 WER improvement?"
  }
}