{
  "id": 625645,
  "title": "**🔍 Anyone else struggling with repetition + short outputs? Looking for ideas**",
  "url": "/competitions/brain-to-text-25/discussion/625645",
  "author_name": "UmutUygurr",
  "post_date": "2025-11-16T09:47:51.263000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I’m getting decent performance from the baseline RNN decoder, but I keep hitting the same issue on the test split:\nthe model collapses into very short phrases like “a / the / a of”, with extremely low diversity.</p>\n<p>Even with:</p>\n<p>Larger beam width</p>\n<p>Lower LM weight</p>\n<p>Temperature scaling</p>\n<p>…it still gravitates toward the most frequent tokens.</p>\n<p>I’m currently experimenting with:</p>\n<p>RNN ensemble (averaged logits)</p>\n<p>5-gram LM (KenLM) instead of unigram</p>\n<p>Length normalization during beam search</p>\n<p>Early results look better, but I’m still trying to understand why the collapse happens so strongly on some sessions and not others.</p>\n<p>❓ Questions for others</p>\n<p>Are you seeing the same collapse?</p>\n<p>Did anything besides LM tuning help?</p>\n<p>Does diversity correlate with speaking strategy / corpus type?</p>\n<p>Has anyone tried a simple heuristic to detect when the beam is collapsing (e.g., entropy check) and recover from it?</p>\n<p>Any thoughts, even small observations, are welcome — I feel like we’re all hitting similar walls and sharing might unblock us.</p>\n<p>Good luck to everyone and thanks in advance 🙌</p>",
  "messages": [
    {
      "id": 3330652,
      "postDate": "2025-11-16T09:47:51.263Z",
      "content": "<p>Hi all,</p>\n<p>I’m getting decent performance from the baseline RNN decoder, but I keep hitting the same issue on the test split:\nthe model collapses into very short phrases like “a / the / a of”, with extremely low diversity.</p>\n<p>Even with:</p>\n<p>Larger beam width</p>\n<p>Lower LM weight</p>\n<p>Temperature scaling</p>\n<p>…it still gravitates toward the most frequent tokens.</p>\n<p>I’m currently experimenting with:</p>\n<p>RNN ensemble (averaged logits)</p>\n<p>5-gram LM (KenLM) instead of unigram</p>\n<p>Length normalization during beam search</p>\n<p>Early results look better, but I’m still trying to understand why the collapse happens so strongly on some sessions and not others.</p>\n<p>❓ Questions for others</p>\n<p>Are you seeing the same collapse?</p>\n<p>Did anything besides LM tuning help?</p>\n<p>Does diversity correlate with speaking strategy / corpus type?</p>\n<p>Has anyone tried a simple heuristic to detect when the beam is collapsing (e.g., entropy check) and recover from it?</p>\n<p>Any thoughts, even small observations, are welcome — I feel like we’re all hitting similar walls and sharing might unblock us.</p>\n<p>Good luck to everyone and thanks in advance 🙌</p>",
      "rawMarkdown": "Hi all,\n\nI’m getting decent performance from the baseline RNN decoder, but I keep hitting the same issue on the test split:\nthe model collapses into very short phrases like “a / the / a of”, with extremely low diversity.\n\nEven with:\n\nLarger beam width\n\nLower LM weight\n\nTemperature scaling\n\n…it still gravitates toward the most frequent tokens.\n\nI’m currently experimenting with:\n\nRNN ensemble (averaged logits)\n\n5-gram LM (KenLM) instead of unigram\n\nLength normalization during beam search\n\nEarly results look better, but I’m still trying to understand why the collapse happens so strongly on some sessions and not others.\n\n❓ Questions for others\n\nAre you seeing the same collapse?\n\nDid anything besides LM tuning help?\n\nDoes diversity correlate with speaking strategy / corpus type?\n\nHas anyone tried a simple heuristic to detect when the beam is collapsing (e.g., entropy check) and recover from it?\n\nAny thoughts, even small observations, are welcome — I feel like we’re all hitting similar walls and sharing might unblock us.\n\nGood luck to everyone and thanks in advance 🙌"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3330652": "Hi all,\n\nI’m getting decent performance from the baseline RNN decoder, but I keep hitting the same issue on the test split:\nthe model collapses into very short phrases like “a / the / a of”, with extremely low diversity.\n\nEven with:\n\nLarger beam width\n\nLower LM weight\n\nTemperature scaling\n\n…it still gravitates toward the most frequent tokens.\n\nI’m currently experimenting with:\n\nRNN ensemble (averaged logits)\n\n5-gram LM (KenLM) instead of unigram\n\nLength normalization during beam search\n\nEarly results look better, but I’m still trying to understand why the collapse happens so strongly on some sessions and not others.\n\n❓ Questions for others\n\nAre you seeing the same collapse?\n\nDid anything besides LM tuning help?\n\nDoes diversity correlate with speaking strategy / corpus type?\n\nHas anyone tried a simple heuristic to detect when the beam is collapsing (e.g., entropy check) and recover from it?\n\nAny thoughts, even small observations, are welcome — I feel like we’re all hitting similar walls and sharing might unblock us.\n\nGood luck to everyone and thanks in advance 🙌"
  }
}