{
  "id": 610005,
  "title": "Corpus-dependent decoder?",
  "url": "/competitions/brain-to-text-25/discussion/610005",
  "author_name": "",
  "post_date": "2025-10-01T01:29:59.393054200Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I was reading the rules, and this seems like a gray area to me:</p>\n<p>Would it be allowed for us to incorporate a corpus-dependent decoder profile? The most obvious application here is with the random words corpus, as using any ngram or LLM rescoring method would make the results on this corpus worse, since the sentences are nonsense.</p>\n<p>I'm inclined to interpret the rules to mean that this is not allowed, but if it is, I would like to take advantage.</p>",
  "messages": [
    {
      "id": "3296475",
      "postDate": "10/01/2025 01:29:59",
      "content": "<p>I was reading the rules, and this seems like a gray area to me:</p>\n<p>Would it be allowed for us to incorporate a corpus-dependent decoder profile? The most obvious application here is with the random words corpus, as using any ngram or LLM rescoring method would make the results on this corpus worse, since the sentences are nonsense.</p>\n<p>I'm inclined to interpret the rules to mean that this is not allowed, but if it is, I would like to take advantage.</p>",
      "rawMarkdown": "I was reading the rules, and this seems like a gray area to me:\n\nWould it be allowed for us to incorporate a corpus-dependent decoder profile? The most obvious application here is with the random words corpus, as using any ngram or LLM rescoring method would make the results on this corpus worse, since the sentences are nonsense.\n\nI'm inclined to interpret the rules to mean that this is not allowed, but if it is, I would like to take advantage.",
      "votes": null
    },
    {
      "id": "3299044",
      "postDate": "10/07/2025 04:17:37",
      "content": "<p>From the rules:</p>\n<blockquote>\n  <p>The decoding should be based solely on code/algorithms that <strong><em>would not rely on human intervention and can be expected to generalize to new text.</em></strong></p>\n</blockquote>\n<p>If your approach can generalize to text without requiring knowledge of the appropriate corpus to use, it would be fine. But if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement.</p>",
      "rawMarkdown": "From the rules:\n>The decoding should be based solely on code/algorithms that ***would not rely on human intervention and can be expected to generalize to new text.***\n\nIf your approach can generalize to text without requiring knowledge of the appropriate corpus to use, it would be fine. But if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement.",
      "votes": null
    },
    {
      "id": "3303400",
      "postDate": "10/17/2025 22:29:53",
      "content": "<p>Will you have a way to check submissions for this, or eg. for human mind-assisted results? The codes will not be evaluated by the organizers as I understand, so is it up to the participants to play properly, write the one-page long report honestly?<br>\nSimilarly, the date, or post-implant day carries extra information. Is it allowed to train separate models for trial-groups recorded many months away, or that information should not be used as to new text it would not generalize. Again, how to make sure nobody uses then?</p>",
      "rawMarkdown": "Will you have a way to check submissions for this, or eg. for human mind-assisted results? The codes will not be evaluated by the organizers as I understand, so is it up to the participants to play properly, write the one-page long report honestly?\nSimilarly, the date, or post-implant day carries extra information. Is it allowed to train separate models for trial-groups recorded many months away, or that information should not be used as to new text it would not generalize. Again, how to make sure nobody uses then?",
      "votes": null
    },
    {
      "id": "3303425",
      "postDate": "10/18/2025 01:38:34",
      "content": "<p>As stated in the rules:</p>\n<blockquote>\n  <p>To be considered eligible for cash prizes, competitors must submit a one-page summary of their approach <strong>as well as the code required to reproduce it</strong></p>\n</blockquote>",
      "rawMarkdown": "As stated in the rules:\n> To be considered eligible for cash prizes, competitors must submit a one-page summary of their approach **as well as the code required to reproduce it**",
      "votes": null
    },
    {
      "id": "3306913",
      "postDate": "10/25/2025 15:17:21",
      "content": "<p>Please help me to understand this corpus-dependent decoder topic clearly: </p>\n<ul>\n<li>The Overview section encourages us to use different language models for each corpora (\"<em>Using <strong>different language models for sentences from different corpora</strong></em>\"), so our brain-to-text engine would depend on the corpora for each prediction.</li>\n<li>The rules require the solution to generalize to new text, and you say that \"<em>But <strong>if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement</strong></em>\".</li>\n</ul>\n<p>I think there's only a thin line between these, some might misunderstand. It can lead to big advantages, therefore important to clearly draw the line. Do I understand well, that:</p>\n<ul>\n<li><p>creating a detector for Switchboard training files and using it to predict Switchboard Testsets, and<br>\ncreating a detector for \"Freq words\" training files and using it to predict \"Freq words\" Testsets, <br>\nis only allowed if the algorithm is able to detect based the contents of a test .hdf5 file (not filename) that it belongs to the Switchboard or \"Freq words\" corpus? Basically the corpus data within 't15_copyTaskData_descriptions.csv' is not allowed to be used while processing the testset .hdf5 files, right?</p></li>\n<li><p>Word list, word frequency statistics or n-grams for each corpus separately can not be created and then used for testset prediction. Eg. it's cheating to download the original switchboard dataset and only allow the words in that to be predicted, because it does not generalize to new text.</p></li>\n</ul>\n<p>Also, if you feel like, please add your explanation of the two quoted sentences to make it clear what's allowed and what's not.</p>",
      "rawMarkdown": "Please help me to understand this corpus-dependent decoder topic clearly: \n- The Overview section encourages us to use different language models for each corpora (\"*Using **different language models for sentences from different corpora***\"), so our brain-to-text engine would depend on the corpora for each prediction.\n- The rules require the solution to generalize to new text, and you say that \"*But **if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement***\".\n\nI think there's only a thin line between these, some might misunderstand. It can lead to big advantages, therefore important to clearly draw the line. Do I understand well, that:\n\n- creating a detector for Switchboard training files and using it to predict Switchboard Testsets, and\ncreating a detector for \"Freq words\" training files and using it to predict \"Freq words\" Testsets, \nis only allowed if the algorithm is able to detect based the contents of a test .hdf5 file (not filename) that it belongs to the Switchboard or \"Freq words\" corpus? Basically the corpus data within 't15_copyTaskData_descriptions.csv' is not allowed to be used while processing the testset .hdf5 files, right?\n\n- Word list, word frequency statistics or n-grams for each corpus separately can not be created and then used for testset prediction. Eg. it's cheating to download the original switchboard dataset and only allow the words in that to be predicted, because it does not generalize to new text.\n\nAlso, if you feel like, please add your explanation of the two quoted sentences to make it clear what's allowed and what's not.",
      "votes": null
    },
    {
      "id": "3306990",
      "postDate": "10/25/2025 21:42:01",
      "content": "<p>You need to use a model and a decoding pipeline that doesn't care about the differences between sources (corpora) at the point of inference.</p>\n<p>The core requirement from the competition host is generalization.</p>",
      "rawMarkdown": "You need to use a model and a decoding pipeline that doesn't care about the differences between sources (corpora) at the point of inference.\n\nThe core requirement from the competition host is generalization.",
      "votes": null
    },
    {
      "id": "3307142",
      "postDate": "10/26/2025 09:41:30",
      "content": "<p>In the same time, we are encouraged to try using difference models for each corpora. Of course it only makes sense if you do it both at training and testing time, right? It doesn't make sense to train separate models for each corpora, but doesn't care about the corpora at test time. How do you resolve this conflict? Your answer doesn't explain how can my quote from the Overview section be implemented. I think we still miss some clarification or maybe that part of the Overview conflicts with the generalization requirement?<br>\nI'm happy if this corpus-dependent decoder way is not allowed, I think so generalization should be a requirement, but I'm afraid some might build corpus-dependent detector as the Overview section encourages it, and might gain advantage or will be badly disappointed if disqualified at the end.</p>",
      "rawMarkdown": "In the same time, we are encouraged to try using difference models for each corpora. Of course it only makes sense if you do it both at training and testing time, right? It doesn't make sense to train separate models for each corpora, but doesn't care about the corpora at test time. How do you resolve this conflict? Your answer doesn't explain how can my quote from the Overview section be implemented. I think we still miss some clarification or maybe that part of the Overview conflicts with the generalization requirement?\nI'm happy if this corpus-dependent decoder way is not allowed, I think so generalization should be a requirement, but I'm afraid some might build corpus-dependent detector as the Overview section encourages it, and might gain advantage or will be badly disappointed if disqualified at the end.",
      "votes": null
    },
    {
      "id": "3307151",
      "postDate": "10/26/2025 10:01:57",
      "content": "<ol>\n<li><p>Use a non-LM decoder (e.g., Greedy Search or Beam Search without a Language Model).</p></li>\n<li><p>Use dynamic decoding: Use a Language Model (LM) for the Natural Language Processing (NLP)/Coherent corpus sessions, and use a separate, distinct LM (or a non-LM decoder) for the Random corpus sessions. (This usually implies implementing a corpus-dependent switch.)</p></li>\n<li><p>Use a single, unified LM decoder: Treat the Random corpus data as noise during the acoustic model training phase (to increase robustness), but then apply the single Language Model (LM) uniformly to all test sessions (both coherent and random).</p></li>\n</ol>",
      "rawMarkdown": "1. Use a non-LM decoder (e.g., Greedy Search or Beam Search without a Language Model).\n\n2.  Use dynamic decoding: Use a Language Model (LM) for the Natural Language Processing (NLP)/Coherent corpus sessions, and use a separate, distinct LM (or a non-LM decoder) for the Random corpus sessions. (This usually implies implementing a corpus-dependent switch.)\n\n3. Use a single, unified LM decoder: Treat the Random corpus data as noise during the acoustic model training phase (to increase robustness), but then apply the single Language Model (LM) uniformly to all test sessions (both coherent and random).",
      "votes": null
    },
    {
      "id": "3307154",
      "postDate": "10/26/2025 10:16:11",
      "content": "<p>If you only train your model on the Natural Language Processing (NLP) corpus (the coherent sentences) and then test it on the Random words corpus, your model will be much more likely to output a blank space (\"\") for the random sessions.</p>",
      "rawMarkdown": "If you only train your model on the Natural Language Processing (NLP) corpus (the coherent sentences) and then test it on the Random words corpus, your model will be much more likely to output a blank space (\"\") for the random sessions.",
      "votes": null
    },
    {
      "id": "3307437",
      "postDate": "10/27/2025 01:50:07",
      "content": "<p>You are talking about the difference between training your model while knowing what the target corpus is vs. testing on unseen data while knowing what the target corpus is. The former is fine, but with the goal of generalization in mind, we do not want to have to tell the model what corpus to use at test time. In retrospect, I should not have included corpus labels for the test set to simplify this issue. But we will screen for this in final code submissions for prize-eligible submissions.</p>\n<p>A more concrete example of what is ok or not ok:</p>\n<ul>\n<li><strong>OK:</strong> you use the training dataset to train corpus-dependent models, and at test time you automatically detect (based on the input neural data for each test trial) which corpus-dependent decoder to use for inference.</li>\n<li><strong>NOT OK:</strong> you use the training dataset to train corpus-dependent models, and at test time you use the 't15_copyTaskData_descriptions.csv' file to decide which model to use for each test trial.</li>\n</ul>\n<p>Long story short: at test time, you should not have to give your model any information other than (1) the neural data for the test trial, and optionally (2) the session's date (to use the appropriate day-specific layer to account for neural nonstationarities). Telling the inference model what corpus to use, word frequency, word list, etc is <strong>obviously</strong> not in the spirit of a generalizable model that enables a BCI user to talk about whatever the might want to talk about.</p>",
      "rawMarkdown": "You are talking about the difference between training your model while knowing what the target corpus is vs. testing on unseen data while knowing what the target corpus is. The former is fine, but with the goal of generalization in mind, we do not want to have to tell the model what corpus to use at test time. In retrospect, I should not have included corpus labels for the test set to simplify this issue. But we will screen for this in final code submissions for prize-eligible submissions.\n\nA more concrete example of what is ok or not ok:\n- **OK:** you use the training dataset to train corpus-dependent models, and at test time you automatically detect (based on the input neural data for each test trial) which corpus-dependent decoder to use for inference.\n- **NOT OK:** you use the training dataset to train corpus-dependent models, and at test time you use the 't15_copyTaskData_descriptions.csv' file to decide which model to use for each test trial.\n\nLong story short: at test time, you should not have to give your model any information other than (1) the neural data for the test trial, and optionally (2) the session's date (to use the appropriate day-specific layer to account for neural nonstationarities). Telling the inference model what corpus to use, word frequency, word list, etc is **obviously** not in the spirit of a generalizable model that enables a BCI user to talk about whatever the might want to talk about.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3299044,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "10/07/2025 04:17:37",
      "content": "<p>From the rules:</p>\n<blockquote>\n  <p>The decoding should be based solely on code/algorithms that <strong><em>would not rely on human intervention and can be expected to generalize to new text.</em></strong></p>\n</blockquote>\n<p>If your approach can generalize to text without requiring knowledge of the appropriate corpus to use, it would be fine. But if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3303400,
          "author_name": "gorogm",
          "author_url": "",
          "post_date": "10/17/2025 22:29:53",
          "content": "<p>Will you have a way to check submissions for this, or eg. for human mind-assisted results? The codes will not be evaluated by the organizers as I understand, so is it up to the participants to play properly, write the one-page long report honestly?<br>\nSimilarly, the date, or post-implant day carries extra information. Is it allowed to train separate models for trial-groups recorded many months away, or that information should not be used as to new text it would not generalize. Again, how to make sure nobody uses then?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3303425,
              "author_name": "notnickc",
              "author_url": "",
              "post_date": "10/18/2025 01:38:34",
              "content": "<p>As stated in the rules:</p>\n<blockquote>\n  <p>To be considered eligible for cash prizes, competitors must submit a one-page summary of their approach <strong>as well as the code required to reproduce it</strong></p>\n</blockquote>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3306913,
      "author_name": "gorogm",
      "author_url": "",
      "post_date": "10/25/2025 15:17:21",
      "content": "<p>Please help me to understand this corpus-dependent decoder topic clearly: </p>\n<ul>\n<li>The Overview section encourages us to use different language models for each corpora (\"<em>Using <strong>different language models for sentences from different corpora</strong></em>\"), so our brain-to-text engine would depend on the corpora for each prediction.</li>\n<li>The rules require the solution to generalize to new text, and you say that \"<em>But <strong>if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement</strong></em>\".</li>\n</ul>\n<p>I think there's only a thin line between these, some might misunderstand. It can lead to big advantages, therefore important to clearly draw the line. Do I understand well, that:</p>\n<ul>\n<li><p>creating a detector for Switchboard training files and using it to predict Switchboard Testsets, and<br>\ncreating a detector for \"Freq words\" training files and using it to predict \"Freq words\" Testsets, <br>\nis only allowed if the algorithm is able to detect based the contents of a test .hdf5 file (not filename) that it belongs to the Switchboard or \"Freq words\" corpus? Basically the corpus data within 't15_copyTaskData_descriptions.csv' is not allowed to be used while processing the testset .hdf5 files, right?</p></li>\n<li><p>Word list, word frequency statistics or n-grams for each corpus separately can not be created and then used for testset prediction. Eg. it's cheating to download the original switchboard dataset and only allow the words in that to be predicted, because it does not generalize to new text.</p></li>\n</ul>\n<p>Also, if you feel like, please add your explanation of the two quoted sentences to make it clear what's allowed and what's not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3307437,
          "author_name": "notnickc",
          "author_url": "",
          "post_date": "10/27/2025 01:50:07",
          "content": "<p>You are talking about the difference between training your model while knowing what the target corpus is vs. testing on unseen data while knowing what the target corpus is. The former is fine, but with the goal of generalization in mind, we do not want to have to tell the model what corpus to use at test time. In retrospect, I should not have included corpus labels for the test set to simplify this issue. But we will screen for this in final code submissions for prize-eligible submissions.</p>\n<p>A more concrete example of what is ok or not ok:</p>\n<ul>\n<li><strong>OK:</strong> you use the training dataset to train corpus-dependent models, and at test time you automatically detect (based on the input neural data for each test trial) which corpus-dependent decoder to use for inference.</li>\n<li><strong>NOT OK:</strong> you use the training dataset to train corpus-dependent models, and at test time you use the 't15_copyTaskData_descriptions.csv' file to decide which model to use for each test trial.</li>\n</ul>\n<p>Long story short: at test time, you should not have to give your model any information other than (1) the neural data for the test trial, and optionally (2) the session's date (to use the appropriate day-specific layer to account for neural nonstationarities). Telling the inference model what corpus to use, word frequency, word list, etc is <strong>obviously</strong> not in the spirit of a generalizable model that enables a BCI user to talk about whatever the might want to talk about.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3306990,
      "author_name": "nocarbonintelligence",
      "author_url": "",
      "post_date": "10/25/2025 21:42:01",
      "content": "<p>You need to use a model and a decoding pipeline that doesn't care about the differences between sources (corpora) at the point of inference.</p>\n<p>The core requirement from the competition host is generalization.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3307142,
          "author_name": "gorogm",
          "author_url": "",
          "post_date": "10/26/2025 09:41:30",
          "content": "<p>In the same time, we are encouraged to try using difference models for each corpora. Of course it only makes sense if you do it both at training and testing time, right? It doesn't make sense to train separate models for each corpora, but doesn't care about the corpora at test time. How do you resolve this conflict? Your answer doesn't explain how can my quote from the Overview section be implemented. I think we still miss some clarification or maybe that part of the Overview conflicts with the generalization requirement?<br>\nI'm happy if this corpus-dependent decoder way is not allowed, I think so generalization should be a requirement, but I'm afraid some might build corpus-dependent detector as the Overview section encourages it, and might gain advantage or will be badly disappointed if disqualified at the end.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3307151,
              "author_name": "nocarbonintelligence",
              "author_url": "",
              "post_date": "10/26/2025 10:01:57",
              "content": "<ol>\n<li><p>Use a non-LM decoder (e.g., Greedy Search or Beam Search without a Language Model).</p></li>\n<li><p>Use dynamic decoding: Use a Language Model (LM) for the Natural Language Processing (NLP)/Coherent corpus sessions, and use a separate, distinct LM (or a non-LM decoder) for the Random corpus sessions. (This usually implies implementing a corpus-dependent switch.)</p></li>\n<li><p>Use a single, unified LM decoder: Treat the Random corpus data as noise during the acoustic model training phase (to increase robustness), but then apply the single Language Model (LM) uniformly to all test sessions (both coherent and random).</p></li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 3307154,
                  "author_name": "nocarbonintelligence",
                  "author_url": "",
                  "post_date": "10/26/2025 10:16:11",
                  "content": "<p>If you only train your model on the Natural Language Processing (NLP) corpus (the coherent sentences) and then test it on the Random words corpus, your model will be much more likely to output a blank space (\"\") for the random sessions.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3296475": "I was reading the rules, and this seems like a gray area to me:\n\nWould it be allowed for us to incorporate a corpus-dependent decoder profile? The most obvious application here is with the random words corpus, as using any ngram or LLM rescoring method would make the results on this corpus worse, since the sentences are nonsense.\n\nI'm inclined to interpret the rules to mean that this is not allowed, but if it is, I would like to take advantage.",
    "3299044": "From the rules:\n>The decoding should be based solely on code/algorithms that ***would not rely on human intervention and can be expected to generalize to new text.***\n\nIf your approach can generalize to text without requiring knowledge of the appropriate corpus to use, it would be fine. But if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement.",
    "3303400": "Will you have a way to check submissions for this, or eg. for human mind-assisted results? The codes will not be evaluated by the organizers as I understand, so is it up to the participants to play properly, write the one-page long report honestly?\nSimilarly, the date, or post-implant day carries extra information. Is it allowed to train separate models for trial-groups recorded many months away, or that information should not be used as to new text it would not generalize. Again, how to make sure nobody uses then?",
    "3303425": "As stated in the rules:\n> To be considered eligible for cash prizes, competitors must submit a one-page summary of their approach **as well as the code required to reproduce it**",
    "3306913": "Please help me to understand this corpus-dependent decoder topic clearly: \n- The Overview section encourages us to use different language models for each corpora (\"*Using **different language models for sentences from different corpora***\"), so our brain-to-text engine would depend on the corpora for each prediction.\n- The rules require the solution to generalize to new text, and you say that \"*But **if you have to tell the decoder whether to use the switchboard-optimized decoder, or the random word optimized decoder, you would not fit that requirement***\".\n\nI think there's only a thin line between these, some might misunderstand. It can lead to big advantages, therefore important to clearly draw the line. Do I understand well, that:\n\n- creating a detector for Switchboard training files and using it to predict Switchboard Testsets, and\ncreating a detector for \"Freq words\" training files and using it to predict \"Freq words\" Testsets, \nis only allowed if the algorithm is able to detect based the contents of a test .hdf5 file (not filename) that it belongs to the Switchboard or \"Freq words\" corpus? Basically the corpus data within 't15_copyTaskData_descriptions.csv' is not allowed to be used while processing the testset .hdf5 files, right?\n\n- Word list, word frequency statistics or n-grams for each corpus separately can not be created and then used for testset prediction. Eg. it's cheating to download the original switchboard dataset and only allow the words in that to be predicted, because it does not generalize to new text.\n\nAlso, if you feel like, please add your explanation of the two quoted sentences to make it clear what's allowed and what's not.",
    "3306990": "You need to use a model and a decoding pipeline that doesn't care about the differences between sources (corpora) at the point of inference.\n\nThe core requirement from the competition host is generalization.",
    "3307142": "In the same time, we are encouraged to try using difference models for each corpora. Of course it only makes sense if you do it both at training and testing time, right? It doesn't make sense to train separate models for each corpora, but doesn't care about the corpora at test time. How do you resolve this conflict? Your answer doesn't explain how can my quote from the Overview section be implemented. I think we still miss some clarification or maybe that part of the Overview conflicts with the generalization requirement?\nI'm happy if this corpus-dependent decoder way is not allowed, I think so generalization should be a requirement, but I'm afraid some might build corpus-dependent detector as the Overview section encourages it, and might gain advantage or will be badly disappointed if disqualified at the end.",
    "3307151": "1. Use a non-LM decoder (e.g., Greedy Search or Beam Search without a Language Model).\n\n2.  Use dynamic decoding: Use a Language Model (LM) for the Natural Language Processing (NLP)/Coherent corpus sessions, and use a separate, distinct LM (or a non-LM decoder) for the Random corpus sessions. (This usually implies implementing a corpus-dependent switch.)\n\n3. Use a single, unified LM decoder: Treat the Random corpus data as noise during the acoustic model training phase (to increase robustness), but then apply the single Language Model (LM) uniformly to all test sessions (both coherent and random).",
    "3307154": "If you only train your model on the Natural Language Processing (NLP) corpus (the coherent sentences) and then test it on the Random words corpus, your model will be much more likely to output a blank space (\"\") for the random sessions.",
    "3307437": "You are talking about the difference between training your model while knowing what the target corpus is vs. testing on unseen data while knowing what the target corpus is. The former is fine, but with the goal of generalization in mind, we do not want to have to tell the model what corpus to use at test time. In retrospect, I should not have included corpus labels for the test set to simplify this issue. But we will screen for this in final code submissions for prize-eligible submissions.\n\nA more concrete example of what is ok or not ok:\n- **OK:** you use the training dataset to train corpus-dependent models, and at test time you automatically detect (based on the input neural data for each test trial) which corpus-dependent decoder to use for inference.\n- **NOT OK:** you use the training dataset to train corpus-dependent models, and at test time you use the 't15_copyTaskData_descriptions.csv' file to decide which model to use for each test trial.\n\nLong story short: at test time, you should not have to give your model any information other than (1) the neural data for the test trial, and optionally (2) the session's date (to use the appropriate day-specific layer to account for neural nonstationarities). Telling the inference model what corpus to use, word frequency, word list, etc is **obviously** not in the spirit of a generalizable model that enables a BCI user to talk about whatever the might want to talk about."
  },
  "source": "meta"
}