{
  "id": 613722,
  "title": "Question about the predefined train/validation split and its difficulty",
  "url": "/competitions/brain-to-text-25/discussion/613722",
  "author_name": "",
  "post_date": "2025-10-29T08:53:09.913772400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I’d like to know whether the training and validation sets were split in advance on purpose.<br>\nWhen I pooled all the data and ran 5-fold cross-validation, the validation PER turned out to be much lower.<br>\nDoes this mean the official validation set is deliberately harder?<br>\nWas this done to make its distribution closer to that of the test set?</p>",
  "messages": [
    {
      "id": "3308366",
      "postDate": "10/29/2025 08:53:09",
      "content": "<p>I’d like to know whether the training and validation sets were split in advance on purpose.<br>\nWhen I pooled all the data and ran 5-fold cross-validation, the validation PER turned out to be much lower.<br>\nDoes this mean the official validation set is deliberately harder?<br>\nWas this done to make its distribution closer to that of the test set?</p>",
      "rawMarkdown": "I’d like to know whether the training and validation sets were split in advance on purpose.  \nWhen I pooled all the data and ran 5-fold cross-validation, the validation PER turned out to be much lower.  \nDoes this mean the official validation set is deliberately harder?  \nWas this done to make its distribution closer to that of the test set?",
      "votes": null
    },
    {
      "id": "3308743",
      "postDate": "10/30/2025 06:41:51",
      "content": "<p>As you can see from the data split information <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/data/t15_copyTaskData_description.csv\" target=\"_blank\">here</a>, val/test blocks span a range of corpora. Validation data are roughly representative of test data, as they are 50/50 splits of trials from the same blocks.</p>",
      "rawMarkdown": "As you can see from the data split information [here](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/data/t15_copyTaskData_description.csv), val/test blocks span a range of corpora. Validation data are roughly representative of test data, as they are 50/50 splits of trials from the same blocks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3308743,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "10/30/2025 06:41:51",
      "content": "<p>As you can see from the data split information <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/data/t15_copyTaskData_description.csv\" target=\"_blank\">here</a>, val/test blocks span a range of corpora. Validation data are roughly representative of test data, as they are 50/50 splits of trials from the same blocks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3308366": "I’d like to know whether the training and validation sets were split in advance on purpose.  \nWhen I pooled all the data and ran 5-fold cross-validation, the validation PER turned out to be much lower.  \nDoes this mean the official validation set is deliberately harder?  \nWas this done to make its distribution closer to that of the test set?",
    "3308743": "As you can see from the data split information [here](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/data/t15_copyTaskData_description.csv), val/test blocks span a range of corpora. Validation data are roughly representative of test data, as they are 50/50 splits of trials from the same blocks."
  },
  "source": "meta"
}