{
  "id": 589508,
  "title": "t15_copyTask_neuralData vs t15_copyTask",
  "url": "/competitions/brain-to-text-25/discussion/589508",
  "author_name": "",
  "post_date": "2025-07-13T11:40:00.707382100Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks for this exciting dataset &amp; competition!<br>\n1) Are the sentences in <code>t15_copyTask.pkl</code> a subset of the sentences in <code>t15_copyTask_neuralData.zip</code>? What is the relation of the trials in the two sets?<br>\n2) In <code>t15_copyTask_neuralData.zip</code> I found the 45 sessions as expected, but when I sum the number of trials_XXXes in the .hdf5 files, I count 10948 trials instead of the 11300 trials mentioned in the introduction. What might I miss? Is it possible, that the number of trials is less then 11000, but in some trials there are multiple sentences? In this case this is misleading in the README.txt: \"11,000+ Copy Task trials\". Or should the <code>t15_copyTask_neuralData</code> and <code>t15_copyTask</code> be combined to reach 11300? - that seems way more then 11.300 trials.</p>\n<p>Thank in advance!</p>",
  "messages": [
    {
      "id": "3247767",
      "postDate": "07/13/2025 11:40:00",
      "content": "<p>Thanks for this exciting dataset &amp; competition!<br>\n1) Are the sentences in <code>t15_copyTask.pkl</code> a subset of the sentences in <code>t15_copyTask_neuralData.zip</code>? What is the relation of the trials in the two sets?<br>\n2) In <code>t15_copyTask_neuralData.zip</code> I found the 45 sessions as expected, but when I sum the number of trials_XXXes in the .hdf5 files, I count 10948 trials instead of the 11300 trials mentioned in the introduction. What might I miss? Is it possible, that the number of trials is less then 11000, but in some trials there are multiple sentences? In this case this is misleading in the README.txt: \"11,000+ Copy Task trials\". Or should the <code>t15_copyTask_neuralData</code> and <code>t15_copyTask</code> be combined to reach 11300? - that seems way more then 11.300 trials.</p>\n<p>Thank in advance!</p>",
      "rawMarkdown": "Thanks for this exciting dataset & competition!\n1) Are the sentences in `t15_copyTask.pkl` a subset of the sentences in `t15_copyTask_neuralData.zip`? What is the relation of the trials in the two sets?\n2) In `t15_copyTask_neuralData.zip` I found the 45 sessions as expected, but when I sum the number of trials_XXXes in the .hdf5 files, I count 10948 trials instead of the 11300 trials mentioned in the introduction. What might I miss? Is it possible, that the number of trials is less then 11000, but in some trials there are multiple sentences? In this case this is misleading in the README.txt: \"11,000+ Copy Task trials\". Or should the `t15_copyTask_neuralData` and `t15_copyTask` be combined to reach 11300? - that seems way more then 11.300 trials.\n\nThank in advance!",
      "votes": null
    },
    {
      "id": "3248524",
      "postDate": "07/14/2025 18:35:27",
      "content": "<p>Hi there,<br>\n1) You can ignore <code>t15_copyTask.pkl</code>, this is only used to recreate some plots from the NEJM paper. Just use <code>t15_copyTask_neuralData.zip</code>.<br>\n2) You are correct that there are 10,948 total trials (8072 train, 1426 val, 1450 test). I mistakenly listed it as 11,300 trials. I will correct this in the README.<br>\nThanks!</p>",
      "rawMarkdown": "Hi there,\n1) You can ignore `t15_copyTask.pkl`, this is only used to recreate some plots from the NEJM paper. Just use `t15_copyTask_neuralData.zip`.\n2) You are correct that there are 10,948 total trials (8072 train, 1426 val, 1450 test). I mistakenly listed it as 11,300 trials. I will correct this in the README.\nThanks!",
      "votes": null
    },
    {
      "id": "3248591",
      "postDate": "07/14/2025 21:04:17",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3248524,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "07/14/2025 18:35:27",
      "content": "<p>Hi there,<br>\n1) You can ignore <code>t15_copyTask.pkl</code>, this is only used to recreate some plots from the NEJM paper. Just use <code>t15_copyTask_neuralData.zip</code>.<br>\n2) You are correct that there are 10,948 total trials (8072 train, 1426 val, 1450 test). I mistakenly listed it as 11,300 trials. I will correct this in the README.<br>\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3248591,
          "author_name": "gorogm",
          "author_url": "",
          "post_date": "07/14/2025 21:04:17",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3247767": "Thanks for this exciting dataset & competition!\n1) Are the sentences in `t15_copyTask.pkl` a subset of the sentences in `t15_copyTask_neuralData.zip`? What is the relation of the trials in the two sets?\n2) In `t15_copyTask_neuralData.zip` I found the 45 sessions as expected, but when I sum the number of trials_XXXes in the .hdf5 files, I count 10948 trials instead of the 11300 trials mentioned in the introduction. What might I miss? Is it possible, that the number of trials is less then 11000, but in some trials there are multiple sentences? In this case this is misleading in the README.txt: \"11,000+ Copy Task trials\". Or should the `t15_copyTask_neuralData` and `t15_copyTask` be combined to reach 11300? - that seems way more then 11.300 trials.\n\nThank in advance!",
    "3248524": "Hi there,\n1) You can ignore `t15_copyTask.pkl`, this is only used to recreate some plots from the NEJM paper. Just use `t15_copyTask_neuralData.zip`.\n2) You are correct that there are 10,948 total trials (8072 train, 1426 val, 1450 test). I mistakenly listed it as 11,300 trials. I will correct this in the README.\nThanks!",
    "3248591": "Thank you!"
  },
  "source": "meta"
}