{
  "id": 589515,
  "title": "How to match Corpus to Trials?",
  "url": "/competitions/brain-to-text-25/discussion/589515",
  "author_name": "Görög Márton",
  "post_date": "2025-07-13T12:47:17.952000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p><code>t15_copyTaskData_description.csv</code> tells us that during a day (therefore in a single .hdf5 file) there might be sentences with different corpuses (eg. 2023-08-13). But <code>t15_copyTaskData_description.csv</code> does not tell us which trials belong to which block, and the .hdf5 files also doesn't seem to contain block border information. One trial seems to contain multiple sentences in some cases, therefore the \"Number of sentences\" column in the .csv is not easy to translate to trials in the .hdf5 files. Is there a way to match trials and blocks, or trials and corpus information?</p>",
  "messages": [
    {
      "id": 3248528,
      "postDate": "2025-07-14T18:42:32.357Z",
      "content": "<p>Hi,</p>\n<p>To clarify, data is organized as so:</p>\n<ul>\n<li>Each \"session\" (i.e., date) of data has multiple \"blocks\".</li>\n<li>Each \"block\" has multiple \"trials\", all from the same corpus.</li>\n<li>Each \"trial\" is a single sentence.</li>\n<li>So, given the above points, each trial can be described by it's session date, block number, and trial number (e.g., t15.2023.08.13, block 9, trial 4 is from the Switchboard block).</li>\n</ul>\n<p>When you use the example <code>load_h5py_file()</code> <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29\" target=\"_blank\">function from the GitHub repo</a> it will return the session name, block number, and trial number for each trial. You can then use that information to match up with the <code>t15_copyTaskData_description.csv</code> file to figure out the corpus for each trial. I will look into adding example code for this in the GitHub repo.</p>",
      "rawMarkdown": "Hi,\n\nTo clarify, data is organized as so:\n- Each \"session\" (i.e., date) of data has multiple \"blocks\".\n- Each \"block\" has multiple \"trials\", all from the same corpus.\n- Each \"trial\" is a single sentence.\n- So, given the above points, each trial can be described by it's session date, block number, and trial number (e.g., t15.2023.08.13, block 9, trial 4 is from the Switchboard block).\n\nWhen you use the example `load_h5py_file()` [function from the GitHub repo](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29) it will return the session name, block number, and trial number for each trial. You can then use that information to match up with the `t15_copyTaskData_description.csv` file to figure out the corpus for each trial. I will look into adding example code for this in the GitHub repo.",
      "replies": [
        {
          "id": 3248587,
          "postDate": "2025-07-14T21:00:01.240Z",
          "content": "<p>I've now updated the example load_h5py_file() <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29\" target=\"_blank\">function from the GitHub repo</a> so that it also returns the corpus for each trial.</p>",
          "rawMarkdown": "I've now updated the example load_h5py_file() [function from the GitHub repo](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29) so that it also returns the corpus for each trial."
        },
        {
          "id": 3248588,
          "postDate": "2025-07-14T21:01:36.677Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3248589,
          "postDate": "2025-07-14T21:02:31.400Z",
          "content": "<p>Thank you for pointing to the <code>load_h5py_file()</code> ! I did not discover the 'attrs' attributes previously by myself.</p>",
          "rawMarkdown": "Thank you for pointing to the `load_h5py_file()` ! I did not discover the 'attrs' attributes previously by myself."
        }
      ]
    },
    {
      "id": 3247792,
      "postDate": "2025-07-13T12:47:17.953Z",
      "content": "<p><code>t15_copyTaskData_description.csv</code> tells us that during a day (therefore in a single .hdf5 file) there might be sentences with different corpuses (eg. 2023-08-13). But <code>t15_copyTaskData_description.csv</code> does not tell us which trials belong to which block, and the .hdf5 files also doesn't seem to contain block border information. One trial seems to contain multiple sentences in some cases, therefore the \"Number of sentences\" column in the .csv is not easy to translate to trials in the .hdf5 files. Is there a way to match trials and blocks, or trials and corpus information?</p>",
      "rawMarkdown": "`t15_copyTaskData_description.csv` tells us that during a day (therefore in a single .hdf5 file) there might be sentences with different corpuses (eg. 2023-08-13). But `t15_copyTaskData_description.csv` does not tell us which trials belong to which block, and the .hdf5 files also doesn't seem to contain block border information. One trial seems to contain multiple sentences in some cases, therefore the \"Number of sentences\" column in the .csv is not easy to translate to trials in the .hdf5 files. Is there a way to match trials and blocks, or trials and corpus information?"
    }
  ],
  "comments": [
    {
      "id": 3248528,
      "author_name": "Nick Card",
      "author_url": "",
      "post_date": "2025-07-14T18:42:32.357000",
      "content": "<p>Hi,</p>\n<p>To clarify, data is organized as so:</p>\n<ul>\n<li>Each \"session\" (i.e., date) of data has multiple \"blocks\".</li>\n<li>Each \"block\" has multiple \"trials\", all from the same corpus.</li>\n<li>Each \"trial\" is a single sentence.</li>\n<li>So, given the above points, each trial can be described by it's session date, block number, and trial number (e.g., t15.2023.08.13, block 9, trial 4 is from the Switchboard block).</li>\n</ul>\n<p>When you use the example <code>load_h5py_file()</code> <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29\" target=\"_blank\">function from the GitHub repo</a> it will return the session name, block number, and trial number for each trial. You can then use that information to match up with the <code>t15_copyTaskData_description.csv</code> file to figure out the corpus for each trial. I will look into adding example code for this in the GitHub repo.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3248587,
          "author_name": "Nick Card",
          "author_url": "",
          "post_date": "2025-07-14T21:00:01.240000",
          "content": "<p>I've now updated the example load_h5py_file() <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29\" target=\"_blank\">function from the GitHub repo</a> so that it also returns the corpus for each trial.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3248588,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-07-14T21:01:36.677000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3248589,
          "author_name": "Görög Márton",
          "author_url": "",
          "post_date": "2025-07-14T21:02:31.400000",
          "content": "<p>Thank you for pointing to the <code>load_h5py_file()</code> ! I did not discover the 'attrs' attributes previously by myself.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3248528": "Hi,\n\nTo clarify, data is organized as so:\n- Each \"session\" (i.e., date) of data has multiple \"blocks\".\n- Each \"block\" has multiple \"trials\", all from the same corpus.\n- Each \"trial\" is a single sentence.\n- So, given the above points, each trial can be described by it's session date, block number, and trial number (e.g., t15.2023.08.13, block 9, trial 4 is from the Switchboard block).\n\nWhen you use the example `load_h5py_file()` [function from the GitHub repo](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29) it will return the session name, block number, and trial number for each trial. You can then use that information to match up with the `t15_copyTaskData_description.csv` file to figure out the corpus for each trial. I will look into adding example code for this in the GitHub repo.",
    "3247792": "`t15_copyTaskData_description.csv` tells us that during a day (therefore in a single .hdf5 file) there might be sentences with different corpuses (eg. 2023-08-13). But `t15_copyTaskData_description.csv` does not tell us which trials belong to which block, and the .hdf5 files also doesn't seem to contain block border information. One trial seems to contain multiple sentences in some cases, therefore the \"Number of sentences\" column in the .csv is not easy to translate to trials in the .hdf5 files. Is there a way to match trials and blocks, or trials and corpus information?"
  }
}