{
  "id": 589934,
  "title": "Some EDA on the Competition Data",
  "url": "/competitions/brain-to-text-25/discussion/589934",
  "author_name": "",
  "post_date": "2025-07-16T14:32:48.168280700Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am sharing a notebook with some EDA on the Brain-to-text '25 data competition. The data is available <a href=\"https://www.kaggle.com/datasets/paulnkamau/brain-to-text-25-data\" target=\"_blank\">here</a>. Check out the notebook, <a href=\"https://www.kaggle.com/code/paulnkamau/some-eda-brain-to-text-25\" target=\"_blank\">Some EDA _ Brain-to-Text '25</a></p>",
  "messages": [
    {
      "id": "3249465",
      "postDate": "07/16/2025 14:32:48",
      "content": "<p>I am sharing a notebook with some EDA on the Brain-to-text '25 data competition. The data is available <a href=\"https://www.kaggle.com/datasets/paulnkamau/brain-to-text-25-data\" target=\"_blank\">here</a>. Check out the notebook, <a href=\"https://www.kaggle.com/code/paulnkamau/some-eda-brain-to-text-25\" target=\"_blank\">Some EDA _ Brain-to-Text '25</a></p>",
      "rawMarkdown": "I am sharing a notebook with some EDA on the Brain-to-text '25 data competition. The data is available [here](https://www.kaggle.com/datasets/paulnkamau/brain-to-text-25-data). Check out the notebook, [Some EDA _ Brain-to-Text '25](https://www.kaggle.com/code/paulnkamau/some-eda-brain-to-text-25)",
      "votes": null
    },
    {
      "id": "3249536",
      "postDate": "07/16/2025 16:46:10",
      "content": "<p>Hi, thanks for this. Just two notes:</p>\n<ol>\n<li>The dataset is now hosted on this Kaggle competition page as well (but thanks for reuploading it)</li>\n<li>There are additional data properties in the .hdf5 files that you do not appear to be loading in your code, so I just wanted to point that out. For example, you are manually decoding the transcriptions back into sentences, but there is also a data field that just has that sentence label in it. A full data loading example can be found on our GitHub <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29\" target=\"_blank\">here</a>, and pasted below:</li>\n</ol>\n<pre><code> h5py\n\n ():\n    data = {\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n    }\n    \n     h5py.File(file_path, )  f:\n\n        keys = (f.keys())\n\n        \n         key  keys:\n            g = f[key]\n\n            neural_features = g[][:]\n            n_time_steps = g.attrs[]\n            seq_class_ids = g[][:]    g  \n            seq_len = g.attrs[]    g.attrs  \n            transcription = g[][:]    g  \n            sentence_label = g.attrs[][:]    g.attrs  \n            session = g.attrs[]\n            block_num = g.attrs[]\n            trial_num = g.attrs[]\n\n            data[].append(neural_features)\n            data[].append(n_time_steps)\n            data[].append(seq_class_ids)\n            data[].append(seq_len)\n            data[].append(transcription)\n            data[].append(sentence_label)\n            data[].append(session)\n            data[].append(block_num)\n            data[].append(trial_num)\n     data\n</code></pre>",
      "rawMarkdown": "Hi, thanks for this. Just two notes:\n1. The dataset is now hosted on this Kaggle competition page as well (but thanks for reuploading it)\n2. There are additional data properties in the .hdf5 files that you do not appear to be loading in your code, so I just wanted to point that out. For example, you are manually decoding the transcriptions back into sentences, but there is also a data field that just has that sentence label in it. A full data loading example can be found on our GitHub [here](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29), and pasted below:\n```python\nimport h5py\n\ndef load_h5py_file(file_path):\n    data = {\n        'neural_features': [],\n        'n_time_steps': [],\n        'seq_class_ids': [],\n        'seq_len': [],\n        'transcriptions': [],\n        'sentence_label': [],\n        'session': [],\n        'block_num': [],\n        'trial_num': [],\n    }\n    # Open the hdf5 file for that day\n    with h5py.File(file_path, 'r') as f:\n\n        keys = list(f.keys())\n\n        # For each trial in the selected trials in that day\n        for key in keys:\n            g = f[key]\n\n            neural_features = g['input_features'][:]\n            n_time_steps = g.attrs['n_time_steps']\n            seq_class_ids = g['seq_class_ids'][:] if 'seq_class_ids' in g else None\n            seq_len = g.attrs['seq_len'] if 'seq_len' in g.attrs else None\n            transcription = g['transcription'][:] if 'transcription' in g else None\n            sentence_label = g.attrs['sentence_label'][:] if 'sentence_label' in g.attrs else None\n            session = g.attrs['session']\n            block_num = g.attrs['block_num']\n            trial_num = g.attrs['trial_num']\n\n            data['neural_features'].append(neural_features)\n            data['n_time_steps'].append(n_time_steps)\n            data['seq_class_ids'].append(seq_class_ids)\n            data['seq_len'].append(seq_len)\n            data['transcriptions'].append(transcription)\n            data['sentence_label'].append(sentence_label)\n            data['session'].append(session)\n            data['block_num'].append(block_num)\n            data['trial_num'].append(trial_num)\n    return data\n```",
      "votes": null
    },
    {
      "id": "3249548",
      "postDate": "07/16/2025 17:02:17",
      "content": "<p>Hi Nick, thanks for the feedback. I'll update the notebook. Appreciate the correction.</p>",
      "rawMarkdown": "Hi Nick, thanks for the feedback. I'll update the notebook. Appreciate the correction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3249536,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "07/16/2025 16:46:10",
      "content": "<p>Hi, thanks for this. Just two notes:</p>\n<ol>\n<li>The dataset is now hosted on this Kaggle competition page as well (but thanks for reuploading it)</li>\n<li>There are additional data properties in the .hdf5 files that you do not appear to be loading in your code, so I just wanted to point that out. For example, you are manually decoding the transcriptions back into sentences, but there is also a data field that just has that sentence label in it. A full data loading example can be found on our GitHub <a href=\"https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29\" target=\"_blank\">here</a>, and pasted below:</li>\n</ol>\n<pre><code> h5py\n\n ():\n    data = {\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n        : [],\n    }\n    \n     h5py.File(file_path, )  f:\n\n        keys = (f.keys())\n\n        \n         key  keys:\n            g = f[key]\n\n            neural_features = g[][:]\n            n_time_steps = g.attrs[]\n            seq_class_ids = g[][:]    g  \n            seq_len = g.attrs[]    g.attrs  \n            transcription = g[][:]    g  \n            sentence_label = g.attrs[][:]    g.attrs  \n            session = g.attrs[]\n            block_num = g.attrs[]\n            trial_num = g.attrs[]\n\n            data[].append(neural_features)\n            data[].append(n_time_steps)\n            data[].append(seq_class_ids)\n            data[].append(seq_len)\n            data[].append(transcription)\n            data[].append(sentence_label)\n            data[].append(session)\n            data[].append(block_num)\n            data[].append(trial_num)\n     data\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3249548,
          "author_name": "paulnkamau",
          "author_url": "",
          "post_date": "07/16/2025 17:02:17",
          "content": "<p>Hi Nick, thanks for the feedback. I'll update the notebook. Appreciate the correction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3249465": "I am sharing a notebook with some EDA on the Brain-to-text '25 data competition. The data is available [here](https://www.kaggle.com/datasets/paulnkamau/brain-to-text-25-data). Check out the notebook, [Some EDA _ Brain-to-Text '25](https://www.kaggle.com/code/paulnkamau/some-eda-brain-to-text-25)",
    "3249536": "Hi, thanks for this. Just two notes:\n1. The dataset is now hosted on this Kaggle competition page as well (but thanks for reuploading it)\n2. There are additional data properties in the .hdf5 files that you do not appear to be loading in your code, so I just wanted to point that out. For example, you are manually decoding the transcriptions back into sentences, but there is also a data field that just has that sentence label in it. A full data loading example can be found on our GitHub [here](https://github.com/Neuroprosthetics-Lab/nejm-brain-to-text/blob/main/model_training/evaluate_model_helpers.py#L29), and pasted below:\n```python\nimport h5py\n\ndef load_h5py_file(file_path):\n    data = {\n        'neural_features': [],\n        'n_time_steps': [],\n        'seq_class_ids': [],\n        'seq_len': [],\n        'transcriptions': [],\n        'sentence_label': [],\n        'session': [],\n        'block_num': [],\n        'trial_num': [],\n    }\n    # Open the hdf5 file for that day\n    with h5py.File(file_path, 'r') as f:\n\n        keys = list(f.keys())\n\n        # For each trial in the selected trials in that day\n        for key in keys:\n            g = f[key]\n\n            neural_features = g['input_features'][:]\n            n_time_steps = g.attrs['n_time_steps']\n            seq_class_ids = g['seq_class_ids'][:] if 'seq_class_ids' in g else None\n            seq_len = g.attrs['seq_len'] if 'seq_len' in g.attrs else None\n            transcription = g['transcription'][:] if 'transcription' in g else None\n            sentence_label = g.attrs['sentence_label'][:] if 'sentence_label' in g.attrs else None\n            session = g.attrs['session']\n            block_num = g.attrs['block_num']\n            trial_num = g.attrs['trial_num']\n\n            data['neural_features'].append(neural_features)\n            data['n_time_steps'].append(n_time_steps)\n            data['seq_class_ids'].append(seq_class_ids)\n            data['seq_len'].append(seq_len)\n            data['transcriptions'].append(transcription)\n            data['sentence_label'].append(sentence_label)\n            data['session'].append(session)\n            data['block_num'].append(block_num)\n            data['trial_num'].append(trial_num)\n    return data\n```",
    "3249548": "Hi Nick, thanks for the feedback. I'll update the notebook. Appreciate the correction."
  },
  "source": "meta"
}