{
  "id": 609284,
  "title": "What are the seq_class_ids, transcriptions, and seq_len attributes referring to?",
  "url": "/competitions/brain-to-text-25/discussion/609284",
  "author_name": "El Hönan",
  "post_date": "2025-09-25T15:02:11.721000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When using the provided h5py loader:<br>\n````<br>\ndef load_h5py_file(file_path):</p>\n<pre><code>data = {\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n}\n\n#  the hdf5 file for that day\nwith h5py.(file_path, ) as f:\n\n    keys = list(f.keys())\n\n    #  each trial in the selected trials in that day\n    for key in keys:\n        g = f[key]\n\n        neural_features = g[][:]\n        n_time_steps = g.attrs[]\n        seq_class_ids = g[][:] if  in g else \n        seq_len = g.attrs[] if  in g.attrs else \n        transcription = g[][:] if  in g else \n        sentence_label = g.attrs[][:] if  in g.attrs else \n        session = g.attrs[]\n        block_num = g.attrs[]\n        trial_num = g.attrs[]\n\n        # match this trial up with the csv to get the corpus name\n        year, month, day = session.split()[:]\n\n        data[].append(neural_features)\n        data[].append(n_time_steps)\n        data[].append(seq_class_ids)\n        data[].append(seq_len)\n        data[].append(transcription)\n        data[].append(sentence_label)\n        data[].append(session)\n        data[].append(block_num)\n        data[].append(trial_num)\n\nreturn data\n</code></pre>\n<p>My question is what the transcription, seq_len, and seq_class_ids attributes contain?</p>",
  "messages": [
    {
      "id": 3295383,
      "postDate": "2025-09-28T15:27:05.043Z",
      "content": "<p>Regule is correct:</p>\n<ul>\n<li><code>transcription</code> is ACII code representation of sentence_label</li>\n<li><code>seq_class_ids</code> is phoneme indices of the ground truth label, corresponding to:</li>\n</ul>\n<pre><code>LOGIT_TO_PHONEME: Final[[]] = [\n,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , ,\n,\n]\n</code></pre>\n<ul>\n<li><code>seq_len</code> is the number of phoneme labels per trial (<code>seq_class_ids</code> is zero-padded to 500 values).</li>\n</ul>\n<p>I will add some of these details to the data description.</p>",
      "rawMarkdown": "Regule is correct:\n- `transcription` is ACII code representation of sentence_label\n- `seq_class_ids` is phoneme indices of the ground truth label, corresponding to:\n```python\nLOGIT_TO_PHONEME: Final[Sequence[str]] = [\n'BLANK',\n'AA', 'AE', 'AH', 'AO', 'AW',\n'AY', 'B', 'CH', 'D', 'DH',\n'EH', 'ER', 'EY', 'F', 'G',\n'HH', 'IH', 'IY', 'JH', 'K',\n'L', 'M', 'N', 'NG', 'OW',\n'OY', 'P', 'R', 'S', 'SH',\n'T', 'TH', 'UH', 'UW', 'V',\n'W', 'Y', 'Z', 'ZH',\n' | ',\n]\n```\n- `seq_len` is the number of phoneme labels per trial (`seq_class_ids` is zero-padded to 500 values).\n\nI will add some of these details to the data description.",
      "votes": 1,
      "replies": [
        {
          "id": 3296821,
          "postDate": "2025-10-01T17:16:47.920Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 3294193,
      "postDate": "2025-09-25T15:02:11.723Z",
      "content": "<p>When using the provided h5py loader:<br>\n````<br>\ndef load_h5py_file(file_path):</p>\n<pre><code>data = {\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n    : [],\n}\n\n#  the hdf5 file for that day\nwith h5py.(file_path, ) as f:\n\n    keys = list(f.keys())\n\n    #  each trial in the selected trials in that day\n    for key in keys:\n        g = f[key]\n\n        neural_features = g[][:]\n        n_time_steps = g.attrs[]\n        seq_class_ids = g[][:] if  in g else \n        seq_len = g.attrs[] if  in g.attrs else \n        transcription = g[][:] if  in g else \n        sentence_label = g.attrs[][:] if  in g.attrs else \n        session = g.attrs[]\n        block_num = g.attrs[]\n        trial_num = g.attrs[]\n\n        # match this trial up with the csv to get the corpus name\n        year, month, day = session.split()[:]\n\n        data[].append(neural_features)\n        data[].append(n_time_steps)\n        data[].append(seq_class_ids)\n        data[].append(seq_len)\n        data[].append(transcription)\n        data[].append(sentence_label)\n        data[].append(session)\n        data[].append(block_num)\n        data[].append(trial_num)\n\nreturn data\n</code></pre>\n<p>My question is what the transcription, seq_len, and seq_class_ids attributes contain?</p>",
      "rawMarkdown": "When using the provided h5py loader:\n````\ndef load_h5py_file(file_path):\n\n    data = {\n        'neural_features': [],\n        'n_time_steps': [],\n        'seq_class_ids': [],\n        'seq_len': [],\n        'transcriptions': [],\n        'sentence_label': [],\n        'session': [],\n        'block_num': [],\n        'trial_num': [],\n    }\n\n    # Open the hdf5 file for that day\n    with h5py.File(file_path, 'r') as f:\n\n        keys = list(f.keys())\n\n        # For each trial in the selected trials in that day\n        for key in keys:\n            g = f[key]\n\n            neural_features = g['input_features'][:]\n            n_time_steps = g.attrs['n_time_steps']\n            seq_class_ids = g['seq_class_ids'][:] if 'seq_class_ids' in g else None\n            seq_len = g.attrs['seq_len'] if 'seq_len' in g.attrs else None\n            transcription = g['transcription'][:] if 'transcription' in g else None\n            sentence_label = g.attrs['sentence_label'][:] if 'sentence_label' in g.attrs else None\n            session = g.attrs['session']\n            block_num = g.attrs['block_num']\n            trial_num = g.attrs['trial_num']\n\n            # match this trial up with the csv to get the corpus name\n            year, month, day = session.split('.')[1:]\n\n            data['neural_features'].append(neural_features)\n            data['n_time_steps'].append(n_time_steps)\n            data['seq_class_ids'].append(seq_class_ids)\n            data['seq_len'].append(seq_len)\n            data['transcriptions'].append(transcription)\n            data['sentence_label'].append(sentence_label)\n            data['session'].append(session)\n            data['block_num'].append(block_num)\n            data['trial_num'].append(trial_num)\n\n    return data\n\n\nMy question is what the transcription, seq_len, and seq_class_ids attributes contain?\n",
      "votes": 1
    },
    {
      "id": 3295264,
      "postDate": "2025-09-28T08:21:42.717Z",
      "content": "<p>I also think that this is not explained very clearly. After tinkering  a bit I found out what is what.<br>\nFirst of all seq_len is just length of non-zero sequence stored in seq_class_ids, this is probably related to <br>\nhow data was generated as seq_class_ids is always 500 elements long. <br>\nNon zero values stored in seq_class_ids contain codes of phonemes so that if you run </p>\n<blockquote>\n  <p>LOGIT_TO_PHONEME: Final[Sequence[str]] = [<br>\n     'BLANK',<br>\n     'AA', 'AE', 'AH', 'AO', 'AW',<br>\n     'AY', 'B',  'CH', 'D', 'DH',<br>\n     'EH', 'ER', 'EY', 'F', 'G',<br>\n     'HH', 'IH', 'IY', 'JH', 'K',<br>\n     'L', 'M', 'N', 'NG', 'OW',<br>\n     'OY', 'P', 'R', 'S', 'SH',<br>\n     'T', 'TH', 'UH', 'UW', 'V',<br>\n     'W', 'Y', 'Z', 'ZH',<br>\n     ' | ',<br>\n  ]<br>\n  class_seq = np.trim_zeros(class_seq)<br>\n  print(f'Class sequence ({class_seq.shape}) = {[LOGIT_TO_PHONEME[c] for c in class_seq]}')</p>\n</blockquote>\n<p>assuming that you assigned seq_class_ids from specific trial to variable class_seq  you will get something like </p>\n<blockquote>\n  <p>Class sequence ((18,)) = [ 1 28 40 37 34 40  9  3 23 40 36 17 10 40 10 17 29 40]<br>\n  Class sequence ((18,)) = ['AA', 'R', ' | ', 'Y', 'UW', ' | ', 'D', 'AH', 'N', ' | ', 'W', 'IH', 'DH', ' | ', 'DH', 'IH', 'S', ' | ']</p>\n</blockquote>\n<p>As for transcriptions this is simply an ACII code representation of sentence_label, you can see that length of its non-zero values corresponds to length of sentence_label. <br>\nAs this is also always 500 element long I'm sure that this is also related to how data was collected as in Python<br>\nthis entry is unnecessary as you can easily generate an ASCII sequence from string. </p>",
      "rawMarkdown": "I also think that this is not explained very clearly. After tinkering  a bit I found out what is what.\nFirst of all seq_len is just length of non-zero sequence stored in seq_class_ids, this is probably related to \nhow data was generated as seq_class_ids is always 500 elements long. \nNon zero values stored in seq_class_ids contain codes of phonemes so that if you run \n\n>LOGIT_TO_PHONEME: Final[Sequence[str]] = [\n>    'BLANK',\n>    'AA', 'AE', 'AH', 'AO', 'AW',\n>    'AY', 'B',  'CH', 'D', 'DH',\n>    'EH', 'ER', 'EY', 'F', 'G',\n>    'HH', 'IH', 'IY', 'JH', 'K',\n>    'L', 'M', 'N', 'NG', 'OW',\n>    'OY', 'P', 'R', 'S', 'SH',\n>    'T', 'TH', 'UH', 'UW', 'V',\n>    'W', 'Y', 'Z', 'ZH',\n>    ' | ',\n>]\n>class_seq = np.trim_zeros(class_seq)\n>print(f'Class sequence ({class_seq.shape}) = {[LOGIT_TO_PHONEME[c] for c in class_seq]}')\n\n assuming that you assigned seq_class_ids from specific trial to variable class_seq  you will get something like \n>Class sequence ((18,)) = [ 1 28 40 37 34 40  9  3 23 40 36 17 10 40 10 17 29 40]\n>Class sequence ((18,)) = ['AA', 'R', ' | ', 'Y', 'UW', ' | ', 'D', 'AH', 'N', ' | ', 'W', 'IH', 'DH', ' | ', 'DH', 'IH', 'S', ' | ']\n\nAs for transcriptions this is simply an ACII code representation of sentence_label, you can see that length of its non-zero values corresponds to length of sentence_label. \nAs this is also always 500 element long I'm sure that this is also related to how data was collected as in Python\nthis entry is unnecessary as you can easily generate an ASCII sequence from string. \n\n",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 3296820,
          "postDate": "2025-10-01T17:16:42.727Z",
          "content": "<p>This was very helpful, thank you!</p>",
          "rawMarkdown": "This was very helpful, thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3295383,
      "author_name": "Nick Card",
      "author_url": "",
      "post_date": "2025-09-28T15:27:05.043000",
      "content": "<p>Regule is correct:</p>\n<ul>\n<li><code>transcription</code> is ACII code representation of sentence_label</li>\n<li><code>seq_class_ids</code> is phoneme indices of the ground truth label, corresponding to:</li>\n</ul>\n<pre><code>LOGIT_TO_PHONEME: Final[[]] = [\n,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , , ,\n, , , ,\n,\n]\n</code></pre>\n<ul>\n<li><code>seq_len</code> is the number of phoneme labels per trial (<code>seq_class_ids</code> is zero-padded to 500 values).</li>\n</ul>\n<p>I will add some of these details to the data description.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3296821,
          "author_name": "El Hönan",
          "author_url": "",
          "post_date": "2025-10-01T17:16:47.920000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3295264,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-09-28T08:21:42.717000",
      "content": "<p>I also think that this is not explained very clearly. After tinkering  a bit I found out what is what.<br>\nFirst of all seq_len is just length of non-zero sequence stored in seq_class_ids, this is probably related to <br>\nhow data was generated as seq_class_ids is always 500 elements long. <br>\nNon zero values stored in seq_class_ids contain codes of phonemes so that if you run </p>\n<blockquote>\n  <p>LOGIT_TO_PHONEME: Final[Sequence[str]] = [<br>\n     'BLANK',<br>\n     'AA', 'AE', 'AH', 'AO', 'AW',<br>\n     'AY', 'B',  'CH', 'D', 'DH',<br>\n     'EH', 'ER', 'EY', 'F', 'G',<br>\n     'HH', 'IH', 'IY', 'JH', 'K',<br>\n     'L', 'M', 'N', 'NG', 'OW',<br>\n     'OY', 'P', 'R', 'S', 'SH',<br>\n     'T', 'TH', 'UH', 'UW', 'V',<br>\n     'W', 'Y', 'Z', 'ZH',<br>\n     ' | ',<br>\n  ]<br>\n  class_seq = np.trim_zeros(class_seq)<br>\n  print(f'Class sequence ({class_seq.shape}) = {[LOGIT_TO_PHONEME[c] for c in class_seq]}')</p>\n</blockquote>\n<p>assuming that you assigned seq_class_ids from specific trial to variable class_seq  you will get something like </p>\n<blockquote>\n  <p>Class sequence ((18,)) = [ 1 28 40 37 34 40  9  3 23 40 36 17 10 40 10 17 29 40]<br>\n  Class sequence ((18,)) = ['AA', 'R', ' | ', 'Y', 'UW', ' | ', 'D', 'AH', 'N', ' | ', 'W', 'IH', 'DH', ' | ', 'DH', 'IH', 'S', ' | ']</p>\n</blockquote>\n<p>As for transcriptions this is simply an ACII code representation of sentence_label, you can see that length of its non-zero values corresponds to length of sentence_label. <br>\nAs this is also always 500 element long I'm sure that this is also related to how data was collected as in Python<br>\nthis entry is unnecessary as you can easily generate an ASCII sequence from string. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3296820,
          "author_name": "El Hönan",
          "author_url": "",
          "post_date": "2025-10-01T17:16:42.727000",
          "content": "<p>This was very helpful, thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3295383": "Regule is correct:\n- `transcription` is ACII code representation of sentence_label\n- `seq_class_ids` is phoneme indices of the ground truth label, corresponding to:\n```python\nLOGIT_TO_PHONEME: Final[Sequence[str]] = [\n'BLANK',\n'AA', 'AE', 'AH', 'AO', 'AW',\n'AY', 'B', 'CH', 'D', 'DH',\n'EH', 'ER', 'EY', 'F', 'G',\n'HH', 'IH', 'IY', 'JH', 'K',\n'L', 'M', 'N', 'NG', 'OW',\n'OY', 'P', 'R', 'S', 'SH',\n'T', 'TH', 'UH', 'UW', 'V',\n'W', 'Y', 'Z', 'ZH',\n' | ',\n]\n```\n- `seq_len` is the number of phoneme labels per trial (`seq_class_ids` is zero-padded to 500 values).\n\nI will add some of these details to the data description.",
    "3294193": "When using the provided h5py loader:\n````\ndef load_h5py_file(file_path):\n\n    data = {\n        'neural_features': [],\n        'n_time_steps': [],\n        'seq_class_ids': [],\n        'seq_len': [],\n        'transcriptions': [],\n        'sentence_label': [],\n        'session': [],\n        'block_num': [],\n        'trial_num': [],\n    }\n\n    # Open the hdf5 file for that day\n    with h5py.File(file_path, 'r') as f:\n\n        keys = list(f.keys())\n\n        # For each trial in the selected trials in that day\n        for key in keys:\n            g = f[key]\n\n            neural_features = g['input_features'][:]\n            n_time_steps = g.attrs['n_time_steps']\n            seq_class_ids = g['seq_class_ids'][:] if 'seq_class_ids' in g else None\n            seq_len = g.attrs['seq_len'] if 'seq_len' in g.attrs else None\n            transcription = g['transcription'][:] if 'transcription' in g else None\n            sentence_label = g.attrs['sentence_label'][:] if 'sentence_label' in g.attrs else None\n            session = g.attrs['session']\n            block_num = g.attrs['block_num']\n            trial_num = g.attrs['trial_num']\n\n            # match this trial up with the csv to get the corpus name\n            year, month, day = session.split('.')[1:]\n\n            data['neural_features'].append(neural_features)\n            data['n_time_steps'].append(n_time_steps)\n            data['seq_class_ids'].append(seq_class_ids)\n            data['seq_len'].append(seq_len)\n            data['transcriptions'].append(transcription)\n            data['sentence_label'].append(sentence_label)\n            data['session'].append(session)\n            data['block_num'].append(block_num)\n            data['trial_num'].append(trial_num)\n\n    return data\n\n\nMy question is what the transcription, seq_len, and seq_class_ids attributes contain?\n",
    "3295264": "I also think that this is not explained very clearly. After tinkering  a bit I found out what is what.\nFirst of all seq_len is just length of non-zero sequence stored in seq_class_ids, this is probably related to \nhow data was generated as seq_class_ids is always 500 elements long. \nNon zero values stored in seq_class_ids contain codes of phonemes so that if you run \n\n>LOGIT_TO_PHONEME: Final[Sequence[str]] = [\n>    'BLANK',\n>    'AA', 'AE', 'AH', 'AO', 'AW',\n>    'AY', 'B',  'CH', 'D', 'DH',\n>    'EH', 'ER', 'EY', 'F', 'G',\n>    'HH', 'IH', 'IY', 'JH', 'K',\n>    'L', 'M', 'N', 'NG', 'OW',\n>    'OY', 'P', 'R', 'S', 'SH',\n>    'T', 'TH', 'UH', 'UW', 'V',\n>    'W', 'Y', 'Z', 'ZH',\n>    ' | ',\n>]\n>class_seq = np.trim_zeros(class_seq)\n>print(f'Class sequence ({class_seq.shape}) = {[LOGIT_TO_PHONEME[c] for c in class_seq]}')\n\n assuming that you assigned seq_class_ids from specific trial to variable class_seq  you will get something like \n>Class sequence ((18,)) = [ 1 28 40 37 34 40  9  3 23 40 36 17 10 40 10 17 29 40]\n>Class sequence ((18,)) = ['AA', 'R', ' | ', 'Y', 'UW', ' | ', 'D', 'AH', 'N', ' | ', 'W', 'IH', 'DH', ' | ', 'DH', 'IH', 'S', ' | ']\n\nAs for transcriptions this is simply an ACII code representation of sentence_label, you can see that length of its non-zero values corresponds to length of sentence_label. \nAs this is also always 500 element long I'm sure that this is also related to how data was collected as in Python\nthis entry is unnecessary as you can easily generate an ASCII sequence from string. \n\n"
  }
}