{
  "id": 596283,
  "title": "Question about the dataset columns",
  "url": "/competitions/brain-to-text-25/discussion/596283",
  "author_name": "",
  "post_date": "2025-08-02T17:15:30.527624Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello, <br>\nI wanted to ask what data the 'transcriptions' and 'seq_class_ids' columns represent? I initially thought that one column represented the value of each word in the sentence and the other column represented each letter in the sentence, but this doesn't work out in terms of the length of the arrays and the values in the arrays. I would be happy to receive an answer, thanks in advance :)</p>",
  "messages": [
    {
      "id": "3261988",
      "postDate": "08/02/2025 17:15:30",
      "content": "<p>Hello, <br>\nI wanted to ask what data the 'transcriptions' and 'seq_class_ids' columns represent? I initially thought that one column represented the value of each word in the sentence and the other column represented each letter in the sentence, but this doesn't work out in terms of the length of the arrays and the values in the arrays. I would be happy to receive an answer, thanks in advance :)</p>",
      "rawMarkdown": "Hello, \nI wanted to ask what data the 'transcriptions' and 'seq_class_ids' columns represent? I initially thought that one column represented the value of each word in the sentence and the other column represented each letter in the sentence, but this doesn't work out in terms of the length of the arrays and the values in the arrays. I would be happy to receive an answer, thanks in advance :)",
      "votes": null
    },
    {
      "id": "3263786",
      "postDate": "08/05/2025 17:56:17",
      "content": "<p>Hi, this information is available in the \"Data format\" section of the \"Data\" tab on this Kaggle competition. There's also some example python code there to load the files into a python dictionary. After doing so, here is a description of what each dictionary key is:</p>\n<ul>\n<li>neural_features: list of neural data for a given trial, of shape TxF, where F is the number of neural features (512) and T is the number of 20 ms time steps in this trial</li>\n<li>n_time_steps: number of 10 ms time steps in each trial</li>\n<li>seq_class_ids: phoneme indexes for each trial's phoneme label. Numbers correspond to the phoneme index in this list:</li>\n</ul>\n<pre><code>LOGIT_TO_PHONEME = [\n    ,\n    , , , , ,\n    , ,  , , ,\n    , , , , ,\n    , , , , ,\n    , , , , ,\n    , , , , ,\n    , , , , ,\n    , , , ,\n    ,\n]\n</code></pre>\n<ul>\n<li>seq_len: number of phonemes for each trial</li>\n<li>transcriptions: character indexes for the sentence labels of each trial</li>\n<li>sentence_label: plain text sentence label for each trial</li>\n<li>session: session date that this trial was collected on</li>\n<li>block_num: block number for each trial</li>\n<li>trial_num: trial number for each trial</li>\n</ul>",
      "rawMarkdown": "Hi, this information is available in the \"Data format\" section of the \"Data\" tab on this Kaggle competition. There's also some example python code there to load the files into a python dictionary. After doing so, here is a description of what each dictionary key is:\n\n- neural_features: list of neural data for a given trial, of shape TxF, where F is the number of neural features (512) and T is the number of 20 ms time steps in this trial\n- n_time_steps: number of 10 ms time steps in each trial\n- seq_class_ids: phoneme indexes for each trial's phoneme label. Numbers correspond to the phoneme index in this list:\n```python\nLOGIT_TO_PHONEME = [\n    'BLANK',\n    'AA', 'AE', 'AH', 'AO', 'AW',\n    'AY', 'B',  'CH', 'D', 'DH',\n    'EH', 'ER', 'EY', 'F', 'G',\n    'HH', 'IH', 'IY', 'JH', 'K',\n    'L', 'M', 'N', 'NG', 'OW',\n    'OY', 'P', 'R', 'S', 'SH',\n    'T', 'TH', 'UH', 'UW', 'V',\n    'W', 'Y', 'Z', 'ZH',\n    ' | ',\n]\n```\n- seq_len: number of phonemes for each trial\n- transcriptions: character indexes for the sentence labels of each trial\n- sentence_label: plain text sentence label for each trial\n- session: session date that this trial was collected on\n- block_num: block number for each trial\n- trial_num: trial number for each trial",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3263786,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "08/05/2025 17:56:17",
      "content": "<p>Hi, this information is available in the \"Data format\" section of the \"Data\" tab on this Kaggle competition. There's also some example python code there to load the files into a python dictionary. After doing so, here is a description of what each dictionary key is:</p>\n<ul>\n<li>neural_features: list of neural data for a given trial, of shape TxF, where F is the number of neural features (512) and T is the number of 20 ms time steps in this trial</li>\n<li>n_time_steps: number of 10 ms time steps in each trial</li>\n<li>seq_class_ids: phoneme indexes for each trial's phoneme label. Numbers correspond to the phoneme index in this list:</li>\n</ul>\n<pre><code>LOGIT_TO_PHONEME = [\n    ,\n    , , , , ,\n    , ,  , , ,\n    , , , , ,\n    , , , , ,\n    , , , , ,\n    , , , , ,\n    , , , , ,\n    , , , ,\n    ,\n]\n</code></pre>\n<ul>\n<li>seq_len: number of phonemes for each trial</li>\n<li>transcriptions: character indexes for the sentence labels of each trial</li>\n<li>sentence_label: plain text sentence label for each trial</li>\n<li>session: session date that this trial was collected on</li>\n<li>block_num: block number for each trial</li>\n<li>trial_num: trial number for each trial</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3261988": "Hello, \nI wanted to ask what data the 'transcriptions' and 'seq_class_ids' columns represent? I initially thought that one column represented the value of each word in the sentence and the other column represented each letter in the sentence, but this doesn't work out in terms of the length of the arrays and the values in the arrays. I would be happy to receive an answer, thanks in advance :)",
    "3263786": "Hi, this information is available in the \"Data format\" section of the \"Data\" tab on this Kaggle competition. There's also some example python code there to load the files into a python dictionary. After doing so, here is a description of what each dictionary key is:\n\n- neural_features: list of neural data for a given trial, of shape TxF, where F is the number of neural features (512) and T is the number of 20 ms time steps in this trial\n- n_time_steps: number of 10 ms time steps in each trial\n- seq_class_ids: phoneme indexes for each trial's phoneme label. Numbers correspond to the phoneme index in this list:\n```python\nLOGIT_TO_PHONEME = [\n    'BLANK',\n    'AA', 'AE', 'AH', 'AO', 'AW',\n    'AY', 'B',  'CH', 'D', 'DH',\n    'EH', 'ER', 'EY', 'F', 'G',\n    'HH', 'IH', 'IY', 'JH', 'K',\n    'L', 'M', 'N', 'NG', 'OW',\n    'OY', 'P', 'R', 'S', 'SH',\n    'T', 'TH', 'UH', 'UW', 'V',\n    'W', 'Y', 'Z', 'ZH',\n    ' | ',\n]\n```\n- seq_len: number of phonemes for each trial\n- transcriptions: character indexes for the sentence labels of each trial\n- sentence_label: plain text sentence label for each trial\n- session: session date that this trial was collected on\n- block_num: block number for each trial\n- trial_num: trial number for each trial"
  },
  "source": "meta"
}