{
  "id": 596460,
  "title": "Neural Data Preprocessing",
  "url": "/competitions/brain-to-text-25/discussion/596460",
  "author_name": "",
  "post_date": "2025-08-03T18:50:31.469378800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, my name is Krishna. I had a question about feature engineering and preprocessing of the neural data. I'd also like to share a result if its helpful to understand the correlation between the neural data and sentences.</p>\n<p>Questions: Is it necessary to perform any normalization/preprocessing to the neural data? Moreso, would any processing help in model predictions? I see the iEEG features are quite small, but was wondering if any feaure engineering should be applied such as noise reduction.</p>\n<p>Using linear interpolation, I aligned sentences and neural features over time. When looking at all neural fature groups (64), similar words yield similar neural features by using cosine similarity between BERT embeddings of word tokens and neural features.</p>\n<p>Using a linear regression ML model using sklearn, I used subsets of the neural features to predict BERT embeddings of whole sentences so see if there was any more important feature set. When doing this, r^2 for each subset was almost the same at 0.18, while using all features was 0.84.</p>\n<p><strong>Key takeaways:</strong></p>\n<ul>\n<li>Similar word semantic structure yield similar neural features</li>\n<li>All feature subsets [n, n+64] are almost equal in importance and do correlate linearly to semantic embeddings.</li>\n</ul>\n<p>Hope this helps!</p>",
  "messages": [
    {
      "id": "3262514",
      "postDate": "08/03/2025 18:50:31",
      "content": "<p>Hello, my name is Krishna. I had a question about feature engineering and preprocessing of the neural data. I'd also like to share a result if its helpful to understand the correlation between the neural data and sentences.</p>\n<p>Questions: Is it necessary to perform any normalization/preprocessing to the neural data? Moreso, would any processing help in model predictions? I see the iEEG features are quite small, but was wondering if any feaure engineering should be applied such as noise reduction.</p>\n<p>Using linear interpolation, I aligned sentences and neural features over time. When looking at all neural fature groups (64), similar words yield similar neural features by using cosine similarity between BERT embeddings of word tokens and neural features.</p>\n<p>Using a linear regression ML model using sklearn, I used subsets of the neural features to predict BERT embeddings of whole sentences so see if there was any more important feature set. When doing this, r^2 for each subset was almost the same at 0.18, while using all features was 0.84.</p>\n<p><strong>Key takeaways:</strong></p>\n<ul>\n<li>Similar word semantic structure yield similar neural features</li>\n<li>All feature subsets [n, n+64] are almost equal in importance and do correlate linearly to semantic embeddings.</li>\n</ul>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hello, my name is Krishna. I had a question about feature engineering and preprocessing of the neural data. I'd also like to share a result if its helpful to understand the correlation between the neural data and sentences.\n\nQuestions: Is it necessary to perform any normalization/preprocessing to the neural data? Moreso, would any processing help in model predictions? I see the iEEG features are quite small, but was wondering if any feaure engineering should be applied such as noise reduction.\n\nUsing linear interpolation, I aligned sentences and neural features over time. When looking at all neural fature groups (64), similar words yield similar neural features by using cosine similarity between BERT embeddings of word tokens and neural features.\n\nUsing a linear regression ML model using sklearn, I used subsets of the neural features to predict BERT embeddings of whole sentences so see if there was any more important feature set. When doing this, r^2 for each subset was almost the same at 0.18, while using all features was 0.84.\n\n**Key takeaways:**\n- Similar word semantic structure yield similar neural features\n- All feature subsets [n, n+64] are almost equal in importance and do correlate linearly to semantic embeddings.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "3263773",
      "postDate": "08/05/2025 17:47:51",
      "content": "<p>Hi there,</p>\n<p>The neural data has already been preprocessed (bandpass filtered, denoised via linear regression referencing, feature-extracted, and temporally binned to 20 ms bins) and normalized (z-scored based on the prior 10 sentences in chronological order).</p>\n<p>For the current baseline RNN model, there is a little bit more preprocessing (temporal smoothing) and data augmentation (adding white noise and constant offsets to training data) that occurs. You may find that additional preprocessing or data augmentation steps are helpful, depending on what kind of decoding approach that you use.</p>",
      "rawMarkdown": "Hi there,\n\nThe neural data has already been preprocessed (bandpass filtered, denoised via linear regression referencing, feature-extracted, and temporally binned to 20 ms bins) and normalized (z-scored based on the prior 10 sentences in chronological order).\n\nFor the current baseline RNN model, there is a little bit more preprocessing (temporal smoothing) and data augmentation (adding white noise and constant offsets to training data) that occurs. You may find that additional preprocessing or data augmentation steps are helpful, depending on what kind of decoding approach that you use.",
      "votes": null
    },
    {
      "id": "3273085",
      "postDate": "08/22/2025 05:49:31",
      "content": "<p>Hi Nick, </p>\n<p>I was wondering — in the preprocessing pipeline, you mention that the data was feature-extracted. Could you explain what kind of features were extracted? </p>",
      "rawMarkdown": "Hi Nick, \n\nI was wondering — in the preprocessing pipeline, you mention that the data was feature-extracted. Could you explain what kind of features were extracted?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3263773,
      "author_name": "notnickc",
      "author_url": "",
      "post_date": "08/05/2025 17:47:51",
      "content": "<p>Hi there,</p>\n<p>The neural data has already been preprocessed (bandpass filtered, denoised via linear regression referencing, feature-extracted, and temporally binned to 20 ms bins) and normalized (z-scored based on the prior 10 sentences in chronological order).</p>\n<p>For the current baseline RNN model, there is a little bit more preprocessing (temporal smoothing) and data augmentation (adding white noise and constant offsets to training data) that occurs. You may find that additional preprocessing or data augmentation steps are helpful, depending on what kind of decoding approach that you use.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3273085,
          "author_name": "mertcokun",
          "author_url": "",
          "post_date": "08/22/2025 05:49:31",
          "content": "<p>Hi Nick, </p>\n<p>I was wondering — in the preprocessing pipeline, you mention that the data was feature-extracted. Could you explain what kind of features were extracted? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3262514": "Hello, my name is Krishna. I had a question about feature engineering and preprocessing of the neural data. I'd also like to share a result if its helpful to understand the correlation between the neural data and sentences.\n\nQuestions: Is it necessary to perform any normalization/preprocessing to the neural data? Moreso, would any processing help in model predictions? I see the iEEG features are quite small, but was wondering if any feaure engineering should be applied such as noise reduction.\n\nUsing linear interpolation, I aligned sentences and neural features over time. When looking at all neural fature groups (64), similar words yield similar neural features by using cosine similarity between BERT embeddings of word tokens and neural features.\n\nUsing a linear regression ML model using sklearn, I used subsets of the neural features to predict BERT embeddings of whole sentences so see if there was any more important feature set. When doing this, r^2 for each subset was almost the same at 0.18, while using all features was 0.84.\n\n**Key takeaways:**\n- Similar word semantic structure yield similar neural features\n- All feature subsets [n, n+64] are almost equal in importance and do correlate linearly to semantic embeddings.\n\nHope this helps!",
    "3263773": "Hi there,\n\nThe neural data has already been preprocessed (bandpass filtered, denoised via linear regression referencing, feature-extracted, and temporally binned to 20 ms bins) and normalized (z-scored based on the prior 10 sentences in chronological order).\n\nFor the current baseline RNN model, there is a little bit more preprocessing (temporal smoothing) and data augmentation (adding white noise and constant offsets to training data) that occurs. You may find that additional preprocessing or data augmentation steps are helpful, depending on what kind of decoding approach that you use.",
    "3273085": "Hi Nick, \n\nI was wondering — in the preprocessing pipeline, you mention that the data was feature-extracted. Could you explain what kind of features were extracted?"
  },
  "source": "meta"
}