{
  "id": 477021,
  "title": "Raw EEG have been already preprocessed?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477021",
  "author_name": "Ángel Jacinto Sánchez Ruiz",
  "post_date": "2024-02-14T12:21:02.860000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi. Is my first time working with eeg and I was looking a reasonable way to normalize train eeg to  replicate some papper pipelines I'd like to try and I've been checking the histograms. I was especting different non normalized distributions for each electrode but all of them looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fd9a2c6fa9bff82e259669ecefeecb69b%2Fca915245-fa14-4c45-8c25-7a6ee0830e26.png?generation=1707913122280115&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F2ee486fd5ddbe901163e475511ed2254%2F26db3aa7-2e19-4e29-8ec1-5828bc4ebd10.png?generation=1707913498875179&amp;alt=media\"><br>\nThere is always same pattern with peaks at -10000, 0, 10000 and 35000. Has been data already normalized. What I'm missing? May be my code is wrong.</p>\n<pre><code> electrode  electrodes:\n    c = {}\n     eeg_id  tqdm(np(df.eeg_id)):\n        eeg = pd(f)\n        eeg = eeg\n        values,counts = np(eeg,return_counts=True)\n           ((values)):\n            try:\n                c] += counts\n            except:\n                c] = counts \n         values,counts,eeg\n    plt(electrode)\n    plt(c(),c(),)\n    plt()\n     c \n</code></pre>",
  "messages": [
    {
      "id": 2651879,
      "postDate": "2024-02-14T12:21:02.860Z",
      "content": "<p>Hi. Is my first time working with eeg and I was looking a reasonable way to normalize train eeg to  replicate some papper pipelines I'd like to try and I've been checking the histograms. I was especting different non normalized distributions for each electrode but all of them looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fd9a2c6fa9bff82e259669ecefeecb69b%2Fca915245-fa14-4c45-8c25-7a6ee0830e26.png?generation=1707913122280115&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F2ee486fd5ddbe901163e475511ed2254%2F26db3aa7-2e19-4e29-8ec1-5828bc4ebd10.png?generation=1707913498875179&amp;alt=media\"><br>\nThere is always same pattern with peaks at -10000, 0, 10000 and 35000. Has been data already normalized. What I'm missing? May be my code is wrong.</p>\n<pre><code> electrode  electrodes:\n    c = {}\n     eeg_id  tqdm(np(df.eeg_id)):\n        eeg = pd(f)\n        eeg = eeg\n        values,counts = np(eeg,return_counts=True)\n           ((values)):\n            try:\n                c] += counts\n            except:\n                c] = counts \n         values,counts,eeg\n    plt(electrode)\n    plt(c(),c(),)\n    plt()\n     c \n</code></pre>",
      "rawMarkdown": "Hi. Is my first time working with eeg and I was looking a reasonable way to normalize train eeg to  replicate some papper pipelines I'd like to try and I've been checking the histograms. I was especting different non normalized distributions for each electrode but all of them looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fd9a2c6fa9bff82e259669ecefeecb69b%2Fca915245-fa14-4c45-8c25-7a6ee0830e26.png?generation=1707913122280115&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F2ee486fd5ddbe901163e475511ed2254%2F26db3aa7-2e19-4e29-8ec1-5828bc4ebd10.png?generation=1707913498875179&alt=media)\nThere is always same pattern with peaks at -10000, 0, 10000 and 35000. Has been data already normalized. What I'm missing? May be my code is wrong.\n\n    \n    for electrode in electrodes:\n        c = {}\n        for eeg_id in tqdm.tqdm(np.unique(df.eeg_id)):\n            eeg = pd.read_parquet(f'train_eegs/{eeg_id}.parquet')[electrode].values\n            eeg = eeg[~np.isnan(eeg)]\n            values,counts = np.unique(eeg,return_counts=True)\n            for i in range(len(values)):\n                try:\n                    c[values[i]] += counts[i]\n                except:\n                    c[values[i]] = counts[i] \n            del values,counts,eeg\n        plt.title(electrode)\n        plt.plot(c.keys(),c.values(),'.')\n        plt.show()\n        del c ",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2651879": "Hi. Is my first time working with eeg and I was looking a reasonable way to normalize train eeg to  replicate some papper pipelines I'd like to try and I've been checking the histograms. I was especting different non normalized distributions for each electrode but all of them looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fd9a2c6fa9bff82e259669ecefeecb69b%2Fca915245-fa14-4c45-8c25-7a6ee0830e26.png?generation=1707913122280115&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F2ee486fd5ddbe901163e475511ed2254%2F26db3aa7-2e19-4e29-8ec1-5828bc4ebd10.png?generation=1707913498875179&alt=media)\nThere is always same pattern with peaks at -10000, 0, 10000 and 35000. Has been data already normalized. What I'm missing? May be my code is wrong.\n\n    \n    for electrode in electrodes:\n        c = {}\n        for eeg_id in tqdm.tqdm(np.unique(df.eeg_id)):\n            eeg = pd.read_parquet(f'train_eegs/{eeg_id}.parquet')[electrode].values\n            eeg = eeg[~np.isnan(eeg)]\n            values,counts = np.unique(eeg,return_counts=True)\n            for i in range(len(values)):\n                try:\n                    c[values[i]] += counts[i]\n                except:\n                    c[values[i]] = counts[i] \n            del values,counts,eeg\n        plt.title(electrode)\n        plt.plot(c.keys(),c.values(),'.')\n        plt.show()\n        del c "
  }
}