{
  "id": 467110,
  "title": "Problems with understanding how to get the correct subsets",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467110",
  "author_name": "",
  "post_date": "2024-01-11T06:52:29.262571900Z",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi to all! <br>\nWhen I read how we should get annotated subsets, I can't imagine it, here is an example:<br>\ndf=pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')<br>\ndf[df['eeg_id']==1628180742]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F3cb5482645d8ca0c92f74c2db9cce87c%2F1.jpeg?generation=1704954836026280&amp;alt=media\"></p>\n<p>As I understand for the first sample we need to take part of his eeg from 0 to 200x10 [0:200x10] because we take 10 second segments starting with spectrogram_label_offset_seconds and each second includes 200 lines from the train_eegs file (By the way, this file is 90 seconds long, not the declared 50)</p>\n<p>But I have a question about how to get segments for spectrograms: <br>\ndata=pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/353733.parquet')<br>\ndata<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fa17b2122f746691fea108c45dbdd701f%2F2.jpeg?generation=1704955530094289&amp;alt=media\"></p>\n<p>In this sample, we have a 639 second long spectrogram recording with a 2 second step, what segments are we taking here? Do we ignore only the beginning of the recording, or do we only take 10 second windows of 5 elements, I can't understand it. Thanks for the comments.</p>",
  "messages": [
    {
      "id": "2596560",
      "postDate": "01/11/2024 06:52:29",
      "content": "<p>Hi to all! <br>\nWhen I read how we should get annotated subsets, I can't imagine it, here is an example:<br>\ndf=pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')<br>\ndf[df['eeg_id']==1628180742]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F3cb5482645d8ca0c92f74c2db9cce87c%2F1.jpeg?generation=1704954836026280&amp;alt=media\"></p>\n<p>As I understand for the first sample we need to take part of his eeg from 0 to 200x10 [0:200x10] because we take 10 second segments starting with spectrogram_label_offset_seconds and each second includes 200 lines from the train_eegs file (By the way, this file is 90 seconds long, not the declared 50)</p>\n<p>But I have a question about how to get segments for spectrograms: <br>\ndata=pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/353733.parquet')<br>\ndata<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fa17b2122f746691fea108c45dbdd701f%2F2.jpeg?generation=1704955530094289&amp;alt=media\"></p>\n<p>In this sample, we have a 639 second long spectrogram recording with a 2 second step, what segments are we taking here? Do we ignore only the beginning of the recording, or do we only take 10 second windows of 5 elements, I can't understand it. Thanks for the comments.</p>",
      "rawMarkdown": "Hi to all! \nWhen I read how we should get annotated subsets, I can't imagine it, here is an example:\ndf=pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ndf[df['eeg_id']==1628180742]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F3cb5482645d8ca0c92f74c2db9cce87c%2F1.jpeg?generation=1704954836026280&alt=media)\n\nAs I understand for the first sample we need to take part of his eeg from 0 to 200x10 [0:200x10] because we take 10 second segments starting with spectrogram_label_offset_seconds and each second includes 200 lines from the train_eegs file (By the way, this file is 90 seconds long, not the declared 50)\n\nBut I have a question about how to get segments for spectrograms: \ndata=pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/353733.parquet')\ndata\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fa17b2122f746691fea108c45dbdd701f%2F2.jpeg?generation=1704955530094289&alt=media)\n\nIn this sample, we have a 639 second long spectrogram recording with a 2 second step, what segments are we taking here? Do we ignore only the beginning of the recording, or do we only take 10 second windows of 5 elements, I can't understand it. Thanks for the comments.",
      "votes": null
    },
    {
      "id": "2596607",
      "postDate": "01/11/2024 07:31:15",
      "content": "<p>In the data description, it's mentioned that the eeg and spectrograms are centered at the same time. So if the eeg is 90 seconds, could it be the center 90 seconds of the spectrograms? Then you would take the beginning 10 seconds of that center piece, and it would account to the first subsample.</p>\n<p>I'm not really sure either yet.</p>\n<p>One thing I was wondering about and I made a topic on was how come some spectrograms have multiple eeg ids? I thought each spectrogram was associated with a single eeg?</p>",
      "rawMarkdown": "In the data description, it's mentioned that the eeg and spectrograms are centered at the same time. So if the eeg is 90 seconds, could it be the center 90 seconds of the spectrograms? Then you would take the beginning 10 seconds of that center piece, and it would account to the first subsample.\n\nI'm not really sure either yet.\n\nOne thing I was wondering about and I made a topic on was how come some spectrograms have multiple eeg ids? I thought each spectrogram was associated with a single eeg?",
      "votes": null
    },
    {
      "id": "2596657",
      "postDate": "01/11/2024 08:39:57",
      "content": "<p>Damn this data is hard. I also started a topic about how to merge targets into eegs.</p>",
      "rawMarkdown": "Damn this data is hard. I also started a topic about how to merge targets into eegs.",
      "votes": null
    },
    {
      "id": "2596764",
      "postDate": "01/11/2024 10:00:06",
      "content": "<p>Here is code to get the spectrogram:</p>\n<pre><code>    train = pd.read_csv()\n     = train.iloc[]\n    spectrogram = pd.read_parquet(+str(.spectrogram_id)+)\n    start = (.spectrogram_label_offset_seconds)\n     %==:  += \n    end =  + \n    spectrogram = spectrogram.loc[(spectrogram.time&gt;=)&amp;(spectrogram.time&lt;=)]\n</code></pre>\n<p>After running this code you have the spectrogram for train row <code>ROW</code> in the dataframe named <code>spectrogram</code>. It has shape <code>(300,401)</code>.</p>",
      "rawMarkdown": "Here is code to get the spectrogram:\n\n        train = pd.read_csv('train.csv')\n        row = train.iloc[ROW]\n        spectrogram = pd.read_parquet(PATH+str(row.spectrogram_id)+'.parquet')\n        start = int(row.spectrogram_label_offset_seconds)\n        if start%2==0: start += 1\n        end = start + 598\n        spectrogram = spectrogram.loc[(spectrogram.time>=start)&(spectrogram.time<=end)]\n\nAfter running this code you have the spectrogram for train row `ROW` in the dataframe named `spectrogram`. It has shape `(300,401)`.",
      "votes": null
    },
    {
      "id": "2596814",
      "postDate": "01/11/2024 10:48:22",
      "content": "<p>Thanks Chris! Great to see you in the competition! <br>\nSo, if I understand correctly, under such conditions, each of our 106,800 samples will have 7 or 8 samples that will intersect with it.<br>\nBy the way:<br>\n    train_eeg_id = train[['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']]<br>\n    unique_rows_count = train_eeg_id.nunique()<br>\n    print('unique_rows_count:',unique_rows_count)<br>\n    unique_eeg_id_count=train_eeg_id['eeg_id'].unique().shape[0]<br>\n    print('unique_eeg_id_count:',unique_eeg_id_count)<br>\n​<br>\n     unique_rows_count: eeg_id 17089<br>\n                                       seizure_vote 18<br>\n                                       lpd_vote 19<br>\n                                       gpd_vote 17<br>\n                                       lrda_vote 16<br>\n                                       grda_vote 16<br>\n                                       other_vote 26<br>\n                                       dtype: int64<br>\n      unique_eeg_id_count: 17089</p>\n<p>We have a total of 17089 trials and it seems that the predictions are the same for all sample segments, in fact we can safely take a 50 second training segment, or divide it not only into 10 second segments, but into any suitable segments.</p>",
      "rawMarkdown": "Thanks Chris! Great to see you in the competition! \nSo, if I understand correctly, under such conditions, each of our 106,800 samples will have 7 or 8 samples that will intersect with it.\nBy the way:\n    train_eeg_id = train[['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']]\n    unique_rows_count = train_eeg_id.nunique()\n    print('unique_rows_count:',unique_rows_count)\n    unique_eeg_id_count=train_eeg_id['eeg_id'].unique().shape[0]\n    print('unique_eeg_id_count:',unique_eeg_id_count)\n​\n     unique_rows_count: eeg_id 17089\n                                       seizure_vote 18\n                                       lpd_vote 19\n                                       gpd_vote 17\n                                       lrda_vote 16\n                                       grda_vote 16\n                                       other_vote 26\n                                       dtype: int64\n      unique_eeg_id_count: 17089\n\nWe have a total of 17089 trials and it seems that the predictions are the same for all sample segments, in fact we can safely take a 50 second training segment, or divide it not only into 10 second segments, but into any suitable segments.",
      "votes": null
    },
    {
      "id": "2597933",
      "postDate": "01/12/2024 03:25:46",
      "content": "<p>It doesn't look like all the subsamples within an eeg_id have the same labels, but most do (only 1807 out of 17089 don't).</p>\n<pre><code>mean_labels = train()]()\n\neeg_ids_with_diff_labels_in_sub_samples = \n\n idx, row  train():\n     not ((row](np.float32)) ==\n         mean_labels](np.float32))():\n        eeg_ids_with_diff_labels_in_sub_samples(row.eeg_id)\n\n))\n</code></pre>",
      "rawMarkdown": "It doesn't look like all the subsamples within an eeg_id have the same labels, but most do (only 1807 out of 17089 don't).\n\n```\nmean_labels = train.groupby('eeg_id')[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].mean()\n\neeg_ids_with_diff_labels_in_sub_samples = []\n\nfor idx, row in train.iterrows():\n    if not ((row.loc[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].values.astype(np.float32)) ==\n         mean_labels.loc[row.eeg_id, ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].values.astype(np.float32)).all():\n        eeg_ids_with_diff_labels_in_sub_samples.append(row.eeg_id)\n\nprint(len(np.unique(eeg_ids_with_diff_labels_in_sub_samples)))\n```",
      "votes": null
    },
    {
      "id": "2601719",
      "postDate": "01/14/2024 16:17:19",
      "content": "<p>Sorry, I have a question: Each eeg_id has multiple eeg_sub_id. Does each eeg_sub_id represent 50s of EEG recordings? This 50s EEG record is part of eeg_id. I don’t know if my understanding is correct.</p>",
      "rawMarkdown": "Sorry, I have a question: Each eeg_id has multiple eeg_sub_id. Does each eeg_sub_id represent 50s of EEG recordings? This 50s EEG record is part of eeg_id. I don’t know if my understanding is correct.",
      "votes": null
    },
    {
      "id": "2602562",
      "postDate": "01/15/2024 07:35:45",
      "content": "<blockquote>\n  <p>Does each eeg_sub_id represent 50s of EEG recordings?</p>\n</blockquote>\n<p>Yes.</p>",
      "rawMarkdown": "> Does each eeg_sub_id represent 50s of EEG recordings?\n\nYes.",
      "votes": null
    },
    {
      "id": "2604396",
      "postDate": "01/16/2024 11:58:42",
      "content": "<p>check this topic maybe helps: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010</a></p>",
      "rawMarkdown": "check this topic maybe helps: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2596607,
      "author_name": "yousof9",
      "author_url": "",
      "post_date": "01/11/2024 07:31:15",
      "content": "<p>In the data description, it's mentioned that the eeg and spectrograms are centered at the same time. So if the eeg is 90 seconds, could it be the center 90 seconds of the spectrograms? Then you would take the beginning 10 seconds of that center piece, and it would account to the first subsample.</p>\n<p>I'm not really sure either yet.</p>\n<p>One thing I was wondering about and I made a topic on was how come some spectrograms have multiple eeg ids? I thought each spectrogram was associated with a single eeg?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2596657,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "01/11/2024 08:39:57",
          "content": "<p>Damn this data is hard. I also started a topic about how to merge targets into eegs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2596764,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/11/2024 10:00:06",
      "content": "<p>Here is code to get the spectrogram:</p>\n<pre><code>    train = pd.read_csv()\n     = train.iloc[]\n    spectrogram = pd.read_parquet(+str(.spectrogram_id)+)\n    start = (.spectrogram_label_offset_seconds)\n     %==:  += \n    end =  + \n    spectrogram = spectrogram.loc[(spectrogram.time&gt;=)&amp;(spectrogram.time&lt;=)]\n</code></pre>\n<p>After running this code you have the spectrogram for train row <code>ROW</code> in the dataframe named <code>spectrogram</code>. It has shape <code>(300,401)</code>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2596814,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "01/11/2024 10:48:22",
          "content": "<p>Thanks Chris! Great to see you in the competition! <br>\nSo, if I understand correctly, under such conditions, each of our 106,800 samples will have 7 or 8 samples that will intersect with it.<br>\nBy the way:<br>\n    train_eeg_id = train[['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']]<br>\n    unique_rows_count = train_eeg_id.nunique()<br>\n    print('unique_rows_count:',unique_rows_count)<br>\n    unique_eeg_id_count=train_eeg_id['eeg_id'].unique().shape[0]<br>\n    print('unique_eeg_id_count:',unique_eeg_id_count)<br>\n​<br>\n     unique_rows_count: eeg_id 17089<br>\n                                       seizure_vote 18<br>\n                                       lpd_vote 19<br>\n                                       gpd_vote 17<br>\n                                       lrda_vote 16<br>\n                                       grda_vote 16<br>\n                                       other_vote 26<br>\n                                       dtype: int64<br>\n      unique_eeg_id_count: 17089</p>\n<p>We have a total of 17089 trials and it seems that the predictions are the same for all sample segments, in fact we can safely take a 50 second training segment, or divide it not only into 10 second segments, but into any suitable segments.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2597933,
              "author_name": "yousof9",
              "author_url": "",
              "post_date": "01/12/2024 03:25:46",
              "content": "<p>It doesn't look like all the subsamples within an eeg_id have the same labels, but most do (only 1807 out of 17089 don't).</p>\n<pre><code>mean_labels = train()]()\n\neeg_ids_with_diff_labels_in_sub_samples = \n\n idx, row  train():\n     not ((row](np.float32)) ==\n         mean_labels](np.float32))():\n        eeg_ids_with_diff_labels_in_sub_samples(row.eeg_id)\n\n))\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2601719,
      "author_name": "gentlezdh",
      "author_url": "",
      "post_date": "01/14/2024 16:17:19",
      "content": "<p>Sorry, I have a question: Each eeg_id has multiple eeg_sub_id. Does each eeg_sub_id represent 50s of EEG recordings? This 50s EEG record is part of eeg_id. I don’t know if my understanding is correct.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2602562,
          "author_name": "zrongchu",
          "author_url": "",
          "post_date": "01/15/2024 07:35:45",
          "content": "<blockquote>\n  <p>Does each eeg_sub_id represent 50s of EEG recordings?</p>\n</blockquote>\n<p>Yes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2604396,
      "author_name": "icozma",
      "author_url": "",
      "post_date": "01/16/2024 11:58:42",
      "content": "<p>check this topic maybe helps: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2596560": "Hi to all! \nWhen I read how we should get annotated subsets, I can't imagine it, here is an example:\ndf=pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ndf[df['eeg_id']==1628180742]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F3cb5482645d8ca0c92f74c2db9cce87c%2F1.jpeg?generation=1704954836026280&alt=media)\n\nAs I understand for the first sample we need to take part of his eeg from 0 to 200x10 [0:200x10] because we take 10 second segments starting with spectrogram_label_offset_seconds and each second includes 200 lines from the train_eegs file (By the way, this file is 90 seconds long, not the declared 50)\n\nBut I have a question about how to get segments for spectrograms: \ndata=pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/353733.parquet')\ndata\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2Fa17b2122f746691fea108c45dbdd701f%2F2.jpeg?generation=1704955530094289&alt=media)\n\nIn this sample, we have a 639 second long spectrogram recording with a 2 second step, what segments are we taking here? Do we ignore only the beginning of the recording, or do we only take 10 second windows of 5 elements, I can't understand it. Thanks for the comments.",
    "2596607": "In the data description, it's mentioned that the eeg and spectrograms are centered at the same time. So if the eeg is 90 seconds, could it be the center 90 seconds of the spectrograms? Then you would take the beginning 10 seconds of that center piece, and it would account to the first subsample.\n\nI'm not really sure either yet.\n\nOne thing I was wondering about and I made a topic on was how come some spectrograms have multiple eeg ids? I thought each spectrogram was associated with a single eeg?",
    "2596657": "Damn this data is hard. I also started a topic about how to merge targets into eegs.",
    "2596764": "Here is code to get the spectrogram:\n\n        train = pd.read_csv('train.csv')\n        row = train.iloc[ROW]\n        spectrogram = pd.read_parquet(PATH+str(row.spectrogram_id)+'.parquet')\n        start = int(row.spectrogram_label_offset_seconds)\n        if start%2==0: start += 1\n        end = start + 598\n        spectrogram = spectrogram.loc[(spectrogram.time>=start)&(spectrogram.time<=end)]\n\nAfter running this code you have the spectrogram for train row `ROW` in the dataframe named `spectrogram`. It has shape `(300,401)`.",
    "2596814": "Thanks Chris! Great to see you in the competition! \nSo, if I understand correctly, under such conditions, each of our 106,800 samples will have 7 or 8 samples that will intersect with it.\nBy the way:\n    train_eeg_id = train[['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']]\n    unique_rows_count = train_eeg_id.nunique()\n    print('unique_rows_count:',unique_rows_count)\n    unique_eeg_id_count=train_eeg_id['eeg_id'].unique().shape[0]\n    print('unique_eeg_id_count:',unique_eeg_id_count)\n​\n     unique_rows_count: eeg_id 17089\n                                       seizure_vote 18\n                                       lpd_vote 19\n                                       gpd_vote 17\n                                       lrda_vote 16\n                                       grda_vote 16\n                                       other_vote 26\n                                       dtype: int64\n      unique_eeg_id_count: 17089\n\nWe have a total of 17089 trials and it seems that the predictions are the same for all sample segments, in fact we can safely take a 50 second training segment, or divide it not only into 10 second segments, but into any suitable segments.",
    "2597933": "It doesn't look like all the subsamples within an eeg_id have the same labels, but most do (only 1807 out of 17089 don't).\n\n```\nmean_labels = train.groupby('eeg_id')[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].mean()\n\neeg_ids_with_diff_labels_in_sub_samples = []\n\nfor idx, row in train.iterrows():\n    if not ((row.loc[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].values.astype(np.float32)) ==\n         mean_labels.loc[row.eeg_id, ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].values.astype(np.float32)).all():\n        eeg_ids_with_diff_labels_in_sub_samples.append(row.eeg_id)\n\nprint(len(np.unique(eeg_ids_with_diff_labels_in_sub_samples)))\n```",
    "2601719": "Sorry, I have a question: Each eeg_id has multiple eeg_sub_id. Does each eeg_sub_id represent 50s of EEG recordings? This 50s EEG record is part of eeg_id. I don’t know if my understanding is correct.",
    "2602562": "> Does each eeg_sub_id represent 50s of EEG recordings?\n\nYes.",
    "2604396": "check this topic maybe helps: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010"
  },
  "source": "meta"
}