{
  "id": 487386,
  "title": "Question on Retrieving EEG and Spectrogram",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/487386",
  "author_name": "",
  "post_date": "2024-03-28T22:38:15.968862200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I know we are approaching the end of this competition and the topic of retrieving corresponding EEG and Spectrogram has been discussed at length. However, please indulge me a little.</p>\n<p>In Chris Kaggle discussion file <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">here</a>, he explained at length how to go about it. However, I do have a question.</p>\n<p>In the train_egg parquet files, when loaded into a dataframe, each row has an eeg_id number. Also, in the train_spectrogram parquet files, when loaded into a dataframe, each row has a spectrogram_id number. Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file? </p>\n<p>If one matches these numbers together, would it still be necessary to calculate corresponding EEG and Spectrum. Keep in mind for example the following: </p>\n<ul>\n<li>eeg_label_offset_seconds - is the time between the beginning of the consolidated EEG and this <strong>subsample</strong></li>\n<li>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to</li>\n</ul>\n<p>I guess my underlying question is this - why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?</p>",
  "messages": [
    {
      "id": "2721240",
      "postDate": "03/28/2024 22:38:15",
      "content": "<p>I know we are approaching the end of this competition and the topic of retrieving corresponding EEG and Spectrogram has been discussed at length. However, please indulge me a little.</p>\n<p>In Chris Kaggle discussion file <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">here</a>, he explained at length how to go about it. However, I do have a question.</p>\n<p>In the train_egg parquet files, when loaded into a dataframe, each row has an eeg_id number. Also, in the train_spectrogram parquet files, when loaded into a dataframe, each row has a spectrogram_id number. Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file? </p>\n<p>If one matches these numbers together, would it still be necessary to calculate corresponding EEG and Spectrum. Keep in mind for example the following: </p>\n<ul>\n<li>eeg_label_offset_seconds - is the time between the beginning of the consolidated EEG and this <strong>subsample</strong></li>\n<li>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to</li>\n</ul>\n<p>I guess my underlying question is this - why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?</p>",
      "rawMarkdown": "I know we are approaching the end of this competition and the topic of retrieving corresponding EEG and Spectrogram has been discussed at length. However, please indulge me a little.\n\nIn Chris Kaggle discussion file [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010), he explained at length how to go about it. However, I do have a question.\n\nIn the train_egg parquet files, when loaded into a dataframe, each row has an eeg_id number. Also, in the train_spectrogram parquet files, when loaded into a dataframe, each row has a spectrogram_id number. Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file? \n\nIf one matches these numbers together, would it still be necessary to calculate corresponding EEG and Spectrum. Keep in mind for example the following: \n- eeg_label_offset_seconds - is the time between the beginning of the consolidated EEG and this **subsample**\n- eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to\n\nI guess my underlying question is this - why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?",
      "votes": null
    },
    {
      "id": "2724080",
      "postDate": "03/30/2024 17:49:57",
      "content": "<blockquote>\n  <p>Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?</p>\n</blockquote>\n<p>It is not necessary, you could load the 3 files (train.csv, EEG parquets and Spectrogram parquets), then match across the dataframes to get the subsamples.</p>\n<p>There are 106,800 subsamples with unique label_id<br>\nThe EEG of a subsample = eeg_label_offset_seconds + 50 seconds<br>\nThe Spectrogram of a subsample = spectrogram_label_offset_seconds + 10 minutes</p>",
      "rawMarkdown": "> Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file?\n\nYes.\n\n>  why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?\n\nIt is not necessary, you could load the 3 files (train.csv, EEG parquets and Spectrogram parquets), then match across the dataframes to get the subsamples.\n\nThere are 106,800 subsamples with unique label_id\nThe EEG of a subsample = eeg_label_offset_seconds + 50 seconds\nThe Spectrogram of a subsample = spectrogram_label_offset_seconds + 10 minutes",
      "votes": null
    },
    {
      "id": "2724653",
      "postDate": "03/31/2024 02:46:28",
      "content": "<p><a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a> - Thanks much for clarifying. I thought as much that it was not the only path to take to calculate corresponding EEG and Spectrogram. However, from reading the discussion posts it appeared most, if not all, were following the calculation path.</p>",
      "rawMarkdown": "nartaa - Thanks much for clarifying. I thought as much that it was not the only path to take to calculate corresponding EEG and Spectrogram. However, from reading the discussion posts it appeared most, if not all, were following the calculation path.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2724080,
      "author_name": "nartaa",
      "author_url": "",
      "post_date": "03/30/2024 17:49:57",
      "content": "<blockquote>\n  <p>Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?</p>\n</blockquote>\n<p>It is not necessary, you could load the 3 files (train.csv, EEG parquets and Spectrogram parquets), then match across the dataframes to get the subsamples.</p>\n<p>There are 106,800 subsamples with unique label_id<br>\nThe EEG of a subsample = eeg_label_offset_seconds + 50 seconds<br>\nThe Spectrogram of a subsample = spectrogram_label_offset_seconds + 10 minutes</p>",
      "votes": null,
      "replies": [
        {
          "id": 2724653,
          "author_name": "michaelwaynetreasure",
          "author_url": "",
          "post_date": "03/31/2024 02:46:28",
          "content": "<p><a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a> - Thanks much for clarifying. I thought as much that it was not the only path to take to calculate corresponding EEG and Spectrogram. However, from reading the discussion posts it appeared most, if not all, were following the calculation path.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2721240": "I know we are approaching the end of this competition and the topic of retrieving corresponding EEG and Spectrogram has been discussed at length. However, please indulge me a little.\n\nIn Chris Kaggle discussion file [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010), he explained at length how to go about it. However, I do have a question.\n\nIn the train_egg parquet files, when loaded into a dataframe, each row has an eeg_id number. Also, in the train_spectrogram parquet files, when loaded into a dataframe, each row has a spectrogram_id number. Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file? \n\nIf one matches these numbers together, would it still be necessary to calculate corresponding EEG and Spectrum. Keep in mind for example the following: \n- eeg_label_offset_seconds - is the time between the beginning of the consolidated EEG and this **subsample**\n- eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to\n\nI guess my underlying question is this - why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?",
    "2724080": "> Are the eeg_id numbers and the spectrogram_id numbers here the same set of numbers in the train.csv file?\n\nYes.\n\n>  why is it necessary to calculate corresponding EEG and Spectrogram if one simply loads the 3 files into a dataframe and then match across the dataframes on the eeg_id and spectrogram_id columns?\n\nIt is not necessary, you could load the 3 files (train.csv, EEG parquets and Spectrogram parquets), then match across the dataframes to get the subsamples.\n\nThere are 106,800 subsamples with unique label_id\nThe EEG of a subsample = eeg_label_offset_seconds + 50 seconds\nThe Spectrogram of a subsample = spectrogram_label_offset_seconds + 10 minutes",
    "2724653": "nartaa - Thanks much for clarifying. I thought as much that it was not the only path to take to calculate corresponding EEG and Spectrogram. However, from reading the discussion posts it appeared most, if not all, were following the calculation path."
  },
  "source": "meta"
}