{
  "id": 468506,
  "title": "Relationship between patient_id vs eeg_id vs spectrogram_id",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/468506",
  "author_name": "",
  "post_date": "2024-01-16T21:38:54.367407Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>unique patient_id = 1950<br>\nfor each patient_id , we have more then 1 eeg_id and spectrogram_id is present.</p>\n<p>let me take patient_id == 30631 ,</p>\n<pre><code>pat_id=df[df[]==]\n(,(pat_id))\n\nspec_count=pat_id[].nunique()\neeg_count=pat_id[].nunique()\n(,eeg_count)\n(,spec_count)\n\neeg_id =pat_id[pat_id[]==]\nunique_spec_id_wrt_eegid=eeg_id[].nunique()\n(,unique_spec_id_wrt_eegid)\n\nspec_id =pat_id[pat_id[]==]\nunique_eeg_id_wrt_specid=spec_id[].nunique()\n(,unique_eeg_id_wrt_specid)\n</code></pre>\n<p>The response:</p>\n<p>the length: 2215<br>\neeg_count: 270<br>\nspec_count: 66<br>\ntotal counts of spectogram_id for a unique eeg_id: 1<br>\ntotal counts of eeg_id for a unique spectogram_id: 107</p>\n<p><strong>Note:</strong><br>\nthe above code shows that , for every unique patient_id , have more then 1 eeg_id and spectrogram_id are available.<br>\nthe relationship between eeg_id and spectrogram_id for a particular patient_id are as follows:</p>\n<ol>\n<li>for each eeg_id we have only one spectrogram_id is there but for every spectrogram_id , we have more then 1 eeg_id are present</li>\n</ol>",
  "messages": [
    {
      "id": "2605138",
      "postDate": "01/16/2024 21:38:54",
      "content": "<p>unique patient_id = 1950<br>\nfor each patient_id , we have more then 1 eeg_id and spectrogram_id is present.</p>\n<p>let me take patient_id == 30631 ,</p>\n<pre><code>pat_id=df[df[]==]\n(,(pat_id))\n\nspec_count=pat_id[].nunique()\neeg_count=pat_id[].nunique()\n(,eeg_count)\n(,spec_count)\n\neeg_id =pat_id[pat_id[]==]\nunique_spec_id_wrt_eegid=eeg_id[].nunique()\n(,unique_spec_id_wrt_eegid)\n\nspec_id =pat_id[pat_id[]==]\nunique_eeg_id_wrt_specid=spec_id[].nunique()\n(,unique_eeg_id_wrt_specid)\n</code></pre>\n<p>The response:</p>\n<p>the length: 2215<br>\neeg_count: 270<br>\nspec_count: 66<br>\ntotal counts of spectogram_id for a unique eeg_id: 1<br>\ntotal counts of eeg_id for a unique spectogram_id: 107</p>\n<p><strong>Note:</strong><br>\nthe above code shows that , for every unique patient_id , have more then 1 eeg_id and spectrogram_id are available.<br>\nthe relationship between eeg_id and spectrogram_id for a particular patient_id are as follows:</p>\n<ol>\n<li>for each eeg_id we have only one spectrogram_id is there but for every spectrogram_id , we have more then 1 eeg_id are present</li>\n</ol>",
      "rawMarkdown": "unique patient_id = 1950\nfor each patient_id , we have more then 1 eeg_id and spectrogram_id is present.\n\nlet me take patient_id == 30631 ,\n\n\n```python\npat_id=df[df['patient_id']==30631]\nprint('the length :',len(pat_id))\n\nspec_count=pat_id['spectrogram_id'].nunique()\neeg_count=pat_id['eeg_id'].nunique()\nprint('eeg_count:',eeg_count)\nprint('spec_count:',spec_count)\n\neeg_id =pat_id[pat_id['eeg_id']==1270973624]\nunique_spec_id_wrt_eegid=eeg_id['spectrogram_id'].nunique()\nprint('total counts of spectogram_id for a unique eeg_id:',unique_spec_id_wrt_eegid)\n\nspec_id =pat_id[pat_id['spectrogram_id']==764146759]\nunique_eeg_id_wrt_specid=spec_id['eeg_id'].nunique()\nprint('total counts of eeg_id for a unique spectogram_id:',unique_eeg_id_wrt_specid)\n```\n\nThe response:\n\nthe length: 2215\neeg_count: 270\nspec_count: 66\ntotal counts of spectogram_id for a unique eeg_id: 1\ntotal counts of eeg_id for a unique spectogram_id: 107\n\n**Note:**\nthe above code shows that , for every unique patient_id , have more then 1 eeg_id and spectrogram_id are available.\nthe relationship between eeg_id and spectrogram_id for a particular patient_id are as follows:\n1. for each eeg_id we have only one spectrogram_id is there but for every spectrogram_id , we have more then 1 eeg_id are present",
      "votes": null
    },
    {
      "id": "2655474",
      "postDate": "02/16/2024 22:48:04",
      "content": "<p>I observed the same. According to the data description both <em>eeg_id</em> and <em>spectrogram_id</em> should be \"A unique identifier for the entire EEG recording.\", however, as noted by <a href=\"https://www.kaggle.com/santusrk\" target=\"_blank\">@santusrk</a>, the same <em>spectrogram_id</em> appears in the rows with different <em>eeg_id values</em>   Would it be possible to get a clarification from the organizers?</p>",
      "rawMarkdown": "I observed the same. According to the data description both *eeg_id* and *spectrogram_id* should be \"A unique identifier for the entire EEG recording.\", however, as noted by @santusrk, the same *spectrogram_id* appears in the rows with different *eeg_id values*   Would it be possible to get a clarification from the organizers?",
      "votes": null
    },
    {
      "id": "2655487",
      "postDate": "02/16/2024 23:11:40",
      "content": "<p>We have data on 1950 patients. Kaggle could have created 1950 EEG parquets and 1950 Spectrogram parquets. But then the files would be too large. So, Kaggle split the 1950 patients' EEG into 17089 EEG parquet files. And Kaggle split the 1950 patients' Spectrograms into 11138 Spectrogram files. </p>\n<p>Each row of the <code>train.csv</code> extracts a portion (10_000 rows) of one EEG parquet and a portion (300 rows) of one Spectrogram parquet. Since two different <code>eeg_id</code> can refer to the same patient, then sometimes both of those <code>eeg_id</code> extract Spectrogram from the same Spectrogram parquet.</p>\n<p>Imagine if Kaggle wrote all Spectrograms to <strong>only one</strong> Spectrogram parquet. Then <strong>every eeg id</strong> would access their Spectrogram from the same Spectrogram parquet.</p>",
      "rawMarkdown": "We have data on 1950 patients. Kaggle could have created 1950 EEG parquets and 1950 Spectrogram parquets. But then the files would be too large. So, Kaggle split the 1950 patients' EEG into 17089 EEG parquet files. And Kaggle split the 1950 patients' Spectrograms into 11138 Spectrogram files. \n\nEach row of the `train.csv` extracts a portion (10_000 rows) of one EEG parquet and a portion (300 rows) of one Spectrogram parquet. Since two different `eeg_id` can refer to the same patient, then sometimes both of those `eeg_id` extract Spectrogram from the same Spectrogram parquet.\n\nImagine if Kaggle wrote all Spectrograms to **only one** Spectrogram parquet. Then **every eeg id** would access their Spectrogram from the same Spectrogram parquet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2655474,
      "author_name": "tomyellowhammer",
      "author_url": "",
      "post_date": "02/16/2024 22:48:04",
      "content": "<p>I observed the same. According to the data description both <em>eeg_id</em> and <em>spectrogram_id</em> should be \"A unique identifier for the entire EEG recording.\", however, as noted by <a href=\"https://www.kaggle.com/santusrk\" target=\"_blank\">@santusrk</a>, the same <em>spectrogram_id</em> appears in the rows with different <em>eeg_id values</em>   Would it be possible to get a clarification from the organizers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655487,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2024 23:11:40",
          "content": "<p>We have data on 1950 patients. Kaggle could have created 1950 EEG parquets and 1950 Spectrogram parquets. But then the files would be too large. So, Kaggle split the 1950 patients' EEG into 17089 EEG parquet files. And Kaggle split the 1950 patients' Spectrograms into 11138 Spectrogram files. </p>\n<p>Each row of the <code>train.csv</code> extracts a portion (10_000 rows) of one EEG parquet and a portion (300 rows) of one Spectrogram parquet. Since two different <code>eeg_id</code> can refer to the same patient, then sometimes both of those <code>eeg_id</code> extract Spectrogram from the same Spectrogram parquet.</p>\n<p>Imagine if Kaggle wrote all Spectrograms to <strong>only one</strong> Spectrogram parquet. Then <strong>every eeg id</strong> would access their Spectrogram from the same Spectrogram parquet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2605138": "unique patient_id = 1950\nfor each patient_id , we have more then 1 eeg_id and spectrogram_id is present.\n\nlet me take patient_id == 30631 ,\n\n\n```python\npat_id=df[df['patient_id']==30631]\nprint('the length :',len(pat_id))\n\nspec_count=pat_id['spectrogram_id'].nunique()\neeg_count=pat_id['eeg_id'].nunique()\nprint('eeg_count:',eeg_count)\nprint('spec_count:',spec_count)\n\neeg_id =pat_id[pat_id['eeg_id']==1270973624]\nunique_spec_id_wrt_eegid=eeg_id['spectrogram_id'].nunique()\nprint('total counts of spectogram_id for a unique eeg_id:',unique_spec_id_wrt_eegid)\n\nspec_id =pat_id[pat_id['spectrogram_id']==764146759]\nunique_eeg_id_wrt_specid=spec_id['eeg_id'].nunique()\nprint('total counts of eeg_id for a unique spectogram_id:',unique_eeg_id_wrt_specid)\n```\n\nThe response:\n\nthe length: 2215\neeg_count: 270\nspec_count: 66\ntotal counts of spectogram_id for a unique eeg_id: 1\ntotal counts of eeg_id for a unique spectogram_id: 107\n\n**Note:**\nthe above code shows that , for every unique patient_id , have more then 1 eeg_id and spectrogram_id are available.\nthe relationship between eeg_id and spectrogram_id for a particular patient_id are as follows:\n1. for each eeg_id we have only one spectrogram_id is there but for every spectrogram_id , we have more then 1 eeg_id are present",
    "2655474": "I observed the same. According to the data description both *eeg_id* and *spectrogram_id* should be \"A unique identifier for the entire EEG recording.\", however, as noted by @santusrk, the same *spectrogram_id* appears in the rows with different *eeg_id values*   Would it be possible to get a clarification from the organizers?",
    "2655487": "We have data on 1950 patients. Kaggle could have created 1950 EEG parquets and 1950 Spectrogram parquets. But then the files would be too large. So, Kaggle split the 1950 patients' EEG into 17089 EEG parquet files. And Kaggle split the 1950 patients' Spectrograms into 11138 Spectrogram files. \n\nEach row of the `train.csv` extracts a portion (10_000 rows) of one EEG parquet and a portion (300 rows) of one Spectrogram parquet. Since two different `eeg_id` can refer to the same patient, then sometimes both of those `eeg_id` extract Spectrogram from the same Spectrogram parquet.\n\nImagine if Kaggle wrote all Spectrograms to **only one** Spectrogram parquet. Then **every eeg id** would access their Spectrogram from the same Spectrogram parquet."
  },
  "source": "meta"
}