{
  "id": 491078,
  "title": "Multiple Expert Consensus Diagnoses for Single EEG Recordings!",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/491078",
  "author_name": "",
  "post_date": "2024-04-04T14:57:05.258443700Z",
  "votes": 7,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Greetings, everyone.</p>\n<p>I have encountered a peculiar situation in the training dataset, and I would appreciate if someone could provide some insight. <br>\nAmong the EEG recordings, I have identified 783 instances where a single EEG_ID has been assigned multiple expert consensus diagnoses.</p>\n<p>The concerning aspect is that each EEG recording expected to originate from a single patient. Consequently, each EEG_ID should logically correspond to a unique diagnosis, as it is highly unlikely for an individual patient to have multiple distinct conditions simultaneously.</p>\n<p>If this scenario is not an anomaly and there is a valid explanation for the presence of multiple diagnoses for a single EEG_ID, I would greatly appreciate your insights.</p>\n<p><code>(train.groupby('eeg_id')['expert_consensus'].nunique() &gt; 1).sum()</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F1e63b1329117b5949a5e0c90c378cc57%2FScreenshot%20from%202024-04-04%2015-01-48.png?generation=1712242615039186&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2735147",
      "postDate": "04/04/2024 14:57:05",
      "content": "<p>Greetings, everyone.</p>\n<p>I have encountered a peculiar situation in the training dataset, and I would appreciate if someone could provide some insight. <br>\nAmong the EEG recordings, I have identified 783 instances where a single EEG_ID has been assigned multiple expert consensus diagnoses.</p>\n<p>The concerning aspect is that each EEG recording expected to originate from a single patient. Consequently, each EEG_ID should logically correspond to a unique diagnosis, as it is highly unlikely for an individual patient to have multiple distinct conditions simultaneously.</p>\n<p>If this scenario is not an anomaly and there is a valid explanation for the presence of multiple diagnoses for a single EEG_ID, I would greatly appreciate your insights.</p>\n<p><code>(train.groupby('eeg_id')['expert_consensus'].nunique() &gt; 1).sum()</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F1e63b1329117b5949a5e0c90c378cc57%2FScreenshot%20from%202024-04-04%2015-01-48.png?generation=1712242615039186&amp;alt=media\"></p>",
      "rawMarkdown": "Greetings, everyone.\n\nI have encountered a peculiar situation in the training dataset, and I would appreciate if someone could provide some insight. \nAmong the EEG recordings, I have identified 783 instances where a single EEG_ID has been assigned multiple expert consensus diagnoses.\n\nThe concerning aspect is that each EEG recording expected to originate from a single patient. Consequently, each EEG_ID should logically correspond to a unique diagnosis, as it is highly unlikely for an individual patient to have multiple distinct conditions simultaneously.\n\nIf this scenario is not an anomaly and there is a valid explanation for the presence of multiple diagnoses for a single EEG_ID, I would greatly appreciate your insights.\n\n`(train.groupby('eeg_id')['expert_consensus'].nunique() > 1).sum()`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F1e63b1329117b5949a5e0c90c378cc57%2FScreenshot%20from%202024-04-04%2015-01-48.png?generation=1712242615039186&alt=media)",
      "votes": null
    },
    {
      "id": "2735209",
      "postDate": "04/04/2024 15:39:23",
      "content": "<p>Not necessarily. There is various subsamples in each eeg_id. Each one with different 10 seconds to be labeled. Even beign the same patient can have different labels. Even more if each one have been labeled independently. The model has to predict more various experts vote distribution than a true label.</p>",
      "rawMarkdown": "Not necessarily. There is various subsamples in each eeg_id. Each one with different 10 seconds to be labeled. Even beign the same patient can have different labels. Even more if each one have been labeled independently. The model has to predict more various experts vote distribution than a true label.",
      "votes": null
    },
    {
      "id": "2735220",
      "postDate": "04/04/2024 15:42:40",
      "content": "<p>see the Data page - </p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.</p>\n  <ul>\n  <li><p>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.</p></li>\n  \n  <li><p>label_id - An ID for this set of labels.</p></li>\n  </ul>\n</blockquote>\n<p>Have been working on this after finding issues and considering all the discussions about &gt;= 10 votes.<br>\nReally don't think that was made clear about how train data was put together. But probably some figured it out earlier?<br>\nThe labels are not consistently for a 50s or even 10s period, so deciding how to split out<br>\nHave a csv file could put in a dataset, not sure if too late for sharing though.  Value counts below for train expert consensus counts and the 5 entry of 234 is just for one eeg id.  Possibly these could/should be left out of training.</p>\n<blockquote>\n  <p>exp_cons_cnts<br>\n  1    93956<br>\n  2     9134<br>\n  3     2187<br>\n  4     1289<br>\n  5      234`</p>\n</blockquote>\n<p>Considering the test data is meant to be 50s only per eeg id and one spectrogram no overlapping, not sure if this will also happen in test labels.  And think is meant to be only the centre 10s that was labeled not sure.</p>",
      "rawMarkdown": "see the Data page - \n>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\n\n\n>- eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.\n\n\n\n>- label_id - An ID for this set of labels.\n\nHave been working on this after finding issues and considering all the discussions about >= 10 votes.\nReally don't think that was made clear about how train data was put together. But probably some figured it out earlier?\nThe labels are not consistently for a 50s or even 10s period, so deciding how to split out\nHave a csv file could put in a dataset, not sure if too late for sharing though.  Value counts below for train expert consensus counts and the 5 entry of 234 is just for one eeg id.  Possibly these could/should be left out of training.\n\n>exp_cons_cnts\n1    93956\n2     9134\n3     2187\n4     1289\n5      234`\n\nConsidering the test data is meant to be 50s only per eeg id and one spectrogram no overlapping, not sure if this will also happen in test labels.  And think is meant to be only the centre 10s that was labeled not sure.",
      "votes": null
    },
    {
      "id": "2736346",
      "postDate": "04/05/2024 06:35:36",
      "content": "<p>If you plot the consensuses (?) against time, you'll see that they evolve in a reasonable manner: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fbbfd41171920532c3d3ff34d4e4bc517%2Fconsensus_evolve.png?generation=1712298663242029&amp;alt=media\"></p>\n<p>If the top vote is less than 60%, I plotted the top two labels. Otherwise, I plotted the top label. Most of the transitions make sense (e.g. LPD to LRDA or GPD to GRDA; also seizures and LPD seem to mix. I read that patients with LPD are more likely to have seizures, although I don't remember on what time scale).</p>",
      "rawMarkdown": "If you plot the consensuses (?) against time, you'll see that they evolve in a reasonable manner: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fbbfd41171920532c3d3ff34d4e4bc517%2Fconsensus_evolve.png?generation=1712298663242029&alt=media)\n\nIf the top vote is less than 60%, I plotted the top two labels. Otherwise, I plotted the top label. Most of the transitions make sense (e.g. LPD to LRDA or GPD to GRDA; also seizures and LPD seem to mix. I read that patients with LPD are more likely to have seizures, although I don't remember on what time scale).",
      "votes": null
    },
    {
      "id": "2736381",
      "postDate": "04/05/2024 06:54:57",
      "content": "<p>Think the issue is with time - acc to Data page, the labels are for 50 sec subsamples from the eeg_label_offset_seconds, a bit vague if the central 10 secs applies to eeg data, spectrograms or both.  So taking a centre 10 secs from a 50 sec window of eeg data does not always (if ever) equate to the labels. Perhaps windowing by 2-4 secs might be closer. And train eegs have a variety of time lengths, well beyond 50 secs.  All of this probably accounts for the high number of Other in train data, maybe why &gt;=10 seems to work well.  Many of the public notebooks are not really considering eeg_label_offset_seconds from what I can tell.</p>\n<p>In contrast,  the test data is meant to be 50s only per eeg id and one spectrogram no overlapping. Somewhat apples and oranges.</p>",
      "rawMarkdown": "Think the issue is with time - acc to Data page, the labels are for 50 sec subsamples from the eeg_label_offset_seconds, a bit vague if the central 10 secs applies to eeg data, spectrograms or both.  So taking a centre 10 secs from a 50 sec window of eeg data does not always (if ever) equate to the labels. Perhaps windowing by 2-4 secs might be closer. And train eegs have a variety of time lengths, well beyond 50 secs.  All of this probably accounts for the high number of Other in train data, maybe why >=10 seems to work well.  Many of the public notebooks are not really considering eeg_label_offset_seconds from what I can tell.\n\nIn contrast,  the test data is meant to be 50s only per eeg id and one spectrogram no overlapping. Somewhat apples and oranges.",
      "votes": null
    },
    {
      "id": "2736429",
      "postDate": "04/05/2024 07:39:35",
      "content": "<p>For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fb29caca1b03a8a31eab2ca75b734ac68%2F21379701_eeg.png?generation=1712302677695583&amp;alt=media\"><br>\nEach 10s region is the center of a 50s region (which isn't shown). (Note: the colors here aren't the same as the colors on the previous plot)</p>\n<p>These are the labels and their locations:</p>\n<p>onset    duration    description<br>\n20.0      10.0    ('other', 'lrda')<br>\n30.0      10.0    ('other', 'lrda')<br>\n32.0      10.0    ('other', 'lrda')<br>\n50.0      10.0    ('seizure', 'other')<br>\n64.0      10.0    seizure<br>\n68.0      10.0    seizure<br>\n82.0      10.0    lpd<br>\n96.0      10.0    lpd<br>\n98.0      10.0    lpd<br>\n100.0      10.0    lpd<br>\n102.0      10.0    lpd<br>\n106.0      10.0    lpd<br>\n110.0      10.0    lpd<br>\n124.0      10.0    lpd<br>\n126.0      10.0    lpd<br>\n128.0      10.0    lpd<br>\n130.0      10.0    lpd<br>\n178.0      10.0    other</p>\n<p>The first offset was 0s, so the 10s scored region starts at 20s and ends at 30s, etc.</p>\n<p>This is the first 50s window for spec (cropped to the central 50s of the 10 minute window) and eeg, with the middle 10s highlighted:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2F30769663cb2ccd66f92193a2087fd355%2F21379701_eeg_spec.png?generation=1712303735165968&amp;alt=media\"><br>\nYou can see that there is a high frequency response in the spectrogram that corresponds with the EEG data in time. So I think the offsets are accurately described (although a bit cryptic, some of the other posts discussing the offsets were helpful.)</p>",
      "rawMarkdown": "For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fb29caca1b03a8a31eab2ca75b734ac68%2F21379701_eeg.png?generation=1712302677695583&alt=media)\nEach 10s region is the center of a 50s region (which isn't shown). (Note: the colors here aren't the same as the colors on the previous plot)\n\nThese are the labels and their locations:\n\nonset\tduration\tdescription\n20.0 \t 10.0 \t ('other', 'lrda')\n30.0 \t 10.0 \t ('other', 'lrda')\n32.0 \t 10.0 \t ('other', 'lrda')\n50.0 \t 10.0 \t ('seizure', 'other')\n64.0 \t 10.0 \t seizure\n68.0 \t 10.0 \t seizure\n82.0 \t 10.0 \t lpd\n96.0 \t 10.0 \t lpd\n98.0 \t 10.0 \t lpd\n100.0 \t 10.0 \t lpd\n102.0 \t 10.0 \t lpd\n106.0 \t 10.0 \t lpd\n110.0 \t 10.0 \t lpd\n124.0 \t 10.0 \t lpd\n126.0 \t 10.0 \t lpd\n128.0 \t 10.0 \t lpd\n130.0 \t 10.0 \t lpd\n178.0 \t 10.0 \t other\n\nThe first offset was 0s, so the 10s scored region starts at 20s and ends at 30s, etc.\n\nThis is the first 50s window for spec (cropped to the central 50s of the 10 minute window) and eeg, with the middle 10s highlighted:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2F30769663cb2ccd66f92193a2087fd355%2F21379701_eeg_spec.png?generation=1712303735165968&alt=media)\nYou can see that there is a high frequency response in the spectrogram that corresponds with the EEG data in time. So I think the offsets are accurately described (although a bit cryptic, some of the other posts discussing the offsets were helpful.)",
      "votes": null
    },
    {
      "id": "2736489",
      "postDate": "04/05/2024 08:13:29",
      "content": "<blockquote>\n  <p>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.<br>\n  eeg_label_offset_seconds - The time between the beginning of the consolidated EEG and this subsample.</p>\n</blockquote>\n<p>From train, the most votes for (21379701) are Other and Seizure and LPD. Using train eeg_label_offset_seconds starting at 0 for 50s will have Other (2 or 1)  and Seizure (1 or 2)  and LRDA (1 or 0) which consumes the entries up to 48s.  The next train at 62s for 50s will have LPD 75% of votes LRDA 25% of votes.  All these are low vote counts 1-3.  This eeg data has 208s or 41600 rows.<br>\n.<br>\nIt is difficult to read the text corresponding to labels, and not sure where your 50s region starts. Can only go by what is provided in train in terms of where eeg data and labels are meant to be in the parquet files.</p>\n<p>EDIT: your post came after was doing this.  <br>\nfrom train Other corr to 0, 10, 12 secs and 158s.  Seizure 30, 34, 48. LPD 62 to 106 at 2-4s intervals</p>\n<p>What is  the source of your plots?  Know a notebook is not possible now but maybe a code snippet, looks like using mne?</p>",
      "rawMarkdown": ">eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.\neeg_label_offset_seconds - The time between the beginning of the consolidated EEG and this subsample.\n\nFrom train, the most votes for (21379701) are Other and Seizure and LPD. Using train eeg_label_offset_seconds starting at 0 for 50s will have Other (2 or 1)  and Seizure (1 or 2)  and LRDA (1 or 0) which consumes the entries up to 48s.  The next train at 62s for 50s will have LPD 75% of votes LRDA 25% of votes.  All these are low vote counts 1-3.  This eeg data has 208s or 41600 rows.\n.\nIt is difficult to read the text corresponding to labels, and not sure where your 50s region starts. Can only go by what is provided in train in terms of where eeg data and labels are meant to be in the parquet files.\n\nEDIT: your post came after was doing this.  \nfrom train Other corr to 0, 10, 12 secs and 158s.  Seizure 30, 34, 48. LPD 62 to 106 at 2-4s intervals\n\nWhat is  the source of your plots?  Know a notebook is not possible now but maybe a code snippet, looks like using mne?",
      "votes": null
    },
    {
      "id": "2736492",
      "postDate": "04/05/2024 08:15:31",
      "content": "<p>If it is expected that a patient's state or condition can transition from one to another over time. To accurately capture these transitions, the <code>eeg_label_offset_seconds</code> should be utilized for each specific state or diagnosis.</p>\n<p>I have been applying the same <code>eeg_label_offset_seconds</code> value for a single <code>eeg_id</code>, under the assumption that each <code>eeg_id</code> would have only one associated diagnosis or state. This approach might not be appropriate, as some patients have multiple diagnoses associated with their EEG recordings. Incorporating the correct <code>eeg_label_offset_seconds</code> for each diagnosis or state associated with an <code>eeg_id</code>, the  model captures the temporal dynamics and transitions within the EEG data.</p>\n<p>While the number of samples with multiple diagnoses might not be substantial, considering this issue and accounting for potential transitions between different states or diagnoses could potentially improve the model's performance.</p>",
      "rawMarkdown": "If it is expected that a patient's state or condition can transition from one to another over time. To accurately capture these transitions, the `eeg_label_offset_seconds` should be utilized for each specific state or diagnosis.\n\nI have been applying the same `eeg_label_offset_seconds` value for a single `eeg_id`, under the assumption that each `eeg_id` would have only one associated diagnosis or state. This approach might not be appropriate, as some patients have multiple diagnoses associated with their EEG recordings. Incorporating the correct `eeg_label_offset_seconds` for each diagnosis or state associated with an `eeg_id`, the  model captures the temporal dynamics and transitions within the EEG data.\n\nWhile the number of samples with multiple diagnoses might not be substantial, considering this issue and accounting for potential transitions between different states or diagnoses could potentially improve the model's performance.",
      "votes": null
    },
    {
      "id": "2736497",
      "postDate": "04/05/2024 08:27:17",
      "content": "<p>Not entirely sure the train eeg_label_offset_seconds correctly corresponds to the row labels. And it is not just about multiple labels per eeg id, but getting the vote percentages somewhere in the ballpark. <br>\nThink the wording of a 50s subsample that the labels applied to was either misleading or incorrect.  </p>",
      "rawMarkdown": "Not entirely sure the train eeg_label_offset_seconds correctly corresponds to the row labels. And it is not just about multiple labels per eeg id, but getting the vote percentages somewhere in the ballpark. \nThink the wording of a 50s subsample that the labels applied to was either misleading or incorrect.",
      "votes": null
    },
    {
      "id": "2736532",
      "postDate": "04/05/2024 08:54:36",
      "content": "<p>Apologies in advance for this giant \"snippet\"… it would have been a lot of work to make it shorter.</p>\n<pre><code>def plot:\n    fig = plt.figure(figsize=(, ), layout=None)\n    gs = fig.add\n\n    axs0 = )  i  range()]\n\n    sub_spec = )\n    plot\n\n    ax1 = fig.add\n    raw_single_ann = eeg.raw.copy.set)\n    fig_raw = raw_single_ann.plot(\n        duration=,\n        start=eeg.eeg_offsets,\n        show_scrollbars=False,\n        show_scalebars=False,\n        **mne_plot_kwargs,\n    )\n\n    canvas = \n    canvas.draw\n    img = np.asarray(canvas.buffer)\n    im = ax1.imshow(img)\n    ax1.set\n</code></pre>\n<p>This is the code snippet. The rest of the code you need to run it is:</p>\n<pre><code>train_meta = pd.read_csv(data_path / ).set_index([, ]).sort_index()\nvote_cols = [col  col  train_meta.columns  str(col).endswith()]\ntrain_meta[] = train_meta[vote_cols].sum(axis=)\n\neeg_ids = train_meta.reset_index()[].unique()\neeg_id_index = {eeg_ids[i]: i  i  range(len(eeg_ids))}\n\n\nvote_probs = train_meta[vote_cols].div(train_meta[], axis=)\n\n\ndef get_vote_case str:\n    (k2, v2), (k1, v1) = x.sort_values()[-:].items()\n     v1 == :\n         \n     v1 &gt;  + tol:\n         \n       k1    k2:\n         \n     \n\n\ndef get_vote_case_str str:\n    (k2, v2), (k1, v1) = x.sort_values()[-:].items()\n     v1 == :\n         k1.split()[]\n     v1 &gt;  + tol:\n         k1.split()[]\n     (k1.split()[], k2.split()[])\n\n\ndef get_vote_cases_df pd.DataFrame:\n    vote_cases_file = output_path / \n\n      vote_cases_file.exists():\n        vote_cases = vote_probs.apply(get_vote_case, axis=)\n        vote_cases_str = vote_probs.apply(get_vote_case_str, axis=)\n        df = pd.DataFrame({: vote_cases, : vote_cases_str})\n        df.to_csv(vote_cases_file)\n\n     pd.read_csv(vote_cases_file).set_index()\n\n\nvote_cases_df = get_vote_cases_df()\nvote_cases = vote_cases_df[]\nvote_cases_str = vote_cases_df[]\n\n\n\n\n\nmne_plot_kwargs = dict(\n    scalings={: , : }, clipping=None, remove_dc=False, proj=False, show=False\n)\n\n\n\n\n\n\ndef set_double_banana_ref None:\n    \n    LL = [, , , , ]\n    LP = [, , , , ]\n    RL = [, , , , ]\n    RP = [, , , , ]\n    central = [, , ]\n\n    chains = [LL, LP, RL, RP, central]\n\n     chain  chains:\n        mne.set_bipolar_reference(raw, chain[:-], chain[:], copy=False, drop_refs=False)\n\n\n\n\n\n\n\ndef pick_bananas mne.io.Raw:\n    picks = [ch_name  ch_name  raw.info[]    ch_name]\n     raw.pick(picks)\n\n\ndef make_mne_raw(\n    eeg: pd.DataFrame,\n    notch: Optional[tuple[float]] = (,),\n    average_ref: bool = False,\n    banana_ref: bool = False,\n    ekg: bool = False,\n    high_pass: Optional[float] = None,\n    low_pass: Optional[float] = None,\n    ref_before_filt: bool = False,\n    ffill_na: bool = False,\n    fill_na_zero: bool = False,\n    fill_na_mean: bool = False,\n    detrend: bool = False,\n    debug: bool = False,\n    clip: Optional[float] = None,\n) -&gt; mne.io.Raw:\n      ekg:\n        eeg = eeg.drop(columns=)\n\n     ffill_na:\n        eeg = eeg.ffill(axis=, limit=)\n\n     fill_na_zero:\n        eeg = eeg.fillna()\n\n     fill_na_mean:\n        eeg = eeg.fillna(eeg.mean(axis=))\n\n    eeg_arr = eeg.values.T *   \n\n     debug:\n        n_nan = np.sum(np.isnan(eeg_arr))\n         n_nan &gt; :\n            (f)\n\n      ekg:\n        info = mne.create_info((eeg.columns), sfreq=, ch_types=)\n    else:\n        info = mne.create_info(\n            (eeg.columns), sfreq=, ch_types=[] * (len(eeg.columns) - ) + []\n        )\n    raw = mne.io.RawArray(eeg_arr, info=info)\n\n    raw.set_montage()\n\n     ref_before_filt:\n         average_ref:\n            raw.set_eeg_reference(ref_channels=)\n         banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n     notch:\n        raw.notch_filter(freqs=notch)\n\n     high_pass  low_pass:\n        raw.filter(l_freq=high_pass, h_freq=low_pass)\n\n     debug:\n        n_nan = raw.to_data_frame().isna().sum().sum()\n         n_nan &gt; :\n            (f)\n\n     clip:\n        clip = abs(clip) * \n        raw.apply_function(lambda x: np.clip(x, -clip, clip))\n\n     detrend:\n        raw.apply_function(lambda x: signal.detrend(x, type=))\n\n      ref_before_filt:\n         average_ref:\n            raw.set_eeg_reference(ref_channels=)\n         banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n     raw\n\n\n :\n    spec_regions = [, , , ]\n\n    def __init__(self, eeg_num: int, raw_kwargs: Optional[dict] = None):\n        self.eeg_id = eeg_ids[eeg_num]\n        self.eeg_path = data_path /  / (str(self.eeg_id) + )\n        self._data = None  \n        self._raw = None  \n        self._raw_kwargs = raw_kwargs  raw_kwargs   None  {}\n\n        meta = train_meta.loc[self.eeg_id]\n        self.spec_path = data_path /  / (str(meta[][]) + )\n        self._spec = None\n\n    @classmethod\n    def from_id(cls, eeg_id: int, raw_kwargs: Optional[dict] = None):\n        eeg_num = eeg_id_index[eeg_id]\n         cls(eeg_num, raw_kwargs)\n\n    @property\n    def meta pd.DataFrame:\n         train_meta.loc[self.eeg_id]\n\n    @property\n    def data pd.DataFrame:\n         self._data  None:\n            eeg = pd.read_parquet(self.eeg_path)\n            eeg = eeg.set_index(pd.to_timedelta( * eeg.index, unit=))\n            self._data = eeg\n         self._data\n\n    @property\n    def spec pd.DataFrame:\n         self._spec  None:\n            spec = pd.read_parquet(self.spec_path)\n            spec = spec.set_index(pd.to_timedelta(spec.time, unit=)).drop(columns=)\n            spec.columns = pd.MultiIndex.from_tuples(\n                [col.split()  col  spec.columns], names=(, )\n            )\n            self._spec = spec\n         self._spec\n\n    @property\n    def raw mne.io.Raw:\n         self._raw  None:\n            self._raw = make_mne_raw(self.data, **self._raw_kwargs)\n            ann_dict = {\n                : self.eeg_offsets + ,\n                : [] * len(self.eeg_offsets),\n                : vote_cases_str.loc[self.eeg_id],\n            }\n            self._raw.set_annotations(mne.Annotations(**ann_dict))\n         self._raw\n\n    @property\n    def eeg_offsets(self):\n         self.meta[]\n\n    @property\n    def spec_offsets(self):\n         self.meta[]\n\n    @property\n    def offsets(self):\n         pd.DataFrame({: self.eeg_offsets, : self.spec_offsets})\n\n    def sub_eeg pd.DataFrame:\n        offset = self.eeg_offsets[sub_eeg_id]\n         self.data.loc[slice(f, f)]\n\n    def sub_spec pd.DataFrame:\n        offset = self.spec_offsets[sub_eeg_id]\n\n        \n        \n        \n         match_eeg_window:\n            window = slice(f, f)\n        else:\n            window = slice(f, f)\n\n         self.spec.loc[window]\n\n\ndef spec_freq_medians_region_agg pd.Series:\n    total_medians0 = spec.T.unstack(level=).sort_index(key=lambda x: x.astype(float)).median(axis=)\n    total_medians = pd.DataFrame({region: total_medians0  region  EEG.spec_regions}).unstack()\n    total_medians.index.names = (, )\n     total_medians\n\n\ndef to_decibel pd.DataFrame:\n      isinstance(baseline, pd.Series):\n         baseline == :\n            baseline = spec.median(axis=)\n        elif baseline == :\n            baseline = spec_freq_medians_region_agg(spec)\n        else:\n            raise ValueError()\n    spec = spec.div(baseline, axis=).fillna()  \n\n    spec =  * np.log10(spec.clip(, 5))\n     spec\n\n\n\ndef plot_spec(\n    spec: pd.DataFrame,\n    show_label_window: bool = True,\n    axs: Optional[[mpl.axes.Axes]] = None,\n    cbar: bool = True,\n    **kwargs,\n) -&gt; None:\n     axs  None:\n        fig, axs = plt.subplots(, , figsize=(, ), sharex=True)\n\n    vmin, vmax = spec.min().min(), spec.max().max()\n\n    time_ticks_stride = len(spec.index)  ]\n        td = pd.Timedelta()\n        label_times = times[(times &gt;= mid_point - td) &amp; (times &lt;= mid_point + td)].seconds\n\n     region, ax  zip(EEG.spec_regions, axs):\n        \n        data = spec.loc[:, region].T.sort_index(ascending=False, key=lambda x: x.astype(float))\n\n        \n        data.columns = data.columns.seconds\n\n         region != EEG.spec_regions[-]:\n            sns.heatmap(\n                data=data, ax=ax, xticklabels=False, vmin=vmin, vmax=vmax, cmap=, cbar=cbar, **kwargs\n            )\n\n             show_label_window:\n                label_data =  * data\n                label_data.loc[:, label_times] = \n                sns.heatmap(\n                    label_data,\n                    alpha=,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=False,\n                    mask=label_data.(bool),\n                    **kwargs,\n                )\n\n            ax.set(xlabel=)\n        else:\n            sns.heatmap(\n                data=data,\n                ax=ax,\n                xticklabels=time_ticks_stride,\n                vmin=vmin,\n                vmax=vmax,\n                cmap=,\n                cbar=cbar,\n                **kwargs,\n            )\n\n             show_label_window:\n                label_data =  * data\n                label_data.loc[:, label_times] = \n                sns.heatmap(\n                    label_data,\n                    alpha=,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=time_ticks_stride,\n                    mask=label_data.(bool),\n                    **kwargs,\n                )\n\n        ax.set(ylabel=f)\n        \n</code></pre>\n<p>To make the spec/eeg plot I ran:</p>\n<pre><code>eeg = from})\nplot\nplt.show\n</code></pre>\n<p>To make the \"evolution\" plot, I did</p>\n<pre><code> functools import reduce\n matplotlib as mpl\n matplotlib.pyplot as plt\n\n\n = train_meta.reset_index()[].apply(lambda x: set([x])).groupby()\n = g.apply(lambda x: len(reduce(lambda a, b: a.union(b), x)))\n\n = n_consensuses &gt; \n\n = vote_cases_str.loc[eeg_ids[ncon_filt]]\n\n = mixed_con_vote_str.groupby().apply(lambda x: list(x))\n\n = mixed_con_votes_evolve.apply(len) &lt; \n\n = {: , : , : , : , : , : }\n\n plot_votes_evolve_row(row_num, votes_list, ax, x_spacing=., y_spacing=.):\n     i, votes in enumerate(votes_list):\n         isinstance(votes, tuple):\n             offset, vote in zip([-., .], votes[::-]):\n                .scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote], alpha=.)\n         isinstance(votes, str) and  in votes:\n            \n             = votes[:-].split()\n             offset, vote in zip([-., .], votes[::-]):\n                .scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote.strip()], alpha=.)\n        :\n            .scatter(x_spacing * i, y_spacing * row_num, color=cat_colors[votes])\n\n\n, ax = plt.subplots(figsize=(, ))\n\n =\n\n = \n = .\n\n = mixed_con_votes_evolve.loc[len_filt]\n\n i in range(n_patients):\n    (i, mcve_filtered.iloc[i], ax, y_spacing=y_spacing)\n.legend(handles=patches)\n.set_yticks(ticks=(y_spacing * np.arange(n_patients)), labels=mcve_filtered.index[:n_patients])\n</code></pre>",
      "rawMarkdown": "Apologies in advance for this giant \"snippet\"... it would have been a lot of work to make it shorter.\n\n```\ndef plot_spec_eeg(eeg: EEG, offset: int = 0):\n    fig = plt.figure(figsize=(20, 10), layout=None)\n    gs = fig.add_gridspec(nrows=4, ncols=2, hspace=0.1, wspace=0.0)\n\n    axs0 = [fig.add_subplot(gs[i, 0]) for i in range(4)]\n\n    sub_spec = to_decibel(eeg.sub_spec(offset, match_eeg_window=True))\n    plot_spec(sub_spec, axs=axs0)\n\n    ax1 = fig.add_subplot(gs[:, 1])\n    raw_single_ann = eeg.raw.copy().set_annotations(mne.Annotations(**eeg.raw.annotations[offset]))\n    fig_raw = raw_single_ann.plot(\n        duration=50,\n        start=eeg.eeg_offsets[offset],\n        show_scrollbars=False,\n        show_scalebars=False,\n        **mne_plot_kwargs,\n    )\n\n    canvas = FigureCanvasAgg(fig_raw)\n    canvas.draw()\n    img = np.asarray(canvas.buffer_rgba())\n    im = ax1.imshow(img)\n    ax1.set_axis_off()\n```\n\nThis is the code snippet. The rest of the code you need to run it is:\n\n```\ntrain_meta = pd.read_csv(data_path / \"train.csv\").set_index([\"eeg_id\", \"eeg_sub_id\"]).sort_index()\nvote_cols = [col for col in train_meta.columns if str(col).endswith(\"_vote\")]\ntrain_meta[\"vote_total\"] = train_meta[vote_cols].sum(axis=1)\n\neeg_ids = train_meta.reset_index()[\"eeg_id\"].unique()\neeg_id_index = {eeg_ids[i]: i for i in range(len(eeg_ids))}\n\n# extra target info\nvote_probs = train_meta[vote_cols].div(train_meta[\"vote_total\"], axis=0)\n\n\ndef get_vote_case(x: pd.Series, tol: float = 0.05) -> str:\n    (k2, v2), (k1, v1) = x.sort_values()[-2:].items()\n    if v1 == 1.0:\n        return \"perfect\"\n    if v1 > 0.5 + tol:\n        return \"ideal\"\n    if \"other\" in k1 or \"other\" in k2:\n        return \"proto\"\n    return \"edge\"\n\n\ndef get_vote_case_str(x: pd.Series, tol: float = 0.05) -> str:\n    (k2, v2), (k1, v1) = x.sort_values()[-2:].items()\n    if v1 == 1.0:\n        return k1.split(\"_\")[0]\n    if v1 > 0.5 + tol:\n        return k1.split(\"_\")[0]\n    return (k1.split(\"_\")[0], k2.split(\"_\")[0])\n\n\ndef get_vote_cases_df() -> pd.DataFrame:\n    vote_cases_file = output_path / \"vote_cases.csv\"\n\n    if not vote_cases_file.exists():\n        vote_cases = vote_probs.apply(get_vote_case, axis=1)\n        vote_cases_str = vote_probs.apply(get_vote_case_str, axis=1)\n        df = pd.DataFrame({\"vote_case\": vote_cases, \"vote_case_str\": vote_cases_str})\n        df.to_csv(vote_cases_file)\n\n    return pd.read_csv(vote_cases_file).set_index(\"eeg_id\")\n\n\nvote_cases_df = get_vote_cases_df()\nvote_cases = vote_cases_df[\"vote_case\"]\nvote_cases_str = vote_cases_df[\"vote_case_str\"]\n\n\n# MNE code and EEG class\n\n# kwargs for mne.io.Raw.plot(); need to specify duration to show more than 10s\nmne_plot_kwargs = dict(\n    scalings={\"eeg\": 100e-6, \"ecg\": 1e-4}, clipping=None, remove_dc=False, proj=False, show=False\n)\n\n\n# Bipolar reference code\n# TODO: integrate this into EEG class?\n\n\ndef set_double_banana_ref(raw: mne.io.Raw) -> None:\n    \"\"\"Add bipolar reference (\"double banana\") channels to\n    mne.io.Raw object.\n\n    Modifies in place.\n    \"\"\"\n    LL = [\"Fp1\", \"F7\", \"T3\", \"T5\", \"O1\"]\n    LP = [\"Fp1\", \"F3\", \"C3\", \"P3\", \"O1\"]\n    RL = [\"Fp2\", \"F8\", \"T4\", \"T6\", \"O2\"]\n    RP = [\"Fp2\", \"F4\", \"C4\", \"P4\", \"O2\"]\n    central = [\"Fz\", \"Cz\", \"Pz\"]\n\n    chains = [LL, LP, RL, RP, central]\n\n    for chain in chains:\n        mne.set_bipolar_reference(raw, chain[:-1], chain[1:], copy=False, drop_refs=False)\n\n\n# There are some issues with this... the locations of these \"virtual\" channels is set as the same location as the first item in the channel tuple that is differenced. For instance, `(\"Fp1\", \"F7\")` yields the virtual channel `Fp1-F7`, which is assigned the same location as `Fp1`.\n#\n# Also, we didn't drop the original channels because Fp1, Fp2, O1, and O2 are each used twice, so we can't drop them as we go. The following code allows us to pick the bipolar channels:\n\n\ndef pick_bananas(raw: mne.io.Raw) -> mne.io.Raw:\n    picks = [ch_name for ch_name in raw.info[\"ch_names\"] if \"-\" in ch_name]\n    return raw.pick(picks)\n\n\ndef make_mne_raw(\n    eeg: pd.DataFrame,\n    notch: Optional[tuple[float]] = (60,),\n    average_ref: bool = False,\n    banana_ref: bool = False,\n    ekg: bool = False,\n    high_pass: Optional[float] = None,\n    low_pass: Optional[float] = None,\n    ref_before_filt: bool = False,\n    ffill_na: bool = False,\n    fill_na_zero: bool = False,\n    fill_na_mean: bool = False,\n    detrend: bool = False,\n    debug: bool = False,\n    clip: Optional[float] = None,\n) -> mne.io.Raw:\n    if not ekg:\n        eeg = eeg.drop(columns=\"EKG\")\n\n    if ffill_na:\n        eeg = eeg.ffill(axis=0, limit=5)\n\n    if fill_na_zero:\n        eeg = eeg.fillna(0.0)\n\n    if fill_na_mean:\n        eeg = eeg.fillna(eeg.mean(axis=0))\n\n    eeg_arr = eeg.values.T * 1e-6  # units = microvolts\n\n    if debug:\n        n_nan = np.sum(np.isnan(eeg_arr))\n        if n_nan > 0:\n            print(f\"Number NaN: {n_nan}\")\n\n    if not ekg:\n        info = mne.create_info(list(eeg.columns), sfreq=200.0, ch_types=\"eeg\")\n    else:\n        info = mne.create_info(\n            list(eeg.columns), sfreq=200.0, ch_types=[\"eeg\"] * (len(eeg.columns) - 1) + [\"ecg\"]\n        )\n    raw = mne.io.RawArray(eeg_arr, info=info)\n\n    raw.set_montage(\"standard_1020\")\n\n    if ref_before_filt:\n        if average_ref:\n            raw.set_eeg_reference(ref_channels=\"average\")\n        if banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n    if notch:\n        raw.notch_filter(freqs=notch)\n\n    if high_pass or low_pass:\n        raw.filter(l_freq=high_pass, h_freq=low_pass)\n\n    if debug:\n        n_nan = raw.to_data_frame().isna().sum().sum()\n        if n_nan > 0:\n            print(f\"Number NaN: {n_nan}\")\n\n    if clip:\n        clip = abs(clip) * 1e-6\n        raw.apply_function(lambda x: np.clip(x, -clip, clip))\n\n    if detrend:\n        raw.apply_function(lambda x: signal.detrend(x, type=\"linear\"))\n\n    if not ref_before_filt:\n        if average_ref:\n            raw.set_eeg_reference(ref_channels=\"average\")\n        if banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n    return raw\n\n\nclass EEG:\n    spec_regions = [\"LL\", \"LP\", \"RP\", \"RL\"]\n\n    def __init__(self, eeg_num: int, raw_kwargs: Optional[dict] = None):\n        self.eeg_id = eeg_ids[eeg_num]\n        self.eeg_path = data_path / \"train_eegs\" / (str(self.eeg_id) + \".parquet\")\n        self._data = None  # lazy load data\n        self._raw = None  # MNE raw EEG\n        self._raw_kwargs = raw_kwargs if raw_kwargs is not None else {}\n\n        meta = train_meta.loc[self.eeg_id]\n        self.spec_path = data_path / \"train_spectrograms\" / (str(meta[\"spectrogram_id\"][0]) + \".parquet\")\n        self._spec = None\n\n    @classmethod\n    def from_id(cls, eeg_id: int, raw_kwargs: Optional[dict] = None):\n        eeg_num = eeg_id_index[eeg_id]\n        return cls(eeg_num, raw_kwargs)\n\n    @property\n    def meta(self) -> pd.DataFrame:\n        return train_meta.loc[self.eeg_id]\n\n    @property\n    def data(self) -> pd.DataFrame:\n        if self._data is None:\n            eeg = pd.read_parquet(self.eeg_path)\n            eeg = eeg.set_index(pd.to_timedelta(5 * eeg.index, unit=\"ms\"))\n            self._data = eeg\n        return self._data\n\n    @property\n    def spec(self) -> pd.DataFrame:\n        if self._spec is None:\n            spec = pd.read_parquet(self.spec_path)\n            spec = spec.set_index(pd.to_timedelta(spec.time, unit=\"s\")).drop(columns=\"time\")\n            spec.columns = pd.MultiIndex.from_tuples(\n                [col.split(\"_\") for col in spec.columns], names=(\"region\", \"frequency\")\n            )\n            self._spec = spec\n        return self._spec\n\n    @property\n    def raw(self) -> mne.io.Raw:\n        if self._raw is None:\n            self._raw = make_mne_raw(self.data, **self._raw_kwargs)\n            ann_dict = {\n                \"onset\": self.eeg_offsets + 20.0,\n                \"duration\": [10.0] * len(self.eeg_offsets),\n                \"description\": vote_cases_str.loc[self.eeg_id],\n            }\n            self._raw.set_annotations(mne.Annotations(**ann_dict))\n        return self._raw\n\n    @property\n    def eeg_offsets(self):\n        return self.meta[\"eeg_label_offset_seconds\"]\n\n    @property\n    def spec_offsets(self):\n        return self.meta[\"spectrogram_label_offset_seconds\"]\n\n    @property\n    def offsets(self):\n        return pd.DataFrame({\"eeg_offset\": self.eeg_offsets, \"spec_offset\": self.spec_offsets})\n\n    def sub_eeg(self, sub_eeg_id: int) -> pd.DataFrame:\n        offset = self.eeg_offsets[sub_eeg_id]\n        return self.data.loc[slice(f\"{offset}s\", f\"{49 + offset}s\")]\n\n    def sub_spec(self, sub_eeg_id: int, match_eeg_window: bool = False) -> pd.DataFrame:\n        offset = self.spec_offsets[sub_eeg_id]\n\n        # NOTE: we're adding 0.1s to the end of the slices to make sure they don't end\n        # with 0 seconds, in which case the units are inferred as minutes, and we get an\n        # extra 59 seconds.\n        if match_eeg_window:\n            window = slice(f\"4 minutes {59 + offset - 25 + 0.1}s\", f\"4 minutes {59 + offset + 25 + 0.1}s\")\n        else:\n            window = slice(f\"{offset + 0.1}s\", f\"10 minutes {offset + 0.1}s\")\n\n        return self.spec.loc[window]\n\n\ndef spec_freq_medians_region_agg(spec: pd.DataFrame) -> pd.Series:\n    total_medians0 = spec.T.unstack(level=\"region\").sort_index(key=lambda x: x.astype(float)).median(axis=1)\n    total_medians = pd.DataFrame({region: total_medians0 for region in EEG.spec_regions}).unstack()\n    total_medians.index.names = (\"region\", \"frequency\")\n    return total_medians\n\n\ndef to_decibel(spec: pd.DataFrame, baseline: Union[pd.Series, Literal[\"agg\", \"sep\"]] = \"agg\") -> pd.DataFrame:\n    if not isinstance(baseline, pd.Series):\n        if baseline == \"sep\":\n            baseline = spec.median(axis=0)\n        elif baseline == \"agg\":\n            baseline = spec_freq_medians_region_agg(spec)\n        else:\n            raise ValueError(\"baseline must be 'agg' or 'sep' or pd.Series\")\n    spec = spec.div(baseline, axis=1).fillna(0.0)  # if median was 0, then row show be zero\n\n    spec = 10 * np.log10(spec.clip(1e-5, 1e5))\n    return spec\n\n\n# Plotting code\ndef plot_spec(\n    spec: pd.DataFrame,\n    show_label_window: bool = True,\n    axs: Optional[list[mpl.axes.Axes]] = None,\n    cbar: bool = True,\n    **kwargs,\n) -> None:\n    if axs is None:\n        fig, axs = plt.subplots(4, 1, figsize=(8, 10), sharex=True)\n\n    vmin, vmax = spec.min().min(), spec.max().max()\n\n    time_ticks_stride = len(spec.index) // 25\n\n    if show_label_window:\n        times = spec.index\n        fmin, fmax = 0, len(spec.loc[:, EEG.spec_regions[0]].columns)\n        mid_point = times[len(times) // 2]\n        td = pd.Timedelta(\"5s\")\n        label_times = times[(times >= mid_point - td) & (times <= mid_point + td)].seconds\n\n    for region, ax in zip(EEG.spec_regions, axs):\n        # sort frequencies in descending order\n        data = spec.loc[:, region].T.sort_index(ascending=False, key=lambda x: x.astype(float))\n\n        # convert times to seconds (else seaborn converts to nanoseconds)\n        data.columns = data.columns.seconds\n\n        if region != EEG.spec_regions[-1]:\n            sns.heatmap(\n                data=data, ax=ax, xticklabels=False, vmin=vmin, vmax=vmax, cmap=\"jet\", cbar=cbar, **kwargs\n            )\n\n            if show_label_window:\n                label_data = 0.0 * data\n                label_data.loc[:, label_times] = 1.0\n                sns.heatmap(\n                    label_data,\n                    alpha=0.4,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=False,\n                    mask=label_data.map(bool),\n                    **kwargs,\n                )\n\n            ax.set(xlabel=\"\")\n        else:\n            sns.heatmap(\n                data=data,\n                ax=ax,\n                xticklabels=time_ticks_stride,\n                vmin=vmin,\n                vmax=vmax,\n                cmap=\"jet\",\n                cbar=cbar,\n                **kwargs,\n            )\n\n            if show_label_window:\n                label_data = 0.0 * data\n                label_data.loc[:, label_times] = 1.0\n                sns.heatmap(\n                    label_data,\n                    alpha=0.4,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=time_ticks_stride,\n                    mask=label_data.map(bool),\n                    **kwargs,\n                )\n\n        ax.set(ylabel=f\"{region} frequency (Hz)\")\n        # ax.set_title(region, loc=\"left\")\n\n```\n\nTo make the spec/eeg plot I ran:\n```\neeg = EEG.from_id(21379701, raw_kwargs={\"banana_ref\": True, \"ffill_na\": True, \"notch\": (60.0,)})\nplot_spec_eeg(eeg, offset=13)\nplt.show()\n```\n\nTo make the \"evolution\" plot, I did\n\n```\nfrom functools import reduce\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\n\n\ng = train_meta.reset_index(\"eeg_sub_id\")[\"expert_consensus\"].apply(lambda x: set([x])).groupby(\"eeg_id\")\nn_consensuses = g.apply(lambda x: len(reduce(lambda a, b: a.union(b), x)))\n\nncon_filt = n_consensuses > 1\n\nmixed_con_vote_str = vote_cases_str.loc[eeg_ids[ncon_filt]]\n\nmixed_con_votes_evolve = mixed_con_vote_str.groupby(\"eeg_id\").apply(lambda x: list(x))\n\nlen_filt = mixed_con_votes_evolve.apply(len) < 20\n\ncat_colors = {\"other\": \"lightslategrey\", \"lpd\": \"tomato\", \"gpd\": \"mediumseagreen\", \"lrda\": \"palevioletred\", \"grda\": \"greenyellow\", \"seizure\": \"dodgerblue\"}\n\ndef plot_votes_evolve_row(row_num, votes_list, ax, x_spacing=0.5, y_spacing=0.5):\n    for i, votes in enumerate(votes_list):\n        if isinstance(votes, tuple):\n            for offset, vote in zip([-0.1, 0.1], votes[::-1]):\n                ax.scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote], alpha=0.7)\n        elif isinstance(votes, str) and \", \" in votes:\n            # saving vote cases to csv means tuples of strings are loaded as a single string\n            votes = votes[1:-1].split(\", \")\n            for offset, vote in zip([-0.1, 0.1], votes[::-1]):\n                ax.scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote.strip(\"'\")], alpha=0.7)\n        else:\n            ax.scatter(x_spacing * i, y_spacing * row_num, color=cat_colors[votes])\n\n            \nfig, ax = plt.subplots(figsize=(10, 15))\n\npatches = [mpl.patches.Patch(color=v, label=k) for k, v in cat_colors.items()]\n\nn_patients = 50\ny_spacing = 0.2\n\nmcve_filtered = mixed_con_votes_evolve.loc[len_filt]\n\nfor i in range(n_patients):\n    plot_votes_evolve_row(i, mcve_filtered.iloc[i], ax, y_spacing=y_spacing)\nax.legend(handles=patches)\nax.set_yticks(ticks=(y_spacing * np.arange(n_patients)), labels=mcve_filtered.index[:n_patients])\n\n```",
      "votes": null
    },
    {
      "id": "2736634",
      "postDate": "04/05/2024 10:04:33",
      "content": "<p>Many thanks for all the snippets.  Can reproduce evolution plot and spectrogram.  Was really interested in the plot you showed with this section, if you have a chance to comment:</p>\n<blockquote>\n  <p>For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:</p>\n</blockquote>",
      "rawMarkdown": "Many thanks for all the snippets.  Can reproduce evolution plot and spectrogram.  Was really interested in the plot you showed with this section, if you have a chance to comment:\n\n>For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:",
      "votes": null
    },
    {
      "id": "2736744",
      "postDate": "04/05/2024 11:27:21",
      "content": "<p>That is mne's plotting, I think for the eeg plot I posted, it was  <code>eeg.raw.plot(duration=240, **mne_plot_kwargs)</code> (you need at least the eeg scale from my mne plot kwargs, or else it will look really weird). MNE automatically uses the annotations I made with the labels</p>",
      "rawMarkdown": "That is mne's plotting, I think for the eeg plot I posted, it was  `eeg.raw.plot(duration=240, **mne_plot_kwargs)` (you need at least the eeg scale from my mne plot kwargs, or else it will look really weird). MNE automatically uses the annotations I made with the labels",
      "votes": null
    },
    {
      "id": "2736797",
      "postDate": "04/05/2024 12:06:35",
      "content": "<p>Thanks!  Should have realised by the axis labels. (getting tired)    </p>",
      "rawMarkdown": "Thanks!  Should have realised by the axis labels. (getting tired)",
      "votes": null
    },
    {
      "id": "2737778",
      "postDate": "04/05/2024 23:48:53",
      "content": "<p>Tks for sharing 👊</p>",
      "rawMarkdown": "Tks for sharing 👊",
      "votes": null
    },
    {
      "id": "2738862",
      "postDate": "04/06/2024 17:01:01",
      "content": "<p>Thanks for bring this up! I also noticed this problem. Currently I am using Chris's method to extract EEG signal from the raw parquet, which always takes the middle 50s of a signal, regardless of the <code>eeg_label_offset_seconds</code>. To get non-overlap samples, I applied groupby on <code>eeg_id + targets</code>, which gives each unique <code>eeg_id + vote</code> pattern a single entry. In this case, the offset of the signal segment indeed does not match with the actual <code>eeg_label_offset_seconds</code>. That would lead to the fact that the same EEG segment is paired with multiple vote patterns, and thus confuse the model. </p>\n<p>Any solution yet to fix it? </p>\n<p>Attached is a plot of eeg_id = 21379170. Black Solid lines in the plot indicates the middle 50sec segment. Color represent different expert consensus. </p>\n<p>Code snippet:</p>\n<pre><code> =  [,,,,,,,]\nfig, axes = plt.subplots(, , figsize=(, ), sharex=)\nfor i, ax in enumerate(axes):\n    ax.plot(np.arange(len_seq), eeg_signal[:, i]) # eeg_signal (, ) the signal data from parquet\n    ax.text(, , [i], transform=ax.transAxes, ha=, va=)\n\ndef plot_timespan(ax, t1, t2, color, label=):\n    vspan = ax.axvspan(t1, t2, alpha=, color=color, label=label)\n    return ax, vspan\n\ncolors = [, , , , , ]\nfor i, row in train_csv[(train_csv[]==)].iterrows():\n    target_color = colors[row[targets].argmax()]\n    target_label = row[]\n    t_middle = row[] + \n    for ax in axes:\n        ax, vspan = plot_timespan(ax, (t_middle)*, (t_middle+)*, target_color, target_label)\n        ax.axvline((t_total//)*, color=, ls=, lw=)\n        ax.axvline((t_total//+)*, color=, ls=, lw=)\n\nax.legend(bbox_to_anchor=(, ), loc=)\nfig.tight_layout()\nplt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F278125a3c40ea5158c2271e3ab47b4c9%2Foutput.png?generation=1712422551573864&amp;alt=media\"></p>",
      "rawMarkdown": "Thanks for bring this up! I also noticed this problem. Currently I am using Chris's method to extract EEG signal from the raw parquet, which always takes the middle 50s of a signal, regardless of the `eeg_label_offset_seconds`. To get non-overlap samples, I applied groupby on `eeg_id + targets`, which gives each unique `eeg_id + vote` pattern a single entry. In this case, the offset of the signal segment indeed does not match with the actual `eeg_label_offset_seconds`. That would lead to the fact that the same EEG segment is paired with multiple vote patterns, and thus confuse the model. \n\nAny solution yet to fix it? \n\nAttached is a plot of eeg_id = 21379170. Black Solid lines in the plot indicates the middle 50sec segment. Color represent different expert consensus. \n\nCode snippet:\n\n```\nEEG_FEAT_USE =  ['Fp1','T3','C3','O1','Fp2','C4','T4','O2']\nfig, axes = plt.subplots(8, 1, figsize=(12, 8), sharex=True)\nfor i, ax in enumerate(axes):\n    ax.plot(np.arange(len_seq), eeg_signal[:, i]) # eeg_signal (L, C) the signal data from parquet\n    ax.text(0.01, 0.05, EEG_FEAT_USE[i], transform=ax.transAxes, ha='left', va='bottom')\n\ndef plot_timespan(ax, t1, t2, color, label=None):\n    vspan = ax.axvspan(t1, t2, alpha=0.2, color=color, label=label)\n    return ax, vspan\n\ncolors = ['r', 'g', 'b', 'c', 'm', 'y']\nfor i, row in train_csv[(train_csv['eeg_id']==21379701)].iterrows():\n    target_color = colors[row[targets].argmax()]\n    target_label = row['expert_consensus']\n    t_middle = row['eeg_label_offset_seconds'] + 25\n    for ax in axes:\n        ax, vspan = plot_timespan(ax, (t_middle-5)*200, (t_middle+5)*200, target_color, target_label)\n        ax.axvline((t_total//2-25)*200, color='k', ls='-', lw=2)\n        ax.axvline((t_total//2+25)*200, color='k', ls='-', lw=2)\n\nax.legend(bbox_to_anchor=(1.15, 4.5), loc='center right')\nfig.tight_layout()\nplt.show()\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F278125a3c40ea5158c2271e3ab47b4c9%2Foutput.png?generation=1712422551573864&alt=media)",
      "votes": null
    },
    {
      "id": "2738939",
      "postDate": "04/06/2024 18:01:35",
      "content": "<p>You can add an offset of 20 secs to eeg_label_offset_seconds, in theory this is the start of the 50s for the row label and they looked at the centre 10s  for it.  So offset add of 20s should get the centre 10 secs.  But there is overlap sometimes.</p>\n<blockquote>\n  <p>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.</p>\n</blockquote>",
      "rawMarkdown": "You can add an offset of 20 secs to eeg_label_offset_seconds, in theory this is the start of the 50s for the row label and they looked at the centre 10s  for it.  So offset add of 20s should get the centre 10 secs.  But there is overlap sometimes.\n\n>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2735209,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "04/04/2024 15:39:23",
      "content": "<p>Not necessarily. There is various subsamples in each eeg_id. Each one with different 10 seconds to be labeled. Even beign the same patient can have different labels. Even more if each one have been labeled independently. The model has to predict more various experts vote distribution than a true label.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2735220,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "04/04/2024 15:42:40",
      "content": "<p>see the Data page - </p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.</p>\n  <ul>\n  <li><p>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.</p></li>\n  \n  <li><p>label_id - An ID for this set of labels.</p></li>\n  </ul>\n</blockquote>\n<p>Have been working on this after finding issues and considering all the discussions about &gt;= 10 votes.<br>\nReally don't think that was made clear about how train data was put together. But probably some figured it out earlier?<br>\nThe labels are not consistently for a 50s or even 10s period, so deciding how to split out<br>\nHave a csv file could put in a dataset, not sure if too late for sharing though.  Value counts below for train expert consensus counts and the 5 entry of 234 is just for one eeg id.  Possibly these could/should be left out of training.</p>\n<blockquote>\n  <p>exp_cons_cnts<br>\n  1    93956<br>\n  2     9134<br>\n  3     2187<br>\n  4     1289<br>\n  5      234`</p>\n</blockquote>\n<p>Considering the test data is meant to be 50s only per eeg id and one spectrogram no overlapping, not sure if this will also happen in test labels.  And think is meant to be only the centre 10s that was labeled not sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2736346,
      "author_name": "brendandanger",
      "author_url": "",
      "post_date": "04/05/2024 06:35:36",
      "content": "<p>If you plot the consensuses (?) against time, you'll see that they evolve in a reasonable manner: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fbbfd41171920532c3d3ff34d4e4bc517%2Fconsensus_evolve.png?generation=1712298663242029&amp;alt=media\"></p>\n<p>If the top vote is less than 60%, I plotted the top two labels. Otherwise, I plotted the top label. Most of the transitions make sense (e.g. LPD to LRDA or GPD to GRDA; also seizures and LPD seem to mix. I read that patients with LPD are more likely to have seizures, although I don't remember on what time scale).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2736381,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "04/05/2024 06:54:57",
          "content": "<p>Think the issue is with time - acc to Data page, the labels are for 50 sec subsamples from the eeg_label_offset_seconds, a bit vague if the central 10 secs applies to eeg data, spectrograms or both.  So taking a centre 10 secs from a 50 sec window of eeg data does not always (if ever) equate to the labels. Perhaps windowing by 2-4 secs might be closer. And train eegs have a variety of time lengths, well beyond 50 secs.  All of this probably accounts for the high number of Other in train data, maybe why &gt;=10 seems to work well.  Many of the public notebooks are not really considering eeg_label_offset_seconds from what I can tell.</p>\n<p>In contrast,  the test data is meant to be 50s only per eeg id and one spectrogram no overlapping. Somewhat apples and oranges.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2736429,
              "author_name": "brendandanger",
              "author_url": "",
              "post_date": "04/05/2024 07:39:35",
              "content": "<p>For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fb29caca1b03a8a31eab2ca75b734ac68%2F21379701_eeg.png?generation=1712302677695583&amp;alt=media\"><br>\nEach 10s region is the center of a 50s region (which isn't shown). (Note: the colors here aren't the same as the colors on the previous plot)</p>\n<p>These are the labels and their locations:</p>\n<p>onset    duration    description<br>\n20.0      10.0    ('other', 'lrda')<br>\n30.0      10.0    ('other', 'lrda')<br>\n32.0      10.0    ('other', 'lrda')<br>\n50.0      10.0    ('seizure', 'other')<br>\n64.0      10.0    seizure<br>\n68.0      10.0    seizure<br>\n82.0      10.0    lpd<br>\n96.0      10.0    lpd<br>\n98.0      10.0    lpd<br>\n100.0      10.0    lpd<br>\n102.0      10.0    lpd<br>\n106.0      10.0    lpd<br>\n110.0      10.0    lpd<br>\n124.0      10.0    lpd<br>\n126.0      10.0    lpd<br>\n128.0      10.0    lpd<br>\n130.0      10.0    lpd<br>\n178.0      10.0    other</p>\n<p>The first offset was 0s, so the 10s scored region starts at 20s and ends at 30s, etc.</p>\n<p>This is the first 50s window for spec (cropped to the central 50s of the 10 minute window) and eeg, with the middle 10s highlighted:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2F30769663cb2ccd66f92193a2087fd355%2F21379701_eeg_spec.png?generation=1712303735165968&amp;alt=media\"><br>\nYou can see that there is a high frequency response in the spectrogram that corresponds with the EEG data in time. So I think the offsets are accurately described (although a bit cryptic, some of the other posts discussing the offsets were helpful.)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2736489,
                  "author_name": "something4kag",
                  "author_url": "",
                  "post_date": "04/05/2024 08:13:29",
                  "content": "<blockquote>\n  <p>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.<br>\n  eeg_label_offset_seconds - The time between the beginning of the consolidated EEG and this subsample.</p>\n</blockquote>\n<p>From train, the most votes for (21379701) are Other and Seizure and LPD. Using train eeg_label_offset_seconds starting at 0 for 50s will have Other (2 or 1)  and Seizure (1 or 2)  and LRDA (1 or 0) which consumes the entries up to 48s.  The next train at 62s for 50s will have LPD 75% of votes LRDA 25% of votes.  All these are low vote counts 1-3.  This eeg data has 208s or 41600 rows.<br>\n.<br>\nIt is difficult to read the text corresponding to labels, and not sure where your 50s region starts. Can only go by what is provided in train in terms of where eeg data and labels are meant to be in the parquet files.</p>\n<p>EDIT: your post came after was doing this.  <br>\nfrom train Other corr to 0, 10, 12 secs and 158s.  Seizure 30, 34, 48. LPD 62 to 106 at 2-4s intervals</p>\n<p>What is  the source of your plots?  Know a notebook is not possible now but maybe a code snippet, looks like using mne?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2736532,
                      "author_name": "brendandanger",
                      "author_url": "",
                      "post_date": "04/05/2024 08:54:36",
                      "content": "<p>Apologies in advance for this giant \"snippet\"… it would have been a lot of work to make it shorter.</p>\n<pre><code>def plot:\n    fig = plt.figure(figsize=(, ), layout=None)\n    gs = fig.add\n\n    axs0 = )  i  range()]\n\n    sub_spec = )\n    plot\n\n    ax1 = fig.add\n    raw_single_ann = eeg.raw.copy.set)\n    fig_raw = raw_single_ann.plot(\n        duration=,\n        start=eeg.eeg_offsets,\n        show_scrollbars=False,\n        show_scalebars=False,\n        **mne_plot_kwargs,\n    )\n\n    canvas = \n    canvas.draw\n    img = np.asarray(canvas.buffer)\n    im = ax1.imshow(img)\n    ax1.set\n</code></pre>\n<p>This is the code snippet. The rest of the code you need to run it is:</p>\n<pre><code>train_meta = pd.read_csv(data_path / ).set_index([, ]).sort_index()\nvote_cols = [col  col  train_meta.columns  str(col).endswith()]\ntrain_meta[] = train_meta[vote_cols].sum(axis=)\n\neeg_ids = train_meta.reset_index()[].unique()\neeg_id_index = {eeg_ids[i]: i  i  range(len(eeg_ids))}\n\n\nvote_probs = train_meta[vote_cols].div(train_meta[], axis=)\n\n\ndef get_vote_case str:\n    (k2, v2), (k1, v1) = x.sort_values()[-:].items()\n     v1 == :\n         \n     v1 &gt;  + tol:\n         \n       k1    k2:\n         \n     \n\n\ndef get_vote_case_str str:\n    (k2, v2), (k1, v1) = x.sort_values()[-:].items()\n     v1 == :\n         k1.split()[]\n     v1 &gt;  + tol:\n         k1.split()[]\n     (k1.split()[], k2.split()[])\n\n\ndef get_vote_cases_df pd.DataFrame:\n    vote_cases_file = output_path / \n\n      vote_cases_file.exists():\n        vote_cases = vote_probs.apply(get_vote_case, axis=)\n        vote_cases_str = vote_probs.apply(get_vote_case_str, axis=)\n        df = pd.DataFrame({: vote_cases, : vote_cases_str})\n        df.to_csv(vote_cases_file)\n\n     pd.read_csv(vote_cases_file).set_index()\n\n\nvote_cases_df = get_vote_cases_df()\nvote_cases = vote_cases_df[]\nvote_cases_str = vote_cases_df[]\n\n\n\n\n\nmne_plot_kwargs = dict(\n    scalings={: , : }, clipping=None, remove_dc=False, proj=False, show=False\n)\n\n\n\n\n\n\ndef set_double_banana_ref None:\n    \n    LL = [, , , , ]\n    LP = [, , , , ]\n    RL = [, , , , ]\n    RP = [, , , , ]\n    central = [, , ]\n\n    chains = [LL, LP, RL, RP, central]\n\n     chain  chains:\n        mne.set_bipolar_reference(raw, chain[:-], chain[:], copy=False, drop_refs=False)\n\n\n\n\n\n\n\ndef pick_bananas mne.io.Raw:\n    picks = [ch_name  ch_name  raw.info[]    ch_name]\n     raw.pick(picks)\n\n\ndef make_mne_raw(\n    eeg: pd.DataFrame,\n    notch: Optional[tuple[float]] = (,),\n    average_ref: bool = False,\n    banana_ref: bool = False,\n    ekg: bool = False,\n    high_pass: Optional[float] = None,\n    low_pass: Optional[float] = None,\n    ref_before_filt: bool = False,\n    ffill_na: bool = False,\n    fill_na_zero: bool = False,\n    fill_na_mean: bool = False,\n    detrend: bool = False,\n    debug: bool = False,\n    clip: Optional[float] = None,\n) -&gt; mne.io.Raw:\n      ekg:\n        eeg = eeg.drop(columns=)\n\n     ffill_na:\n        eeg = eeg.ffill(axis=, limit=)\n\n     fill_na_zero:\n        eeg = eeg.fillna()\n\n     fill_na_mean:\n        eeg = eeg.fillna(eeg.mean(axis=))\n\n    eeg_arr = eeg.values.T *   \n\n     debug:\n        n_nan = np.sum(np.isnan(eeg_arr))\n         n_nan &gt; :\n            (f)\n\n      ekg:\n        info = mne.create_info((eeg.columns), sfreq=, ch_types=)\n    else:\n        info = mne.create_info(\n            (eeg.columns), sfreq=, ch_types=[] * (len(eeg.columns) - ) + []\n        )\n    raw = mne.io.RawArray(eeg_arr, info=info)\n\n    raw.set_montage()\n\n     ref_before_filt:\n         average_ref:\n            raw.set_eeg_reference(ref_channels=)\n         banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n     notch:\n        raw.notch_filter(freqs=notch)\n\n     high_pass  low_pass:\n        raw.filter(l_freq=high_pass, h_freq=low_pass)\n\n     debug:\n        n_nan = raw.to_data_frame().isna().sum().sum()\n         n_nan &gt; :\n            (f)\n\n     clip:\n        clip = abs(clip) * \n        raw.apply_function(lambda x: np.clip(x, -clip, clip))\n\n     detrend:\n        raw.apply_function(lambda x: signal.detrend(x, type=))\n\n      ref_before_filt:\n         average_ref:\n            raw.set_eeg_reference(ref_channels=)\n         banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n     raw\n\n\n :\n    spec_regions = [, , , ]\n\n    def __init__(self, eeg_num: int, raw_kwargs: Optional[dict] = None):\n        self.eeg_id = eeg_ids[eeg_num]\n        self.eeg_path = data_path /  / (str(self.eeg_id) + )\n        self._data = None  \n        self._raw = None  \n        self._raw_kwargs = raw_kwargs  raw_kwargs   None  {}\n\n        meta = train_meta.loc[self.eeg_id]\n        self.spec_path = data_path /  / (str(meta[][]) + )\n        self._spec = None\n\n    @classmethod\n    def from_id(cls, eeg_id: int, raw_kwargs: Optional[dict] = None):\n        eeg_num = eeg_id_index[eeg_id]\n         cls(eeg_num, raw_kwargs)\n\n    @property\n    def meta pd.DataFrame:\n         train_meta.loc[self.eeg_id]\n\n    @property\n    def data pd.DataFrame:\n         self._data  None:\n            eeg = pd.read_parquet(self.eeg_path)\n            eeg = eeg.set_index(pd.to_timedelta( * eeg.index, unit=))\n            self._data = eeg\n         self._data\n\n    @property\n    def spec pd.DataFrame:\n         self._spec  None:\n            spec = pd.read_parquet(self.spec_path)\n            spec = spec.set_index(pd.to_timedelta(spec.time, unit=)).drop(columns=)\n            spec.columns = pd.MultiIndex.from_tuples(\n                [col.split()  col  spec.columns], names=(, )\n            )\n            self._spec = spec\n         self._spec\n\n    @property\n    def raw mne.io.Raw:\n         self._raw  None:\n            self._raw = make_mne_raw(self.data, **self._raw_kwargs)\n            ann_dict = {\n                : self.eeg_offsets + ,\n                : [] * len(self.eeg_offsets),\n                : vote_cases_str.loc[self.eeg_id],\n            }\n            self._raw.set_annotations(mne.Annotations(**ann_dict))\n         self._raw\n\n    @property\n    def eeg_offsets(self):\n         self.meta[]\n\n    @property\n    def spec_offsets(self):\n         self.meta[]\n\n    @property\n    def offsets(self):\n         pd.DataFrame({: self.eeg_offsets, : self.spec_offsets})\n\n    def sub_eeg pd.DataFrame:\n        offset = self.eeg_offsets[sub_eeg_id]\n         self.data.loc[slice(f, f)]\n\n    def sub_spec pd.DataFrame:\n        offset = self.spec_offsets[sub_eeg_id]\n\n        \n        \n        \n         match_eeg_window:\n            window = slice(f, f)\n        else:\n            window = slice(f, f)\n\n         self.spec.loc[window]\n\n\ndef spec_freq_medians_region_agg pd.Series:\n    total_medians0 = spec.T.unstack(level=).sort_index(key=lambda x: x.astype(float)).median(axis=)\n    total_medians = pd.DataFrame({region: total_medians0  region  EEG.spec_regions}).unstack()\n    total_medians.index.names = (, )\n     total_medians\n\n\ndef to_decibel pd.DataFrame:\n      isinstance(baseline, pd.Series):\n         baseline == :\n            baseline = spec.median(axis=)\n        elif baseline == :\n            baseline = spec_freq_medians_region_agg(spec)\n        else:\n            raise ValueError()\n    spec = spec.div(baseline, axis=).fillna()  \n\n    spec =  * np.log10(spec.clip(, 5))\n     spec\n\n\n\ndef plot_spec(\n    spec: pd.DataFrame,\n    show_label_window: bool = True,\n    axs: Optional[[mpl.axes.Axes]] = None,\n    cbar: bool = True,\n    **kwargs,\n) -&gt; None:\n     axs  None:\n        fig, axs = plt.subplots(, , figsize=(, ), sharex=True)\n\n    vmin, vmax = spec.min().min(), spec.max().max()\n\n    time_ticks_stride = len(spec.index)  ]\n        td = pd.Timedelta()\n        label_times = times[(times &gt;= mid_point - td) &amp; (times &lt;= mid_point + td)].seconds\n\n     region, ax  zip(EEG.spec_regions, axs):\n        \n        data = spec.loc[:, region].T.sort_index(ascending=False, key=lambda x: x.astype(float))\n\n        \n        data.columns = data.columns.seconds\n\n         region != EEG.spec_regions[-]:\n            sns.heatmap(\n                data=data, ax=ax, xticklabels=False, vmin=vmin, vmax=vmax, cmap=, cbar=cbar, **kwargs\n            )\n\n             show_label_window:\n                label_data =  * data\n                label_data.loc[:, label_times] = \n                sns.heatmap(\n                    label_data,\n                    alpha=,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=False,\n                    mask=label_data.(bool),\n                    **kwargs,\n                )\n\n            ax.set(xlabel=)\n        else:\n            sns.heatmap(\n                data=data,\n                ax=ax,\n                xticklabels=time_ticks_stride,\n                vmin=vmin,\n                vmax=vmax,\n                cmap=,\n                cbar=cbar,\n                **kwargs,\n            )\n\n             show_label_window:\n                label_data =  * data\n                label_data.loc[:, label_times] = \n                sns.heatmap(\n                    label_data,\n                    alpha=,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=time_ticks_stride,\n                    mask=label_data.(bool),\n                    **kwargs,\n                )\n\n        ax.set(ylabel=f)\n        \n</code></pre>\n<p>To make the spec/eeg plot I ran:</p>\n<pre><code>eeg = from})\nplot\nplt.show\n</code></pre>\n<p>To make the \"evolution\" plot, I did</p>\n<pre><code> functools import reduce\n matplotlib as mpl\n matplotlib.pyplot as plt\n\n\n = train_meta.reset_index()[].apply(lambda x: set([x])).groupby()\n = g.apply(lambda x: len(reduce(lambda a, b: a.union(b), x)))\n\n = n_consensuses &gt; \n\n = vote_cases_str.loc[eeg_ids[ncon_filt]]\n\n = mixed_con_vote_str.groupby().apply(lambda x: list(x))\n\n = mixed_con_votes_evolve.apply(len) &lt; \n\n = {: , : , : , : , : , : }\n\n plot_votes_evolve_row(row_num, votes_list, ax, x_spacing=., y_spacing=.):\n     i, votes in enumerate(votes_list):\n         isinstance(votes, tuple):\n             offset, vote in zip([-., .], votes[::-]):\n                .scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote], alpha=.)\n         isinstance(votes, str) and  in votes:\n            \n             = votes[:-].split()\n             offset, vote in zip([-., .], votes[::-]):\n                .scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote.strip()], alpha=.)\n        :\n            .scatter(x_spacing * i, y_spacing * row_num, color=cat_colors[votes])\n\n\n, ax = plt.subplots(figsize=(, ))\n\n =\n\n = \n = .\n\n = mixed_con_votes_evolve.loc[len_filt]\n\n i in range(n_patients):\n    (i, mcve_filtered.iloc[i], ax, y_spacing=y_spacing)\n.legend(handles=patches)\n.set_yticks(ticks=(y_spacing * np.arange(n_patients)), labels=mcve_filtered.index[:n_patients])\n</code></pre>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2736634,
                          "author_name": "something4kag",
                          "author_url": "",
                          "post_date": "04/05/2024 10:04:33",
                          "content": "<p>Many thanks for all the snippets.  Can reproduce evolution plot and spectrogram.  Was really interested in the plot you showed with this section, if you have a chance to comment:</p>\n<blockquote>\n  <p>For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:</p>\n</blockquote>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2736744,
                              "author_name": "brendandanger",
                              "author_url": "",
                              "post_date": "04/05/2024 11:27:21",
                              "content": "<p>That is mne's plotting, I think for the eeg plot I posted, it was  <code>eeg.raw.plot(duration=240, **mne_plot_kwargs)</code> (you need at least the eeg scale from my mne plot kwargs, or else it will look really weird). MNE automatically uses the annotations I made with the labels</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2736797,
                                  "author_name": "something4kag",
                                  "author_url": "",
                                  "post_date": "04/05/2024 12:06:35",
                                  "content": "<p>Thanks!  Should have realised by the axis labels. (getting tired)    </p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2736492,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "04/05/2024 08:15:31",
          "content": "<p>If it is expected that a patient's state or condition can transition from one to another over time. To accurately capture these transitions, the <code>eeg_label_offset_seconds</code> should be utilized for each specific state or diagnosis.</p>\n<p>I have been applying the same <code>eeg_label_offset_seconds</code> value for a single <code>eeg_id</code>, under the assumption that each <code>eeg_id</code> would have only one associated diagnosis or state. This approach might not be appropriate, as some patients have multiple diagnoses associated with their EEG recordings. Incorporating the correct <code>eeg_label_offset_seconds</code> for each diagnosis or state associated with an <code>eeg_id</code>, the  model captures the temporal dynamics and transitions within the EEG data.</p>\n<p>While the number of samples with multiple diagnoses might not be substantial, considering this issue and accounting for potential transitions between different states or diagnoses could potentially improve the model's performance.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2736497,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "04/05/2024 08:27:17",
              "content": "<p>Not entirely sure the train eeg_label_offset_seconds correctly corresponds to the row labels. And it is not just about multiple labels per eeg id, but getting the vote percentages somewhere in the ballpark. <br>\nThink the wording of a 50s subsample that the labels applied to was either misleading or incorrect.  </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2737778,
      "author_name": "rafaelzimmermann1",
      "author_url": "",
      "post_date": "04/05/2024 23:48:53",
      "content": "<p>Tks for sharing 👊</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2738862,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "04/06/2024 17:01:01",
      "content": "<p>Thanks for bring this up! I also noticed this problem. Currently I am using Chris's method to extract EEG signal from the raw parquet, which always takes the middle 50s of a signal, regardless of the <code>eeg_label_offset_seconds</code>. To get non-overlap samples, I applied groupby on <code>eeg_id + targets</code>, which gives each unique <code>eeg_id + vote</code> pattern a single entry. In this case, the offset of the signal segment indeed does not match with the actual <code>eeg_label_offset_seconds</code>. That would lead to the fact that the same EEG segment is paired with multiple vote patterns, and thus confuse the model. </p>\n<p>Any solution yet to fix it? </p>\n<p>Attached is a plot of eeg_id = 21379170. Black Solid lines in the plot indicates the middle 50sec segment. Color represent different expert consensus. </p>\n<p>Code snippet:</p>\n<pre><code> =  [,,,,,,,]\nfig, axes = plt.subplots(, , figsize=(, ), sharex=)\nfor i, ax in enumerate(axes):\n    ax.plot(np.arange(len_seq), eeg_signal[:, i]) # eeg_signal (, ) the signal data from parquet\n    ax.text(, , [i], transform=ax.transAxes, ha=, va=)\n\ndef plot_timespan(ax, t1, t2, color, label=):\n    vspan = ax.axvspan(t1, t2, alpha=, color=color, label=label)\n    return ax, vspan\n\ncolors = [, , , , , ]\nfor i, row in train_csv[(train_csv[]==)].iterrows():\n    target_color = colors[row[targets].argmax()]\n    target_label = row[]\n    t_middle = row[] + \n    for ax in axes:\n        ax, vspan = plot_timespan(ax, (t_middle)*, (t_middle+)*, target_color, target_label)\n        ax.axvline((t_total//)*, color=, ls=, lw=)\n        ax.axvline((t_total//+)*, color=, ls=, lw=)\n\nax.legend(bbox_to_anchor=(, ), loc=)\nfig.tight_layout()\nplt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F278125a3c40ea5158c2271e3ab47b4c9%2Foutput.png?generation=1712422551573864&amp;alt=media\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2738939,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "04/06/2024 18:01:35",
          "content": "<p>You can add an offset of 20 secs to eeg_label_offset_seconds, in theory this is the start of the 50s for the row label and they looked at the centre 10s  for it.  So offset add of 20s should get the centre 10 secs.  But there is overlap sometimes.</p>\n<blockquote>\n  <p>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2735147": "Greetings, everyone.\n\nI have encountered a peculiar situation in the training dataset, and I would appreciate if someone could provide some insight. \nAmong the EEG recordings, I have identified 783 instances where a single EEG_ID has been assigned multiple expert consensus diagnoses.\n\nThe concerning aspect is that each EEG recording expected to originate from a single patient. Consequently, each EEG_ID should logically correspond to a unique diagnosis, as it is highly unlikely for an individual patient to have multiple distinct conditions simultaneously.\n\nIf this scenario is not an anomaly and there is a valid explanation for the presence of multiple diagnoses for a single EEG_ID, I would greatly appreciate your insights.\n\n`(train.groupby('eeg_id')['expert_consensus'].nunique() > 1).sum()`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F1e63b1329117b5949a5e0c90c378cc57%2FScreenshot%20from%202024-04-04%2015-01-48.png?generation=1712242615039186&alt=media)",
    "2735209": "Not necessarily. There is various subsamples in each eeg_id. Each one with different 10 seconds to be labeled. Even beign the same patient can have different labels. Even more if each one have been labeled independently. The model has to predict more various experts vote distribution than a true label.",
    "2735220": "see the Data page - \n>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\n\n\n>- eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.\n\n\n\n>- label_id - An ID for this set of labels.\n\nHave been working on this after finding issues and considering all the discussions about >= 10 votes.\nReally don't think that was made clear about how train data was put together. But probably some figured it out earlier?\nThe labels are not consistently for a 50s or even 10s period, so deciding how to split out\nHave a csv file could put in a dataset, not sure if too late for sharing though.  Value counts below for train expert consensus counts and the 5 entry of 234 is just for one eeg id.  Possibly these could/should be left out of training.\n\n>exp_cons_cnts\n1    93956\n2     9134\n3     2187\n4     1289\n5      234`\n\nConsidering the test data is meant to be 50s only per eeg id and one spectrogram no overlapping, not sure if this will also happen in test labels.  And think is meant to be only the centre 10s that was labeled not sure.",
    "2736346": "If you plot the consensuses (?) against time, you'll see that they evolve in a reasonable manner: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fbbfd41171920532c3d3ff34d4e4bc517%2Fconsensus_evolve.png?generation=1712298663242029&alt=media)\n\nIf the top vote is less than 60%, I plotted the top two labels. Otherwise, I plotted the top label. Most of the transitions make sense (e.g. LPD to LRDA or GPD to GRDA; also seizures and LPD seem to mix. I read that patients with LPD are more likely to have seizures, although I don't remember on what time scale).",
    "2736381": "Think the issue is with time - acc to Data page, the labels are for 50 sec subsamples from the eeg_label_offset_seconds, a bit vague if the central 10 secs applies to eeg data, spectrograms or both.  So taking a centre 10 secs from a 50 sec window of eeg data does not always (if ever) equate to the labels. Perhaps windowing by 2-4 secs might be closer. And train eegs have a variety of time lengths, well beyond 50 secs.  All of this probably accounts for the high number of Other in train data, maybe why >=10 seems to work well.  Many of the public notebooks are not really considering eeg_label_offset_seconds from what I can tell.\n\nIn contrast,  the test data is meant to be 50s only per eeg id and one spectrogram no overlapping. Somewhat apples and oranges.",
    "2736429": "For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2Fb29caca1b03a8a31eab2ca75b734ac68%2F21379701_eeg.png?generation=1712302677695583&alt=media)\nEach 10s region is the center of a 50s region (which isn't shown). (Note: the colors here aren't the same as the colors on the previous plot)\n\nThese are the labels and their locations:\n\nonset\tduration\tdescription\n20.0 \t 10.0 \t ('other', 'lrda')\n30.0 \t 10.0 \t ('other', 'lrda')\n32.0 \t 10.0 \t ('other', 'lrda')\n50.0 \t 10.0 \t ('seizure', 'other')\n64.0 \t 10.0 \t seizure\n68.0 \t 10.0 \t seizure\n82.0 \t 10.0 \t lpd\n96.0 \t 10.0 \t lpd\n98.0 \t 10.0 \t lpd\n100.0 \t 10.0 \t lpd\n102.0 \t 10.0 \t lpd\n106.0 \t 10.0 \t lpd\n110.0 \t 10.0 \t lpd\n124.0 \t 10.0 \t lpd\n126.0 \t 10.0 \t lpd\n128.0 \t 10.0 \t lpd\n130.0 \t 10.0 \t lpd\n178.0 \t 10.0 \t other\n\nThe first offset was 0s, so the 10s scored region starts at 20s and ends at 30s, etc.\n\nThis is the first 50s window for spec (cropped to the central 50s of the 10 minute window) and eeg, with the middle 10s highlighted:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F394185%2F30769663cb2ccd66f92193a2087fd355%2F21379701_eeg_spec.png?generation=1712303735165968&alt=media)\nYou can see that there is a high frequency response in the spectrogram that corresponds with the EEG data in time. So I think the offsets are accurately described (although a bit cryptic, some of the other posts discussing the offsets were helpful.)",
    "2736489": ">eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.\neeg_label_offset_seconds - The time between the beginning of the consolidated EEG and this subsample.\n\nFrom train, the most votes for (21379701) are Other and Seizure and LPD. Using train eeg_label_offset_seconds starting at 0 for 50s will have Other (2 or 1)  and Seizure (1 or 2)  and LRDA (1 or 0) which consumes the entries up to 48s.  The next train at 62s for 50s will have LPD 75% of votes LRDA 25% of votes.  All these are low vote counts 1-3.  This eeg data has 208s or 41600 rows.\n.\nIt is difficult to read the text corresponding to labels, and not sure where your 50s region starts. Can only go by what is provided in train in terms of where eeg data and labels are meant to be in the parquet files.\n\nEDIT: your post came after was doing this.  \nfrom train Other corr to 0, 10, 12 secs and 158s.  Seizure 30, 34, 48. LPD 62 to 106 at 2-4s intervals\n\nWhat is  the source of your plots?  Know a notebook is not possible now but maybe a code snippet, looks like using mne?",
    "2736492": "If it is expected that a patient's state or condition can transition from one to another over time. To accurately capture these transitions, the `eeg_label_offset_seconds` should be utilized for each specific state or diagnosis.\n\nI have been applying the same `eeg_label_offset_seconds` value for a single `eeg_id`, under the assumption that each `eeg_id` would have only one associated diagnosis or state. This approach might not be appropriate, as some patients have multiple diagnoses associated with their EEG recordings. Incorporating the correct `eeg_label_offset_seconds` for each diagnosis or state associated with an `eeg_id`, the  model captures the temporal dynamics and transitions within the EEG data.\n\nWhile the number of samples with multiple diagnoses might not be substantial, considering this issue and accounting for potential transitions between different states or diagnoses could potentially improve the model's performance.",
    "2736497": "Not entirely sure the train eeg_label_offset_seconds correctly corresponds to the row labels. And it is not just about multiple labels per eeg id, but getting the vote percentages somewhere in the ballpark. \nThink the wording of a 50s subsample that the labels applied to was either misleading or incorrect.",
    "2736532": "Apologies in advance for this giant \"snippet\"... it would have been a lot of work to make it shorter.\n\n```\ndef plot_spec_eeg(eeg: EEG, offset: int = 0):\n    fig = plt.figure(figsize=(20, 10), layout=None)\n    gs = fig.add_gridspec(nrows=4, ncols=2, hspace=0.1, wspace=0.0)\n\n    axs0 = [fig.add_subplot(gs[i, 0]) for i in range(4)]\n\n    sub_spec = to_decibel(eeg.sub_spec(offset, match_eeg_window=True))\n    plot_spec(sub_spec, axs=axs0)\n\n    ax1 = fig.add_subplot(gs[:, 1])\n    raw_single_ann = eeg.raw.copy().set_annotations(mne.Annotations(**eeg.raw.annotations[offset]))\n    fig_raw = raw_single_ann.plot(\n        duration=50,\n        start=eeg.eeg_offsets[offset],\n        show_scrollbars=False,\n        show_scalebars=False,\n        **mne_plot_kwargs,\n    )\n\n    canvas = FigureCanvasAgg(fig_raw)\n    canvas.draw()\n    img = np.asarray(canvas.buffer_rgba())\n    im = ax1.imshow(img)\n    ax1.set_axis_off()\n```\n\nThis is the code snippet. The rest of the code you need to run it is:\n\n```\ntrain_meta = pd.read_csv(data_path / \"train.csv\").set_index([\"eeg_id\", \"eeg_sub_id\"]).sort_index()\nvote_cols = [col for col in train_meta.columns if str(col).endswith(\"_vote\")]\ntrain_meta[\"vote_total\"] = train_meta[vote_cols].sum(axis=1)\n\neeg_ids = train_meta.reset_index()[\"eeg_id\"].unique()\neeg_id_index = {eeg_ids[i]: i for i in range(len(eeg_ids))}\n\n# extra target info\nvote_probs = train_meta[vote_cols].div(train_meta[\"vote_total\"], axis=0)\n\n\ndef get_vote_case(x: pd.Series, tol: float = 0.05) -> str:\n    (k2, v2), (k1, v1) = x.sort_values()[-2:].items()\n    if v1 == 1.0:\n        return \"perfect\"\n    if v1 > 0.5 + tol:\n        return \"ideal\"\n    if \"other\" in k1 or \"other\" in k2:\n        return \"proto\"\n    return \"edge\"\n\n\ndef get_vote_case_str(x: pd.Series, tol: float = 0.05) -> str:\n    (k2, v2), (k1, v1) = x.sort_values()[-2:].items()\n    if v1 == 1.0:\n        return k1.split(\"_\")[0]\n    if v1 > 0.5 + tol:\n        return k1.split(\"_\")[0]\n    return (k1.split(\"_\")[0], k2.split(\"_\")[0])\n\n\ndef get_vote_cases_df() -> pd.DataFrame:\n    vote_cases_file = output_path / \"vote_cases.csv\"\n\n    if not vote_cases_file.exists():\n        vote_cases = vote_probs.apply(get_vote_case, axis=1)\n        vote_cases_str = vote_probs.apply(get_vote_case_str, axis=1)\n        df = pd.DataFrame({\"vote_case\": vote_cases, \"vote_case_str\": vote_cases_str})\n        df.to_csv(vote_cases_file)\n\n    return pd.read_csv(vote_cases_file).set_index(\"eeg_id\")\n\n\nvote_cases_df = get_vote_cases_df()\nvote_cases = vote_cases_df[\"vote_case\"]\nvote_cases_str = vote_cases_df[\"vote_case_str\"]\n\n\n# MNE code and EEG class\n\n# kwargs for mne.io.Raw.plot(); need to specify duration to show more than 10s\nmne_plot_kwargs = dict(\n    scalings={\"eeg\": 100e-6, \"ecg\": 1e-4}, clipping=None, remove_dc=False, proj=False, show=False\n)\n\n\n# Bipolar reference code\n# TODO: integrate this into EEG class?\n\n\ndef set_double_banana_ref(raw: mne.io.Raw) -> None:\n    \"\"\"Add bipolar reference (\"double banana\") channels to\n    mne.io.Raw object.\n\n    Modifies in place.\n    \"\"\"\n    LL = [\"Fp1\", \"F7\", \"T3\", \"T5\", \"O1\"]\n    LP = [\"Fp1\", \"F3\", \"C3\", \"P3\", \"O1\"]\n    RL = [\"Fp2\", \"F8\", \"T4\", \"T6\", \"O2\"]\n    RP = [\"Fp2\", \"F4\", \"C4\", \"P4\", \"O2\"]\n    central = [\"Fz\", \"Cz\", \"Pz\"]\n\n    chains = [LL, LP, RL, RP, central]\n\n    for chain in chains:\n        mne.set_bipolar_reference(raw, chain[:-1], chain[1:], copy=False, drop_refs=False)\n\n\n# There are some issues with this... the locations of these \"virtual\" channels is set as the same location as the first item in the channel tuple that is differenced. For instance, `(\"Fp1\", \"F7\")` yields the virtual channel `Fp1-F7`, which is assigned the same location as `Fp1`.\n#\n# Also, we didn't drop the original channels because Fp1, Fp2, O1, and O2 are each used twice, so we can't drop them as we go. The following code allows us to pick the bipolar channels:\n\n\ndef pick_bananas(raw: mne.io.Raw) -> mne.io.Raw:\n    picks = [ch_name for ch_name in raw.info[\"ch_names\"] if \"-\" in ch_name]\n    return raw.pick(picks)\n\n\ndef make_mne_raw(\n    eeg: pd.DataFrame,\n    notch: Optional[tuple[float]] = (60,),\n    average_ref: bool = False,\n    banana_ref: bool = False,\n    ekg: bool = False,\n    high_pass: Optional[float] = None,\n    low_pass: Optional[float] = None,\n    ref_before_filt: bool = False,\n    ffill_na: bool = False,\n    fill_na_zero: bool = False,\n    fill_na_mean: bool = False,\n    detrend: bool = False,\n    debug: bool = False,\n    clip: Optional[float] = None,\n) -> mne.io.Raw:\n    if not ekg:\n        eeg = eeg.drop(columns=\"EKG\")\n\n    if ffill_na:\n        eeg = eeg.ffill(axis=0, limit=5)\n\n    if fill_na_zero:\n        eeg = eeg.fillna(0.0)\n\n    if fill_na_mean:\n        eeg = eeg.fillna(eeg.mean(axis=0))\n\n    eeg_arr = eeg.values.T * 1e-6  # units = microvolts\n\n    if debug:\n        n_nan = np.sum(np.isnan(eeg_arr))\n        if n_nan > 0:\n            print(f\"Number NaN: {n_nan}\")\n\n    if not ekg:\n        info = mne.create_info(list(eeg.columns), sfreq=200.0, ch_types=\"eeg\")\n    else:\n        info = mne.create_info(\n            list(eeg.columns), sfreq=200.0, ch_types=[\"eeg\"] * (len(eeg.columns) - 1) + [\"ecg\"]\n        )\n    raw = mne.io.RawArray(eeg_arr, info=info)\n\n    raw.set_montage(\"standard_1020\")\n\n    if ref_before_filt:\n        if average_ref:\n            raw.set_eeg_reference(ref_channels=\"average\")\n        if banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n    if notch:\n        raw.notch_filter(freqs=notch)\n\n    if high_pass or low_pass:\n        raw.filter(l_freq=high_pass, h_freq=low_pass)\n\n    if debug:\n        n_nan = raw.to_data_frame().isna().sum().sum()\n        if n_nan > 0:\n            print(f\"Number NaN: {n_nan}\")\n\n    if clip:\n        clip = abs(clip) * 1e-6\n        raw.apply_function(lambda x: np.clip(x, -clip, clip))\n\n    if detrend:\n        raw.apply_function(lambda x: signal.detrend(x, type=\"linear\"))\n\n    if not ref_before_filt:\n        if average_ref:\n            raw.set_eeg_reference(ref_channels=\"average\")\n        if banana_ref:\n            set_double_banana_ref(raw)\n            raw = pick_bananas(raw)\n\n    return raw\n\n\nclass EEG:\n    spec_regions = [\"LL\", \"LP\", \"RP\", \"RL\"]\n\n    def __init__(self, eeg_num: int, raw_kwargs: Optional[dict] = None):\n        self.eeg_id = eeg_ids[eeg_num]\n        self.eeg_path = data_path / \"train_eegs\" / (str(self.eeg_id) + \".parquet\")\n        self._data = None  # lazy load data\n        self._raw = None  # MNE raw EEG\n        self._raw_kwargs = raw_kwargs if raw_kwargs is not None else {}\n\n        meta = train_meta.loc[self.eeg_id]\n        self.spec_path = data_path / \"train_spectrograms\" / (str(meta[\"spectrogram_id\"][0]) + \".parquet\")\n        self._spec = None\n\n    @classmethod\n    def from_id(cls, eeg_id: int, raw_kwargs: Optional[dict] = None):\n        eeg_num = eeg_id_index[eeg_id]\n        return cls(eeg_num, raw_kwargs)\n\n    @property\n    def meta(self) -> pd.DataFrame:\n        return train_meta.loc[self.eeg_id]\n\n    @property\n    def data(self) -> pd.DataFrame:\n        if self._data is None:\n            eeg = pd.read_parquet(self.eeg_path)\n            eeg = eeg.set_index(pd.to_timedelta(5 * eeg.index, unit=\"ms\"))\n            self._data = eeg\n        return self._data\n\n    @property\n    def spec(self) -> pd.DataFrame:\n        if self._spec is None:\n            spec = pd.read_parquet(self.spec_path)\n            spec = spec.set_index(pd.to_timedelta(spec.time, unit=\"s\")).drop(columns=\"time\")\n            spec.columns = pd.MultiIndex.from_tuples(\n                [col.split(\"_\") for col in spec.columns], names=(\"region\", \"frequency\")\n            )\n            self._spec = spec\n        return self._spec\n\n    @property\n    def raw(self) -> mne.io.Raw:\n        if self._raw is None:\n            self._raw = make_mne_raw(self.data, **self._raw_kwargs)\n            ann_dict = {\n                \"onset\": self.eeg_offsets + 20.0,\n                \"duration\": [10.0] * len(self.eeg_offsets),\n                \"description\": vote_cases_str.loc[self.eeg_id],\n            }\n            self._raw.set_annotations(mne.Annotations(**ann_dict))\n        return self._raw\n\n    @property\n    def eeg_offsets(self):\n        return self.meta[\"eeg_label_offset_seconds\"]\n\n    @property\n    def spec_offsets(self):\n        return self.meta[\"spectrogram_label_offset_seconds\"]\n\n    @property\n    def offsets(self):\n        return pd.DataFrame({\"eeg_offset\": self.eeg_offsets, \"spec_offset\": self.spec_offsets})\n\n    def sub_eeg(self, sub_eeg_id: int) -> pd.DataFrame:\n        offset = self.eeg_offsets[sub_eeg_id]\n        return self.data.loc[slice(f\"{offset}s\", f\"{49 + offset}s\")]\n\n    def sub_spec(self, sub_eeg_id: int, match_eeg_window: bool = False) -> pd.DataFrame:\n        offset = self.spec_offsets[sub_eeg_id]\n\n        # NOTE: we're adding 0.1s to the end of the slices to make sure they don't end\n        # with 0 seconds, in which case the units are inferred as minutes, and we get an\n        # extra 59 seconds.\n        if match_eeg_window:\n            window = slice(f\"4 minutes {59 + offset - 25 + 0.1}s\", f\"4 minutes {59 + offset + 25 + 0.1}s\")\n        else:\n            window = slice(f\"{offset + 0.1}s\", f\"10 minutes {offset + 0.1}s\")\n\n        return self.spec.loc[window]\n\n\ndef spec_freq_medians_region_agg(spec: pd.DataFrame) -> pd.Series:\n    total_medians0 = spec.T.unstack(level=\"region\").sort_index(key=lambda x: x.astype(float)).median(axis=1)\n    total_medians = pd.DataFrame({region: total_medians0 for region in EEG.spec_regions}).unstack()\n    total_medians.index.names = (\"region\", \"frequency\")\n    return total_medians\n\n\ndef to_decibel(spec: pd.DataFrame, baseline: Union[pd.Series, Literal[\"agg\", \"sep\"]] = \"agg\") -> pd.DataFrame:\n    if not isinstance(baseline, pd.Series):\n        if baseline == \"sep\":\n            baseline = spec.median(axis=0)\n        elif baseline == \"agg\":\n            baseline = spec_freq_medians_region_agg(spec)\n        else:\n            raise ValueError(\"baseline must be 'agg' or 'sep' or pd.Series\")\n    spec = spec.div(baseline, axis=1).fillna(0.0)  # if median was 0, then row show be zero\n\n    spec = 10 * np.log10(spec.clip(1e-5, 1e5))\n    return spec\n\n\n# Plotting code\ndef plot_spec(\n    spec: pd.DataFrame,\n    show_label_window: bool = True,\n    axs: Optional[list[mpl.axes.Axes]] = None,\n    cbar: bool = True,\n    **kwargs,\n) -> None:\n    if axs is None:\n        fig, axs = plt.subplots(4, 1, figsize=(8, 10), sharex=True)\n\n    vmin, vmax = spec.min().min(), spec.max().max()\n\n    time_ticks_stride = len(spec.index) // 25\n\n    if show_label_window:\n        times = spec.index\n        fmin, fmax = 0, len(spec.loc[:, EEG.spec_regions[0]].columns)\n        mid_point = times[len(times) // 2]\n        td = pd.Timedelta(\"5s\")\n        label_times = times[(times >= mid_point - td) & (times <= mid_point + td)].seconds\n\n    for region, ax in zip(EEG.spec_regions, axs):\n        # sort frequencies in descending order\n        data = spec.loc[:, region].T.sort_index(ascending=False, key=lambda x: x.astype(float))\n\n        # convert times to seconds (else seaborn converts to nanoseconds)\n        data.columns = data.columns.seconds\n\n        if region != EEG.spec_regions[-1]:\n            sns.heatmap(\n                data=data, ax=ax, xticklabels=False, vmin=vmin, vmax=vmax, cmap=\"jet\", cbar=cbar, **kwargs\n            )\n\n            if show_label_window:\n                label_data = 0.0 * data\n                label_data.loc[:, label_times] = 1.0\n                sns.heatmap(\n                    label_data,\n                    alpha=0.4,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=False,\n                    mask=label_data.map(bool),\n                    **kwargs,\n                )\n\n            ax.set(xlabel=\"\")\n        else:\n            sns.heatmap(\n                data=data,\n                ax=ax,\n                xticklabels=time_ticks_stride,\n                vmin=vmin,\n                vmax=vmax,\n                cmap=\"jet\",\n                cbar=cbar,\n                **kwargs,\n            )\n\n            if show_label_window:\n                label_data = 0.0 * data\n                label_data.loc[:, label_times] = 1.0\n                sns.heatmap(\n                    label_data,\n                    alpha=0.4,\n                    ax=ax,\n                    cbar=False,\n                    xticklabels=time_ticks_stride,\n                    mask=label_data.map(bool),\n                    **kwargs,\n                )\n\n        ax.set(ylabel=f\"{region} frequency (Hz)\")\n        # ax.set_title(region, loc=\"left\")\n\n```\n\nTo make the spec/eeg plot I ran:\n```\neeg = EEG.from_id(21379701, raw_kwargs={\"banana_ref\": True, \"ffill_na\": True, \"notch\": (60.0,)})\nplot_spec_eeg(eeg, offset=13)\nplt.show()\n```\n\nTo make the \"evolution\" plot, I did\n\n```\nfrom functools import reduce\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\n\n\ng = train_meta.reset_index(\"eeg_sub_id\")[\"expert_consensus\"].apply(lambda x: set([x])).groupby(\"eeg_id\")\nn_consensuses = g.apply(lambda x: len(reduce(lambda a, b: a.union(b), x)))\n\nncon_filt = n_consensuses > 1\n\nmixed_con_vote_str = vote_cases_str.loc[eeg_ids[ncon_filt]]\n\nmixed_con_votes_evolve = mixed_con_vote_str.groupby(\"eeg_id\").apply(lambda x: list(x))\n\nlen_filt = mixed_con_votes_evolve.apply(len) < 20\n\ncat_colors = {\"other\": \"lightslategrey\", \"lpd\": \"tomato\", \"gpd\": \"mediumseagreen\", \"lrda\": \"palevioletred\", \"grda\": \"greenyellow\", \"seizure\": \"dodgerblue\"}\n\ndef plot_votes_evolve_row(row_num, votes_list, ax, x_spacing=0.5, y_spacing=0.5):\n    for i, votes in enumerate(votes_list):\n        if isinstance(votes, tuple):\n            for offset, vote in zip([-0.1, 0.1], votes[::-1]):\n                ax.scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote], alpha=0.7)\n        elif isinstance(votes, str) and \", \" in votes:\n            # saving vote cases to csv means tuples of strings are loaded as a single string\n            votes = votes[1:-1].split(\", \")\n            for offset, vote in zip([-0.1, 0.1], votes[::-1]):\n                ax.scatter(x_spacing * i, y_spacing * (row_num + offset), color=cat_colors[vote.strip(\"'\")], alpha=0.7)\n        else:\n            ax.scatter(x_spacing * i, y_spacing * row_num, color=cat_colors[votes])\n\n            \nfig, ax = plt.subplots(figsize=(10, 15))\n\npatches = [mpl.patches.Patch(color=v, label=k) for k, v in cat_colors.items()]\n\nn_patients = 50\ny_spacing = 0.2\n\nmcve_filtered = mixed_con_votes_evolve.loc[len_filt]\n\nfor i in range(n_patients):\n    plot_votes_evolve_row(i, mcve_filtered.iloc[i], ax, y_spacing=y_spacing)\nax.legend(handles=patches)\nax.set_yticks(ticks=(y_spacing * np.arange(n_patients)), labels=mcve_filtered.index[:n_patients])\n\n```",
    "2736634": "Many thanks for all the snippets.  Can reproduce evolution plot and spectrogram.  Was really interested in the plot you showed with this section, if you have a chance to comment:\n\n>For the last EEG id on the previous plot (21379701), here are the annotations over the raw EEGs, showing the 10s corresponding to the label:",
    "2736744": "That is mne's plotting, I think for the eeg plot I posted, it was  `eeg.raw.plot(duration=240, **mne_plot_kwargs)` (you need at least the eeg scale from my mne plot kwargs, or else it will look really weird). MNE automatically uses the annotations I made with the labels",
    "2736797": "Thanks!  Should have realised by the axis labels. (getting tired)",
    "2737778": "Tks for sharing 👊",
    "2738862": "Thanks for bring this up! I also noticed this problem. Currently I am using Chris's method to extract EEG signal from the raw parquet, which always takes the middle 50s of a signal, regardless of the `eeg_label_offset_seconds`. To get non-overlap samples, I applied groupby on `eeg_id + targets`, which gives each unique `eeg_id + vote` pattern a single entry. In this case, the offset of the signal segment indeed does not match with the actual `eeg_label_offset_seconds`. That would lead to the fact that the same EEG segment is paired with multiple vote patterns, and thus confuse the model. \n\nAny solution yet to fix it? \n\nAttached is a plot of eeg_id = 21379170. Black Solid lines in the plot indicates the middle 50sec segment. Color represent different expert consensus. \n\nCode snippet:\n\n```\nEEG_FEAT_USE =  ['Fp1','T3','C3','O1','Fp2','C4','T4','O2']\nfig, axes = plt.subplots(8, 1, figsize=(12, 8), sharex=True)\nfor i, ax in enumerate(axes):\n    ax.plot(np.arange(len_seq), eeg_signal[:, i]) # eeg_signal (L, C) the signal data from parquet\n    ax.text(0.01, 0.05, EEG_FEAT_USE[i], transform=ax.transAxes, ha='left', va='bottom')\n\ndef plot_timespan(ax, t1, t2, color, label=None):\n    vspan = ax.axvspan(t1, t2, alpha=0.2, color=color, label=label)\n    return ax, vspan\n\ncolors = ['r', 'g', 'b', 'c', 'm', 'y']\nfor i, row in train_csv[(train_csv['eeg_id']==21379701)].iterrows():\n    target_color = colors[row[targets].argmax()]\n    target_label = row['expert_consensus']\n    t_middle = row['eeg_label_offset_seconds'] + 25\n    for ax in axes:\n        ax, vspan = plot_timespan(ax, (t_middle-5)*200, (t_middle+5)*200, target_color, target_label)\n        ax.axvline((t_total//2-25)*200, color='k', ls='-', lw=2)\n        ax.axvline((t_total//2+25)*200, color='k', ls='-', lw=2)\n\nax.legend(bbox_to_anchor=(1.15, 4.5), loc='center right')\nfig.tight_layout()\nplt.show()\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F278125a3c40ea5158c2271e3ab47b4c9%2Foutput.png?generation=1712422551573864&alt=media)",
    "2738939": "You can add an offset of 20 secs to eeg_label_offset_seconds, in theory this is the start of the 50s for the row label and they looked at the centre 10s  for it.  So offset add of 20s should get the centre 10 secs.  But there is overlap sometimes.\n\n>eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to."
  },
  "source": "meta"
}