{
  "id": 467127,
  "title": "Correct way to merge targets",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467127",
  "author_name": "",
  "post_date": "2024-01-11T08:37:20.296401100Z",
  "votes": 26,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I'm trying to figure out correct way of merging targets so I need some help.</p>\n<p>We have this dataframe for eeg_id 1628180742<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F28045613639783a9cff40b842da71f99%2FScreenshot%20from%202024-01-11%2011-29-46.png?generation=1704961820794192&amp;alt=media\" alt=\"eeg_id\"></p>\n<p>and we load the eeg parquet file<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb39daaaa97f2cde4df76019e7e0d2f42%2FScreenshot%20from%202024-01-11%2011-30-44.png?generation=1704961866211132&amp;alt=media\" alt=\"eeg_df\"></p>\n<p>we know that sample per second is 200, so we create a second column on the eeg dataframe</p>\n<pre><code>sample_per_second = \nseconds = df_eeg.shape[] / sample_per_second\ndf_eeg[] = np.repeat(np.arange(seconds), )\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F128932ddfe0c5265fe3b92592e44ecc6%2FScreenshot%20from%202024-01-11%2011-32-38.png?generation=1704961968625087&amp;alt=media\" alt=\"eeg seconds\"></p>\n<p>since we have the seconds, we can merge the targets on that column</p>\n<pre><code>target_columns = [, , , , , ]\ndf_eeg_merged = df_eeg.merge(df_train.loc[df_train[] == , target_columns + []].rename(columns={: }), on=, how=)\n</code></pre>\n<p>I'm not sure if I'm doing this correctly. Are we supposed merge targets on those exact seconds? EEGs look like this when visualized. Red areas are where seizure_vote is greater than 0.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3fce7e9535aad74eeb3da945d1435f37%2F1628180742.png?generation=1704962100705649&amp;alt=media\" alt=\"eeg\"></p>",
  "messages": [
    {
      "id": "2596653",
      "postDate": "01/11/2024 08:37:20",
      "content": "<p>I'm trying to figure out correct way of merging targets so I need some help.</p>\n<p>We have this dataframe for eeg_id 1628180742<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F28045613639783a9cff40b842da71f99%2FScreenshot%20from%202024-01-11%2011-29-46.png?generation=1704961820794192&amp;alt=media\" alt=\"eeg_id\"></p>\n<p>and we load the eeg parquet file<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb39daaaa97f2cde4df76019e7e0d2f42%2FScreenshot%20from%202024-01-11%2011-30-44.png?generation=1704961866211132&amp;alt=media\" alt=\"eeg_df\"></p>\n<p>we know that sample per second is 200, so we create a second column on the eeg dataframe</p>\n<pre><code>sample_per_second = \nseconds = df_eeg.shape[] / sample_per_second\ndf_eeg[] = np.repeat(np.arange(seconds), )\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F128932ddfe0c5265fe3b92592e44ecc6%2FScreenshot%20from%202024-01-11%2011-32-38.png?generation=1704961968625087&amp;alt=media\" alt=\"eeg seconds\"></p>\n<p>since we have the seconds, we can merge the targets on that column</p>\n<pre><code>target_columns = [, , , , , ]\ndf_eeg_merged = df_eeg.merge(df_train.loc[df_train[] == , target_columns + []].rename(columns={: }), on=, how=)\n</code></pre>\n<p>I'm not sure if I'm doing this correctly. Are we supposed merge targets on those exact seconds? EEGs look like this when visualized. Red areas are where seizure_vote is greater than 0.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3fce7e9535aad74eeb3da945d1435f37%2F1628180742.png?generation=1704962100705649&amp;alt=media\" alt=\"eeg\"></p>",
      "rawMarkdown": "I'm trying to figure out correct way of merging targets so I need some help.\n\nWe have this dataframe for eeg_id 1628180742\n![eeg_id](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F28045613639783a9cff40b842da71f99%2FScreenshot%20from%202024-01-11%2011-29-46.png?generation=1704961820794192&alt=media)\n\nand we load the eeg parquet file\n![eeg_df](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb39daaaa97f2cde4df76019e7e0d2f42%2FScreenshot%20from%202024-01-11%2011-30-44.png?generation=1704961866211132&alt=media)\n\nwe know that sample per second is 200, so we create a second column on the eeg dataframe\n```python\nsample_per_second = 200\nseconds = df_eeg.shape[0] / sample_per_second\ndf_eeg['second'] = np.repeat(np.arange(seconds), 200)\n```\n\n![eeg seconds](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F128932ddfe0c5265fe3b92592e44ecc6%2FScreenshot%20from%202024-01-11%2011-32-38.png?generation=1704961968625087&alt=media)\n\nsince we have the seconds, we can merge the targets on that column\n```python\ntarget_columns = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\ndf_eeg_merged = df_eeg.merge(df_train.loc[df_train['eeg_id'] == 1628180742, target_columns + ['eeg_label_offset_seconds']].rename(columns={'eeg_label_offset_seconds': 'second'}), on='second', how='left')\n```\n\nI'm not sure if I'm doing this correctly. Are we supposed merge targets on those exact seconds? EEGs look like this when visualized. Red areas are where seizure_vote is greater than 0.\n\n![eeg](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3fce7e9535aad74eeb3da945d1435f37%2F1628180742.png?generation=1704962100705649&alt=media)",
      "votes": null
    },
    {
      "id": "2597097",
      "postDate": "01/11/2024 14:16:27",
      "content": "<p><code>eeg_label_offset_seconds</code> - The time between the beginning of the consolidated EEG and this subsample.</p>\n<p>This explanation is so confusing. Isn't the time between the beginning of the EGG should always be <code>eeg_label_offset_seconds</code> since beginning is always 0?</p>\n<p>There is another thing I found. We are predicting a single target for each eeg_id but some of the eegs have more than 1 unique target. Here is the distribution of unique targets per eeg_id</p>\n<pre><code>df_train.groupby()[].nunique().value_counts()\n\nexpert_consensus\n    \n      \n       \n       \n        \nName: count, dtype: int64\n</code></pre>",
      "rawMarkdown": "`eeg_label_offset_seconds` - The time between the beginning of the consolidated EEG and this subsample.\n\nThis explanation is so confusing. Isn't the time between the beginning of the EGG should always be `eeg_label_offset_seconds` since beginning is always 0?\n\nThere is another thing I found. We are predicting a single target for each eeg_id but some of the eegs have more than 1 unique target. Here is the distribution of unique targets per eeg_id\n\n```python\ndf_train.groupby('eeg_id')['expert_consensus'].nunique().value_counts()\n\nexpert_consensus\n1    16306\n2      666\n3       94\n4       22\n5        1\nName: count, dtype: int64\n```",
      "votes": null
    },
    {
      "id": "2597245",
      "postDate": "01/11/2024 15:39:22",
      "content": "<p>Each eeg_id can have multiple eeg_sub_id, and each of these can have a different expert_consensus</p>",
      "rawMarkdown": "Each eeg_id can have multiple eeg_sub_id, and each of these can have a different expert_consensus",
      "votes": null
    },
    {
      "id": "2597257",
      "postDate": "01/11/2024 15:48:13",
      "content": "<p>Thanks for the clarification. If I understood correctly, each eeg_sub_id can have a single expert consensus and they are 50 seconds. Test set eegs are exactly like this so we have to extract similar samples from training set. </p>",
      "rawMarkdown": "Thanks for the clarification. If I understood correctly, each eeg_sub_id can have a single expert consensus and they are 50 seconds. Test set eegs are exactly like this so we have to extract similar samples from training set.",
      "votes": null
    },
    {
      "id": "2597406",
      "postDate": "01/11/2024 16:50:47",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa49d99bde59ef8aa39d01af910fbe9c1%2FUnknown.png?generation=1704991551238060&amp;alt=media\"></p>\n<p>-- <strong>common sequence of all sub-sequences is 40secs-50secs</strong><br>\n-- target votes are do not changed, so my guess is we can ignore the 0-40secs and 50secs-90secs sequence as it not given any importance to the targets ( as expert decided based on sub-sequence spike )</p>\n<p>or 0-40secs, 50secs-90secs is <strong>weak labels approach</strong> (its my understanding, correct me if it is wrong)</p>\n<p>Unique target votes for each egg_id<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F8a5ac4a30d4fcb816cb8fad1ad181f00%2F__results___7_0.png?generation=1704991738200947&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85ce9b71721384fcc665282e71f639ad%2F__results___8_0.png?generation=1704995430359845&amp;alt=media\"></p>\n<p>-- Similar patterns we can observe in all egg_ids where only have same target votes majorly and few have target votes changed.</p>\n<p><a href=\"https://www.kaggle.com/code/seshurajup/eegs-target-analysis/notebook\" target=\"_blank\">Notebook - Eegs Target Analysis</a></p>\n<p>(<a href=\"https://www.youtube.com/watch?v=S9NLrhj0x-M&amp;t\" target=\"_blank\">Part 1</a>) - from this video as reference of my understanding</p>",
      "rawMarkdown": "gunesevitan \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa49d99bde59ef8aa39d01af910fbe9c1%2FUnknown.png?generation=1704991551238060&alt=media)\n\n-- **common sequence of all sub-sequences is 40secs-50secs**\n-- target votes are do not changed, so my guess is we can ignore the 0-40secs and 50secs-90secs sequence as it not given any importance to the targets ( as expert decided based on sub-sequence spike )\n\nor 0-40secs, 50secs-90secs is **weak labels approach** (its my understanding, correct me if it is wrong)\n\nUnique target votes for each egg_id\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F8a5ac4a30d4fcb816cb8fad1ad181f00%2F__results___7_0.png?generation=1704991738200947&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85ce9b71721384fcc665282e71f639ad%2F__results___8_0.png?generation=1704995430359845&alt=media)\n\n-- Similar patterns we can observe in all egg_ids where only have same target votes majorly and few have target votes changed.\n\n[Notebook - Eegs Target Analysis](https://www.kaggle.com/code/seshurajup/eegs-target-analysis/notebook)\n\n([Part 1](https://www.youtube.com/watch?v=S9NLrhj0x-M&t)) - from this video as reference of my understanding",
      "votes": null
    },
    {
      "id": "2597451",
      "postDate": "01/11/2024 17:07:02",
      "content": "<p>Hmm weak labels explains a lot so for each row, expert consensus could happened at any timestep between <code>eeg_label_offset_seconds + 0</code> and <code>eeg_label_offset_seconds + (50 * 200)</code>. Is that what you meant? I wonder if there are any sub sequences with more than 1 unique labels.</p>",
      "rawMarkdown": "Hmm weak labels explains a lot so for each row, expert consensus could happened at any timestep between `eeg_label_offset_seconds + 0` and `eeg_label_offset_seconds + (50 * 200)`. Is that what you meant? I wonder if there are any sub sequences with more than 1 unique labels.",
      "votes": null
    },
    {
      "id": "2597473",
      "postDate": "01/11/2024 17:11:48",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  Similar to bird competition, we have complete 10secs of audio ( with target as primary bird, secondary bird sings ), in which bird sign happen only on 5sec-7sec, but 0-5secs and 7secs-10secs is the audio won't play useful for the primary/secondary target. in this perspective I mention as weak label.</p>\n<p>But, in this competition we have overlapping segments which help to eliminate the 0-5secs and 7secs-10secs is my intention for selecting targets</p>\n<blockquote>\n  <p>Is that what you meant? <br>\n  <strong>Yes</strong></p>\n  <p>I wonder if there are any sub sequences with more than 1 unique labels.<br>\n  <strong>Yes</strong> as per the unique target count graph, we have few egg sub-sequences have different target counts</p>\n</blockquote>",
      "rawMarkdown": "gunesevitan  Similar to bird competition, we have complete 10secs of audio ( with target as primary bird, secondary bird sings ), in which bird sign happen only on 5sec-7sec, but 0-5secs and 7secs-10secs is the audio won't play useful for the primary/secondary target. in this perspective I mention as weak label.\n\nBut, in this competition we have overlapping segments which help to eliminate the 0-5secs and 7secs-10secs is my intention for selecting targets\n\n>Is that what you meant? \n**Yes**\n\n\n\n> I wonder if there are any sub sequences with more than 1 unique labels.\n**Yes** as per the unique target count graph, we have few egg sub-sequences have different target counts",
      "votes": null
    },
    {
      "id": "2597533",
      "postDate": "01/11/2024 17:28:21",
      "content": "<p>I found that 656 eegs have more than 1 unique target values for overlapping regions</p>\n<pre><code>df_train[] = df_train.groupby()[].diff()\ndf_uniques = df_train.loc[df_train[] &lt; ].groupby()[].nunique().value_counts()\n\nexpert_consensus\n    \n      \n       \n       \n        \nName: count, dtype: int64\n</code></pre>\n<p>and this is an example<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F661e6c5cb95d5ac3d847640831d7d490%2FScreenshot%20from%202024-01-11%2020-24-48.png?generation=1704994059097978&amp;alt=media\" alt=\"eeg\"></p>\n<p>Any idea about how to deal with them? </p>",
      "rawMarkdown": "I found that 656 eegs have more than 1 unique target values for overlapping regions\n```python\ndf_train['eeg_label_offset_seconds_diff'] = df_train.groupby('eeg_id')['eeg_label_offset_seconds'].diff()\ndf_uniques = df_train.loc[df_train['eeg_label_offset_seconds_diff'] < 50].groupby('eeg_id')['expert_consensus'].nunique().value_counts()\n\nexpert_consensus\n1    10109\n2      554\n3       81\n4       20\n5        1\nName: count, dtype: int64\n```\n\nand this is an example\n![eeg](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F661e6c5cb95d5ac3d847640831d7d490%2FScreenshot%20from%202024-01-11%2020-24-48.png?generation=1704994059097978&alt=media)\n\nAny idea about how to deal with them?",
      "votes": null
    },
    {
      "id": "2597566",
      "postDate": "01/11/2024 17:46:32",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6b6f83ee672a60d76b77974efb14f73a%2F__results___8_0.png?generation=1704994875486323&amp;alt=media\"></p>\n<p>Exactly I also changed my notebook to get exact sense, </p>\n<pre><code>{: ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : }\n</code></pre>\n<p>89.4% have primary targets and 10.6% have secondary targets is what I want to try or we can ignore those is other approach. Maybe after experiments we will change our opinion.</p>\n<p>1st approach is better since these votes are came from experts while label is my opinion at this point (same approach used in bird competitions)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6b6f83ee672a60d76b77974efb14f73a%2F__results___8_0.png?generation=1704994875486323&alt=media)\n\nExactly I also changed my notebook to get exact sense, \n```python\n{1: 15282,\n 2: 1348,\n 3: 200,\n 4: 98,\n 5: 66,\n 6: 35,\n 7: 17,\n 9: 15,\n 10: 7,\n 8: 4,\n 11: 3,\n 17: 3,\n 12: 3,\n 13: 1,\n 22: 1,\n 26: 1,\n 19: 1,\n 29: 1,\n 43: 1,\n 15: 1,\n 30: 1}\n```\n\n89.4% have primary targets and 10.6% have secondary targets is what I want to try or we can ignore those is other approach. Maybe after experiments we will change our opinion.\n\n1st approach is better since these votes are came from experts while label is my opinion at this point (same approach used in bird competitions)",
      "votes": null
    },
    {
      "id": "2597589",
      "postDate": "01/11/2024 18:07:26",
      "content": "<p>Another approach is aggregating votes of overlapping rows and create a new expert consensus.</p>",
      "rawMarkdown": "Another approach is aggregating votes of overlapping rows and create a new expert consensus.",
      "votes": null
    },
    {
      "id": "2597597",
      "postDate": "01/11/2024 18:10:25",
      "content": "<p>I will try while training, Thanks for sharing <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
      "rawMarkdown": "I will try while training, Thanks for sharing @gunesevitan",
      "votes": null
    },
    {
      "id": "2597828",
      "postDate": "01/12/2024 00:17:58",
      "content": "<p>Note that the labels only apply to the central 10 seconds.</p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. </p>\n</blockquote>\n<p>Is it possible that a patient has multiple targets over the course of a day?</p>",
      "rawMarkdown": "Note that the labels only apply to the central 10 seconds.\n> The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. \n\nIs it possible that a patient has multiple targets over the course of a day?",
      "votes": null
    },
    {
      "id": "2597829",
      "postDate": "01/12/2024 00:26:19",
      "content": "<p>Or multiple targets in the 50 second long EEG sample?</p>\n<p>Specifically the 20 seconds before and after the central 10 seconds.</p>",
      "rawMarkdown": "Or multiple targets in the 50 second long EEG sample?\n\nSpecifically the 20 seconds before and after the central 10 seconds.",
      "votes": null
    },
    {
      "id": "2598155",
      "postDate": "01/12/2024 07:01:02",
      "content": "<p>There are lots of eegs with multiple targets in their whole 50 seconds long sequences which is perfectly fine. There are cases that satisfies <code>eeg_label_offset_seconds_diff &lt; 10</code> condition which means their central 10 seconds are overlapping but they have more than 1 unique target values. </p>\n<p>Check the highlighted row here. Last subsample with GRDA starts at 26th second so its central 10 seconds are between 46 and 56. First Other starts at 34th second so its central 10 seconds are between 54 and 64. There is only 2 seconds of overlap here but I wonder if there are any cases with more overlapping seconds.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fdc1cf01f596f1c63a41b600a598046dd%2FScreenshot%20from%202024-01-12%2009-58-13.png?generation=1705042705049577&amp;alt=media\" alt=\"eeg\"></p>\n<p><strong>Edit:</strong> There are 594 subsamples that have overlapping central 10 seconds with different expert consensus. I checked it with this code</p>\n<pre><code>df_train[] = df_train[].({\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n})\n\ndf_train[] = df_train.groupby()[].diff()\ndf_train[] = df_train.groupby()[].diff()\n\ncondition = (df_train[] &lt; ) &amp; (df_train[] != )\ndf_train.loc[condition, ].value_counts()\n\neeg_label_offset_seconds_diff\n    \n    \n    \n    \nName: count, dtype: int64\n</code></pre>\n<p>2.0 diff value occurs 109 times which means 8 seconds of overlap in the central 10 second. I guess there are very few of them and they can be filtered.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe2ee21305f9cfd14d148d7b904438af0%2FScreenshot%20from%202024-01-12%2010-36-53.png?generation=1705045031179377&amp;alt=media\" alt=\"eeg2\"></p>",
      "rawMarkdown": "There are lots of eegs with multiple targets in their whole 50 seconds long sequences which is perfectly fine. There are cases that satisfies `eeg_label_offset_seconds_diff < 10` condition which means their central 10 seconds are overlapping but they have more than 1 unique target values. \n\nCheck the highlighted row here. Last subsample with GRDA starts at 26th second so its central 10 seconds are between 46 and 56. First Other starts at 34th second so its central 10 seconds are between 54 and 64. There is only 2 seconds of overlap here but I wonder if there are any cases with more overlapping seconds.\n\n![eeg](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fdc1cf01f596f1c63a41b600a598046dd%2FScreenshot%20from%202024-01-12%2009-58-13.png?generation=1705042705049577&alt=media)\n\n**Edit:** There are 594 subsamples that have overlapping central 10 seconds with different expert consensus. I checked it with this code\n\n```python\ndf_train['expert_consensus_encoded'] = df_train['expert_consensus'].map({\n    'Seizure': 1,\n    'LPD': 2,\n    'GPD': 3,\n    'LRDA': 4,\n    'GRDA': 5,\n    'Other': 6,\n})\n\ndf_train['eeg_label_offset_seconds_diff'] = df_train.groupby('eeg_id')['eeg_label_offset_seconds'].diff()\ndf_train['expert_consensus_encoded_diff'] = df_train.groupby('eeg_id')['expert_consensus_encoded'].diff()\n\ncondition = (df_train['eeg_label_offset_seconds_diff'] < 10) & (df_train['expert_consensus_encoded_diff'] != 0)\ndf_train.loc[condition, 'eeg_label_offset_seconds_diff'].value_counts()\n\neeg_label_offset_seconds_diff\n6.0    172\n8.0    167\n4.0    146\n2.0    109\nName: count, dtype: int64\n```\n\n2.0 diff value occurs 109 times which means 8 seconds of overlap in the central 10 second. I guess there are very few of them and they can be filtered.\n\n![eeg2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe2ee21305f9cfd14d148d7b904438af0%2FScreenshot%20from%202024-01-12%2010-36-53.png?generation=1705045031179377&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2597097,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/11/2024 14:16:27",
      "content": "<p><code>eeg_label_offset_seconds</code> - The time between the beginning of the consolidated EEG and this subsample.</p>\n<p>This explanation is so confusing. Isn't the time between the beginning of the EGG should always be <code>eeg_label_offset_seconds</code> since beginning is always 0?</p>\n<p>There is another thing I found. We are predicting a single target for each eeg_id but some of the eegs have more than 1 unique target. Here is the distribution of unique targets per eeg_id</p>\n<pre><code>df_train.groupby()[].nunique().value_counts()\n\nexpert_consensus\n    \n      \n       \n       \n        \nName: count, dtype: int64\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2597245,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "01/11/2024 15:39:22",
          "content": "<p>Each eeg_id can have multiple eeg_sub_id, and each of these can have a different expert_consensus</p>",
          "votes": null,
          "replies": [
            {
              "id": 2597257,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "01/11/2024 15:48:13",
              "content": "<p>Thanks for the clarification. If I understood correctly, each eeg_sub_id can have a single expert consensus and they are 50 seconds. Test set eegs are exactly like this so we have to extract similar samples from training set. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2597406,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/11/2024 16:50:47",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa49d99bde59ef8aa39d01af910fbe9c1%2FUnknown.png?generation=1704991551238060&amp;alt=media\"></p>\n<p>-- <strong>common sequence of all sub-sequences is 40secs-50secs</strong><br>\n-- target votes are do not changed, so my guess is we can ignore the 0-40secs and 50secs-90secs sequence as it not given any importance to the targets ( as expert decided based on sub-sequence spike )</p>\n<p>or 0-40secs, 50secs-90secs is <strong>weak labels approach</strong> (its my understanding, correct me if it is wrong)</p>\n<p>Unique target votes for each egg_id<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F8a5ac4a30d4fcb816cb8fad1ad181f00%2F__results___7_0.png?generation=1704991738200947&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85ce9b71721384fcc665282e71f639ad%2F__results___8_0.png?generation=1704995430359845&amp;alt=media\"></p>\n<p>-- Similar patterns we can observe in all egg_ids where only have same target votes majorly and few have target votes changed.</p>\n<p><a href=\"https://www.kaggle.com/code/seshurajup/eegs-target-analysis/notebook\" target=\"_blank\">Notebook - Eegs Target Analysis</a></p>\n<p>(<a href=\"https://www.youtube.com/watch?v=S9NLrhj0x-M&amp;t\" target=\"_blank\">Part 1</a>) - from this video as reference of my understanding</p>",
      "votes": null,
      "replies": [
        {
          "id": 2597451,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "01/11/2024 17:07:02",
          "content": "<p>Hmm weak labels explains a lot so for each row, expert consensus could happened at any timestep between <code>eeg_label_offset_seconds + 0</code> and <code>eeg_label_offset_seconds + (50 * 200)</code>. Is that what you meant? I wonder if there are any sub sequences with more than 1 unique labels.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2597473,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/11/2024 17:11:48",
              "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  Similar to bird competition, we have complete 10secs of audio ( with target as primary bird, secondary bird sings ), in which bird sign happen only on 5sec-7sec, but 0-5secs and 7secs-10secs is the audio won't play useful for the primary/secondary target. in this perspective I mention as weak label.</p>\n<p>But, in this competition we have overlapping segments which help to eliminate the 0-5secs and 7secs-10secs is my intention for selecting targets</p>\n<blockquote>\n  <p>Is that what you meant? <br>\n  <strong>Yes</strong></p>\n  <p>I wonder if there are any sub sequences with more than 1 unique labels.<br>\n  <strong>Yes</strong> as per the unique target count graph, we have few egg sub-sequences have different target counts</p>\n</blockquote>",
              "votes": null,
              "replies": [
                {
                  "id": 2597533,
                  "author_name": "gunesevitan",
                  "author_url": "",
                  "post_date": "01/11/2024 17:28:21",
                  "content": "<p>I found that 656 eegs have more than 1 unique target values for overlapping regions</p>\n<pre><code>df_train[] = df_train.groupby()[].diff()\ndf_uniques = df_train.loc[df_train[] &lt; ].groupby()[].nunique().value_counts()\n\nexpert_consensus\n    \n      \n       \n       \n        \nName: count, dtype: int64\n</code></pre>\n<p>and this is an example<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F661e6c5cb95d5ac3d847640831d7d490%2FScreenshot%20from%202024-01-11%2020-24-48.png?generation=1704994059097978&amp;alt=media\" alt=\"eeg\"></p>\n<p>Any idea about how to deal with them? </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2597566,
                      "author_name": "seshurajup",
                      "author_url": "",
                      "post_date": "01/11/2024 17:46:32",
                      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6b6f83ee672a60d76b77974efb14f73a%2F__results___8_0.png?generation=1704994875486323&amp;alt=media\"></p>\n<p>Exactly I also changed my notebook to get exact sense, </p>\n<pre><code>{: ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : ,\n : }\n</code></pre>\n<p>89.4% have primary targets and 10.6% have secondary targets is what I want to try or we can ignore those is other approach. Maybe after experiments we will change our opinion.</p>\n<p>1st approach is better since these votes are came from experts while label is my opinion at this point (same approach used in bird competitions)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2597589,
                          "author_name": "gunesevitan",
                          "author_url": "",
                          "post_date": "01/11/2024 18:07:26",
                          "content": "<p>Another approach is aggregating votes of overlapping rows and create a new expert consensus.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2597597,
                              "author_name": "seshurajup",
                              "author_url": "",
                              "post_date": "01/11/2024 18:10:25",
                              "content": "<p>I will try while training, Thanks for sharing <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 2597828,
                              "author_name": "cdeotte",
                              "author_url": "",
                              "post_date": "01/12/2024 00:17:58",
                              "content": "<p>Note that the labels only apply to the central 10 seconds.</p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. </p>\n</blockquote>\n<p>Is it possible that a patient has multiple targets over the course of a day?</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2597829,
                                  "author_name": "brendanartley",
                                  "author_url": "",
                                  "post_date": "01/12/2024 00:26:19",
                                  "content": "<p>Or multiple targets in the 50 second long EEG sample?</p>\n<p>Specifically the 20 seconds before and after the central 10 seconds.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2598155,
                                      "author_name": "gunesevitan",
                                      "author_url": "",
                                      "post_date": "01/12/2024 07:01:02",
                                      "content": "<p>There are lots of eegs with multiple targets in their whole 50 seconds long sequences which is perfectly fine. There are cases that satisfies <code>eeg_label_offset_seconds_diff &lt; 10</code> condition which means their central 10 seconds are overlapping but they have more than 1 unique target values. </p>\n<p>Check the highlighted row here. Last subsample with GRDA starts at 26th second so its central 10 seconds are between 46 and 56. First Other starts at 34th second so its central 10 seconds are between 54 and 64. There is only 2 seconds of overlap here but I wonder if there are any cases with more overlapping seconds.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fdc1cf01f596f1c63a41b600a598046dd%2FScreenshot%20from%202024-01-12%2009-58-13.png?generation=1705042705049577&amp;alt=media\" alt=\"eeg\"></p>\n<p><strong>Edit:</strong> There are 594 subsamples that have overlapping central 10 seconds with different expert consensus. I checked it with this code</p>\n<pre><code>df_train[] = df_train[].({\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n})\n\ndf_train[] = df_train.groupby()[].diff()\ndf_train[] = df_train.groupby()[].diff()\n\ncondition = (df_train[] &lt; ) &amp; (df_train[] != )\ndf_train.loc[condition, ].value_counts()\n\neeg_label_offset_seconds_diff\n    \n    \n    \n    \nName: count, dtype: int64\n</code></pre>\n<p>2.0 diff value occurs 109 times which means 8 seconds of overlap in the central 10 second. I guess there are very few of them and they can be filtered.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe2ee21305f9cfd14d148d7b904438af0%2FScreenshot%20from%202024-01-12%2010-36-53.png?generation=1705045031179377&amp;alt=media\" alt=\"eeg2\"></p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2596653": "I'm trying to figure out correct way of merging targets so I need some help.\n\nWe have this dataframe for eeg_id 1628180742\n![eeg_id](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F28045613639783a9cff40b842da71f99%2FScreenshot%20from%202024-01-11%2011-29-46.png?generation=1704961820794192&alt=media)\n\nand we load the eeg parquet file\n![eeg_df](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb39daaaa97f2cde4df76019e7e0d2f42%2FScreenshot%20from%202024-01-11%2011-30-44.png?generation=1704961866211132&alt=media)\n\nwe know that sample per second is 200, so we create a second column on the eeg dataframe\n```python\nsample_per_second = 200\nseconds = df_eeg.shape[0] / sample_per_second\ndf_eeg['second'] = np.repeat(np.arange(seconds), 200)\n```\n\n![eeg seconds](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F128932ddfe0c5265fe3b92592e44ecc6%2FScreenshot%20from%202024-01-11%2011-32-38.png?generation=1704961968625087&alt=media)\n\nsince we have the seconds, we can merge the targets on that column\n```python\ntarget_columns = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\ndf_eeg_merged = df_eeg.merge(df_train.loc[df_train['eeg_id'] == 1628180742, target_columns + ['eeg_label_offset_seconds']].rename(columns={'eeg_label_offset_seconds': 'second'}), on='second', how='left')\n```\n\nI'm not sure if I'm doing this correctly. Are we supposed merge targets on those exact seconds? EEGs look like this when visualized. Red areas are where seizure_vote is greater than 0.\n\n![eeg](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3fce7e9535aad74eeb3da945d1435f37%2F1628180742.png?generation=1704962100705649&alt=media)",
    "2597097": "`eeg_label_offset_seconds` - The time between the beginning of the consolidated EEG and this subsample.\n\nThis explanation is so confusing. Isn't the time between the beginning of the EGG should always be `eeg_label_offset_seconds` since beginning is always 0?\n\nThere is another thing I found. We are predicting a single target for each eeg_id but some of the eegs have more than 1 unique target. Here is the distribution of unique targets per eeg_id\n\n```python\ndf_train.groupby('eeg_id')['expert_consensus'].nunique().value_counts()\n\nexpert_consensus\n1    16306\n2      666\n3       94\n4       22\n5        1\nName: count, dtype: int64\n```",
    "2597245": "Each eeg_id can have multiple eeg_sub_id, and each of these can have a different expert_consensus",
    "2597257": "Thanks for the clarification. If I understood correctly, each eeg_sub_id can have a single expert consensus and they are 50 seconds. Test set eegs are exactly like this so we have to extract similar samples from training set.",
    "2597406": "gunesevitan \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa49d99bde59ef8aa39d01af910fbe9c1%2FUnknown.png?generation=1704991551238060&alt=media)\n\n-- **common sequence of all sub-sequences is 40secs-50secs**\n-- target votes are do not changed, so my guess is we can ignore the 0-40secs and 50secs-90secs sequence as it not given any importance to the targets ( as expert decided based on sub-sequence spike )\n\nor 0-40secs, 50secs-90secs is **weak labels approach** (its my understanding, correct me if it is wrong)\n\nUnique target votes for each egg_id\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F8a5ac4a30d4fcb816cb8fad1ad181f00%2F__results___7_0.png?generation=1704991738200947&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85ce9b71721384fcc665282e71f639ad%2F__results___8_0.png?generation=1704995430359845&alt=media)\n\n-- Similar patterns we can observe in all egg_ids where only have same target votes majorly and few have target votes changed.\n\n[Notebook - Eegs Target Analysis](https://www.kaggle.com/code/seshurajup/eegs-target-analysis/notebook)\n\n([Part 1](https://www.youtube.com/watch?v=S9NLrhj0x-M&t)) - from this video as reference of my understanding",
    "2597451": "Hmm weak labels explains a lot so for each row, expert consensus could happened at any timestep between `eeg_label_offset_seconds + 0` and `eeg_label_offset_seconds + (50 * 200)`. Is that what you meant? I wonder if there are any sub sequences with more than 1 unique labels.",
    "2597473": "gunesevitan  Similar to bird competition, we have complete 10secs of audio ( with target as primary bird, secondary bird sings ), in which bird sign happen only on 5sec-7sec, but 0-5secs and 7secs-10secs is the audio won't play useful for the primary/secondary target. in this perspective I mention as weak label.\n\nBut, in this competition we have overlapping segments which help to eliminate the 0-5secs and 7secs-10secs is my intention for selecting targets\n\n>Is that what you meant? \n**Yes**\n\n\n\n> I wonder if there are any sub sequences with more than 1 unique labels.\n**Yes** as per the unique target count graph, we have few egg sub-sequences have different target counts",
    "2597533": "I found that 656 eegs have more than 1 unique target values for overlapping regions\n```python\ndf_train['eeg_label_offset_seconds_diff'] = df_train.groupby('eeg_id')['eeg_label_offset_seconds'].diff()\ndf_uniques = df_train.loc[df_train['eeg_label_offset_seconds_diff'] < 50].groupby('eeg_id')['expert_consensus'].nunique().value_counts()\n\nexpert_consensus\n1    10109\n2      554\n3       81\n4       20\n5        1\nName: count, dtype: int64\n```\n\nand this is an example\n![eeg](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F661e6c5cb95d5ac3d847640831d7d490%2FScreenshot%20from%202024-01-11%2020-24-48.png?generation=1704994059097978&alt=media)\n\nAny idea about how to deal with them?",
    "2597566": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6b6f83ee672a60d76b77974efb14f73a%2F__results___8_0.png?generation=1704994875486323&alt=media)\n\nExactly I also changed my notebook to get exact sense, \n```python\n{1: 15282,\n 2: 1348,\n 3: 200,\n 4: 98,\n 5: 66,\n 6: 35,\n 7: 17,\n 9: 15,\n 10: 7,\n 8: 4,\n 11: 3,\n 17: 3,\n 12: 3,\n 13: 1,\n 22: 1,\n 26: 1,\n 19: 1,\n 29: 1,\n 43: 1,\n 15: 1,\n 30: 1}\n```\n\n89.4% have primary targets and 10.6% have secondary targets is what I want to try or we can ignore those is other approach. Maybe after experiments we will change our opinion.\n\n1st approach is better since these votes are came from experts while label is my opinion at this point (same approach used in bird competitions)",
    "2597589": "Another approach is aggregating votes of overlapping rows and create a new expert consensus.",
    "2597597": "I will try while training, Thanks for sharing @gunesevitan",
    "2597828": "Note that the labels only apply to the central 10 seconds.\n> The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. \n\nIs it possible that a patient has multiple targets over the course of a day?",
    "2597829": "Or multiple targets in the 50 second long EEG sample?\n\nSpecifically the 20 seconds before and after the central 10 seconds.",
    "2598155": "There are lots of eegs with multiple targets in their whole 50 seconds long sequences which is perfectly fine. There are cases that satisfies `eeg_label_offset_seconds_diff < 10` condition which means their central 10 seconds are overlapping but they have more than 1 unique target values. \n\nCheck the highlighted row here. Last subsample with GRDA starts at 26th second so its central 10 seconds are between 46 and 56. First Other starts at 34th second so its central 10 seconds are between 54 and 64. There is only 2 seconds of overlap here but I wonder if there are any cases with more overlapping seconds.\n\n![eeg](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fdc1cf01f596f1c63a41b600a598046dd%2FScreenshot%20from%202024-01-12%2009-58-13.png?generation=1705042705049577&alt=media)\n\n**Edit:** There are 594 subsamples that have overlapping central 10 seconds with different expert consensus. I checked it with this code\n\n```python\ndf_train['expert_consensus_encoded'] = df_train['expert_consensus'].map({\n    'Seizure': 1,\n    'LPD': 2,\n    'GPD': 3,\n    'LRDA': 4,\n    'GRDA': 5,\n    'Other': 6,\n})\n\ndf_train['eeg_label_offset_seconds_diff'] = df_train.groupby('eeg_id')['eeg_label_offset_seconds'].diff()\ndf_train['expert_consensus_encoded_diff'] = df_train.groupby('eeg_id')['expert_consensus_encoded'].diff()\n\ncondition = (df_train['eeg_label_offset_seconds_diff'] < 10) & (df_train['expert_consensus_encoded_diff'] != 0)\ndf_train.loc[condition, 'eeg_label_offset_seconds_diff'].value_counts()\n\neeg_label_offset_seconds_diff\n6.0    172\n8.0    167\n4.0    146\n2.0    109\nName: count, dtype: int64\n```\n\n2.0 diff value occurs 109 times which means 8 seconds of overlap in the central 10 second. I guess there are very few of them and they can be filtered.\n\n![eeg2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe2ee21305f9cfd14d148d7b904438af0%2FScreenshot%20from%202024-01-12%2010-36-53.png?generation=1705045031179377&alt=media)"
  },
  "source": "meta"
}