{
  "id": 468208,
  "title": "How to understand the target?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/468208",
  "author_name": "",
  "post_date": "2024-01-15T19:07:50.086665300Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Different row in train.csv corresponds to different time chunk of the EEG. Why different chunk of the same EEG (eeg_id) have different target? My guess is that different time period can lead to different conclusion and expert are asked to rate each chunk separately. Can someone confirm?</p>\n<p>I noticed something weird about the target, e.g. eeg_id==11127485. This table is sorted by time. Somehow, for time two, there are two more ratings than all other time.</p>\n<table>\n<thead>\n<tr>\n<th>eeg_label_offset_seconds</th>\n<th>seizure_vote</th>\n<th>lpd_vote</th>\n<th>gpd_vote</th>\n<th>lrda_vote</th>\n<th>grda_vote</th>\n<th>other_vote</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>10.0</td>\n<td>4</td>\n<td>0</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>26.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>30.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>32.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>42.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>44.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>50.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>58.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<hr>",
  "messages": [
    {
      "id": "2603429",
      "postDate": "01/15/2024 19:07:50",
      "content": "<p>Different row in train.csv corresponds to different time chunk of the EEG. Why different chunk of the same EEG (eeg_id) have different target? My guess is that different time period can lead to different conclusion and expert are asked to rate each chunk separately. Can someone confirm?</p>\n<p>I noticed something weird about the target, e.g. eeg_id==11127485. This table is sorted by time. Somehow, for time two, there are two more ratings than all other time.</p>\n<table>\n<thead>\n<tr>\n<th>eeg_label_offset_seconds</th>\n<th>seizure_vote</th>\n<th>lpd_vote</th>\n<th>gpd_vote</th>\n<th>lrda_vote</th>\n<th>grda_vote</th>\n<th>other_vote</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>10.0</td>\n<td>4</td>\n<td>0</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>26.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>30.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>32.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>42.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>44.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>50.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>58.0</td>\n<td>3</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<hr>",
      "rawMarkdown": "Different row in train.csv corresponds to different time chunk of the EEG. Why different chunk of the same EEG (eeg_id) have different target? My guess is that different time period can lead to different conclusion and expert are asked to rate each chunk separately. Can someone confirm?\n\nI noticed something weird about the target, e.g. eeg_id==11127485. This table is sorted by time. Somehow, for time two, there are two more ratings than all other time.\n\n| eeg_label_offset_seconds | seizure_vote | lpd_vote | gpd_vote | lrda_vote | grda_vote | other_vote |\n|--------------------------|--------------|----------|----------|-----------|-----------|------------|\n| 0.0                      | 3            | 0        | 0        | 0         | 0         | 0          |\n| 10.0                     | 4            | 0        | 1        | 0         | 0         | 0          |\n| 26.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 30.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 32.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 42.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 44.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 50.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 58.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n****",
      "votes": null
    },
    {
      "id": "2603439",
      "postDate": "01/15/2024 19:16:59",
      "content": "<p><a href=\"https://www.kaggle.com/zhenlanwang\" target=\"_blank\">@zhenlanwang</a> 16 to 20 experts given voting for 10secs EGG Sequence as explained in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">Discussion</a></p>\n<ul>\n<li>in this case, minimum 5 experts given voting for these sub-sequences </li>\n<li>add egg offset also to this table, it make sense better understanding</li>\n</ul>",
      "rawMarkdown": "zhenlanwang 16 to 20 experts given voting for 10secs EGG Sequence as explained in this [Discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010)\n- in this case, minimum 5 experts given voting for these sub-sequences \n- add egg offset also to this table, it make sense better understanding",
      "votes": null
    },
    {
      "id": "2603563",
      "postDate": "01/15/2024 21:29:26",
      "content": "<p>I added eeg offset. the time is increasing but votes are not.</p>",
      "rawMarkdown": "I added eeg offset. the time is increasing but votes are not.",
      "votes": null
    },
    {
      "id": "2603573",
      "postDate": "01/15/2024 21:50:58",
      "content": "<pre><code>egg = pd.read_parquet()\negg.shape\n\n(, ) =&gt; / =&gt; 108sec sequence\n</code></pre>\n<pre><code>train[] = train[].apply( x: )\ntrain[] = train[].apply( x: )\ntrain[] = (egg.shape[]/)\ntrain[train[] == ][[,,,,,,,,,,,]].to_markdown()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fad8d8c0b230547a214cc7399f483845e%2FScreenshot%202024-01-16%20at%203.25.43AM.png?generation=1705355775031469&amp;alt=media\"><br>\nFor the sub sequence <strong>30sec-40sec</strong> experts voted seizure_vote = 4 experts and gpd_vote = 1 experts<br>\nAll sub sequences from <strong>20secs to 88secs</strong> except <strong>30secs-40secs</strong> experts voted  seizure_vote = 3 experts<br>\nAll sub sequences from <strong>0secs-20secs</strong> &amp; <strong>88secs to 108secs</strong> don't have target information except it is used as reference while voting for the sequence <strong>20secs-30secs</strong>, <strong>78secs-88secs</strong> respectively </p>\n<p>Since, we have  108secs minimum and major votes for Seizure. We can strong consider subsquence from 20secs to 88secs have Seizure with probability (3<em>8+4)/(3</em>8+4 + 1) ~ <strong>0.965</strong></p>\n<p>Since sub sequences are short in this case 0secs-108secs used where Seizure target is well visible in 20secs-88secs as per voting counts</p>\n<p><strong>Note</strong><br>\nBased on this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606655\" target=\"_blank\">discussion</a>, and data tab<br>\n<em>test_agg</em>: <strong>Exactly 50 seconds of EEG data</strong>. We no need to aggregate votes over each eeg_id, we can consider one sub-sequence at a time.</p>",
      "rawMarkdown": "```python\negg = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/11127485.parquet\")\negg.shape\n\n(21600, 20) => 21600/200 => 108sec sequence\n```\n\n```python\ntrain['range'] = train['eeg_label_offset_seconds'].apply(lambda x: f\"{int(x)}-{int(x)+50}\")\ntrain['target_range'] = train['eeg_label_offset_seconds'].apply(lambda x: f\"{int(x)+20}-{int(x)+50-20}\")\ntrain['eeg_len'] = int(egg.shape[0]/200)\ntrain[train['eeg_id'] == 11127485][['eeg_id','eeg_label_offset_seconds','expert_consensus','seizure_vote','lpd_vote','gpd_vote','lrda_vote','grda_vote','other_vote','range','target_range','eeg_len']].to_markdown()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fad8d8c0b230547a214cc7399f483845e%2FScreenshot%202024-01-16%20at%203.25.43AM.png?generation=1705355775031469&alt=media)\nFor the sub sequence **30sec-40sec** experts voted seizure_vote = 4 experts and gpd_vote = 1 experts\nAll sub sequences from **20secs to 88secs** except **30secs-40secs** experts voted  seizure_vote = 3 experts\nAll sub sequences from **0secs-20secs** & **88secs to 108secs** don't have target information except it is used as reference while voting for the sequence **20secs-30secs**, **78secs-88secs** respectively \n\nSince, we have  108secs minimum and major votes for Seizure. We can strong consider subsquence from 20secs to 88secs have Seizure with probability (3*8+4)/(3*8+4 + 1) ~ **0.965**\n\nSince sub sequences are short in this case 0secs-108secs used where Seizure target is well visible in 20secs-88secs as per voting counts\n\n**Note**\nBased on this [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606655), and data tab\n*test_agg*: **Exactly 50 seconds of EEG data**. We no need to aggregate votes over each eeg_id, we can consider one sub-sequence at a time.",
      "votes": null
    },
    {
      "id": "2604518",
      "postDate": "01/16/2024 13:17:39",
      "content": "<p>Thank you for the explanation. It is very helpful.</p>",
      "rawMarkdown": "Thank you for the explanation. It is very helpful.",
      "votes": null
    },
    {
      "id": "2604559",
      "postDate": "01/16/2024 13:54:57",
      "content": "<p><a href=\"https://www.kaggle.com/zhenlanwang\" target=\"_blank\">@zhenlanwang</a> you are welcome</p>",
      "rawMarkdown": "zhenlanwang you are welcome",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2603439,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/15/2024 19:16:59",
      "content": "<p><a href=\"https://www.kaggle.com/zhenlanwang\" target=\"_blank\">@zhenlanwang</a> 16 to 20 experts given voting for 10secs EGG Sequence as explained in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">Discussion</a></p>\n<ul>\n<li>in this case, minimum 5 experts given voting for these sub-sequences </li>\n<li>add egg offset also to this table, it make sense better understanding</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2603563,
          "author_name": "zhenlanwang",
          "author_url": "",
          "post_date": "01/15/2024 21:29:26",
          "content": "<p>I added eeg offset. the time is increasing but votes are not.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2603573,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/15/2024 21:50:58",
              "content": "<pre><code>egg = pd.read_parquet()\negg.shape\n\n(, ) =&gt; / =&gt; 108sec sequence\n</code></pre>\n<pre><code>train[] = train[].apply( x: )\ntrain[] = train[].apply( x: )\ntrain[] = (egg.shape[]/)\ntrain[train[] == ][[,,,,,,,,,,,]].to_markdown()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fad8d8c0b230547a214cc7399f483845e%2FScreenshot%202024-01-16%20at%203.25.43AM.png?generation=1705355775031469&amp;alt=media\"><br>\nFor the sub sequence <strong>30sec-40sec</strong> experts voted seizure_vote = 4 experts and gpd_vote = 1 experts<br>\nAll sub sequences from <strong>20secs to 88secs</strong> except <strong>30secs-40secs</strong> experts voted  seizure_vote = 3 experts<br>\nAll sub sequences from <strong>0secs-20secs</strong> &amp; <strong>88secs to 108secs</strong> don't have target information except it is used as reference while voting for the sequence <strong>20secs-30secs</strong>, <strong>78secs-88secs</strong> respectively </p>\n<p>Since, we have  108secs minimum and major votes for Seizure. We can strong consider subsquence from 20secs to 88secs have Seizure with probability (3<em>8+4)/(3</em>8+4 + 1) ~ <strong>0.965</strong></p>\n<p>Since sub sequences are short in this case 0secs-108secs used where Seizure target is well visible in 20secs-88secs as per voting counts</p>\n<p><strong>Note</strong><br>\nBased on this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606655\" target=\"_blank\">discussion</a>, and data tab<br>\n<em>test_agg</em>: <strong>Exactly 50 seconds of EEG data</strong>. We no need to aggregate votes over each eeg_id, we can consider one sub-sequence at a time.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2604518,
                  "author_name": "zhenlanwang",
                  "author_url": "",
                  "post_date": "01/16/2024 13:17:39",
                  "content": "<p>Thank you for the explanation. It is very helpful.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2604559,
                      "author_name": "seshurajup",
                      "author_url": "",
                      "post_date": "01/16/2024 13:54:57",
                      "content": "<p><a href=\"https://www.kaggle.com/zhenlanwang\" target=\"_blank\">@zhenlanwang</a> you are welcome</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2603429": "Different row in train.csv corresponds to different time chunk of the EEG. Why different chunk of the same EEG (eeg_id) have different target? My guess is that different time period can lead to different conclusion and expert are asked to rate each chunk separately. Can someone confirm?\n\nI noticed something weird about the target, e.g. eeg_id==11127485. This table is sorted by time. Somehow, for time two, there are two more ratings than all other time.\n\n| eeg_label_offset_seconds | seizure_vote | lpd_vote | gpd_vote | lrda_vote | grda_vote | other_vote |\n|--------------------------|--------------|----------|----------|-----------|-----------|------------|\n| 0.0                      | 3            | 0        | 0        | 0         | 0         | 0          |\n| 10.0                     | 4            | 0        | 1        | 0         | 0         | 0          |\n| 26.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 30.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 32.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 42.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 44.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 50.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n| 58.0                     | 3            | 0        | 0        | 0         | 0         | 0          |\n****",
    "2603439": "zhenlanwang 16 to 20 experts given voting for 10secs EGG Sequence as explained in this [Discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010)\n- in this case, minimum 5 experts given voting for these sub-sequences \n- add egg offset also to this table, it make sense better understanding",
    "2603563": "I added eeg offset. the time is increasing but votes are not.",
    "2603573": "```python\negg = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/11127485.parquet\")\negg.shape\n\n(21600, 20) => 21600/200 => 108sec sequence\n```\n\n```python\ntrain['range'] = train['eeg_label_offset_seconds'].apply(lambda x: f\"{int(x)}-{int(x)+50}\")\ntrain['target_range'] = train['eeg_label_offset_seconds'].apply(lambda x: f\"{int(x)+20}-{int(x)+50-20}\")\ntrain['eeg_len'] = int(egg.shape[0]/200)\ntrain[train['eeg_id'] == 11127485][['eeg_id','eeg_label_offset_seconds','expert_consensus','seizure_vote','lpd_vote','gpd_vote','lrda_vote','grda_vote','other_vote','range','target_range','eeg_len']].to_markdown()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fad8d8c0b230547a214cc7399f483845e%2FScreenshot%202024-01-16%20at%203.25.43AM.png?generation=1705355775031469&alt=media)\nFor the sub sequence **30sec-40sec** experts voted seizure_vote = 4 experts and gpd_vote = 1 experts\nAll sub sequences from **20secs to 88secs** except **30secs-40secs** experts voted  seizure_vote = 3 experts\nAll sub sequences from **0secs-20secs** & **88secs to 108secs** don't have target information except it is used as reference while voting for the sequence **20secs-30secs**, **78secs-88secs** respectively \n\nSince, we have  108secs minimum and major votes for Seizure. We can strong consider subsquence from 20secs to 88secs have Seizure with probability (3*8+4)/(3*8+4 + 1) ~ **0.965**\n\nSince sub sequences are short in this case 0secs-108secs used where Seizure target is well visible in 20secs-88secs as per voting counts\n\n**Note**\nBased on this [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606655), and data tab\n*test_agg*: **Exactly 50 seconds of EEG data**. We no need to aggregate votes over each eeg_id, we can consider one sub-sequence at a time.",
    "2604518": "Thank you for the explanation. It is very helpful.",
    "2604559": "zhenlanwang you are welcome"
  },
  "source": "meta"
}