{
  "id": 492631,
  "title": "One eeg_id should have only 1 label or should have different labels?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492631",
  "author_name": "",
  "post_date": "2024-04-10T11:13:58.408483200Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>As in the notebook Cris shared, he used the mean label for each eeg id, is it  correct as we also see from 1st team sharing, they also use this strategy.<br>\nBut we found many eeg id with quite different labels, does it means just nosie, by default each eeg id should have 1 unique label?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F9809f18100e728488c77012dba714a08%2FAFC9C994D781E6A713D9549AF14CE779.png?generation=1712747652430688&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2745091",
      "postDate": "04/10/2024 11:13:58",
      "content": "<p>As in the notebook Cris shared, he used the mean label for each eeg id, is it  correct as we also see from 1st team sharing, they also use this strategy.<br>\nBut we found many eeg id with quite different labels, does it means just nosie, by default each eeg id should have 1 unique label?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F9809f18100e728488c77012dba714a08%2FAFC9C994D781E6A713D9549AF14CE779.png?generation=1712747652430688&amp;alt=media\"></p>",
      "rawMarkdown": "As in the notebook Cris shared, he used the mean label for each eeg id, is it  correct as we also see from 1st team sharing, they also use this strategy.\nBut we found many eeg id with quite different labels, does it means just nosie, by default each eeg id should have 1 unique label?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F9809f18100e728488c77012dba714a08%2FAFC9C994D781E6A713D9549AF14CE779.png?generation=1712747652430688&alt=media)",
      "votes": null
    },
    {
      "id": "2745208",
      "postDate": "04/10/2024 13:12:25",
      "content": "<p>I have done some experiments based on my previous single model, which randomly choose 1 sub id for each eeg id, and I used the label of that eeg sub id, now I change to use the mean label of the whole eeg id.<br>\nNotice I did not modify my cv strategy, so cv is still calc on all eeg subid based all there local offset, local label and for per eed all eeg subid total weights is 1.</p>\n<p>For model using torchaudio.stft + efficientvit_b2</p>\n<table>\n<thead>\n<tr>\n<th>strategy</th>\n<th>cv</th>\n<th>pb</th>\n<th>lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>using local offset, local label</td>\n<td>0.230</td>\n<td>0.2933</td>\n<td>0.2360</td>\n</tr>\n<tr>\n<td>using local offset, mean label</td>\n<td>0.2286</td>\n<td>0.2951</td>\n<td>0.2364</td>\n</tr>\n<tr>\n<td>using middel offset, mean label</td>\n<td>0.2343</td>\n<td>0.2908</td>\n<td>0.234108</td>\n</tr>\n<tr>\n<td>For model using scipy.spectrogram + efficientvit_b2</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>strategy</th>\n<th>cv</th>\n<th>pb</th>\n<th>lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>using local offset, local label</td>\n<td>0.2291</td>\n<td>0.289261</td>\n<td>0.2333</td>\n</tr>\n<tr>\n<td>using middel offset, mean label</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>So my cv strategy is complex yet not correct, I should have used Chris's eval and train method and submit before trying more complex strategy…</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I have done some experiments based on my previous single model, which randomly choose 1 sub id for each eeg id, and I used the label of that eeg sub id, now I change to use the mean label of the whole eeg id.\nNotice I did not modify my cv strategy, so cv is still calc on all eeg subid based all there local offset, local label and for per eed all eeg subid total weights is 1.\n\nFor model using torchaudio.stft + efficientvit_b2\n| strategy |cv  | pb| lb |\n| --- | --- | --- | --- |\n| using local offset, local label |  0.230 | 0.2933 | 0.2360 |\n| using local offset, mean label |  0.2286 | 0.2951 | 0.2364 |\n| using middel offset, mean label |  0.2343 | 0.2908 | 0.234108 |\nFor model using scipy.spectrogram + efficientvit_b2 \n\n| strategy |cv  | pb| lb |  \n| --- | --- | --- | --- |\n| using local offset, local label |  0.2291 | 0.289261 | 0.2333 |\n| using middel offset, mean label | - | -  | -  |\nSo my cv strategy is complex yet not correct, I should have used Chris's eval and train method and submit before trying more complex strategy...",
      "votes": null
    },
    {
      "id": "2745367",
      "postDate": "04/10/2024 15:19:16",
      "content": "<p>Sure I believe one <code>eeg_id</code> can have multiple labels. For example at 3:05pm, a patient has <code>LPD</code> and then later at 3:20p, the patient has <code>Seizure</code>.</p>\n<p>However, many of the train labels are <strong>fake</strong> as described in the research paper. The most likely <strong>real</strong> label is the middle label. And the test data uses all <strong>real</strong> labels.</p>\n<p>In the research paper they explain how the <strong>fake</strong> labels are generated. Imagine that annotators label the 10 seconds from 3:05:00 to 3:05:10. This is a <strong>real</strong> label from annotators. Next they generate more data by assuming the 10 seconds from 3:04:50 to 3:05:00 also has the same label. This is an educated guess (i.e. <strong>fake</strong>). And they assume 3:05:10 to 3:05:20 has the same label this is an educated guess (i.e. <strong>fake</strong>).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/fake.png\"></p>\n<p>Figure from <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/tree/main\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Sure I believe one `eeg_id` can have multiple labels. For example at 3:05pm, a patient has `LPD` and then later at 3:20p, the patient has `Seizure`.\n\nHowever, many of the train labels are **fake** as described in the research paper. The most likely **real** label is the middle label. And the test data uses all **real** labels.\n\nIn the research paper they explain how the **fake** labels are generated. Imagine that annotators label the 10 seconds from 3:05:00 to 3:05:10. This is a **real** label from annotators. Next they generate more data by assuming the 10 seconds from 3:04:50 to 3:05:00 also has the same label. This is an educated guess (i.e. **fake**). And they assume 3:05:10 to 3:05:20 has the same label this is an educated guess (i.e. **fake**).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/fake.png)\n\nFigure from [here][1]\n\n[1]: https://github.com/bdsp-core/IIIC-SPaRCNet/tree/main",
      "votes": null
    },
    {
      "id": "2745370",
      "postDate": "04/10/2024 15:22:36",
      "content": "<p>There are up to 5 different labels for a given eeg_id in the low votes ones.</p>",
      "rawMarkdown": "There are up to 5 different labels for a given eeg_id in the low votes ones.",
      "votes": null
    },
    {
      "id": "2745378",
      "postDate": "04/10/2024 15:27:45",
      "content": "<p>Thanks  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , and the third solution also tried to find the real label, more clear now!</p>",
      "rawMarkdown": "Thanks  @cdeotte , and the third solution also tried to find the real label, more clear now!",
      "votes": null
    },
    {
      "id": "2745578",
      "postDate": "04/10/2024 17:42:48",
      "content": "<p>It's not noise and eeg IDs can have multiple labels. It was mentioned on hosts paper. Subsamples with same labels belong to same stationary period. We used each stationary period as a data point in our pipeline so we had a slightly larger dataset than others. For training, we randomly selected a subsample from the same stationary period by 100% chance, and for validation, we used the middle subsample.</p>",
      "rawMarkdown": "It's not noise and eeg IDs can have multiple labels. It was mentioned on hosts paper. Subsamples with same labels belong to same stationary period. We used each stationary period as a data point in our pipeline so we had a slightly larger dataset than others. For training, we randomly selected a subsample from the same stationary period by 100% chance, and for validation, we used the middle subsample.",
      "votes": null
    },
    {
      "id": "2745895",
      "postDate": "04/10/2024 22:57:49",
      "content": "<p>Yeah, I did not carefully read that paper. And from the solution share of 1st they just use middel 50s per eeg id, though might not best, does not affect much, while I used per eeg subid local offset and label, this strategy hurt performance as the paper mentioned I should have used the middle one at least for eeg sub ids with same label.</p>",
      "rawMarkdown": "Yeah, I did not carefully read that paper. And from the solution share of 1st they just use middel 50s per eeg id, though might not best, does not affect much, while I used per eeg subid local offset and label, this strategy hurt performance as the paper mentioned I should have used the middle one at least for eeg sub ids with same label.",
      "votes": null
    },
    {
      "id": "2747605",
      "postDate": "04/12/2024 02:00:34",
      "content": "<p>Update experiment results, for my pipline, it seems the best to do is per unique eeg_subid + label, choose the middel eeg_sub_id and use it's local offset and label.<br>\nWith this new strategy, best single model PB improve from 0.289 to 0.287 and ensemble 2 models could get better result then what I did using 20 models.</p>",
      "rawMarkdown": "Update experiment results, for my pipline, it seems the best to do is per unique eeg_subid + label, choose the middel eeg_sub_id and use it's local offset and label.\nWith this new strategy, best single model PB improve from 0.289 to 0.287 and ensemble 2 models could get better result then what I did using 20 models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2745208,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "04/10/2024 13:12:25",
      "content": "<p>I have done some experiments based on my previous single model, which randomly choose 1 sub id for each eeg id, and I used the label of that eeg sub id, now I change to use the mean label of the whole eeg id.<br>\nNotice I did not modify my cv strategy, so cv is still calc on all eeg subid based all there local offset, local label and for per eed all eeg subid total weights is 1.</p>\n<p>For model using torchaudio.stft + efficientvit_b2</p>\n<table>\n<thead>\n<tr>\n<th>strategy</th>\n<th>cv</th>\n<th>pb</th>\n<th>lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>using local offset, local label</td>\n<td>0.230</td>\n<td>0.2933</td>\n<td>0.2360</td>\n</tr>\n<tr>\n<td>using local offset, mean label</td>\n<td>0.2286</td>\n<td>0.2951</td>\n<td>0.2364</td>\n</tr>\n<tr>\n<td>using middel offset, mean label</td>\n<td>0.2343</td>\n<td>0.2908</td>\n<td>0.234108</td>\n</tr>\n<tr>\n<td>For model using scipy.spectrogram + efficientvit_b2</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>strategy</th>\n<th>cv</th>\n<th>pb</th>\n<th>lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>using local offset, local label</td>\n<td>0.2291</td>\n<td>0.289261</td>\n<td>0.2333</td>\n</tr>\n<tr>\n<td>using middel offset, mean label</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>So my cv strategy is complex yet not correct, I should have used Chris's eval and train method and submit before trying more complex strategy…</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2745367,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/10/2024 15:19:16",
      "content": "<p>Sure I believe one <code>eeg_id</code> can have multiple labels. For example at 3:05pm, a patient has <code>LPD</code> and then later at 3:20p, the patient has <code>Seizure</code>.</p>\n<p>However, many of the train labels are <strong>fake</strong> as described in the research paper. The most likely <strong>real</strong> label is the middle label. And the test data uses all <strong>real</strong> labels.</p>\n<p>In the research paper they explain how the <strong>fake</strong> labels are generated. Imagine that annotators label the 10 seconds from 3:05:00 to 3:05:10. This is a <strong>real</strong> label from annotators. Next they generate more data by assuming the 10 seconds from 3:04:50 to 3:05:00 also has the same label. This is an educated guess (i.e. <strong>fake</strong>). And they assume 3:05:10 to 3:05:20 has the same label this is an educated guess (i.e. <strong>fake</strong>).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/fake.png\"></p>\n<p>Figure from <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/tree/main\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2745378,
          "author_name": "goldenlock",
          "author_url": "",
          "post_date": "04/10/2024 15:27:45",
          "content": "<p>Thanks  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , and the third solution also tried to find the real label, more clear now!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2745370,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/10/2024 15:22:36",
      "content": "<p>There are up to 5 different labels for a given eeg_id in the low votes ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2745578,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "04/10/2024 17:42:48",
      "content": "<p>It's not noise and eeg IDs can have multiple labels. It was mentioned on hosts paper. Subsamples with same labels belong to same stationary period. We used each stationary period as a data point in our pipeline so we had a slightly larger dataset than others. For training, we randomly selected a subsample from the same stationary period by 100% chance, and for validation, we used the middle subsample.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2745895,
          "author_name": "goldenlock",
          "author_url": "",
          "post_date": "04/10/2024 22:57:49",
          "content": "<p>Yeah, I did not carefully read that paper. And from the solution share of 1st they just use middel 50s per eeg id, though might not best, does not affect much, while I used per eeg subid local offset and label, this strategy hurt performance as the paper mentioned I should have used the middle one at least for eeg sub ids with same label.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2747605,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "04/12/2024 02:00:34",
      "content": "<p>Update experiment results, for my pipline, it seems the best to do is per unique eeg_subid + label, choose the middel eeg_sub_id and use it's local offset and label.<br>\nWith this new strategy, best single model PB improve from 0.289 to 0.287 and ensemble 2 models could get better result then what I did using 20 models.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2745091": "As in the notebook Cris shared, he used the mean label for each eeg id, is it  correct as we also see from 1st team sharing, they also use this strategy.\nBut we found many eeg id with quite different labels, does it means just nosie, by default each eeg id should have 1 unique label?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F9809f18100e728488c77012dba714a08%2FAFC9C994D781E6A713D9549AF14CE779.png?generation=1712747652430688&alt=media)",
    "2745208": "I have done some experiments based on my previous single model, which randomly choose 1 sub id for each eeg id, and I used the label of that eeg sub id, now I change to use the mean label of the whole eeg id.\nNotice I did not modify my cv strategy, so cv is still calc on all eeg subid based all there local offset, local label and for per eed all eeg subid total weights is 1.\n\nFor model using torchaudio.stft + efficientvit_b2\n| strategy |cv  | pb| lb |\n| --- | --- | --- | --- |\n| using local offset, local label |  0.230 | 0.2933 | 0.2360 |\n| using local offset, mean label |  0.2286 | 0.2951 | 0.2364 |\n| using middel offset, mean label |  0.2343 | 0.2908 | 0.234108 |\nFor model using scipy.spectrogram + efficientvit_b2 \n\n| strategy |cv  | pb| lb |  \n| --- | --- | --- | --- |\n| using local offset, local label |  0.2291 | 0.289261 | 0.2333 |\n| using middel offset, mean label | - | -  | -  |\nSo my cv strategy is complex yet not correct, I should have used Chris's eval and train method and submit before trying more complex strategy...",
    "2745367": "Sure I believe one `eeg_id` can have multiple labels. For example at 3:05pm, a patient has `LPD` and then later at 3:20p, the patient has `Seizure`.\n\nHowever, many of the train labels are **fake** as described in the research paper. The most likely **real** label is the middle label. And the test data uses all **real** labels.\n\nIn the research paper they explain how the **fake** labels are generated. Imagine that annotators label the 10 seconds from 3:05:00 to 3:05:10. This is a **real** label from annotators. Next they generate more data by assuming the 10 seconds from 3:04:50 to 3:05:00 also has the same label. This is an educated guess (i.e. **fake**). And they assume 3:05:10 to 3:05:20 has the same label this is an educated guess (i.e. **fake**).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/fake.png)\n\nFigure from [here][1]\n\n[1]: https://github.com/bdsp-core/IIIC-SPaRCNet/tree/main",
    "2745370": "There are up to 5 different labels for a given eeg_id in the low votes ones.",
    "2745378": "Thanks  @cdeotte , and the third solution also tried to find the real label, more clear now!",
    "2745578": "It's not noise and eeg IDs can have multiple labels. It was mentioned on hosts paper. Subsamples with same labels belong to same stationary period. We used each stationary period as a data point in our pipeline so we had a slightly larger dataset than others. For training, we randomly selected a subsample from the same stationary period by 100% chance, and for validation, we used the middle subsample.",
    "2745895": "Yeah, I did not carefully read that paper. And from the solution share of 1st they just use middel 50s per eeg id, though might not best, does not affect much, while I used per eeg subid local offset and label, this strategy hurt performance as the paper mentioned I should have used the middle one at least for eeg sub ids with same label.",
    "2747605": "Update experiment results, for my pipline, it seems the best to do is per unique eeg_subid + label, choose the middel eeg_sub_id and use it's local offset and label.\nWith this new strategy, best single model PB improve from 0.289 to 0.287 and ensemble 2 models could get better result then what I did using 20 models."
  },
  "source": "meta"
}