{
  "id": 479244,
  "title": "Dataset: The Temple University Hospital Seizure Detection Corpus",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/479244",
  "author_name": "",
  "post_date": "2024-02-23T17:23:33.809348200Z",
  "votes": 43,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Recently, I learned about the existence of the TUH EEG Seizure Corpus but could not access it <a href=\"https://isip.piconepress.com/projects/tuh_eeg/html/downloads.shtml\" target=\"_blank\">here</a>. Luckily, the portion of data we desire is hosted on Kaggle <a href=\"https://www.kaggle.com/datasets/psyryuvok/the-tuh-eeg-seizure-corpus-tusz-v152\" target=\"_blank\">here</a>. I located the <a href=\"https://isip.piconepress.com/publications/book_sections/2018/frontiers_neuroscience/tuh_eeg/\" target=\"_blank\">manuscript</a> describing the dataset, and worked out the rest. More information can be found in this <a href=\"https://www.kaggle.com/code/seanbearden/processing-the-tuh-eeg-seizure-corpus\" target=\"_blank\">notebook</a>. </p>\n<p>Here is some info on the dataset:</p>\n<ul>\n<li>There are 637 patient ids.</li>\n<li>There are 1367 recording sessions.</li>\n<li>There are 443 sessions with seizure events.</li>\n<li>3050 events have been labeled seizures.</li>\n<li>8662 events have been labeled baseline/non-interesting events.</li>\n</ul>\n<p>From the <a href=\"https://isip.piconepress.com/publications/book_sections/2018/frontiers_neuroscience/tuh_eeg/\" target=\"_blank\">manuscript</a>, we learn how this data is organized. The seizures are identified by type: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Faf0188f3a8f5d1f5bf58a30ad7cc2ebd%2FScreenshot%202024-02-23%20at%207.39.22AM.png?generation=1708708353881792&amp;alt=media\"></p>\n<p>Note that two of the seizure types are clinically determined. We might want to avoid using these seizures when training models with EEG data.</p>\n<p>EKG signal not always present, check <code>seizure.csv</code> for details. EKG units may differ from competition dataset. </p>\n<p>The dataset can be accessed <a href=\"https://www.kaggle.com/datasets/seanbearden/hms-hba-tuh-tusz-seizures\" target=\"_blank\">here</a>. An example of usage with EfficientNet can be found <a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">here</a></p>\n<p>I have not included the background events, but will upload them in another dataset. They may be useful for the \"Other\" class, but I'm not sure we can exclude the possibility background belongs to any other non-seizure class. </p>\n<p>What are your thoughts on this dataset? Has anyone used it to enhance a model's performance? Should the spectrograms be adjusted?</p>",
  "messages": [
    {
      "id": "2665477",
      "postDate": "02/23/2024 17:23:33",
      "content": "<p>Recently, I learned about the existence of the TUH EEG Seizure Corpus but could not access it <a href=\"https://isip.piconepress.com/projects/tuh_eeg/html/downloads.shtml\" target=\"_blank\">here</a>. Luckily, the portion of data we desire is hosted on Kaggle <a href=\"https://www.kaggle.com/datasets/psyryuvok/the-tuh-eeg-seizure-corpus-tusz-v152\" target=\"_blank\">here</a>. I located the <a href=\"https://isip.piconepress.com/publications/book_sections/2018/frontiers_neuroscience/tuh_eeg/\" target=\"_blank\">manuscript</a> describing the dataset, and worked out the rest. More information can be found in this <a href=\"https://www.kaggle.com/code/seanbearden/processing-the-tuh-eeg-seizure-corpus\" target=\"_blank\">notebook</a>. </p>\n<p>Here is some info on the dataset:</p>\n<ul>\n<li>There are 637 patient ids.</li>\n<li>There are 1367 recording sessions.</li>\n<li>There are 443 sessions with seizure events.</li>\n<li>3050 events have been labeled seizures.</li>\n<li>8662 events have been labeled baseline/non-interesting events.</li>\n</ul>\n<p>From the <a href=\"https://isip.piconepress.com/publications/book_sections/2018/frontiers_neuroscience/tuh_eeg/\" target=\"_blank\">manuscript</a>, we learn how this data is organized. The seizures are identified by type: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Faf0188f3a8f5d1f5bf58a30ad7cc2ebd%2FScreenshot%202024-02-23%20at%207.39.22AM.png?generation=1708708353881792&amp;alt=media\"></p>\n<p>Note that two of the seizure types are clinically determined. We might want to avoid using these seizures when training models with EEG data.</p>\n<p>EKG signal not always present, check <code>seizure.csv</code> for details. EKG units may differ from competition dataset. </p>\n<p>The dataset can be accessed <a href=\"https://www.kaggle.com/datasets/seanbearden/hms-hba-tuh-tusz-seizures\" target=\"_blank\">here</a>. An example of usage with EfficientNet can be found <a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">here</a></p>\n<p>I have not included the background events, but will upload them in another dataset. They may be useful for the \"Other\" class, but I'm not sure we can exclude the possibility background belongs to any other non-seizure class. </p>\n<p>What are your thoughts on this dataset? Has anyone used it to enhance a model's performance? Should the spectrograms be adjusted?</p>",
      "rawMarkdown": "Recently, I learned about the existence of the TUH EEG Seizure Corpus but could not access it [here][1]. Luckily, the portion of data we desire is hosted on Kaggle [here][2]. I located the [manuscript][3] describing the dataset, and worked out the rest. More information can be found in this [notebook][0]. \n\nHere is some info on the dataset:\n\n- There are 637 patient ids.\n- There are 1367 recording sessions.\n- There are 443 sessions with seizure events.\n- 3050 events have been labeled seizures.\n- 8662 events have been labeled baseline/non-interesting events.\n\nFrom the [manuscript][3], we learn how this data is organized. The seizures are identified by type: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Faf0188f3a8f5d1f5bf58a30ad7cc2ebd%2FScreenshot%202024-02-23%20at%207.39.22AM.png?generation=1708708353881792&alt=media)\n\nNote that two of the seizure types are clinically determined. We might want to avoid using these seizures when training models with EEG data.\n\nEKG signal not always present, check `seizure.csv` for details. EKG units may differ from competition dataset. \n\nThe dataset can be accessed [here][6]. An example of usage with EfficientNet can be found [here][7]\n\nI have not included the background events, but will upload them in another dataset. They may be useful for the \"Other\" class, but I'm not sure we can exclude the possibility background belongs to any other non-seizure class. \n\nWhat are your thoughts on this dataset? Has anyone used it to enhance a model's performance? Should the spectrograms be adjusted?\n\n\n[0]: https://www.kaggle.com/code/seanbearden/processing-the-tuh-eeg-seizure-corpus\n[1]: https://isip.piconepress.com/projects/tuh_eeg/html/downloads.shtml\n[2]: https://www.kaggle.com/datasets/psyryuvok/the-tuh-eeg-seizure-corpus-tusz-v152\n[3]: https://isip.piconepress.com/publications/book_sections/2018/frontiers_neuroscience/tuh_eeg/\n[4]: https://www.kaggle.com/code/seanbearden/missing-data-in-spectrograms\n[5]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/478233\n[6]: https://www.kaggle.com/datasets/seanbearden/hms-hba-tuh-tusz-seizures\n[7]: https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39",
      "votes": null
    },
    {
      "id": "2665824",
      "postDate": "02/23/2024 21:01:26",
      "content": "<p>Thanks for sharing yet another great contribution!</p>",
      "rawMarkdown": "Thanks for sharing yet another great contribution!",
      "votes": null
    },
    {
      "id": "2666104",
      "postDate": "02/24/2024 06:37:05",
      "content": "<p>Thanks for sharing! adding new dataset will be helpful. </p>\n<p>and then i have question that in  <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439</a>, author found out eeg dataset, Have you tried experimenting with this dataset before??</p>",
      "rawMarkdown": "Thanks for sharing! adding new dataset will be helpful. \n\nand then i have question that in  https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439, author found out eeg dataset, Have you tried experimenting with this dataset before??",
      "votes": null
    },
    {
      "id": "2666569",
      "postDate": "02/24/2024 12:58:11",
      "content": "<p>Thanks for the info! I haven’t tried experimenting with it, but it’s on my todo list now</p>",
      "rawMarkdown": "Thanks for the info! I haven’t tried experimenting with it, but it’s on my todo list now",
      "votes": null
    },
    {
      "id": "2669943",
      "postDate": "02/26/2024 15:06:43",
      "content": "<p>Thanks for sharing! I conducted a detailed analysis of CV and found that my models tend to recognize \"seizure\" as \"other\". Additional data will be very useful to address this issue.</p>",
      "rawMarkdown": "Thanks for sharing! I conducted a detailed analysis of CV and found that my models tend to recognize \"seizure\" as \"other\". Additional data will be very useful to address this issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2665824,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "02/23/2024 21:01:26",
      "content": "<p>Thanks for sharing yet another great contribution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2666104,
      "author_name": "seoyunje",
      "author_url": "",
      "post_date": "02/24/2024 06:37:05",
      "content": "<p>Thanks for sharing! adding new dataset will be helpful. </p>\n<p>and then i have question that in  <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439</a>, author found out eeg dataset, Have you tried experimenting with this dataset before??</p>",
      "votes": null,
      "replies": [
        {
          "id": 2666569,
          "author_name": "seanbearden",
          "author_url": "",
          "post_date": "02/24/2024 12:58:11",
          "content": "<p>Thanks for the info! I haven’t tried experimenting with it, but it’s on my todo list now</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2669943,
      "author_name": "zijiangyang1116",
      "author_url": "",
      "post_date": "02/26/2024 15:06:43",
      "content": "<p>Thanks for sharing! I conducted a detailed analysis of CV and found that my models tend to recognize \"seizure\" as \"other\". Additional data will be very useful to address this issue.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2665477": "Recently, I learned about the existence of the TUH EEG Seizure Corpus but could not access it [here][1]. Luckily, the portion of data we desire is hosted on Kaggle [here][2]. I located the [manuscript][3] describing the dataset, and worked out the rest. More information can be found in this [notebook][0]. \n\nHere is some info on the dataset:\n\n- There are 637 patient ids.\n- There are 1367 recording sessions.\n- There are 443 sessions with seizure events.\n- 3050 events have been labeled seizures.\n- 8662 events have been labeled baseline/non-interesting events.\n\nFrom the [manuscript][3], we learn how this data is organized. The seizures are identified by type: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Faf0188f3a8f5d1f5bf58a30ad7cc2ebd%2FScreenshot%202024-02-23%20at%207.39.22AM.png?generation=1708708353881792&alt=media)\n\nNote that two of the seizure types are clinically determined. We might want to avoid using these seizures when training models with EEG data.\n\nEKG signal not always present, check `seizure.csv` for details. EKG units may differ from competition dataset. \n\nThe dataset can be accessed [here][6]. An example of usage with EfficientNet can be found [here][7]\n\nI have not included the background events, but will upload them in another dataset. They may be useful for the \"Other\" class, but I'm not sure we can exclude the possibility background belongs to any other non-seizure class. \n\nWhat are your thoughts on this dataset? Has anyone used it to enhance a model's performance? Should the spectrograms be adjusted?\n\n\n[0]: https://www.kaggle.com/code/seanbearden/processing-the-tuh-eeg-seizure-corpus\n[1]: https://isip.piconepress.com/projects/tuh_eeg/html/downloads.shtml\n[2]: https://www.kaggle.com/datasets/psyryuvok/the-tuh-eeg-seizure-corpus-tusz-v152\n[3]: https://isip.piconepress.com/publications/book_sections/2018/frontiers_neuroscience/tuh_eeg/\n[4]: https://www.kaggle.com/code/seanbearden/missing-data-in-spectrograms\n[5]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/478233\n[6]: https://www.kaggle.com/datasets/seanbearden/hms-hba-tuh-tusz-seizures\n[7]: https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39",
    "2665824": "Thanks for sharing yet another great contribution!",
    "2666104": "Thanks for sharing! adding new dataset will be helpful. \n\nand then i have question that in  https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439, author found out eeg dataset, Have you tried experimenting with this dataset before??",
    "2666569": "Thanks for the info! I haven’t tried experimenting with it, but it’s on my todo list now",
    "2669943": "Thanks for sharing! I conducted a detailed analysis of CV and found that my models tend to recognize \"seizure\" as \"other\". Additional data will be very useful to address this issue."
  },
  "source": "meta"
}