{
  "id": 479207,
  "title": "Preprocessing in the Labeling Done by Experts",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/479207",
  "author_name": "",
  "post_date": "2024-02-23T14:30:45.477300400Z",
  "votes": 47,
  "comment_count": 21,
  "views": 0,
  "content": "<p>For the past few weeks, I have been working on deciphering which type of EEG data preprocessing is most aligned with this competition. After much contemplation and in-depth data analysis, I've reached an insight.</p>\n<p>The data were labeling by experts in specific formats. The EEG data were presented to these experts with a certain type of preprocessing, and the spectrograms were also shown with a type of preprocessing. From this, I concluded that the competition's classifications might be biased towards the preprocessing shown to the experts for making the classifications.</p>\n<p>Note: I am not an expert in the field, and this is just an insight. I believe I am correct, but I do not claim to hold the ultimate truth. If anyone has anything to say on this matter, your input is more than welcome.</p>\n<p>Below is an example image from the competition showing how this classification was made from the experts' perspective.</p>\n<p><img src=\"https://i.imgur.com/hq2il2c.png\"></p>\n<p>Viewing it from this perspective led me to the following conclusion, which may be right or wrong:</p>\n<p>The closer my preprocessing is to the preprocessing used by the experts to generate the classification, the closer my algorithm will come to achieving results similar to theirs.<br>\nWith this in mind, I created a notebook where:</p>\n<ol>\n<li>I try to create spectrograms for the competition as similar as possible to those seen by the experts.</li>\n<li>I attempt to create new EEG spectrograms that are as close as possible to the competition's spectrograms.</li>\n<li>I compare these spectrograms and try to identify if the preprocessing I used on the EEGs generated spectrograms similar to those in the competition, thus attempting to discover how similar my preprocessing is to that used in the competition.</li>\n</ol>\n<p>I know it's a time-consuming, subjective task. If anyone has better ideas, I would love to hear them.</p>\n<p>Here is the link to this notebook:<br>\n<a href=\"https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\" target=\"_blank\">https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu</a></p>\n<p>Just to give an example of what I'm talking about, here is an image that turned out quite similar to the competition's preprocessing:</p>\n<p><img src=\"https://i.imgur.com/9jL3bGi.png\"></p>\n<p>The preprocessing used in this case was a 60Hz notch filter and a 5th-order bandpass butter filter - low_cut 0.7, high_cut 20.</p>\n<p>To illustrate, here is an example where the spectrogram did not turn out similar. I believe something is missing; in this case, it seems there might be some treatment of a bad channel here that I couldn't figure out how to perform.</p>\n<p><img src=\"https://i.imgur.com/ub0E9mH.png\"></p>",
  "messages": [
    {
      "id": "2665225",
      "postDate": "02/23/2024 14:30:45",
      "content": "<p>For the past few weeks, I have been working on deciphering which type of EEG data preprocessing is most aligned with this competition. After much contemplation and in-depth data analysis, I've reached an insight.</p>\n<p>The data were labeling by experts in specific formats. The EEG data were presented to these experts with a certain type of preprocessing, and the spectrograms were also shown with a type of preprocessing. From this, I concluded that the competition's classifications might be biased towards the preprocessing shown to the experts for making the classifications.</p>\n<p>Note: I am not an expert in the field, and this is just an insight. I believe I am correct, but I do not claim to hold the ultimate truth. If anyone has anything to say on this matter, your input is more than welcome.</p>\n<p>Below is an example image from the competition showing how this classification was made from the experts' perspective.</p>\n<p><img src=\"https://i.imgur.com/hq2il2c.png\"></p>\n<p>Viewing it from this perspective led me to the following conclusion, which may be right or wrong:</p>\n<p>The closer my preprocessing is to the preprocessing used by the experts to generate the classification, the closer my algorithm will come to achieving results similar to theirs.<br>\nWith this in mind, I created a notebook where:</p>\n<ol>\n<li>I try to create spectrograms for the competition as similar as possible to those seen by the experts.</li>\n<li>I attempt to create new EEG spectrograms that are as close as possible to the competition's spectrograms.</li>\n<li>I compare these spectrograms and try to identify if the preprocessing I used on the EEGs generated spectrograms similar to those in the competition, thus attempting to discover how similar my preprocessing is to that used in the competition.</li>\n</ol>\n<p>I know it's a time-consuming, subjective task. If anyone has better ideas, I would love to hear them.</p>\n<p>Here is the link to this notebook:<br>\n<a href=\"https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\" target=\"_blank\">https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu</a></p>\n<p>Just to give an example of what I'm talking about, here is an image that turned out quite similar to the competition's preprocessing:</p>\n<p><img src=\"https://i.imgur.com/9jL3bGi.png\"></p>\n<p>The preprocessing used in this case was a 60Hz notch filter and a 5th-order bandpass butter filter - low_cut 0.7, high_cut 20.</p>\n<p>To illustrate, here is an example where the spectrogram did not turn out similar. I believe something is missing; in this case, it seems there might be some treatment of a bad channel here that I couldn't figure out how to perform.</p>\n<p><img src=\"https://i.imgur.com/ub0E9mH.png\"></p>",
      "rawMarkdown": "For the past few weeks, I have been working on deciphering which type of EEG data preprocessing is most aligned with this competition. After much contemplation and in-depth data analysis, I've reached an insight.\n\nThe data were labeling by experts in specific formats. The EEG data were presented to these experts with a certain type of preprocessing, and the spectrograms were also shown with a type of preprocessing. From this, I concluded that the competition's classifications might be biased towards the preprocessing shown to the experts for making the classifications.\n\nNote: I am not an expert in the field, and this is just an insight. I believe I am correct, but I do not claim to hold the ultimate truth. If anyone has anything to say on this matter, your input is more than welcome.\n\nBelow is an example image from the competition showing how this classification was made from the experts' perspective.\n\n![](https://i.imgur.com/hq2il2c.png)\n\nViewing it from this perspective led me to the following conclusion, which may be right or wrong:\n\nThe closer my preprocessing is to the preprocessing used by the experts to generate the classification, the closer my algorithm will come to achieving results similar to theirs.\nWith this in mind, I created a notebook where:\n\n1. I try to create spectrograms for the competition as similar as possible to those seen by the experts.\n1. I attempt to create new EEG spectrograms that are as close as possible to the competition's spectrograms.\n1. I compare these spectrograms and try to identify if the preprocessing I used on the EEGs generated spectrograms similar to those in the competition, thus attempting to discover how similar my preprocessing is to that used in the competition.\n\nI know it's a time-consuming, subjective task. If anyone has better ideas, I would love to hear them.\n\nHere is the link to this notebook:\nhttps://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\n\nJust to give an example of what I'm talking about, here is an image that turned out quite similar to the competition's preprocessing:\n\n![](https://i.imgur.com/9jL3bGi.png)\n\nThe preprocessing used in this case was a 60Hz notch filter and a 5th-order bandpass butter filter - low_cut 0.7, high_cut 20.\n\nTo illustrate, here is an example where the spectrogram did not turn out similar. I believe something is missing; in this case, it seems there might be some treatment of a bad channel here that I couldn't figure out how to perform.\n\n![](https://i.imgur.com/ub0E9mH.png)",
      "votes": null
    },
    {
      "id": "2665263",
      "postDate": "02/23/2024 15:03:24",
      "content": "<p>Oh. I thought the experts received the exact subsamples we extract from eegs and spectograms. Without data enginnering. Good to know and I definitely agree.</p>",
      "rawMarkdown": "Oh. I thought the experts received the exact subsamples we extract from eegs and spectograms. Without data enginnering. Good to know and I definitely agree.",
      "votes": null
    },
    {
      "id": "2665533",
      "postDate": "02/23/2024 17:49:33",
      "content": "<p>Hello, Ángel, how are you? Thanks for the comment.</p>\n<p>It seems that at least the eeg data receives a notch filter and a bandpass, because if you look at the eeg data in the example images of the competition, you can see that there appears to be a filter at least on the higher frequency data and at 60Hz, This seems to be confirmed when comparing the data from the spectogram provided by the competition and the spectogram created from EEGs, apparently there seems to be more pre-processing, in the case of null values, bad channels, artifacts and EKG.</p>",
      "rawMarkdown": "Hello, Ángel, how are you? Thanks for the comment.\n\nIt seems that at least the eeg data receives a notch filter and a bandpass, because if you look at the eeg data in the example images of the competition, you can see that there appears to be a filter at least on the higher frequency data and at 60Hz, This seems to be confirmed when comparing the data from the spectogram provided by the competition and the spectogram created from EEGs, apparently there seems to be more pre-processing, in the case of null values, bad channels, artifacts and EKG.",
      "votes": null
    },
    {
      "id": "2665616",
      "postDate": "02/23/2024 18:44:20",
      "content": "<p><a href=\"https://journals.lww.com/clinicalneurophys/Fulltext/2016/08000/American_Clinical_Neurophysiology_Society.3.aspx\" target=\"_blank\">https://journals.lww.com/clinicalneurophys/Fulltext/2016/08000/American_Clinical_Neurophysiology_Society.3.aspx</a></p>\n<p>The document describes the 'standards' that should have been followed for our data.</p>",
      "rawMarkdown": "https://journals.lww.com/clinicalneurophys/Fulltext/2016/08000/American_Clinical_Neurophysiology_Society.3.aspx\n\nThe document describes the 'standards' that should have been followed for our data.",
      "votes": null
    },
    {
      "id": "2665631",
      "postDate": "02/23/2024 18:55:13",
      "content": "<p>\"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"<br>\nAre you sure that examples are not just that, examples to give us ideas? From data description I understand experts reviewed data as we extract it from eegs and spectogram folders.<br>\nExamples eeg covers only the 10 middle labeled seconds.</p>",
      "rawMarkdown": "\"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"\nAre you sure that examples are not just that, examples to give us ideas? From data description I understand experts reviewed data as we extract it from eegs and spectogram folders.\nExamples eeg covers only the 10 middle labeled seconds.",
      "votes": null
    },
    {
      "id": "2665686",
      "postDate": "02/23/2024 19:40:49",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>. This is an awesome way to visualize and inspect the data</p>",
      "rawMarkdown": "Great work @rafaelzimmermann1. This is an awesome way to visualize and inspect the data",
      "votes": null
    },
    {
      "id": "2665693",
      "postDate": "02/23/2024 19:47:17",
      "content": "<p>Tks Shane </p>",
      "rawMarkdown": "Tks Shane",
      "votes": null
    },
    {
      "id": "2665712",
      "postDate": "02/23/2024 19:58:34",
      "content": "<p>Good question, unfortunately I don't have the answer, I believe the organizers can answer this…</p>\n<p>My main doubt is the preprocessing,<br>\n1- Whether standard preprocessing was done to show the data to experts or not?<br>\n2- Each expert performed their preprocessing and then labeled the 10 seconds.<br>\n3- No pre-processing was done, the data was shown in tools like eeglab and each expert used their preferred filter parameters etc in the tool?</p>\n<p>I believe these are important questions, but I'm not sure of any answers, one of the main reasons for the discussion is to see if anyone knows the answer :)</p>",
      "rawMarkdown": "Good question, unfortunately I don't have the answer, I believe the organizers can answer this...\n\nMy main doubt is the preprocessing,\n1- Whether standard preprocessing was done to show the data to experts or not?\n2- Each expert performed their preprocessing and then labeled the 10 seconds.\n3- No pre-processing was done, the data was shown in tools like eeglab and each expert used their preferred filter parameters etc in the tool?\n\nI believe these are important questions, but I'm not sure of any answers, one of the main reasons for the discussion is to see if anyone knows the answer :)",
      "votes": null
    },
    {
      "id": "2666093",
      "postDate": "02/24/2024 06:30:37",
      "content": "<p>First of all, Thanks for sharing! it's good approach</p>\n<p>And as you know, eeg spectrogram was more effective than competition spectrogram.</p>\n<p>So, i'm worrying about overfitting. but it's not sure haha </p>",
      "rawMarkdown": "First of all, Thanks for sharing! it's good approach\n\nAnd as you know, eeg spectrogram was more effective than competition spectrogram.\n\nSo, i'm worrying about overfitting. but it's not sure haha",
      "votes": null
    },
    {
      "id": "2666860",
      "postDate": "02/24/2024 17:19:22",
      "content": "<p>Thanks for sharing your knowledge PC Jimmy,<br>\nJust to be clear, the document contains what kind of information:</p>\n<p>-- The preprocessing used by the competition organizers?</p>\n<p>Or</p>\n<p>-- What preprocessing is recommended in this type of study?</p>\n<p>Or</p>\n<p>-- Is the minimum preprocessing recommended in this type of study?</p>",
      "rawMarkdown": "Thanks for sharing your knowledge PC Jimmy,\nJust to be clear, the document contains what kind of information:\n\n-- The preprocessing used by the competition organizers?\n\nOr\n\n  -- What preprocessing is recommended in this type of study?\n\nOr\n\n-- Is the minimum preprocessing recommended in this type of study?",
      "votes": null
    },
    {
      "id": "2666865",
      "postDate": "02/24/2024 17:24:52",
      "content": "<p>Indeed, overfitting is an interesting issue in this competition, given that the leaderboard only has 35% of the test data.</p>\n<p>One thing that many are betting on is the blend of various models, and various types of inputs….</p>\n<p>I don't know what to think about this, maybe it reduces the chance of overfitting, maybe not</p>",
      "rawMarkdown": "Indeed, overfitting is an interesting issue in this competition, given that the leaderboard only has 35% of the test data.\n\nOne thing that many are betting on is the blend of various models, and various types of inputs....\n\nI don't know what to think about this, maybe it reduces the chance of overfitting, maybe not",
      "votes": null
    },
    {
      "id": "2666911",
      "postDate": "02/24/2024 18:18:38",
      "content": "<p>Among other things I believe the document concurs with your idea - how the raw data is processed and presented will change the label.  When you watch video 3 you might arrive at the idea that the order of presentation is also important.</p>",
      "rawMarkdown": "Among other things I believe the document concurs with your idea - how the raw data is processed and presented will change the label.  When you watch video 3 you might arrive at the idea that the order of presentation is also important.",
      "votes": null
    },
    {
      "id": "2668676",
      "postDate": "02/25/2024 19:42:19",
      "content": "<p>I asked something like this of the Kaggle coordinator, Ashley Chow\".  She wrote back that I should post it in the discussions.  I just sent her a message pointing to your post.</p>",
      "rawMarkdown": "I asked something like this of the Kaggle coordinator, Ashley Chow\".  She wrote back that I should post it in the discussions.  I just sent her a message pointing to your post.",
      "votes": null
    },
    {
      "id": "2669610",
      "postDate": "02/26/2024 11:28:16",
      "content": "<p>Maybe the key missing part is a amplitude filter:</p>\n<blockquote>\n  <p>A healthy human EEG will show certain patterns of activity that correlate with how awake a person is. The range of frequencies one observes are between 1 and 30 Hz, <strong>and amplitudes will vary between 20 and 100 μV</strong></p>\n</blockquote>\n<p>Can you try this idea and publish a update in the topic ?</p>",
      "rawMarkdown": "Maybe the key missing part is a amplitude filter:\n\n>A healthy human EEG will show certain patterns of activity that correlate with how awake a person is. The range of frequencies one observes are between 1 and 30 Hz, **and amplitudes will vary between 20 and 100 μV**\n\nCan you try this idea and publish a update in the topic ?",
      "votes": null
    },
    {
      "id": "2669680",
      "postDate": "02/26/2024 12:22:14",
      "content": "<p>Thanks for the tip, Sergio! I will implement the amplitude filter and see the impact. Looking forward to sharing the findings with you all soon.</p>",
      "rawMarkdown": "Thanks for the tip, Sergio! I will implement the amplitude filter and see the impact. Looking forward to sharing the findings with you all soon.",
      "votes": null
    },
    {
      "id": "2669684",
      "postDate": "02/26/2024 12:23:26",
      "content": "<p>Thank you Abra Kadabra, let's wait for the answer</p>",
      "rawMarkdown": "Thank you Abra Kadabra, let's wait for the answer",
      "votes": null
    },
    {
      "id": "2681215",
      "postDate": "03/04/2024 14:45:52",
      "content": "<p>I guess no one is coming to save us :)</p>",
      "rawMarkdown": "I guess no one is coming to save us :)",
      "votes": null
    },
    {
      "id": "2685677",
      "postDate": "03/07/2024 10:55:06",
      "content": "<p>Have wondered about the differences in using mne for preprocessing and looking at some of that documentation and tutorials. <br>\nThings like artifact detection and correction, setting_eeg_reference - </p>\n<blockquote>\n  <p>Even in cases where no electrode is specifically designated as the reference, EEG recording hardware will still treat one of the scalp electrodes as the reference, and the recording software may or may not display it to you (it might appear as a completely flat channel, or the software might subtract out the average of all signals before displaying, making it look like there is no reference).</p>\n</blockquote>\n<p>A lot to consider but best results may depend on how good preprocessing is with the data provided.  And it could be that the same preprocessing for all does not yield the best results.  Not sure if patients are all from the same locations, using the same equipment? Probably unlikely.  Most of the information relates to the voting for the 6 patterns, unfortunately not much on the data.</p>\n<p>Have also wondered if different sensors ( or their combinations) are more relevant for the 5 out of 6 patterns here and if so how to address that.  Still researching and looking at papers…  </p>",
      "rawMarkdown": "Have wondered about the differences in using mne for preprocessing and looking at some of that documentation and tutorials. \nThings like artifact detection and correction, setting_eeg_reference - \n>Even in cases where no electrode is specifically designated as the reference, EEG recording hardware will still treat one of the scalp electrodes as the reference, and the recording software may or may not display it to you (it might appear as a completely flat channel, or the software might subtract out the average of all signals before displaying, making it look like there is no reference).\n\n\nA lot to consider but best results may depend on how good preprocessing is with the data provided.  And it could be that the same preprocessing for all does not yield the best results.  Not sure if patients are all from the same locations, using the same equipment? Probably unlikely.  Most of the information relates to the voting for the 6 patterns, unfortunately not much on the data.\n\nHave also wondered if different sensors ( or their combinations) are more relevant for the 5 out of 6 patterns here and if so how to address that.  Still researching and looking at papers...",
      "votes": null
    },
    {
      "id": "2686469",
      "postDate": "03/07/2024 20:35:03",
      "content": "<p>Very interesting questions, to which I also don't have the answers, but I have pondered some of them myself, such as why there are differences in results using MNE and the signal functions, for example…</p>\n<p>Considering your ideas, I thought that perhaps the preprocessing is not specified because it's likely that they are not all similar, as this would limit the study. That is, the data comes from various different sources, and therefore the preprocessing might be different, not to mention the individuality of each patient.</p>\n<p>The main point is that this is a research competition, and theoretically, to conduct proper research, we should have access to this kind of information, just as a researcher would have while working on such a problem. Unfortunately, it seems the organizers do not think the same way. :(</p>",
      "rawMarkdown": "Very interesting questions, to which I also don't have the answers, but I have pondered some of them myself, such as why there are differences in results using MNE and the signal functions, for example...\n\nConsidering your ideas, I thought that perhaps the preprocessing is not specified because it's likely that they are not all similar, as this would limit the study. That is, the data comes from various different sources, and therefore the preprocessing might be different, not to mention the individuality of each patient.\n\nThe main point is that this is a research competition, and theoretically, to conduct proper research, we should have access to this kind of information, just as a researcher would have while working on such a problem. Unfortunately, it seems the organizers do not think the same way. :(",
      "votes": null
    },
    {
      "id": "2686912",
      "postDate": "03/08/2024 06:18:54",
      "content": "<p>Would have thought the EEG data (not spectrograms) was raw data and to be preprocessed here in some way like standards noted in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479207#2665616\" target=\"_blank\">this post</a>.  And to some extent considering other papers and code would have thought it likely mne is used because others have done so.  Perhaps explains some of differences when creating own spectrograms and as you pointed out competition spectrograms could have been created by different sources and handled artifacts or not. </p>\n<p>As a research competition though and because datasets like this can scarce, it would seem beneficial to have consistency in preprocessing.  like e.g.,<br>\n<a href=\"https://github.com/braindecode/braindecode\" target=\"_blank\">Braindecode</a> </p>\n<blockquote>\n  <p>is an open-source Python toolbox for decoding raw electrophysiological brain data with deep learning models. It includes dataset fetchers, data preprocessing and visualization tools, as well as implementations of several deep learning architectures and data augmentations for analysis of EEG, ECoG and MEG.</p>\n</blockquote>\n<p><a href=\"https://github.com/NeuroTechX/moabb\" target=\"_blank\">Mother of all BCI Benchmarks</a></p>\n<blockquote>\n  <p>In this project, we focus mostly on electroencephalographic signals (EEG), that is a very active research domain, with worldwide scientific contributions. Still:</p>\n</blockquote>\n<ul>\n<li>Reproducible Research in BCI has a long way to go.</li>\n<li>While many BCI datasets are made freely available, researchers do not publish code, and reproducing results required to benchmark new algorithms turns out to be trickier than it should be.</li>\n<li>Performances can be significantly impacted by parameters of the preprocessing steps, toolboxes used and implementation “tricks” that are almost never reported in the literature.</li>\n</ul>\n<blockquote>\n  <p>As a result, there is no comprehensive benchmark of BCI algorithms, and newcomers are spending a tremendous amount of time browsing literature to find out what algorithm works best and on which dataset.</p>\n</blockquote>",
      "rawMarkdown": "Would have thought the EEG data (not spectrograms) was raw data and to be preprocessed here in some way like standards noted in [this post](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479207#2665616).  And to some extent considering other papers and code would have thought it likely mne is used because others have done so.  Perhaps explains some of differences when creating own spectrograms and as you pointed out competition spectrograms could have been created by different sources and handled artifacts or not. \n\nAs a research competition though and because datasets like this can scarce, it would seem beneficial to have consistency in preprocessing.  like e.g.,\n[Braindecode](https://github.com/braindecode/braindecode) \n>is an open-source Python toolbox for decoding raw electrophysiological brain data with deep learning models. It includes dataset fetchers, data preprocessing and visualization tools, as well as implementations of several deep learning architectures and data augmentations for analysis of EEG, ECoG and MEG.\n\n[Mother of all BCI Benchmarks](https://github.com/NeuroTechX/moabb)\n>In this project, we focus mostly on electroencephalographic signals (EEG), that is a very active research domain, with worldwide scientific contributions. Still:\n\n   -  Reproducible Research in BCI has a long way to go.\n   - While many BCI datasets are made freely available, researchers do not publish code, and reproducing results required to benchmark new algorithms turns out to be trickier than it should be.\n   - Performances can be significantly impacted by parameters of the preprocessing steps, toolboxes used and implementation “tricks” that are almost never reported in the literature.\n\n>As a result, there is no comprehensive benchmark of BCI algorithms, and newcomers are spending a tremendous amount of time browsing literature to find out what algorithm works best and on which dataset.",
      "votes": null
    },
    {
      "id": "2688840",
      "postDate": "03/09/2024 13:57:55",
      "content": "<p>Your comment really has a lot of value, it is interesting to realize that this type of process does not lead to reproduction, at the same time the studies cannot be tested by others…</p>\n<p>My doubt is, would this be an obscurantist way in which the academy is dealing with the BCI? Or is pre-processing so relevant that it is the secret sauce?</p>",
      "rawMarkdown": "Your comment really has a lot of value, it is interesting to realize that this type of process does not lead to reproduction, at the same time the studies cannot be tested by others...\n\nMy doubt is, would this be an obscurantist way in which the academy is dealing with the BCI? Or is pre-processing so relevant that it is the secret sauce?",
      "votes": null
    },
    {
      "id": "2702296",
      "postDate": "03/17/2024 14:33:30",
      "content": "<p>Hi everyone in this thread.    This is sort of related to preprocessing in terms of data labeling quality.  I can't reconcile what I'm seeing for labeled \"periodic discharches\", with the example or the data.  </p>\n<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/484433\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/484433</a></p>",
      "rawMarkdown": "Hi everyone in this thread.    This is sort of related to preprocessing in terms of data labeling quality.  I can't reconcile what I'm seeing for labeled \"periodic discharches\", with the example or the data.  \n\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/484433",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2665263,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/23/2024 15:03:24",
      "content": "<p>Oh. I thought the experts received the exact subsamples we extract from eegs and spectograms. Without data enginnering. Good to know and I definitely agree.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2665533,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "02/23/2024 17:49:33",
          "content": "<p>Hello, Ángel, how are you? Thanks for the comment.</p>\n<p>It seems that at least the eeg data receives a notch filter and a bandpass, because if you look at the eeg data in the example images of the competition, you can see that there appears to be a filter at least on the higher frequency data and at 60Hz, This seems to be confirmed when comparing the data from the spectogram provided by the competition and the spectogram created from EEGs, apparently there seems to be more pre-processing, in the case of null values, bad channels, artifacts and EKG.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2665631,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "02/23/2024 18:55:13",
              "content": "<p>\"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"<br>\nAre you sure that examples are not just that, examples to give us ideas? From data description I understand experts reviewed data as we extract it from eegs and spectogram folders.<br>\nExamples eeg covers only the 10 middle labeled seconds.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2665712,
                  "author_name": "rafaelzimmermann1",
                  "author_url": "",
                  "post_date": "02/23/2024 19:58:34",
                  "content": "<p>Good question, unfortunately I don't have the answer, I believe the organizers can answer this…</p>\n<p>My main doubt is the preprocessing,<br>\n1- Whether standard preprocessing was done to show the data to experts or not?<br>\n2- Each expert performed their preprocessing and then labeled the 10 seconds.<br>\n3- No pre-processing was done, the data was shown in tools like eeglab and each expert used their preferred filter parameters etc in the tool?</p>\n<p>I believe these are important questions, but I'm not sure of any answers, one of the main reasons for the discussion is to see if anyone knows the answer :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2665616,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/23/2024 18:44:20",
      "content": "<p><a href=\"https://journals.lww.com/clinicalneurophys/Fulltext/2016/08000/American_Clinical_Neurophysiology_Society.3.aspx\" target=\"_blank\">https://journals.lww.com/clinicalneurophys/Fulltext/2016/08000/American_Clinical_Neurophysiology_Society.3.aspx</a></p>\n<p>The document describes the 'standards' that should have been followed for our data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2666860,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "02/24/2024 17:19:22",
          "content": "<p>Thanks for sharing your knowledge PC Jimmy,<br>\nJust to be clear, the document contains what kind of information:</p>\n<p>-- The preprocessing used by the competition organizers?</p>\n<p>Or</p>\n<p>-- What preprocessing is recommended in this type of study?</p>\n<p>Or</p>\n<p>-- Is the minimum preprocessing recommended in this type of study?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2666911,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "02/24/2024 18:18:38",
              "content": "<p>Among other things I believe the document concurs with your idea - how the raw data is processed and presented will change the label.  When you watch video 3 you might arrive at the idea that the order of presentation is also important.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2665686,
      "author_name": "m000sey",
      "author_url": "",
      "post_date": "02/23/2024 19:40:49",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>. This is an awesome way to visualize and inspect the data</p>",
      "votes": null,
      "replies": [
        {
          "id": 2665693,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "02/23/2024 19:47:17",
          "content": "<p>Tks Shane </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2666093,
      "author_name": "seoyunje",
      "author_url": "",
      "post_date": "02/24/2024 06:30:37",
      "content": "<p>First of all, Thanks for sharing! it's good approach</p>\n<p>And as you know, eeg spectrogram was more effective than competition spectrogram.</p>\n<p>So, i'm worrying about overfitting. but it's not sure haha </p>",
      "votes": null,
      "replies": [
        {
          "id": 2666865,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "02/24/2024 17:24:52",
          "content": "<p>Indeed, overfitting is an interesting issue in this competition, given that the leaderboard only has 35% of the test data.</p>\n<p>One thing that many are betting on is the blend of various models, and various types of inputs….</p>\n<p>I don't know what to think about this, maybe it reduces the chance of overfitting, maybe not</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2668676,
      "author_name": "idithhaber",
      "author_url": "",
      "post_date": "02/25/2024 19:42:19",
      "content": "<p>I asked something like this of the Kaggle coordinator, Ashley Chow\".  She wrote back that I should post it in the discussions.  I just sent her a message pointing to your post.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2669684,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "02/26/2024 12:23:26",
          "content": "<p>Thank you Abra Kadabra, let's wait for the answer</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2669610,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "02/26/2024 11:28:16",
      "content": "<p>Maybe the key missing part is a amplitude filter:</p>\n<blockquote>\n  <p>A healthy human EEG will show certain patterns of activity that correlate with how awake a person is. The range of frequencies one observes are between 1 and 30 Hz, <strong>and amplitudes will vary between 20 and 100 μV</strong></p>\n</blockquote>\n<p>Can you try this idea and publish a update in the topic ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2669680,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "02/26/2024 12:22:14",
          "content": "<p>Thanks for the tip, Sergio! I will implement the amplitude filter and see the impact. Looking forward to sharing the findings with you all soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2681215,
      "author_name": "idithhaber",
      "author_url": "",
      "post_date": "03/04/2024 14:45:52",
      "content": "<p>I guess no one is coming to save us :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2685677,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/07/2024 10:55:06",
      "content": "<p>Have wondered about the differences in using mne for preprocessing and looking at some of that documentation and tutorials. <br>\nThings like artifact detection and correction, setting_eeg_reference - </p>\n<blockquote>\n  <p>Even in cases where no electrode is specifically designated as the reference, EEG recording hardware will still treat one of the scalp electrodes as the reference, and the recording software may or may not display it to you (it might appear as a completely flat channel, or the software might subtract out the average of all signals before displaying, making it look like there is no reference).</p>\n</blockquote>\n<p>A lot to consider but best results may depend on how good preprocessing is with the data provided.  And it could be that the same preprocessing for all does not yield the best results.  Not sure if patients are all from the same locations, using the same equipment? Probably unlikely.  Most of the information relates to the voting for the 6 patterns, unfortunately not much on the data.</p>\n<p>Have also wondered if different sensors ( or their combinations) are more relevant for the 5 out of 6 patterns here and if so how to address that.  Still researching and looking at papers…  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2686469,
          "author_name": "rafaelzimmermann1",
          "author_url": "",
          "post_date": "03/07/2024 20:35:03",
          "content": "<p>Very interesting questions, to which I also don't have the answers, but I have pondered some of them myself, such as why there are differences in results using MNE and the signal functions, for example…</p>\n<p>Considering your ideas, I thought that perhaps the preprocessing is not specified because it's likely that they are not all similar, as this would limit the study. That is, the data comes from various different sources, and therefore the preprocessing might be different, not to mention the individuality of each patient.</p>\n<p>The main point is that this is a research competition, and theoretically, to conduct proper research, we should have access to this kind of information, just as a researcher would have while working on such a problem. Unfortunately, it seems the organizers do not think the same way. :(</p>",
          "votes": null,
          "replies": [
            {
              "id": 2686912,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "03/08/2024 06:18:54",
              "content": "<p>Would have thought the EEG data (not spectrograms) was raw data and to be preprocessed here in some way like standards noted in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479207#2665616\" target=\"_blank\">this post</a>.  And to some extent considering other papers and code would have thought it likely mne is used because others have done so.  Perhaps explains some of differences when creating own spectrograms and as you pointed out competition spectrograms could have been created by different sources and handled artifacts or not. </p>\n<p>As a research competition though and because datasets like this can scarce, it would seem beneficial to have consistency in preprocessing.  like e.g.,<br>\n<a href=\"https://github.com/braindecode/braindecode\" target=\"_blank\">Braindecode</a> </p>\n<blockquote>\n  <p>is an open-source Python toolbox for decoding raw electrophysiological brain data with deep learning models. It includes dataset fetchers, data preprocessing and visualization tools, as well as implementations of several deep learning architectures and data augmentations for analysis of EEG, ECoG and MEG.</p>\n</blockquote>\n<p><a href=\"https://github.com/NeuroTechX/moabb\" target=\"_blank\">Mother of all BCI Benchmarks</a></p>\n<blockquote>\n  <p>In this project, we focus mostly on electroencephalographic signals (EEG), that is a very active research domain, with worldwide scientific contributions. Still:</p>\n</blockquote>\n<ul>\n<li>Reproducible Research in BCI has a long way to go.</li>\n<li>While many BCI datasets are made freely available, researchers do not publish code, and reproducing results required to benchmark new algorithms turns out to be trickier than it should be.</li>\n<li>Performances can be significantly impacted by parameters of the preprocessing steps, toolboxes used and implementation “tricks” that are almost never reported in the literature.</li>\n</ul>\n<blockquote>\n  <p>As a result, there is no comprehensive benchmark of BCI algorithms, and newcomers are spending a tremendous amount of time browsing literature to find out what algorithm works best and on which dataset.</p>\n</blockquote>",
              "votes": null,
              "replies": [
                {
                  "id": 2688840,
                  "author_name": "rafaelzimmermann1",
                  "author_url": "",
                  "post_date": "03/09/2024 13:57:55",
                  "content": "<p>Your comment really has a lot of value, it is interesting to realize that this type of process does not lead to reproduction, at the same time the studies cannot be tested by others…</p>\n<p>My doubt is, would this be an obscurantist way in which the academy is dealing with the BCI? Or is pre-processing so relevant that it is the secret sauce?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2702296,
      "author_name": "idithhaber",
      "author_url": "",
      "post_date": "03/17/2024 14:33:30",
      "content": "<p>Hi everyone in this thread.    This is sort of related to preprocessing in terms of data labeling quality.  I can't reconcile what I'm seeing for labeled \"periodic discharches\", with the example or the data.  </p>\n<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/484433\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/484433</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2665225": "For the past few weeks, I have been working on deciphering which type of EEG data preprocessing is most aligned with this competition. After much contemplation and in-depth data analysis, I've reached an insight.\n\nThe data were labeling by experts in specific formats. The EEG data were presented to these experts with a certain type of preprocessing, and the spectrograms were also shown with a type of preprocessing. From this, I concluded that the competition's classifications might be biased towards the preprocessing shown to the experts for making the classifications.\n\nNote: I am not an expert in the field, and this is just an insight. I believe I am correct, but I do not claim to hold the ultimate truth. If anyone has anything to say on this matter, your input is more than welcome.\n\nBelow is an example image from the competition showing how this classification was made from the experts' perspective.\n\n![](https://i.imgur.com/hq2il2c.png)\n\nViewing it from this perspective led me to the following conclusion, which may be right or wrong:\n\nThe closer my preprocessing is to the preprocessing used by the experts to generate the classification, the closer my algorithm will come to achieving results similar to theirs.\nWith this in mind, I created a notebook where:\n\n1. I try to create spectrograms for the competition as similar as possible to those seen by the experts.\n1. I attempt to create new EEG spectrograms that are as close as possible to the competition's spectrograms.\n1. I compare these spectrograms and try to identify if the preprocessing I used on the EEGs generated spectrograms similar to those in the competition, thus attempting to discover how similar my preprocessing is to that used in the competition.\n\nI know it's a time-consuming, subjective task. If anyone has better ideas, I would love to hear them.\n\nHere is the link to this notebook:\nhttps://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\n\nJust to give an example of what I'm talking about, here is an image that turned out quite similar to the competition's preprocessing:\n\n![](https://i.imgur.com/9jL3bGi.png)\n\nThe preprocessing used in this case was a 60Hz notch filter and a 5th-order bandpass butter filter - low_cut 0.7, high_cut 20.\n\nTo illustrate, here is an example where the spectrogram did not turn out similar. I believe something is missing; in this case, it seems there might be some treatment of a bad channel here that I couldn't figure out how to perform.\n\n![](https://i.imgur.com/ub0E9mH.png)",
    "2665263": "Oh. I thought the experts received the exact subsamples we extract from eegs and spectograms. Without data enginnering. Good to know and I definitely agree.",
    "2665533": "Hello, Ángel, how are you? Thanks for the comment.\n\nIt seems that at least the eeg data receives a notch filter and a bandpass, because if you look at the eeg data in the example images of the competition, you can see that there appears to be a filter at least on the higher frequency data and at 60Hz, This seems to be confirmed when comparing the data from the spectogram provided by the competition and the spectogram created from EEGs, apparently there seems to be more pre-processing, in the case of null values, bad channels, artifacts and EKG.",
    "2665616": "https://journals.lww.com/clinicalneurophys/Fulltext/2016/08000/American_Clinical_Neurophysiology_Society.3.aspx\n\nThe document describes the 'standards' that should have been followed for our data.",
    "2665631": "\"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"\nAre you sure that examples are not just that, examples to give us ideas? From data description I understand experts reviewed data as we extract it from eegs and spectogram folders.\nExamples eeg covers only the 10 middle labeled seconds.",
    "2665686": "Great work @rafaelzimmermann1. This is an awesome way to visualize and inspect the data",
    "2665693": "Tks Shane",
    "2665712": "Good question, unfortunately I don't have the answer, I believe the organizers can answer this...\n\nMy main doubt is the preprocessing,\n1- Whether standard preprocessing was done to show the data to experts or not?\n2- Each expert performed their preprocessing and then labeled the 10 seconds.\n3- No pre-processing was done, the data was shown in tools like eeglab and each expert used their preferred filter parameters etc in the tool?\n\nI believe these are important questions, but I'm not sure of any answers, one of the main reasons for the discussion is to see if anyone knows the answer :)",
    "2666093": "First of all, Thanks for sharing! it's good approach\n\nAnd as you know, eeg spectrogram was more effective than competition spectrogram.\n\nSo, i'm worrying about overfitting. but it's not sure haha",
    "2666860": "Thanks for sharing your knowledge PC Jimmy,\nJust to be clear, the document contains what kind of information:\n\n-- The preprocessing used by the competition organizers?\n\nOr\n\n  -- What preprocessing is recommended in this type of study?\n\nOr\n\n-- Is the minimum preprocessing recommended in this type of study?",
    "2666865": "Indeed, overfitting is an interesting issue in this competition, given that the leaderboard only has 35% of the test data.\n\nOne thing that many are betting on is the blend of various models, and various types of inputs....\n\nI don't know what to think about this, maybe it reduces the chance of overfitting, maybe not",
    "2666911": "Among other things I believe the document concurs with your idea - how the raw data is processed and presented will change the label.  When you watch video 3 you might arrive at the idea that the order of presentation is also important.",
    "2668676": "I asked something like this of the Kaggle coordinator, Ashley Chow\".  She wrote back that I should post it in the discussions.  I just sent her a message pointing to your post.",
    "2669610": "Maybe the key missing part is a amplitude filter:\n\n>A healthy human EEG will show certain patterns of activity that correlate with how awake a person is. The range of frequencies one observes are between 1 and 30 Hz, **and amplitudes will vary between 20 and 100 μV**\n\nCan you try this idea and publish a update in the topic ?",
    "2669680": "Thanks for the tip, Sergio! I will implement the amplitude filter and see the impact. Looking forward to sharing the findings with you all soon.",
    "2669684": "Thank you Abra Kadabra, let's wait for the answer",
    "2681215": "I guess no one is coming to save us :)",
    "2685677": "Have wondered about the differences in using mne for preprocessing and looking at some of that documentation and tutorials. \nThings like artifact detection and correction, setting_eeg_reference - \n>Even in cases where no electrode is specifically designated as the reference, EEG recording hardware will still treat one of the scalp electrodes as the reference, and the recording software may or may not display it to you (it might appear as a completely flat channel, or the software might subtract out the average of all signals before displaying, making it look like there is no reference).\n\n\nA lot to consider but best results may depend on how good preprocessing is with the data provided.  And it could be that the same preprocessing for all does not yield the best results.  Not sure if patients are all from the same locations, using the same equipment? Probably unlikely.  Most of the information relates to the voting for the 6 patterns, unfortunately not much on the data.\n\nHave also wondered if different sensors ( or their combinations) are more relevant for the 5 out of 6 patterns here and if so how to address that.  Still researching and looking at papers...",
    "2686469": "Very interesting questions, to which I also don't have the answers, but I have pondered some of them myself, such as why there are differences in results using MNE and the signal functions, for example...\n\nConsidering your ideas, I thought that perhaps the preprocessing is not specified because it's likely that they are not all similar, as this would limit the study. That is, the data comes from various different sources, and therefore the preprocessing might be different, not to mention the individuality of each patient.\n\nThe main point is that this is a research competition, and theoretically, to conduct proper research, we should have access to this kind of information, just as a researcher would have while working on such a problem. Unfortunately, it seems the organizers do not think the same way. :(",
    "2686912": "Would have thought the EEG data (not spectrograms) was raw data and to be preprocessed here in some way like standards noted in [this post](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479207#2665616).  And to some extent considering other papers and code would have thought it likely mne is used because others have done so.  Perhaps explains some of differences when creating own spectrograms and as you pointed out competition spectrograms could have been created by different sources and handled artifacts or not. \n\nAs a research competition though and because datasets like this can scarce, it would seem beneficial to have consistency in preprocessing.  like e.g.,\n[Braindecode](https://github.com/braindecode/braindecode) \n>is an open-source Python toolbox for decoding raw electrophysiological brain data with deep learning models. It includes dataset fetchers, data preprocessing and visualization tools, as well as implementations of several deep learning architectures and data augmentations for analysis of EEG, ECoG and MEG.\n\n[Mother of all BCI Benchmarks](https://github.com/NeuroTechX/moabb)\n>In this project, we focus mostly on electroencephalographic signals (EEG), that is a very active research domain, with worldwide scientific contributions. Still:\n\n   -  Reproducible Research in BCI has a long way to go.\n   - While many BCI datasets are made freely available, researchers do not publish code, and reproducing results required to benchmark new algorithms turns out to be trickier than it should be.\n   - Performances can be significantly impacted by parameters of the preprocessing steps, toolboxes used and implementation “tricks” that are almost never reported in the literature.\n\n>As a result, there is no comprehensive benchmark of BCI algorithms, and newcomers are spending a tremendous amount of time browsing literature to find out what algorithm works best and on which dataset.",
    "2688840": "Your comment really has a lot of value, it is interesting to realize that this type of process does not lead to reproduction, at the same time the studies cannot be tested by others...\n\nMy doubt is, would this be an obscurantist way in which the academy is dealing with the BCI? Or is pre-processing so relevant that it is the secret sauce?",
    "2702296": "Hi everyone in this thread.    This is sort of related to preprocessing in terms of data labeling quality.  I can't reconcile what I'm seeing for labeled \"periodic discharches\", with the example or the data.  \n\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/484433"
  },
  "source": "meta"
}