{
  "id": 476007,
  "title": "Insight on relation of 10-minute spectrograms and 50-second raw EEG to annotations",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/476007",
  "author_name": "",
  "post_date": "2024-02-10T18:53:04.121593900Z",
  "votes": 21,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Recently, I was reviewing the competition overview and data description to reconsider how the data is labeled. The competition training set provides EEG spectrograms that cover a 10-minute window. Based on the performance of public notebooks, it would seem the 10-minute spectrogram is more predictive than the 50-second raw EEG signals. We know the timeframes overlap: \"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\" </p>\n<p>There is an insight to be gained about what the experts are looking at when annotating: In the overview, we can compare the Idealized LPD and Proto LPD. Without medical expertise, we can still reason the pattern on the right-side of the Proto LPD spectrogram might be associated with the overall pattern observed in the Idealized LPD spectrogram. We can also observe a difference in the 50-second EEG signal when comparing Idealized to Proto. This makes sense when we consider the 50-second signal is centered on the 10-minute spectrogram. It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below).</p>\n<p><strong>Idealized LPD</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fffdf457ed15d88fa374ec3bcea89b310%2FScreenshot%202024-02-10%20at%2010.09.06AM.png?generation=1707588640619705&amp;alt=media\"></p>\n<p><strong>Proto LPD</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F94128f76573015c0b7dfce57de5ebf22%2FScreenshot%202024-02-10%20at%2010.08.27AM.png?generation=1707589533787991&amp;alt=media\"></p>\n<p>The examples in the overview indicate that the 50-second EEG signal can vary in importance. My insight: a model that can weight the importance of the 50-second EEG signal relative to the 10-minute spectrogram will presumably predict the annotations of experts more accurately. While this is speculation, I believe it is import to consider the human aspect of manually annotating brain activity, as the goal is to predict the votes of experts rather than determine the true nature of the brain activity.</p>\n<p>What are your thoughts?</p>",
  "messages": [
    {
      "id": "2646280",
      "postDate": "02/10/2024 18:53:04",
      "content": "<p>Recently, I was reviewing the competition overview and data description to reconsider how the data is labeled. The competition training set provides EEG spectrograms that cover a 10-minute window. Based on the performance of public notebooks, it would seem the 10-minute spectrogram is more predictive than the 50-second raw EEG signals. We know the timeframes overlap: \"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\" </p>\n<p>There is an insight to be gained about what the experts are looking at when annotating: In the overview, we can compare the Idealized LPD and Proto LPD. Without medical expertise, we can still reason the pattern on the right-side of the Proto LPD spectrogram might be associated with the overall pattern observed in the Idealized LPD spectrogram. We can also observe a difference in the 50-second EEG signal when comparing Idealized to Proto. This makes sense when we consider the 50-second signal is centered on the 10-minute spectrogram. It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below).</p>\n<p><strong>Idealized LPD</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fffdf457ed15d88fa374ec3bcea89b310%2FScreenshot%202024-02-10%20at%2010.09.06AM.png?generation=1707588640619705&amp;alt=media\"></p>\n<p><strong>Proto LPD</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F94128f76573015c0b7dfce57de5ebf22%2FScreenshot%202024-02-10%20at%2010.08.27AM.png?generation=1707589533787991&amp;alt=media\"></p>\n<p>The examples in the overview indicate that the 50-second EEG signal can vary in importance. My insight: a model that can weight the importance of the 50-second EEG signal relative to the 10-minute spectrogram will presumably predict the annotations of experts more accurately. While this is speculation, I believe it is import to consider the human aspect of manually annotating brain activity, as the goal is to predict the votes of experts rather than determine the true nature of the brain activity.</p>\n<p>What are your thoughts?</p>",
      "rawMarkdown": "Recently, I was reviewing the competition overview and data description to reconsider how the data is labeled. The competition training set provides EEG spectrograms that cover a 10-minute window. Based on the performance of public notebooks, it would seem the 10-minute spectrogram is more predictive than the 50-second raw EEG signals. We know the timeframes overlap: \"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\" \n\nThere is an insight to be gained about what the experts are looking at when annotating: In the overview, we can compare the Idealized LPD and Proto LPD. Without medical expertise, we can still reason the pattern on the right-side of the Proto LPD spectrogram might be associated with the overall pattern observed in the Idealized LPD spectrogram. We can also observe a difference in the 50-second EEG signal when comparing Idealized to Proto. This makes sense when we consider the 50-second signal is centered on the 10-minute spectrogram. It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below).\n\n**Idealized LPD**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fffdf457ed15d88fa374ec3bcea89b310%2FScreenshot%202024-02-10%20at%2010.09.06AM.png?generation=1707588640619705&alt=media)\n\n**Proto LPD**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F94128f76573015c0b7dfce57de5ebf22%2FScreenshot%202024-02-10%20at%2010.08.27AM.png?generation=1707589533787991&alt=media)\n\nThe examples in the overview indicate that the 50-second EEG signal can vary in importance. My insight: a model that can weight the importance of the 50-second EEG signal relative to the 10-minute spectrogram will presumably predict the annotations of experts more accurately. While this is speculation, I believe it is import to consider the human aspect of manually annotating brain activity, as the goal is to predict the votes of experts rather than determine the true nature of the brain activity.\n\nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "2646319",
      "postDate": "02/10/2024 19:43:51",
      "content": "<p>This seems to be totally right. From a certain perspective, we should learn the model to have a human-like bias :)</p>",
      "rawMarkdown": "This seems to be totally right. From a certain perspective, we should learn the model to have a human-like bias :)",
      "votes": null
    },
    {
      "id": "2646537",
      "postDate": "02/11/2024 03:19:15",
      "content": "<p>I have participated in many of the kaggle completions involving medical methods.</p>\n<p>It seems to me a universal rule that we should never seek to model the truth in kaggle medical competitions, but rather the \"expert truth\".  </p>\n<p>In every measurement system there is error and bias.   Its very scary to see in competitions like this one that the medical field has more than their share of measurement issues.</p>",
      "rawMarkdown": "I have participated in many of the kaggle completions involving medical methods.\n\nIt seems to me a universal rule that we should never seek to model the truth in kaggle medical competitions, but rather the \"expert truth\".  \n\nIn every measurement system there is error and bias.   Its very scary to see in competitions like this one that the medical field has more than their share of measurement issues.",
      "votes": null
    },
    {
      "id": "2648593",
      "postDate": "02/12/2024 10:08:50",
      "content": "<blockquote>\n  <p>It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below)</p>\n</blockquote>\n<p>There is no 50-second signal in the example_figures but just 10sec raw EEG data, and 10min spectrogram.</p>\n<p>Also, 10sec raw EEG appears to show sufficient evidence for both of the diagnosis. As the overview page explains for this proto case: <em>'(figure) shows frontal lateralized sharp transients at ~1Hz, but they have a reversed polarity, suggesting they may be coming from a non-cerebral source, thus the split between LPD and “Other” (artifact) makes sense</em>.'<br>\nI think the reverse polarity issue is about the signals at channels Fp1-F3 and F3-C3, which, if only one existed, would be a confident LPD - I guess :-)</p>",
      "rawMarkdown": "> It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below)\n\nThere is no 50-second signal in the example_figures but just 10sec raw EEG data, and 10min spectrogram.\n\nAlso, 10sec raw EEG appears to show sufficient evidence for both of the diagnosis. As the overview page explains for this proto case: *'(figure) shows frontal lateralized sharp transients at ~1Hz, but they have a reversed polarity, suggesting they may be coming from a non-cerebral source, thus the split between LPD and “Other” (artifact) makes sense*.'\nI think the reverse polarity issue is about the signals at channels Fp1-F3 and F3-C3, which, if only one existed, would be a confident LPD - I guess :-)",
      "votes": null
    },
    {
      "id": "2649037",
      "postDate": "02/12/2024 15:22:43",
      "content": "<p>Thank you for the correction! I've noticed that some high scoring notebooks are using 256 of 300 time slices from the spectrograms. I'm curious if \"other\" votes are influenced by other observations in the spectrogram, like the right-side of the spectrogram for the proto LPD example above. Perhaps the \"other\" votes were influenced by the polarity reversal and other features in the spectrogram. If so, only using 256 time slices might drop valuable information.</p>",
      "rawMarkdown": "Thank you for the correction! I've noticed that some high scoring notebooks are using 256 of 300 time slices from the spectrograms. I'm curious if \"other\" votes are influenced by other observations in the spectrogram, like the right-side of the spectrogram for the proto LPD example above. Perhaps the \"other\" votes were influenced by the polarity reversal and other features in the spectrogram. If so, only using 256 time slices might drop valuable information.",
      "votes": null
    },
    {
      "id": "2650994",
      "postDate": "02/13/2024 19:47:43",
      "content": "<p>This! I wish we had information on how each expert voted in each case. We could be better at predicting some experts findings than others</p>",
      "rawMarkdown": "This! I wish we had information on how each expert voted in each case. We could be better at predicting some experts findings than others",
      "votes": null
    },
    {
      "id": "2651000",
      "postDate": "02/13/2024 20:01:21",
      "content": "<p>Agreed <a href=\"https://www.kaggle.com/hedberg503\" target=\"_blank\">@hedberg503</a>! Check out <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>'s EDA <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>! He has provided some valuable insights into the training data.</p>",
      "rawMarkdown": "Agreed @hedberg503! Check out @pcjimmmy's EDA [notebook](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda)! He has provided some valuable insights into the training data.",
      "votes": null
    },
    {
      "id": "2680900",
      "postDate": "03/04/2024 11:34:19",
      "content": "<p>Hello, I am new to this competition and have been reviewing the data description. I'm having difficulty understanding the statement that mentions, \"The expert annotators reviewed 50-second long EEG samples plus matched spectrograms covering a 10-minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.\" Could someone kindly explain this to me? I would greatly appreciate your assistance. Thank you.</p>",
      "rawMarkdown": "Hello, I am new to this competition and have been reviewing the data description. I'm having difficulty understanding the statement that mentions, \"The expert annotators reviewed 50-second long EEG samples plus matched spectrograms covering a 10-minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.\" Could someone kindly explain this to me? I would greatly appreciate your assistance. Thank you.",
      "votes": null
    },
    {
      "id": "2681035",
      "postDate": "03/04/2024 12:59:51",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010</a></p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has been super active in this competition and has published many helpful discussions and starter notebooks. Perhaps reading the above post by him will answer your question.</p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\n\n@cdeotte has been super active in this competition and has published many helpful discussions and starter notebooks. Perhaps reading the above post by him will answer your question.",
      "votes": null
    },
    {
      "id": "2682341",
      "postDate": "03/05/2024 08:59:03",
      "content": "<p>I appreciate your assistance</p>",
      "rawMarkdown": "I appreciate your assistance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2646319,
      "author_name": "victorshlepov",
      "author_url": "",
      "post_date": "02/10/2024 19:43:51",
      "content": "<p>This seems to be totally right. From a certain perspective, we should learn the model to have a human-like bias :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2646537,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/11/2024 03:19:15",
      "content": "<p>I have participated in many of the kaggle completions involving medical methods.</p>\n<p>It seems to me a universal rule that we should never seek to model the truth in kaggle medical competitions, but rather the \"expert truth\".  </p>\n<p>In every measurement system there is error and bias.   Its very scary to see in competitions like this one that the medical field has more than their share of measurement issues.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2648593,
      "author_name": "abdulkadirguner",
      "author_url": "",
      "post_date": "02/12/2024 10:08:50",
      "content": "<blockquote>\n  <p>It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below)</p>\n</blockquote>\n<p>There is no 50-second signal in the example_figures but just 10sec raw EEG data, and 10min spectrogram.</p>\n<p>Also, 10sec raw EEG appears to show sufficient evidence for both of the diagnosis. As the overview page explains for this proto case: <em>'(figure) shows frontal lateralized sharp transients at ~1Hz, but they have a reversed polarity, suggesting they may be coming from a non-cerebral source, thus the split between LPD and “Other” (artifact) makes sense</em>.'<br>\nI think the reverse polarity issue is about the signals at channels Fp1-F3 and F3-C3, which, if only one existed, would be a confident LPD - I guess :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2649037,
          "author_name": "seanbearden",
          "author_url": "",
          "post_date": "02/12/2024 15:22:43",
          "content": "<p>Thank you for the correction! I've noticed that some high scoring notebooks are using 256 of 300 time slices from the spectrograms. I'm curious if \"other\" votes are influenced by other observations in the spectrogram, like the right-side of the spectrogram for the proto LPD example above. Perhaps the \"other\" votes were influenced by the polarity reversal and other features in the spectrogram. If so, only using 256 time slices might drop valuable information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2650994,
      "author_name": "hedberg503",
      "author_url": "",
      "post_date": "02/13/2024 19:47:43",
      "content": "<p>This! I wish we had information on how each expert voted in each case. We could be better at predicting some experts findings than others</p>",
      "votes": null,
      "replies": [
        {
          "id": 2651000,
          "author_name": "seanbearden",
          "author_url": "",
          "post_date": "02/13/2024 20:01:21",
          "content": "<p>Agreed <a href=\"https://www.kaggle.com/hedberg503\" target=\"_blank\">@hedberg503</a>! Check out <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>'s EDA <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>! He has provided some valuable insights into the training data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2680900,
      "author_name": "farwa99",
      "author_url": "",
      "post_date": "03/04/2024 11:34:19",
      "content": "<p>Hello, I am new to this competition and have been reviewing the data description. I'm having difficulty understanding the statement that mentions, \"The expert annotators reviewed 50-second long EEG samples plus matched spectrograms covering a 10-minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.\" Could someone kindly explain this to me? I would greatly appreciate your assistance. Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2681035,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "03/04/2024 12:59:51",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010</a></p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has been super active in this competition and has published many helpful discussions and starter notebooks. Perhaps reading the above post by him will answer your question.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2682341,
              "author_name": "farwa99",
              "author_url": "",
              "post_date": "03/05/2024 08:59:03",
              "content": "<p>I appreciate your assistance</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2646280": "Recently, I was reviewing the competition overview and data description to reconsider how the data is labeled. The competition training set provides EEG spectrograms that cover a 10-minute window. Based on the performance of public notebooks, it would seem the 10-minute spectrogram is more predictive than the 50-second raw EEG signals. We know the timeframes overlap: \"The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\" \n\nThere is an insight to be gained about what the experts are looking at when annotating: In the overview, we can compare the Idealized LPD and Proto LPD. Without medical expertise, we can still reason the pattern on the right-side of the Proto LPD spectrogram might be associated with the overall pattern observed in the Idealized LPD spectrogram. We can also observe a difference in the 50-second EEG signal when comparing Idealized to Proto. This makes sense when we consider the 50-second signal is centered on the 10-minute spectrogram. It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below).\n\n**Idealized LPD**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fffdf457ed15d88fa374ec3bcea89b310%2FScreenshot%202024-02-10%20at%2010.09.06AM.png?generation=1707588640619705&alt=media)\n\n**Proto LPD**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F94128f76573015c0b7dfce57de5ebf22%2FScreenshot%202024-02-10%20at%2010.08.27AM.png?generation=1707589533787991&alt=media)\n\nThe examples in the overview indicate that the 50-second EEG signal can vary in importance. My insight: a model that can weight the importance of the 50-second EEG signal relative to the 10-minute spectrogram will presumably predict the annotations of experts more accurately. While this is speculation, I believe it is import to consider the human aspect of manually annotating brain activity, as the goal is to predict the votes of experts rather than determine the true nature of the brain activity.\n\nWhat are your thoughts?",
    "2646319": "This seems to be totally right. From a certain perspective, we should learn the model to have a human-like bias :)",
    "2646537": "I have participated in many of the kaggle completions involving medical methods.\n\nIt seems to me a universal rule that we should never seek to model the truth in kaggle medical competitions, but rather the \"expert truth\".  \n\nIn every measurement system there is error and bias.   Its very scary to see in competitions like this one that the medical field has more than their share of measurement issues.",
    "2648593": "> It is reasonable to assume that the votes for LPD in the Proto case were more heavily influenced by the ending of the 10-minute spectrogram than the 50-second EEG signal (see below)\n\nThere is no 50-second signal in the example_figures but just 10sec raw EEG data, and 10min spectrogram.\n\nAlso, 10sec raw EEG appears to show sufficient evidence for both of the diagnosis. As the overview page explains for this proto case: *'(figure) shows frontal lateralized sharp transients at ~1Hz, but they have a reversed polarity, suggesting they may be coming from a non-cerebral source, thus the split between LPD and “Other” (artifact) makes sense*.'\nI think the reverse polarity issue is about the signals at channels Fp1-F3 and F3-C3, which, if only one existed, would be a confident LPD - I guess :-)",
    "2649037": "Thank you for the correction! I've noticed that some high scoring notebooks are using 256 of 300 time slices from the spectrograms. I'm curious if \"other\" votes are influenced by other observations in the spectrogram, like the right-side of the spectrogram for the proto LPD example above. Perhaps the \"other\" votes were influenced by the polarity reversal and other features in the spectrogram. If so, only using 256 time slices might drop valuable information.",
    "2650994": "This! I wish we had information on how each expert voted in each case. We could be better at predicting some experts findings than others",
    "2651000": "Agreed @hedberg503! Check out @pcjimmmy's EDA [notebook](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda)! He has provided some valuable insights into the training data.",
    "2680900": "Hello, I am new to this competition and have been reviewing the data description. I'm having difficulty understanding the statement that mentions, \"The expert annotators reviewed 50-second long EEG samples plus matched spectrograms covering a 10-minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.\" Could someone kindly explain this to me? I would greatly appreciate your assistance. Thank you.",
    "2681035": "https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\n\n@cdeotte has been super active in this competition and has published many helpful discussions and starter notebooks. Perhaps reading the above post by him will answer your question.",
    "2682341": "I appreciate your assistance"
  },
  "source": "meta"
}