{
  "id": 478233,
  "title": "Missing Data in HMS-HBA Competition Spectrograms",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/478233",
  "author_name": "",
  "post_date": "2024-02-19T19:17:22.138504300Z",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As you may already know, many of the competition spectrograms have missing data, with some missing more than 50% of data. (Below is an example with more than a third of data missing from the LL region.) When training with unique spectrogram_ids, we sometimes have an option for an offset for the spectrogram. I decided to determine the missing spectrogram data for all 106,800 rows of <code>train.csv</code>, so the offset with the least missing data could be selected. The altered <code>train.csv</code> is output <a href=\"https://www.kaggle.com/seanbearden/missing-data-in-spectrograms\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F50aca4f227580ea30e7d18afa65875aa%2Fmissing_data.png?generation=1708368587898646&amp;alt=media\"></p>\n<p>The results were surprising and might affect training with competition-provided spectrograms:</p>\n<p>7.2% of the label_ids have missing data in the offset spectrogram.<br>\n8.7% of spectrograms have missing data in at least one of the offset spectrograms.</p>\n<p>Missing-data spectrogram distribution with complete-data spectrograms removed:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F76b347d12c5b6ef731487e4970fe4232%2Fmissing_data_label_id.png?generation=1708369753500522&amp;alt=media\"></p>\n<p>Max-missing-data spectrogram distribution with complete-data spectrograms removed:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fcce0c940cb1f2ea16561992b7fee905a%2Fmissing_data_spec_id.png?generation=1708369763664848&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2659340",
      "postDate": "02/19/2024 19:17:22",
      "content": "<p>As you may already know, many of the competition spectrograms have missing data, with some missing more than 50% of data. (Below is an example with more than a third of data missing from the LL region.) When training with unique spectrogram_ids, we sometimes have an option for an offset for the spectrogram. I decided to determine the missing spectrogram data for all 106,800 rows of <code>train.csv</code>, so the offset with the least missing data could be selected. The altered <code>train.csv</code> is output <a href=\"https://www.kaggle.com/seanbearden/missing-data-in-spectrograms\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F50aca4f227580ea30e7d18afa65875aa%2Fmissing_data.png?generation=1708368587898646&amp;alt=media\"></p>\n<p>The results were surprising and might affect training with competition-provided spectrograms:</p>\n<p>7.2% of the label_ids have missing data in the offset spectrogram.<br>\n8.7% of spectrograms have missing data in at least one of the offset spectrograms.</p>\n<p>Missing-data spectrogram distribution with complete-data spectrograms removed:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F76b347d12c5b6ef731487e4970fe4232%2Fmissing_data_label_id.png?generation=1708369753500522&amp;alt=media\"></p>\n<p>Max-missing-data spectrogram distribution with complete-data spectrograms removed:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fcce0c940cb1f2ea16561992b7fee905a%2Fmissing_data_spec_id.png?generation=1708369763664848&amp;alt=media\"></p>",
      "rawMarkdown": "As you may already know, many of the competition spectrograms have missing data, with some missing more than 50% of data. (Below is an example with more than a third of data missing from the LL region.) When training with unique spectrogram_ids, we sometimes have an option for an offset for the spectrogram. I decided to determine the missing spectrogram data for all 106,800 rows of `train.csv`, so the offset with the least missing data could be selected. The altered `train.csv` is output [here](https://www.kaggle.com/seanbearden/missing-data-in-spectrograms).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F50aca4f227580ea30e7d18afa65875aa%2Fmissing_data.png?generation=1708368587898646&alt=media)\n\nThe results were surprising and might affect training with competition-provided spectrograms:\n\n7.2% of the label_ids have missing data in the offset spectrogram.\n8.7% of spectrograms have missing data in at least one of the offset spectrograms.\n\nMissing-data spectrogram distribution with complete-data spectrograms removed:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F76b347d12c5b6ef731487e4970fe4232%2Fmissing_data_label_id.png?generation=1708369753500522&alt=media)\n\nMax-missing-data spectrogram distribution with complete-data spectrograms removed:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fcce0c940cb1f2ea16561992b7fee905a%2Fmissing_data_spec_id.png?generation=1708369763664848&alt=media)",
      "votes": null
    },
    {
      "id": "2659516",
      "postDate": "02/19/2024 23:00:18",
      "content": "<p>I've been thinking how data description says that the label is for the central 10 seconds of the spectrogram and corresponding EEG.  The image you have shows the central 10 secs.</p>",
      "rawMarkdown": "I've been thinking how data description says that the label is for the central 10 seconds of the spectrogram and corresponding EEG.  The image you have shows the central 10 secs.",
      "votes": null
    },
    {
      "id": "2659531",
      "postDate": "02/19/2024 23:58:20",
      "content": "<p>I’ve been wondering how the evaluators are influenced by the seeing the full 10 minute spectrogram. Might not be very relevant, but worth a review. </p>",
      "rawMarkdown": "I’ve been wondering how the evaluators are influenced by the seeing the full 10 minute spectrogram. Might not be very relevant, but worth a review.",
      "votes": null
    },
    {
      "id": "2659624",
      "postDate": "02/20/2024 03:35:46",
      "content": "<p>I wonder how doctors can classify which pattern for each patient depending on incomplete spectrograms, or what is the shortest time that the pattern can be confirmed.</p>",
      "rawMarkdown": "I wonder how doctors can classify which pattern for each patient depending on incomplete spectrograms, or what is the shortest time that the pattern can be confirmed.",
      "votes": null
    },
    {
      "id": "2660641",
      "postDate": "02/20/2024 18:24:35",
      "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> The competition overview indicates the evaluators are voting on the central 10 seconds on the diagram, and some of the reasoning for example vote distributions is explained using the central 10 seconds. It appears 10 seconds is enough data for an evaluator to make a decision, but the evaluator has also seen the whole 10 minute spectrogram. While it seems possible for the evaluator to only review the central 10 seconds, the fact that the evaluators have seen more data would indicate that extra data has an influence on the vote. However, that is pure speculation based on general human behavior.</p>",
      "rawMarkdown": "sweetyheehee The competition overview indicates the evaluators are voting on the central 10 seconds on the diagram, and some of the reasoning for example vote distributions is explained using the central 10 seconds. It appears 10 seconds is enough data for an evaluator to make a decision, but the evaluator has also seen the whole 10 minute spectrogram. While it seems possible for the evaluator to only review the central 10 seconds, the fact that the evaluators have seen more data would indicate that extra data has an influence on the vote. However, that is pure speculation based on general human behavior.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2659516,
      "author_name": "idithhaber",
      "author_url": "",
      "post_date": "02/19/2024 23:00:18",
      "content": "<p>I've been thinking how data description says that the label is for the central 10 seconds of the spectrogram and corresponding EEG.  The image you have shows the central 10 secs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2659531,
          "author_name": "seanbearden",
          "author_url": "",
          "post_date": "02/19/2024 23:58:20",
          "content": "<p>I’ve been wondering how the evaluators are influenced by the seeing the full 10 minute spectrogram. Might not be very relevant, but worth a review. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2659624,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "02/20/2024 03:35:46",
      "content": "<p>I wonder how doctors can classify which pattern for each patient depending on incomplete spectrograms, or what is the shortest time that the pattern can be confirmed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2660641,
          "author_name": "seanbearden",
          "author_url": "",
          "post_date": "02/20/2024 18:24:35",
          "content": "<p><a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a> The competition overview indicates the evaluators are voting on the central 10 seconds on the diagram, and some of the reasoning for example vote distributions is explained using the central 10 seconds. It appears 10 seconds is enough data for an evaluator to make a decision, but the evaluator has also seen the whole 10 minute spectrogram. While it seems possible for the evaluator to only review the central 10 seconds, the fact that the evaluators have seen more data would indicate that extra data has an influence on the vote. However, that is pure speculation based on general human behavior.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2659340": "As you may already know, many of the competition spectrograms have missing data, with some missing more than 50% of data. (Below is an example with more than a third of data missing from the LL region.) When training with unique spectrogram_ids, we sometimes have an option for an offset for the spectrogram. I decided to determine the missing spectrogram data for all 106,800 rows of `train.csv`, so the offset with the least missing data could be selected. The altered `train.csv` is output [here](https://www.kaggle.com/seanbearden/missing-data-in-spectrograms).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F50aca4f227580ea30e7d18afa65875aa%2Fmissing_data.png?generation=1708368587898646&alt=media)\n\nThe results were surprising and might affect training with competition-provided spectrograms:\n\n7.2% of the label_ids have missing data in the offset spectrogram.\n8.7% of spectrograms have missing data in at least one of the offset spectrograms.\n\nMissing-data spectrogram distribution with complete-data spectrograms removed:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2F76b347d12c5b6ef731487e4970fe4232%2Fmissing_data_label_id.png?generation=1708369753500522&alt=media)\n\nMax-missing-data spectrogram distribution with complete-data spectrograms removed:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3223907%2Fcce0c940cb1f2ea16561992b7fee905a%2Fmissing_data_spec_id.png?generation=1708369763664848&alt=media)",
    "2659516": "I've been thinking how data description says that the label is for the central 10 seconds of the spectrogram and corresponding EEG.  The image you have shows the central 10 secs.",
    "2659531": "I’ve been wondering how the evaluators are influenced by the seeing the full 10 minute spectrogram. Might not be very relevant, but worth a review.",
    "2659624": "I wonder how doctors can classify which pattern for each patient depending on incomplete spectrograms, or what is the shortest time that the pattern can be confirmed.",
    "2660641": "sweetyheehee The competition overview indicates the evaluators are voting on the central 10 seconds on the diagram, and some of the reasoning for example vote distributions is explained using the central 10 seconds. It appears 10 seconds is enough data for an evaluator to make a decision, but the evaluator has also seen the whole 10 minute spectrogram. While it seems possible for the evaluator to only review the central 10 seconds, the fact that the evaluators have seen more data would indicate that extra data has an influence on the vote. However, that is pure speculation based on general human behavior."
  },
  "source": "meta"
}