{
  "id": 467909,
  "title": "Faulty spectrogram and EEG data?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467909",
  "author_name": "",
  "post_date": "2024-01-14T15:03:00.567755600Z",
  "votes": 30,
  "comment_count": 14,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> :</p>\n<p>I ran some basic EDA on all the Parquet files for the spectrograms and EEGs. I came across a somewhat high number of problems for both types of files:</p>\n<ul>\n<li><p>For spectrograms, I found 969 Parquet files that contain  rows of <code>NaN</code> values. This can be checked by running this notebook: <a href=\"https://www.kaggle.com/code/patrob/check-nan-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/patrob/check-nan-spectrograms</a>. This represents 8.7% of all spectrograms.</p></li>\n<li><p>For EEGs, the problem is a bit different; there are 824 Parquet files where there are rows (or parts of rows) that contain only <code>NaN</code> values. This can be checked by running this notebook: <a href=\"https://www.kaggle.com/code/patrob/check-nan-rows-eeg\" target=\"_blank\">https://www.kaggle.com/code/patrob/check-nan-rows-eeg</a>. The number of <code>NaN</code> rows varies from 1 (e.g. <code>eeg_id</code> 3625731) to 5197 (<code>eeg_id</code> 3289898692), with an average of 405.8 erroneous rows per file. This represents 4.82% of all EEGs (not including the 211 Parquet files that do not have a corresponding row in <code>train.csv</code>, as per <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467058\" target=\"_blank\">this post</a>).</p></li>\n</ul>\n<p>I have also spotted other data issues with the EEG files (e.g. some implausible datasets, like <code>eeg_id</code> 1457334423), but I will wait to see your initial reply before detailing these issues.</p>\n<p>Considering the high number of affected files, I don't believe that it should be a simple matter of data cleaning for the participants in this competition. Could you please let us know: 1- if this is something you should/can fix; 2- if yes, how long it would take for an update; 3- if not, how we should deal with this issue going forward (e.g. ignoring the faulty files?). </p>\n<p><strong>Update:</strong> I have ascertained the number of samples in <code>train.csv</code> that are affected by the issues above. 8552 samples would have faulty EEG data only; 13110 samples would have faulty spectrogram data only; 2300 samples would have both faulty EEG and spectrogram data. So, <strong>a total of 23952 samples</strong> with faulty data, representing <strong>22.4% of the whole dataset</strong>.</p>\n<p><strong>2nd update:</strong> I have updated the spectrogram notebook to show the number of rows and the number of <code>NaN</code> rows for each spectrogram.</p>",
  "messages": [
    {
      "id": "2601623",
      "postDate": "01/14/2024 15:03:00",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> :</p>\n<p>I ran some basic EDA on all the Parquet files for the spectrograms and EEGs. I came across a somewhat high number of problems for both types of files:</p>\n<ul>\n<li><p>For spectrograms, I found 969 Parquet files that contain  rows of <code>NaN</code> values. This can be checked by running this notebook: <a href=\"https://www.kaggle.com/code/patrob/check-nan-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/patrob/check-nan-spectrograms</a>. This represents 8.7% of all spectrograms.</p></li>\n<li><p>For EEGs, the problem is a bit different; there are 824 Parquet files where there are rows (or parts of rows) that contain only <code>NaN</code> values. This can be checked by running this notebook: <a href=\"https://www.kaggle.com/code/patrob/check-nan-rows-eeg\" target=\"_blank\">https://www.kaggle.com/code/patrob/check-nan-rows-eeg</a>. The number of <code>NaN</code> rows varies from 1 (e.g. <code>eeg_id</code> 3625731) to 5197 (<code>eeg_id</code> 3289898692), with an average of 405.8 erroneous rows per file. This represents 4.82% of all EEGs (not including the 211 Parquet files that do not have a corresponding row in <code>train.csv</code>, as per <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467058\" target=\"_blank\">this post</a>).</p></li>\n</ul>\n<p>I have also spotted other data issues with the EEG files (e.g. some implausible datasets, like <code>eeg_id</code> 1457334423), but I will wait to see your initial reply before detailing these issues.</p>\n<p>Considering the high number of affected files, I don't believe that it should be a simple matter of data cleaning for the participants in this competition. Could you please let us know: 1- if this is something you should/can fix; 2- if yes, how long it would take for an update; 3- if not, how we should deal with this issue going forward (e.g. ignoring the faulty files?). </p>\n<p><strong>Update:</strong> I have ascertained the number of samples in <code>train.csv</code> that are affected by the issues above. 8552 samples would have faulty EEG data only; 13110 samples would have faulty spectrogram data only; 2300 samples would have both faulty EEG and spectrogram data. So, <strong>a total of 23952 samples</strong> with faulty data, representing <strong>22.4% of the whole dataset</strong>.</p>\n<p><strong>2nd update:</strong> I have updated the spectrogram notebook to show the number of rows and the number of <code>NaN</code> rows for each spectrogram.</p>",
      "rawMarkdown": "sohier :\n\nI ran some basic EDA on all the Parquet files for the spectrograms and EEGs. I came across a somewhat high number of problems for both types of files:\n\n- For spectrograms, I found 969 Parquet files that contain ~~only~~ rows of `NaN` values. This can be checked by running this notebook: [https://www.kaggle.com/code/patrob/check-nan-spectrograms](https://www.kaggle.com/code/patrob/check-nan-spectrograms). This represents 8.7% of all spectrograms.\n\n- For EEGs, the problem is a bit different; there are 824 Parquet files where there are rows (or parts of rows) that contain only `NaN` values. This can be checked by running this notebook: [https://www.kaggle.com/code/patrob/check-nan-rows-eeg](https://www.kaggle.com/code/patrob/check-nan-rows-eeg). The number of `NaN` rows varies from 1 (e.g. `eeg_id` 3625731) to 5197 (`eeg_id` 3289898692), with an average of 405.8 erroneous rows per file. This represents 4.82% of all EEGs (not including the 211 Parquet files that do not have a corresponding row in `train.csv`, as per [this post](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467058)).\n\nI have also spotted other data issues with the EEG files (e.g. some implausible datasets, like `eeg_id` 1457334423), but I will wait to see your initial reply before detailing these issues.\n\nConsidering the high number of affected files, I don't believe that it should be a simple matter of data cleaning for the participants in this competition. Could you please let us know: 1- if this is something you should/can fix; 2- if yes, how long it would take for an update; 3- if not, how we should deal with this issue going forward (e.g. ignoring the faulty files?). \n\n**Update:** I have ascertained the number of samples in `train.csv` that are affected by the issues above. 8552 samples would have faulty EEG data only; 13110 samples would have faulty spectrogram data only; 2300 samples would have both faulty EEG and spectrogram data. So, **a total of 23952 samples** with faulty data, representing **22.4% of the whole dataset**.\n\n**2nd update:** I have updated the spectrogram notebook to show the number of rows and the number of `NaN` rows for each spectrogram.",
      "votes": null
    },
    {
      "id": "2601917",
      "postDate": "01/14/2024 19:23:41",
      "content": "<p>Hello Patrick, your link for the notebook is not working.</p>",
      "rawMarkdown": "Hello Patrick, your link for the notebook is not working.",
      "votes": null
    },
    {
      "id": "2602194",
      "postDate": "01/15/2024 00:53:11",
      "content": "<p>Thanks Yan, it should work now.</p>",
      "rawMarkdown": "Thanks Yan, it should work now.",
      "votes": null
    },
    {
      "id": "2605709",
      "postDate": "01/17/2024 07:36:46",
      "content": "<p>Any information regarding this <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>?</p>\n<ul>\n<li>Is missing values an error from Kaggle side?</li>\n<li>Why do they occur and what do they mean?</li>\n<li>Does test set have any missing values? </li>\n</ul>",
      "rawMarkdown": "Any information regarding this @sohier?\n\n* Is missing values an error from Kaggle side?\n* Why do they occur and what do they mean?\n* Does test set have any missing values?",
      "votes": null
    },
    {
      "id": "2613398",
      "postDate": "01/22/2024 04:17:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>It's been more than a week since I raised these potential issues, and we are yet to receive any acknowledgement that this is being looked at and, if required, leading to corrections. Could you please clarify as soon as possible where things are at.</p>",
      "rawMarkdown": "Hi @sohier,\n\nIt's been more than a week since I raised these potential issues, and we are yet to receive any acknowledgement that this is being looked at and, if required, leading to corrections. Could you please clarify as soon as possible where things are at.",
      "votes": null
    },
    {
      "id": "2616798",
      "postDate": "01/23/2024 20:36:12",
      "content": "<blockquote>\n  <p>For spectrograms, I found 969 Parquet files that contain only NaN values. This can be checked by running this notebook: <a href=\"https://www.kaggle.com/code/patrob/check-nan-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/patrob/check-nan-spectrograms</a>. This represents 8.7% of all spectrograms.</p>\n</blockquote>\n<p>There are no NaN in the sub spectrograms present in train.csv. </p>\n<p>This is a notebook that checks for NaN in each 10 minutes slice: <a href=\"https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms</a></p>",
      "rawMarkdown": ">For spectrograms, I found 969 Parquet files that contain only NaN values. This can be checked by running this notebook: https://www.kaggle.com/code/patrob/check-nan-spectrograms. This represents 8.7% of all spectrograms.\n\nThere are no NaN in the sub spectrograms present in train.csv. \n\nThis is a notebook that checks for NaN in each 10 minutes slice: [https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms](https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms)",
      "votes": null
    },
    {
      "id": "2616923",
      "postDate": "01/24/2024 00:26:45",
      "content": "<p>\"We can't find that page\" error. Have you made the notebook public?</p>",
      "rawMarkdown": "\"We can't find that page\" error. Have you made the notebook public?",
      "votes": null
    },
    {
      "id": "2616978",
      "postDate": "01/24/2024 02:14:08",
      "content": "<p>It's public now.</p>",
      "rawMarkdown": "It's public now.",
      "votes": null
    },
    {
      "id": "2617007",
      "postDate": "01/24/2024 02:56:54",
      "content": "<p>I will disagree with you, and I think there is something you misunderstood about the spectrogram data. A spectrogram slice of 10 minutes represents 300 rows of a spectrogram Parquet file. Of the 969 faulty spectrograms I have identified, there are 495 that contain more than 100 NaN rows (highest: 550); there are 94 spectrograms where the number of <code>NaN</code> rows is greater than 50% of the Parquet file. I haven't inspected them all, but I would suspect that these are mostly all consecutive rows. When you have a spectrogram file that contains 295 <code>NaN</code> rows out of 300 (spectrogram_id 314642970), I really doubt that it can serve any purpose.</p>\n<p>Replacing the NaN values in these files would be quite silly, especially when the missing rows are consecutive. For a lot of smaller cases, reconstructing the values using the EEG files would be impossible, as the EEG files focus on the central 50 seconds of a 10-minute spectrogram band; if your missing rows are at the beginning or the end of the spectrogram band, there is not much you can do. One alternative would be to use less than 10 minutes of the spectrogram, but you will find that a lot of the faulty spectrograms have data missing around the middle of the band. That's why I keep urging <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> to let us know what is going on with these files (and other issues I have raised). Unfortunately, the silence is still deafening.</p>",
      "rawMarkdown": "I will disagree with you, and I think there is something you misunderstood about the spectrogram data. A spectrogram slice of 10 minutes represents 300 rows of a spectrogram Parquet file. Of the 969 faulty spectrograms I have identified, there are 495 that contain more than 100 NaN rows (highest: 550); there are 94 spectrograms where the number of `NaN` rows is greater than 50% of the Parquet file. I haven't inspected them all, but I would suspect that these are mostly all consecutive rows. When you have a spectrogram file that contains 295 `NaN` rows out of 300 (spectrogram_id 314642970), I really doubt that it can serve any purpose.\n\nReplacing the NaN values in these files would be quite silly, especially when the missing rows are consecutive. For a lot of smaller cases, reconstructing the values using the EEG files would be impossible, as the EEG files focus on the central 50 seconds of a 10-minute spectrogram band; if your missing rows are at the beginning or the end of the spectrogram band, there is not much you can do. One alternative would be to use less than 10 minutes of the spectrogram, but you will find that a lot of the faulty spectrograms have data missing around the middle of the band. That's why I keep urging @sohier to let us know what is going on with these files (and other issues I have raised). Unfortunately, the silence is still deafening.",
      "votes": null
    },
    {
      "id": "2617105",
      "postDate": "01/24/2024 04:30:16",
      "content": "<p>Thank you for your detailed reply, now I see my mistake. I found 7743 sub eeg's (rows in train.csv) where the sub spectrogram's contain NaNs.</p>\n<p>This is 7743 out of 106800, so for now I will just remove these observations from the training set. I'll try generating my own spectrograms in the future.</p>",
      "rawMarkdown": "Thank you for your detailed reply, now I see my mistake. I found 7743 sub eeg's (rows in train.csv) where the sub spectrogram's contain NaNs.\n\nThis is 7743 out of 106800, so for now I will just remove these observations from the training set. I'll try generating my own spectrograms in the future.",
      "votes": null
    },
    {
      "id": "2638016",
      "postDate": "02/06/2024 03:55:31",
      "content": "<p>I used the following code and found 10829 nan samples in <strong>train.csv</strong></p>\n<pre><code>eeg = eeg.iloc[eeg_offset * :(eeg_offset + ) * ].values\nspectrogram = spectrogram.loc[(spectrogram.time &gt;= spectrogram_offset) &amp; (spectrogram.time &lt; spectrogram_offset + )].values[:, :]\n   np.isnan(eeg)    np.isnan(spectrogram):\n   idx\n</code></pre>",
      "rawMarkdown": "I used the following code and found 10829 nan samples in **train.csv**\n\n```python\neeg = eeg.iloc[eeg_offset * 200:(eeg_offset + 50) * 200].values\nspectrogram = spectrogram.loc[(spectrogram.time >= spectrogram_offset) & (spectrogram.time < spectrogram_offset + 600)].values[:, 1:]\nif True in np.isnan(eeg) or True in np.isnan(spectrogram):\n  return idx\n```",
      "votes": null
    },
    {
      "id": "2638973",
      "postDate": "02/06/2024 15:46:53",
      "content": "<p>I can confirm that there are spectrograms with missing values in test set. My submissions fails when I don't do <code>fillna()</code>. There are cases like this in training set (white areas are NaNs).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe25c85763c1c3f185257c12c7399e4a7%2FScreenshot%20from%202024-02-06%2018-43-43.png?generation=1707234258433345&amp;alt=media\" alt=\"1\"></p>\n<p>I don't think it's possible to predict them accurately so they shouldn't be included in training and validation sets. I tried to remove all spectrogram subsamples with at least 1 missing value but my LB score got worse, so there is a trade-off between keeping samples with too many missing values and LB score. The only way to find out is LB probing.</p>",
      "rawMarkdown": "I can confirm that there are spectrograms with missing values in test set. My submissions fails when I don't do `fillna()`. There are cases like this in training set (white areas are NaNs).\n\n![1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe25c85763c1c3f185257c12c7399e4a7%2FScreenshot%20from%202024-02-06%2018-43-43.png?generation=1707234258433345&alt=media)\n\nI don't think it's possible to predict them accurately so they shouldn't be included in training and validation sets. I tried to remove all spectrogram subsamples with at least 1 missing value but my LB score got worse, so there is a trade-off between keeping samples with too many missing values and LB score. The only way to find out is LB probing.",
      "votes": null
    },
    {
      "id": "2640432",
      "postDate": "02/06/2024 23:36:05",
      "content": "<p>It's a matter of judgment I'm sure. There is a lot of data here, and imputed spectrogram data will not have much training value.</p>\n<p>Does anyone know if the final evaluations are done on higher quality data than the original training, or if we should expect to see missing data in the final evaluation?</p>",
      "rawMarkdown": "It's a matter of judgment I'm sure. There is a lot of data here, and imputed spectrogram data will not have much training value.\n\nDoes anyone know if the final evaluations are done on higher quality data than the original training, or if we should expect to see missing data in the final evaluation?",
      "votes": null
    },
    {
      "id": "2645229",
      "postDate": "02/10/2024 04:50:49",
      "content": "<p>If eegs are less faulty, then we can build our own spectrograms. Most NaN in eegs are isolated rows, so these can be imputed by interpolation. Then I build spectrograms this way: <a href=\"https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms</a></p>",
      "rawMarkdown": "If eegs are less faulty, then we can build our own spectrograms. Most NaN in eegs are isolated rows, so these can be imputed by interpolation. Then I build spectrograms this way: [https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms](https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms)",
      "votes": null
    },
    {
      "id": "2645942",
      "postDate": "02/10/2024 15:00:59",
      "content": "<p>Could it be that in the test base we also have these null spectograms and EEGs with null channels?</p>",
      "rawMarkdown": "Could it be that in the test base we also have these null spectograms and EEGs with null channels?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2601917,
      "author_name": "yantxx",
      "author_url": "",
      "post_date": "01/14/2024 19:23:41",
      "content": "<p>Hello Patrick, your link for the notebook is not working.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2602194,
          "author_name": "patrob",
          "author_url": "",
          "post_date": "01/15/2024 00:53:11",
          "content": "<p>Thanks Yan, it should work now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2605709,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/17/2024 07:36:46",
      "content": "<p>Any information regarding this <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>?</p>\n<ul>\n<li>Is missing values an error from Kaggle side?</li>\n<li>Why do they occur and what do they mean?</li>\n<li>Does test set have any missing values? </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2613398,
      "author_name": "patrob",
      "author_url": "",
      "post_date": "01/22/2024 04:17:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>It's been more than a week since I raised these potential issues, and we are yet to receive any acknowledgement that this is being looked at and, if required, leading to corrections. Could you please clarify as soon as possible where things are at.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2616798,
      "author_name": "elgardo1",
      "author_url": "",
      "post_date": "01/23/2024 20:36:12",
      "content": "<blockquote>\n  <p>For spectrograms, I found 969 Parquet files that contain only NaN values. This can be checked by running this notebook: <a href=\"https://www.kaggle.com/code/patrob/check-nan-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/patrob/check-nan-spectrograms</a>. This represents 8.7% of all spectrograms.</p>\n</blockquote>\n<p>There are no NaN in the sub spectrograms present in train.csv. </p>\n<p>This is a notebook that checks for NaN in each 10 minutes slice: <a href=\"https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2616923,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "01/24/2024 00:26:45",
          "content": "<p>\"We can't find that page\" error. Have you made the notebook public?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2616978,
              "author_name": "elgardo1",
              "author_url": "",
              "post_date": "01/24/2024 02:14:08",
              "content": "<p>It's public now.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2617007,
          "author_name": "patrob",
          "author_url": "",
          "post_date": "01/24/2024 02:56:54",
          "content": "<p>I will disagree with you, and I think there is something you misunderstood about the spectrogram data. A spectrogram slice of 10 minutes represents 300 rows of a spectrogram Parquet file. Of the 969 faulty spectrograms I have identified, there are 495 that contain more than 100 NaN rows (highest: 550); there are 94 spectrograms where the number of <code>NaN</code> rows is greater than 50% of the Parquet file. I haven't inspected them all, but I would suspect that these are mostly all consecutive rows. When you have a spectrogram file that contains 295 <code>NaN</code> rows out of 300 (spectrogram_id 314642970), I really doubt that it can serve any purpose.</p>\n<p>Replacing the NaN values in these files would be quite silly, especially when the missing rows are consecutive. For a lot of smaller cases, reconstructing the values using the EEG files would be impossible, as the EEG files focus on the central 50 seconds of a 10-minute spectrogram band; if your missing rows are at the beginning or the end of the spectrogram band, there is not much you can do. One alternative would be to use less than 10 minutes of the spectrogram, but you will find that a lot of the faulty spectrograms have data missing around the middle of the band. That's why I keep urging <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> to let us know what is going on with these files (and other issues I have raised). Unfortunately, the silence is still deafening.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2617105,
              "author_name": "elgardo1",
              "author_url": "",
              "post_date": "01/24/2024 04:30:16",
              "content": "<p>Thank you for your detailed reply, now I see my mistake. I found 7743 sub eeg's (rows in train.csv) where the sub spectrogram's contain NaNs.</p>\n<p>This is 7743 out of 106800, so for now I will just remove these observations from the training set. I'll try generating my own spectrograms in the future.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2638016,
      "author_name": "zznznb",
      "author_url": "",
      "post_date": "02/06/2024 03:55:31",
      "content": "<p>I used the following code and found 10829 nan samples in <strong>train.csv</strong></p>\n<pre><code>eeg = eeg.iloc[eeg_offset * :(eeg_offset + ) * ].values\nspectrogram = spectrogram.loc[(spectrogram.time &gt;= spectrogram_offset) &amp; (spectrogram.time &lt; spectrogram_offset + )].values[:, :]\n   np.isnan(eeg)    np.isnan(spectrogram):\n   idx\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2638973,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "02/06/2024 15:46:53",
      "content": "<p>I can confirm that there are spectrograms with missing values in test set. My submissions fails when I don't do <code>fillna()</code>. There are cases like this in training set (white areas are NaNs).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe25c85763c1c3f185257c12c7399e4a7%2FScreenshot%20from%202024-02-06%2018-43-43.png?generation=1707234258433345&amp;alt=media\" alt=\"1\"></p>\n<p>I don't think it's possible to predict them accurately so they shouldn't be included in training and validation sets. I tried to remove all spectrogram subsamples with at least 1 missing value but my LB score got worse, so there is a trade-off between keeping samples with too many missing values and LB score. The only way to find out is LB probing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2645229,
          "author_name": "elgardo1",
          "author_url": "",
          "post_date": "02/10/2024 04:50:49",
          "content": "<p>If eegs are less faulty, then we can build our own spectrograms. Most NaN in eegs are isolated rows, so these can be imputed by interpolation. Then I build spectrograms this way: <a href=\"https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms\" target=\"_blank\">https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2640432,
      "author_name": "jrhaberstroh",
      "author_url": "",
      "post_date": "02/06/2024 23:36:05",
      "content": "<p>It's a matter of judgment I'm sure. There is a lot of data here, and imputed spectrogram data will not have much training value.</p>\n<p>Does anyone know if the final evaluations are done on higher quality data than the original training, or if we should expect to see missing data in the final evaluation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2645942,
      "author_name": "rafaelzimmermann1",
      "author_url": "",
      "post_date": "02/10/2024 15:00:59",
      "content": "<p>Could it be that in the test base we also have these null spectograms and EEGs with null channels?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2601623": "sohier :\n\nI ran some basic EDA on all the Parquet files for the spectrograms and EEGs. I came across a somewhat high number of problems for both types of files:\n\n- For spectrograms, I found 969 Parquet files that contain ~~only~~ rows of `NaN` values. This can be checked by running this notebook: [https://www.kaggle.com/code/patrob/check-nan-spectrograms](https://www.kaggle.com/code/patrob/check-nan-spectrograms). This represents 8.7% of all spectrograms.\n\n- For EEGs, the problem is a bit different; there are 824 Parquet files where there are rows (or parts of rows) that contain only `NaN` values. This can be checked by running this notebook: [https://www.kaggle.com/code/patrob/check-nan-rows-eeg](https://www.kaggle.com/code/patrob/check-nan-rows-eeg). The number of `NaN` rows varies from 1 (e.g. `eeg_id` 3625731) to 5197 (`eeg_id` 3289898692), with an average of 405.8 erroneous rows per file. This represents 4.82% of all EEGs (not including the 211 Parquet files that do not have a corresponding row in `train.csv`, as per [this post](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467058)).\n\nI have also spotted other data issues with the EEG files (e.g. some implausible datasets, like `eeg_id` 1457334423), but I will wait to see your initial reply before detailing these issues.\n\nConsidering the high number of affected files, I don't believe that it should be a simple matter of data cleaning for the participants in this competition. Could you please let us know: 1- if this is something you should/can fix; 2- if yes, how long it would take for an update; 3- if not, how we should deal with this issue going forward (e.g. ignoring the faulty files?). \n\n**Update:** I have ascertained the number of samples in `train.csv` that are affected by the issues above. 8552 samples would have faulty EEG data only; 13110 samples would have faulty spectrogram data only; 2300 samples would have both faulty EEG and spectrogram data. So, **a total of 23952 samples** with faulty data, representing **22.4% of the whole dataset**.\n\n**2nd update:** I have updated the spectrogram notebook to show the number of rows and the number of `NaN` rows for each spectrogram.",
    "2601917": "Hello Patrick, your link for the notebook is not working.",
    "2602194": "Thanks Yan, it should work now.",
    "2605709": "Any information regarding this @sohier?\n\n* Is missing values an error from Kaggle side?\n* Why do they occur and what do they mean?\n* Does test set have any missing values?",
    "2613398": "Hi @sohier,\n\nIt's been more than a week since I raised these potential issues, and we are yet to receive any acknowledgement that this is being looked at and, if required, leading to corrections. Could you please clarify as soon as possible where things are at.",
    "2616798": ">For spectrograms, I found 969 Parquet files that contain only NaN values. This can be checked by running this notebook: https://www.kaggle.com/code/patrob/check-nan-spectrograms. This represents 8.7% of all spectrograms.\n\nThere are no NaN in the sub spectrograms present in train.csv. \n\nThis is a notebook that checks for NaN in each 10 minutes slice: [https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms](https://www.kaggle.com/code/elgardo1/hms-eda-nan-in-sub-spectrograms)",
    "2616923": "\"We can't find that page\" error. Have you made the notebook public?",
    "2616978": "It's public now.",
    "2617007": "I will disagree with you, and I think there is something you misunderstood about the spectrogram data. A spectrogram slice of 10 minutes represents 300 rows of a spectrogram Parquet file. Of the 969 faulty spectrograms I have identified, there are 495 that contain more than 100 NaN rows (highest: 550); there are 94 spectrograms where the number of `NaN` rows is greater than 50% of the Parquet file. I haven't inspected them all, but I would suspect that these are mostly all consecutive rows. When you have a spectrogram file that contains 295 `NaN` rows out of 300 (spectrogram_id 314642970), I really doubt that it can serve any purpose.\n\nReplacing the NaN values in these files would be quite silly, especially when the missing rows are consecutive. For a lot of smaller cases, reconstructing the values using the EEG files would be impossible, as the EEG files focus on the central 50 seconds of a 10-minute spectrogram band; if your missing rows are at the beginning or the end of the spectrogram band, there is not much you can do. One alternative would be to use less than 10 minutes of the spectrogram, but you will find that a lot of the faulty spectrograms have data missing around the middle of the band. That's why I keep urging @sohier to let us know what is going on with these files (and other issues I have raised). Unfortunately, the silence is still deafening.",
    "2617105": "Thank you for your detailed reply, now I see my mistake. I found 7743 sub eeg's (rows in train.csv) where the sub spectrogram's contain NaNs.\n\nThis is 7743 out of 106800, so for now I will just remove these observations from the training set. I'll try generating my own spectrograms in the future.",
    "2638016": "I used the following code and found 10829 nan samples in **train.csv**\n\n```python\neeg = eeg.iloc[eeg_offset * 200:(eeg_offset + 50) * 200].values\nspectrogram = spectrogram.loc[(spectrogram.time >= spectrogram_offset) & (spectrogram.time < spectrogram_offset + 600)].values[:, 1:]\nif True in np.isnan(eeg) or True in np.isnan(spectrogram):\n  return idx\n```",
    "2638973": "I can confirm that there are spectrograms with missing values in test set. My submissions fails when I don't do `fillna()`. There are cases like this in training set (white areas are NaNs).\n\n![1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fe25c85763c1c3f185257c12c7399e4a7%2FScreenshot%20from%202024-02-06%2018-43-43.png?generation=1707234258433345&alt=media)\n\nI don't think it's possible to predict them accurately so they shouldn't be included in training and validation sets. I tried to remove all spectrogram subsamples with at least 1 missing value but my LB score got worse, so there is a trade-off between keeping samples with too many missing values and LB score. The only way to find out is LB probing.",
    "2640432": "It's a matter of judgment I'm sure. There is a lot of data here, and imputed spectrogram data will not have much training value.\n\nDoes anyone know if the final evaluations are done on higher quality data than the original training, or if we should expect to see missing data in the final evaluation?",
    "2645229": "If eegs are less faulty, then we can build our own spectrograms. Most NaN in eegs are isolated rows, so these can be imputed by interpolation. Then I build spectrograms this way: [https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms](https://www.kaggle.com/code/elgardo1/hms-how-to-build-spectrograms)",
    "2645942": "Could it be that in the test base we also have these null spectograms and EEGs with null channels?"
  },
  "source": "meta"
}