{
  "id": 467058,
  "title": "is eeg_ids missing in train.csv but available in train_eegs folder? [Solved]",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467058",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2024-01-11T01:43:08.247000",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Total spectrogram_id in train.csv: <strong>11138</strong> &amp; total parquet files are <strong>11138</strong><br>\n Total eeg_ids in train.csv: <strong>17089</strong> &amp; total files are <strong>17300</strong></p>\n<p>missing eeg_ids in train.csv <strong>211</strong> ( <strong>can be used as pseudolabels or is training data missing</strong> ) ?</p>\n<pre><code>missing_eeg_ids = [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ]\n</code></pre>\n<h3><a href=\"https://www.kaggle.com/code/seshurajup/missing-eeg-ids-in-train-csv\" target=\"_blank\">Notebook - Missing Eeg_ids in Train.csv vs train_eegs parquet</a></h3>\n<h2>Those are files without paired spectrograms where I missed some cleanup. by <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></h2>",
  "messages": [
    {
      "id": 2596260,
      "postDate": "2024-01-11T01:43:08.247Z",
      "content": "<p>Total spectrogram_id in train.csv: <strong>11138</strong> &amp; total parquet files are <strong>11138</strong><br>\n Total eeg_ids in train.csv: <strong>17089</strong> &amp; total files are <strong>17300</strong></p>\n<p>missing eeg_ids in train.csv <strong>211</strong> ( <strong>can be used as pseudolabels or is training data missing</strong> ) ?</p>\n<pre><code>missing_eeg_ids = [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ]\n</code></pre>\n<h3><a href=\"https://www.kaggle.com/code/seshurajup/missing-eeg-ids-in-train-csv\" target=\"_blank\">Notebook - Missing Eeg_ids in Train.csv vs train_eegs parquet</a></h3>\n<h2>Those are files without paired spectrograms where I missed some cleanup. by <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></h2>",
      "rawMarkdown": "Total spectrogram_id in train.csv: **11138** & total parquet files are **11138**\n Total eeg_ids in train.csv: **17089** & total files are **17300**\n\nmissing eeg_ids in train.csv **211** ( **can be used as pseudolabels or is training data missing** ) ?\n```python\nmissing_eeg_ids = [100261680, 1006575126, 1046303268, 1054182773, 1080752861, 1090493279, 1102260141, 1106457838, 1114646196, 1124971078, 1129703637, 1138516159, 1152469719, 1164247904, 1210061638, 1213982942, 1218899486, 124842023, 1256854516, 1271026602, 1302255870, 1320529265, 1356605204, 1362609712, 1395846857, 1437515164, 1446608286, 1471761225, 1472250249, 1484488768, 1486503212, 1494048680, 1514837355, 1527948998, 1540422245, 1550636461, 159230712, 1593654895, 1644799805, 1667505910, 166924982, 1685149928, 176751182, 1781690069, 17823653, 1783811564, 1789492555, 1827043672, 1835829757, 1844412082, 1880066416, 1946169694, 1967500472, 1977522925, 1980230677, 2008191275, 2021270543, 2029271373, 2031389372, 2032669824, 2062845812, 2069539903, 2133352173, 2161239256, 2163060838, 2167910584, 2169828710, 217188535, 2177382144, 2187713476, 2215145921, 2229664974, 2232763157, 2250059220, 2256830545, 2360465057, 2370617920, 2402422559, 2437996061, 2465170355, 2495888973, 2512462400, 2514244198, 2518637163, 2524754765, 2537083119, 2546815772, 2564933177, 2566899625, 2585735707, 2619383113, 2644594529, 2704546408, 2722685303, 2734605314, 2753356177, 2770720245, 2780688155, 2827773679, 2847186843, 2883273759, 2890567875, 2915092170, 2928453358, 2940354800, 2945876125, 2966099404, 2980896188, 2989528583, 2995737301, 3029056160, 3107283960, 3110785252, 3130316564, 3159731110, 3184041077, 3184243660, 3219158996, 3231425203, 3248636378, 3277143842, 3279947709, 3287073479, 3296877755, 3302245698, 3334497580, 3374558757, 3395079519, 3395717621, 3425324038, 3427375889, 3462221095, 3463780576, 3469377987, 351886689, 3534133394, 3576774144, 358214261, 3588781445, 3613550741, 3624732686, 3672236418, 3677734263, 367926685, 3681177371, 3688844263, 3713571437, 3726556263, 3729617662, 3746528981, 3765156259, 37691174, 3782539999, 3784293316, 3808352981, 3844299098, 3860943081, 3861560461, 3864775587, 389650354, 3897110777, 3897137837, 390524827, 3932079802, 3963505952, 3964740894, 3965080689, 397155403, 3973402202, 4016243421, 4016271604, 4062219555, 4062860278, 4066543763, 4088014176, 4091126054, 4118489199, 416005567, 4187869644, 4263916735, 4265462352, 4283988928, 438882483, 452392893, 463419419, 495719000, 507760170, 512986705, 525392618, 529809133, 535701085, 581008851, 582052247, 592953741, 595031723, 60524973, 667978667, 710494757, 735478859, 757170164, 771666080, 789759143, 846376179, 875712244, 876702254, 880829874, 888227967, 914125289, 936728284, 968174006, 990852758]\n```\n\n### [Notebook - Missing Eeg_ids in Train.csv vs train_eegs parquet](https://www.kaggle.com/code/seshurajup/missing-eeg-ids-in-train-csv)\n\n## Those are files without paired spectrograms where I missed some cleanup. by @sohier",
      "votes": 15
    },
    {
      "id": 2596267,
      "postDate": "2024-01-11T02:07:00.523Z",
      "content": "<p>I don't think we are missing any data. You can</p>\n<pre><code>train = pd()\n )\n )\n</code></pre>\n<p>This prints 17089 and 11138 respectively which is the number of parquets files for each.</p>\n<p>There are only 1950 unique patients. So for EEG they spread the data over 17089 parquets. And for spectrograms, they spread out the data over 11138. </p>\n<p>When creating a single train sample, we load one EEG parquet and one spectrogram parquet. Then we extract 10_000 consecutive rows for EEG. And extract 300 consecutive rows for spectrogram. For each row in <code>train.csv</code>, the rows to extract from parquet files are indicated.</p>",
      "rawMarkdown": "I don't think we are missing any data. You can\n\n    train = pd.read_csv('train.csv')\n    print( train.eeg_id.nunique() )\n    print( train.spectrogram_id.nunique() )\n\nThis prints 17089 and 11138 respectively which is the number of parquets files for each.\n\nThere are only 1950 unique patients. So for EEG they spread the data over 17089 parquets. And for spectrograms, they spread out the data over 11138. \n\nWhen creating a single train sample, we load one EEG parquet and one spectrogram parquet. Then we extract 10_000 consecutive rows for EEG. And extract 300 consecutive rows for spectrogram. For each row in `train.csv`, the rows to extract from parquet files are indicated.",
      "votes": 2,
      "replies": [
        {
          "id": 2596271,
          "postDate": "2024-01-11T02:12:02.953Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Yes, we have all parquets of the train.csv's eeg_id</p>\n<p>but train_eegs folder have <strong>17_300</strong> parquets. did i calculated wrong way?</p>",
          "rawMarkdown": "@cdeotte Yes, we have all parquets of the train.csv's eeg_id\n\nbut train_eegs folder have **17_300** parquets. did i calculated wrong way?",
          "votes": 3
        },
        {
          "id": 2596386,
          "postDate": "2024-01-11T04:54:11.573Z",
          "content": "<p>I can confirm that there are 17300 parquet files in the <code>train_eegs</code> folder. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, there seems to be 211 rows missing in <code>train.csv</code>; could you please clarify?</p>",
          "rawMarkdown": "I can confirm that there are 17300 parquet files in the `train_eegs` folder. @sohier, there seems to be 211 rows missing in `train.csv`; could you please clarify?",
          "votes": 3
        }
      ]
    },
    {
      "id": 2640887,
      "postDate": "2024-02-07T06:12:18.703Z",
      "content": "<p>First of all, thanks for sharing good information! <br>\nI have a question that we need to drop or pesudolabeling?</p>",
      "rawMarkdown": "First of all, thanks for sharing good information! \nI have a question that we need to drop or pesudolabeling?\n"
    },
    {
      "id": 2597718,
      "postDate": "2024-01-11T20:09:57.057Z",
      "content": "<p>Leaked test data?? I downloaded them, just to make sure 😉</p>",
      "rawMarkdown": "Leaked test data?? I downloaded them, just to make sure 😉",
      "replies": [
        {
          "id": 2597726,
          "postDate": "2024-01-11T20:25:08.657Z",
          "content": "<p>It's not leaked test data. Those are files without paired spectrograms where I missed some cleanup.</p>",
          "rawMarkdown": "It's not leaked test data. Those are files without paired spectrograms where I missed some cleanup.",
          "votes": 9,
          "replies": [
            {
              "id": 2597732,
              "postDate": "2024-01-11T20:30:51.620Z",
              "content": "<p>Thanks for the conformation <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "rawMarkdown": "Thanks for the conformation @sohier "
            },
            {
              "id": 2666204,
              "postDate": "2024-02-24T08:01:47.247Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2596267,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-01-11T02:07:00.523000",
      "content": "<p>I don't think we are missing any data. You can</p>\n<pre><code>train = pd()\n )\n )\n</code></pre>\n<p>This prints 17089 and 11138 respectively which is the number of parquets files for each.</p>\n<p>There are only 1950 unique patients. So for EEG they spread the data over 17089 parquets. And for spectrograms, they spread out the data over 11138. </p>\n<p>When creating a single train sample, we load one EEG parquet and one spectrogram parquet. Then we extract 10_000 consecutive rows for EEG. And extract 300 consecutive rows for spectrogram. For each row in <code>train.csv</code>, the rows to extract from parquet files are indicated.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2596271,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-11T02:12:02.953000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Yes, we have all parquets of the train.csv's eeg_id</p>\n<p>but train_eegs folder have <strong>17_300</strong> parquets. did i calculated wrong way?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2596386,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2024-01-11T04:54:11.573000",
          "content": "<p>I can confirm that there are 17300 parquet files in the <code>train_eegs</code> folder. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, there seems to be 211 rows missing in <code>train.csv</code>; could you please clarify?</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2640887,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2024-02-07T06:12:18.703000",
      "content": "<p>First of all, thanks for sharing good information! <br>\nI have a question that we need to drop or pesudolabeling?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2597718,
      "author_name": "tmoroder",
      "author_url": "",
      "post_date": "2024-01-11T20:09:57.057000",
      "content": "<p>Leaked test data?? I downloaded them, just to make sure 😉</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2597726,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2024-01-11T20:25:08.657000",
          "content": "<p>It's not leaked test data. Those are files without paired spectrograms where I missed some cleanup.</p>",
          "votes": 9,
          "replies": [
            {
              "id": 2597732,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-11T20:30:51.620000",
              "content": "<p>Thanks for the conformation <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2666204,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-24T08:01:47.247000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2596260": "Total spectrogram_id in train.csv: **11138** & total parquet files are **11138**\n Total eeg_ids in train.csv: **17089** & total files are **17300**\n\nmissing eeg_ids in train.csv **211** ( **can be used as pseudolabels or is training data missing** ) ?\n```python\nmissing_eeg_ids = [100261680, 1006575126, 1046303268, 1054182773, 1080752861, 1090493279, 1102260141, 1106457838, 1114646196, 1124971078, 1129703637, 1138516159, 1152469719, 1164247904, 1210061638, 1213982942, 1218899486, 124842023, 1256854516, 1271026602, 1302255870, 1320529265, 1356605204, 1362609712, 1395846857, 1437515164, 1446608286, 1471761225, 1472250249, 1484488768, 1486503212, 1494048680, 1514837355, 1527948998, 1540422245, 1550636461, 159230712, 1593654895, 1644799805, 1667505910, 166924982, 1685149928, 176751182, 1781690069, 17823653, 1783811564, 1789492555, 1827043672, 1835829757, 1844412082, 1880066416, 1946169694, 1967500472, 1977522925, 1980230677, 2008191275, 2021270543, 2029271373, 2031389372, 2032669824, 2062845812, 2069539903, 2133352173, 2161239256, 2163060838, 2167910584, 2169828710, 217188535, 2177382144, 2187713476, 2215145921, 2229664974, 2232763157, 2250059220, 2256830545, 2360465057, 2370617920, 2402422559, 2437996061, 2465170355, 2495888973, 2512462400, 2514244198, 2518637163, 2524754765, 2537083119, 2546815772, 2564933177, 2566899625, 2585735707, 2619383113, 2644594529, 2704546408, 2722685303, 2734605314, 2753356177, 2770720245, 2780688155, 2827773679, 2847186843, 2883273759, 2890567875, 2915092170, 2928453358, 2940354800, 2945876125, 2966099404, 2980896188, 2989528583, 2995737301, 3029056160, 3107283960, 3110785252, 3130316564, 3159731110, 3184041077, 3184243660, 3219158996, 3231425203, 3248636378, 3277143842, 3279947709, 3287073479, 3296877755, 3302245698, 3334497580, 3374558757, 3395079519, 3395717621, 3425324038, 3427375889, 3462221095, 3463780576, 3469377987, 351886689, 3534133394, 3576774144, 358214261, 3588781445, 3613550741, 3624732686, 3672236418, 3677734263, 367926685, 3681177371, 3688844263, 3713571437, 3726556263, 3729617662, 3746528981, 3765156259, 37691174, 3782539999, 3784293316, 3808352981, 3844299098, 3860943081, 3861560461, 3864775587, 389650354, 3897110777, 3897137837, 390524827, 3932079802, 3963505952, 3964740894, 3965080689, 397155403, 3973402202, 4016243421, 4016271604, 4062219555, 4062860278, 4066543763, 4088014176, 4091126054, 4118489199, 416005567, 4187869644, 4263916735, 4265462352, 4283988928, 438882483, 452392893, 463419419, 495719000, 507760170, 512986705, 525392618, 529809133, 535701085, 581008851, 582052247, 592953741, 595031723, 60524973, 667978667, 710494757, 735478859, 757170164, 771666080, 789759143, 846376179, 875712244, 876702254, 880829874, 888227967, 914125289, 936728284, 968174006, 990852758]\n```\n\n### [Notebook - Missing Eeg_ids in Train.csv vs train_eegs parquet](https://www.kaggle.com/code/seshurajup/missing-eeg-ids-in-train-csv)\n\n## Those are files without paired spectrograms where I missed some cleanup. by @sohier",
    "2596267": "I don't think we are missing any data. You can\n\n    train = pd.read_csv('train.csv')\n    print( train.eeg_id.nunique() )\n    print( train.spectrogram_id.nunique() )\n\nThis prints 17089 and 11138 respectively which is the number of parquets files for each.\n\nThere are only 1950 unique patients. So for EEG they spread the data over 17089 parquets. And for spectrograms, they spread out the data over 11138. \n\nWhen creating a single train sample, we load one EEG parquet and one spectrogram parquet. Then we extract 10_000 consecutive rows for EEG. And extract 300 consecutive rows for spectrogram. For each row in `train.csv`, the rows to extract from parquet files are indicated.",
    "2640887": "First of all, thanks for sharing good information! \nI have a question that we need to drop or pesudolabeling?\n",
    "2597718": "Leaked test data?? I downloaded them, just to make sure 😉"
  }
}