{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"The goal of this competition is to detect and classify seizures and other types of harmful brain activity. From electroencephalography (EEG) signals recorded from critically ill hospital patients, we must build a classifier that is able to assign for a given EEG time series the correct class based on the detection and recognition of one of 6 patterns:\n1. seizure (SZ)\n2. generalized periodic discharges (GPD)\n3. lateralized periodic discharges (LPD)\n4. lateralized rhythmic delta activity (LRDA)\n5. generalized rhythmic delta activity (GRDA)\n6. “other” for all signals that do not fit inside above categories\nAll patterns have specific characteristics based on their frequency, the existence of time gap between pulse, their temporal extent and their location. ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"We will:\n1. Look at training metadata and understand what information each column gives\n2. Look at an example EEG data, all the signals contained in it, what are the characteristic and plot them on time axis\n3. Try to filter the EKG signal of this example record, with the help of scipy signal module and fft module","metadata":{}},{"cell_type":"markdown","source":"### setup and imports","metadata":{}},{"cell_type":"code","source":"import os\n\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport scipy.fft as fft\nimport scipy.signal as signal\nimport seaborn as sns\n\nmatplotlib.rcParams['font.family'] = 'sans-serif'\nmatplotlib.rcParams['figure.figsize'] = (10, 6)","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:17.833880Z","iopub.execute_input":"2024-05-14T12:46:17.834243Z","iopub.status.idle":"2024-05-14T12:46:20.313084Z","shell.execute_reply.started":"2024-05-14T12:46:17.834208Z","shell.execute_reply":"2024-05-14T12:46:20.311823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Load metadata","metadata":{}},{"cell_type":"code","source":"dataset_path = \"/kaggle/input/hms-harmful-brain-activity-classification\"\ntraining_metadata_path = os.path.join(dataset_path, \"train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:20.315329Z","iopub.execute_input":"2024-05-14T12:46:20.315919Z","iopub.status.idle":"2024-05-14T12:46:20.321909Z","shell.execute_reply.started":"2024-05-14T12:46:20.315880Z","shell.execute_reply":"2024-05-14T12:46:20.320780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_meta = pd.read_csv(training_metadata_path)\nprint(f\"Training metadata shape: {training_meta.shape}\")\ntraining_meta.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:20.323491Z","iopub.execute_input":"2024-05-14T12:46:20.324163Z","iopub.status.idle":"2024-05-14T12:46:20.677046Z","shell.execute_reply.started":"2024-05-14T12:46:20.324127Z","shell.execute_reply":"2024-05-14T12:46:20.675983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### these are the columns -:\nprint(\"Columns: \")\ntraining_meta.columns.to_list()","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:20.679676Z","iopub.execute_input":"2024-05-14T12:46:20.680641Z","iopub.status.idle":"2024-05-14T12:46:20.688799Z","shell.execute_reply.started":"2024-05-14T12:46:20.680598Z","shell.execute_reply":"2024-05-14T12:46:20.687759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### some important insights from the train data csv file","metadata":{}},{"cell_type":"markdown","source":"1. There are 106_800 rows in our training data, which represents labeled eeg subsample\n \n2. Each row has an id corresponding to the EEG and spectrogram data, which is the data we'll use to make our predictions.\n \n3. lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote are our target columns, which must be probabilities\n \n4. Each row has sub_id and offset_seconds for EEG and spectrogram data. During inference, we'll obtain 50-second-long subsample of EEG and 10-minute-long subsample of the spectrogram to make a prediction. Therefore, target values are applicable for this subsamples, but not for the entire EEG and spectrograms\n \n5. The patient_id column would be benefitial to make a train/val/test split\n \n6. expert_consensus will be the useful indicator of contraversial samples. When consensus is low, the row is, probably, the row is probably going to be dropped.","metadata":{}},{"cell_type":"markdown","source":"### comments","metadata":{}},{"cell_type":"code","source":"print(\"For each eeg and spectrogram, there is a unique patient\")\nprint(training_meta.groupby(\"eeg_id\").patient_id.nunique().value_counts())\nprint(training_meta.groupby(\"spectrogram_id\").patient_id.nunique().value_counts())\n\nprint()\nprint(\"-\" * 80)\nprint(\"But one patient can be recorded several times\")\nprint(training_meta.groupby(\"patient_id\").eeg_id.nunique().sort_values(ascending=False))\nprint(training_meta.groupby(\"patient_id\").spectrogram_id.nunique().sort_values(ascending=False))","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:20.690380Z","iopub.execute_input":"2024-05-14T12:46:20.691023Z","iopub.status.idle":"2024-05-14T12:46:20.745734Z","shell.execute_reply.started":"2024-05-14T12:46:20.690986Z","shell.execute_reply":"2024-05-14T12:46:20.744520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of patterns is rather balanced\nfig = plt.figure()\nax = sns.countplot(\n    x=\"expert_consensus\",\n    data=training_meta,\n    hue=\"expert_consensus\",\n    dodge=False\n)\nax.get_legend().get_frame().set_alpha(0.6)\nplt.title(\"Expert consensus distribution\")\nplt.grid()","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:20.747350Z","iopub.execute_input":"2024-05-14T12:46:20.748066Z","iopub.status.idle":"2024-05-14T12:46:21.521287Z","shell.execute_reply.started":"2024-05-14T12:46:20.747983Z","shell.execute_reply":"2024-05-14T12:46:21.520152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Inspect number of patients","metadata":{}},{"cell_type":"code","source":"print(f\"Number of patients: {training_meta['patient_id'].nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:21.523499Z","iopub.execute_input":"2024-05-14T12:46:21.524014Z","iopub.status.idle":"2024-05-14T12:46:21.533065Z","shell.execute_reply.started":"2024-05-14T12:46:21.523975Z","shell.execute_reply":"2024-05-14T12:46:21.531631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. eeg_id: A unique identifier for the entire EEG recording. Those identifies also the parquet files of the eeg in the train_eegs folder. Note that as each row denominates a particular subsample of an EEG, there can be several rows with same eeg_id.\n2. eeg_sub_id: An ID for the specific 50 second long subsample this row's labels apply to. eeg_sub_id are assigned in chronological order for each eeg_id beginning with 0. The subsample is shifted from the beginning of the eeg_id record by eeg_label_offset_seconds. Note that even if subsample is 50 seconds long, the annotation is done by looking at the central 10 seconds.\n3. eeg_label_offset_seconds: The time between the beginning of the consolidated EEG and this subsample. Time shift.\n4. spectrogram_id: A unique identifier for an entire spectrogram. A spectrogram can regroup several EEGs records. Note that spectrogram_id are name of spectrogram files in folder train_spectrograms. For more information on what is a spectrogram.\n5. spectrogram_sub_id: An ID for the specific 10 minute subsample this row's labels apply to. Same concept as for eeg_sub_id.\n6. spectogram_label_offset_seconds: The time between the beginning of the consolidated spectrogram and this subsample. Simple time shift.\n7. label_id: An ID for this set of labels. Unique for each row.\n8. patient_id: An ID for the patient who donated the data. An eeg_id or spectrogram_id is associated with a single patient_id but a single patient can be associated with several eegs and spectrograms.\n9. expert_consensus: The consensus annotator label. Provided for convenience only. Give the final decision for classifying of the 10 seconds central window of a given subsample in one of the 6 available category.\n10. [seizure/lpd/gpd/lrda/grda/other]_vote: The count of annotator votes for a given brain activity class. The full names of the activity classes are as follows: lpd: lateralized periodic discharges, gpd: generalized periodic discharges, lrd: lateralized rhythmic delta activity, and grda: generalized rhythmic delta activity. Note that total number of annotators can vary for each row, even within the same eeg. Size of cohort of experts range from 1 to 28.","metadata":{}},{"cell_type":"markdown","source":"### Explore train EEGs","metadata":{}},{"cell_type":"code","source":"print(\"Train eeg files: \")\n! ls /kaggle/input/hms-harmful-brain-activity-classification/train_eegs | wc -l\nprint(\"Size of train eegs\")\n! du -sh /kaggle/input/hms-harmful-brain-activity-classification/train_eegs","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:21.534742Z","iopub.execute_input":"2024-05-14T12:46:21.535283Z","iopub.status.idle":"2024-05-14T12:46:57.128460Z","shell.execute_reply.started":"2024-05-14T12:46:21.535243Z","shell.execute_reply":"2024-05-14T12:46:57.127194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_train_eeg = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/1000913311.parquet\")\nsample_train_eeg","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.130125Z","iopub.execute_input":"2024-05-14T12:46:57.130463Z","iopub.status.idle":"2024-05-14T12:46:57.312286Z","shell.execute_reply.started":"2024-05-14T12:46:57.130433Z","shell.execute_reply":"2024-05-14T12:46:57.311346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 10_000 rows and 20 columns. Each column represent value from some specific electrode, which placed on it's specific location. Rows represent values from this electrodes over time. The frequency is 200 samples (rows) per second. So, this is a result of 50-second brain activity tracking.","metadata":{}},{"cell_type":"markdown","source":"### let's look at an example eeg","metadata":{}},{"cell_type":"markdown","source":"### Recording parameters","metadata":{}},{"cell_type":"code","source":"fs = 200  # data sampled at 200 Hz\nwindow_length = 10  # 10 seconds central window\nsubsample_length = 50  # 50 seconds subsampling for each row","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.316228Z","iopub.execute_input":"2024-05-14T12:46:57.316562Z","iopub.status.idle":"2024-05-14T12:46:57.321692Z","shell.execute_reply.started":"2024-05-14T12:46:57.316520Z","shell.execute_reply":"2024-05-14T12:46:57.320466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Choose particular eeg_id and load eeg","metadata":{}},{"cell_type":"code","source":"eeg_path_template = dataset_path + \"/train_eegs/{eeg_id}.parquet\"","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.323236Z","iopub.execute_input":"2024-05-14T12:46:57.323690Z","iopub.status.idle":"2024-05-14T12:46:57.333461Z","shell.execute_reply.started":"2024-05-14T12:46:57.323660Z","shell.execute_reply":"2024-05-14T12:46:57.332434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we chose this id as it is associated with a heavily recognized pattern \neeg_id = 2900632927\ntraining_meta[training_meta.eeg_id == eeg_id]","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.334931Z","iopub.execute_input":"2024-05-14T12:46:57.335406Z","iopub.status.idle":"2024-05-14T12:46:57.366116Z","shell.execute_reply.started":"2024-05-14T12:46:57.335277Z","shell.execute_reply":"2024-05-14T12:46:57.364987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_eeg = pd.read_parquet(eeg_path_template.format(eeg_id=eeg_id))\nprint(\"EEG record shape: \", df_eeg.shape)\ndf_eeg.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.367581Z","iopub.execute_input":"2024-05-14T12:46:57.367998Z","iopub.status.idle":"2024-05-14T12:46:57.424718Z","shell.execute_reply.started":"2024-05-14T12:46:57.367968Z","shell.execute_reply":"2024-05-14T12:46:57.423718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_eeg.columns.to_list())","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.425904Z","iopub.execute_input":"2024-05-14T12:46:57.426177Z","iopub.status.idle":"2024-05-14T12:46:57.431718Z","shell.execute_reply.started":"2024-05-14T12:46:57.426155Z","shell.execute_reply":"2024-05-14T12:46:57.430623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 20 signals, 19 corresponding to the EEG measures, which are voltage fluctuations recorded by electrodes placed on the scalp and a last one corresponding to EKG, which is an electrocardiogram of the patient.\n\nThe EEG measures give a view of the electrical activity of the brain. The different signals names (Fp1, F3, ...) corresponds to standardized locations where the electrodes will be attached. The locations are generally associated with a particular area of the brain. The letters design a particular lobe, such as pre-frontal (Fp), frontal (F), temporal (T), parietal (P), occipital (O), and central (C). Even numbers in the electrode names indicate it is placed on the right side of the head. Odd number indicate it is placed on left side. z (for zero) indicate the electrode is placed in the central plane and serves as reference point. For more information, you can see the 10–20 system (EEG)).\n\nThe unit of measurement is  μV\n .","metadata":{}},{"cell_type":"markdown","source":"## add time","metadata":{}},{"cell_type":"code","source":"df_eeg[\"time\"] = df_eeg.index / fs\ndf_eeg.set_index(\"time\", inplace=True)\ndf_eeg.index","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.432908Z","iopub.execute_input":"2024-05-14T12:46:57.433228Z","iopub.status.idle":"2024-05-14T12:46:57.448911Z","shell.execute_reply.started":"2024-05-14T12:46:57.433195Z","shell.execute_reply":"2024-05-14T12:46:57.447782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_eeg","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.450170Z","iopub.execute_input":"2024-05-14T12:46:57.450490Z","iopub.status.idle":"2024-05-14T12:46:57.481387Z","shell.execute_reply.started":"2024-05-14T12:46:57.450456Z","shell.execute_reply":"2024-05-14T12:46:57.480310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### plot one electrode","metadata":{}},{"cell_type":"code","source":"plt.plot(df_eeg.index, df_eeg[\"Fp1\"])\nplt.xlabel(\"Time (s)\")\nplt.ylabel(\"Amplitude (uV)\")\nplt.grid()\nplt.title(\"Fp1\")","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:57.482637Z","iopub.execute_input":"2024-05-14T12:46:57.482951Z","iopub.status.idle":"2024-05-14T12:46:58.022145Z","shell.execute_reply.started":"2024-05-14T12:46:57.482925Z","shell.execute_reply":"2024-05-14T12:46:58.020946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### We can focus on a particular window of time.","metadata":{}},{"cell_type":"code","source":"x_start = 25.0\nx_end = 35.0\nplt.plot(df_eeg.loc[x_start:x_end].index, df_eeg.loc[x_start:x_end][\"Fp1\"])\nplt.grid()\nplt.title(\"Fp1\")\nplt.xlabel(\"Time (s)\")\nplt.ylabel(\"Amplitude (uV)\")","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:58.023586Z","iopub.execute_input":"2024-05-14T12:46:58.024068Z","iopub.status.idle":"2024-05-14T12:46:58.344682Z","shell.execute_reply.started":"2024-05-14T12:46:58.023948Z","shell.execute_reply":"2024-05-14T12:46:58.343614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### whole set of electrodes","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(20, figsize=(10, 100))\n\n# Generate a line plot for each column in the DataFrame\nfor i, column in enumerate(sample_train_eeg.columns):\n    ax[i].plot(sample_train_eeg.index, sample_train_eeg[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\n# plt.legend()\n# plt.title('Simulated Data Line Chart')\n# plt.xlabel('Index')\n# plt.ylabel('Values')\n# plt.grid(True)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:46:58.346076Z","iopub.execute_input":"2024-05-14T12:46:58.346487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Data has a lot of peaks and troughs, indicating a high variation\n2. The scale of most of the plots ranging from -200 to 100, excluding EKG chart\n3. The line seems to fluctuate above and below a central value, which seems to be around zero value","metadata":{}},{"cell_type":"markdown","source":"### Explore train spectrograms","metadata":{}},{"cell_type":"code","source":"train_spectrogram = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/1000086677.parquet\")\ntrain_spectrogram","metadata":{"execution":{"iopub.status.idle":"2024-05-14T12:47:06.273397Z","shell.execute_reply.started":"2024-05-14T12:47:06.186206Z","shell.execute_reply":"2024-05-14T12:47:06.272444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_spectrogram(spectrogram_path):\n    sample_spect = pd.read_parquet(spectrogram_path)\n    \n    split_spect = {\n        \"LL\": sample_spect.filter(regex='^LL', axis=1),\n        \"RL\": sample_spect.filter(regex='^RL', axis=1),\n        \"RP\": sample_spect.filter(regex='^RP', axis=1),\n        \"LP\": sample_spect.filter(regex='^LP', axis=1),\n    }\n    \n    fig, axes = plt.subplots(nrows=2, ncols=2, figsize=(15, 12))\n    axes = axes.flatten()\n    label_interval = 5\n    for i, split_name in enumerate(split_spect.keys()):\n        ax = axes[i]\n        img = ax.imshow(np.log(split_spect[split_name]).T, cmap='viridis', aspect='auto', origin='lower')\n        cbar = fig.colorbar(img, ax=ax)\n        cbar.set_label('Log(Value)')\n        ax.set_title(split_name)\n        ax.set_ylabel(\"Frequency (Hz)\")\n        ax.set_xlabel(\"Time\")\n\n        ax.set_yticks(np.arange(len(split_spect[split_name].columns)))\n        ax.set_yticklabels([column_name[3:] for column_name in split_spect[split_name].columns])\n        frequencies = [column_name[3:] for column_name in split_spect[split_name].columns]\n        ax.set_yticks(np.arange(0, len(split_spect[split_name].columns), label_interval))\n        ax.set_yticklabels(frequencies[::label_interval])\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:47:06.274822Z","iopub.execute_input":"2024-05-14T12:47:06.275217Z","iopub.status.idle":"2024-05-14T12:47:06.286980Z","shell.execute_reply.started":"2024-05-14T12:47:06.275181Z","shell.execute_reply":"2024-05-14T12:47:06.285872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_spectrogram(\"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/1000189855.parquet\")","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:48:30.083069Z","iopub.execute_input":"2024-05-14T12:48:30.083527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## test eeg","metadata":{}},{"cell_type":"code","source":"print(\"Test eeg files: \")\n! ls /kaggle/input/hms-harmful-brain-activity-classification/test_eegs | wc -l\nprint(\"Size of test eegs\")\n! du -sh /kaggle/input/hms-harmful-brain-activity-classification/test_eegs","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:48:41.972694Z","iopub.execute_input":"2024-05-14T12:48:41.973103Z","iopub.status.idle":"2024-05-14T12:48:44.020084Z","shell.execute_reply.started":"2024-05-14T12:48:41.973070Z","shell.execute_reply":"2024-05-14T12:48:44.018864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### train Spectograms","metadata":{}},{"cell_type":"markdown","source":"## spectpgram data","metadata":{}},{"cell_type":"code","source":"print(\"Train spectogram files: \")\n! ls /kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms | wc -l\nprint(\"Size of train spectograms\")\n! du -sh /kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms","metadata":{"execution":{"iopub.status.busy":"2024-05-14T12:58:00.220711Z","iopub.execute_input":"2024-05-14T12:58:00.221480Z","iopub.status.idle":"2024-05-14T12:58:30.348664Z","shell.execute_reply.started":"2024-05-14T12:58:00.221445Z","shell.execute_reply":"2024-05-14T12:58:30.347566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Test spectogram files: \")\n! ls /kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms | wc -l\nprint(\"Size of test spectograms\")\n! du -sh /kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms","metadata":{"execution":{"iopub.status.busy":"2024-05-14T13:05:21.126280Z","iopub.execute_input":"2024-05-14T13:05:21.126701Z","iopub.status.idle":"2024-05-14T13:05:23.193797Z","shell.execute_reply.started":"2024-05-14T13:05:21.126668Z","shell.execute_reply":"2024-05-14T13:05:23.192645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"print(\"Train spectogram files: \")\n! ls /kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms | wc -l\nprint(\"Size of train spectograms\")\n! du -sh /kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms","metadata":{},"execution_count":null,"outputs":[]}]}