{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30626,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# HMS - Data Inspection\n\nLet's inspect and understand the data first. Navigate to each section for summary of findings.\n\n**Comments welcome!**\n\n\n## Table of Contents\n- [train.csv](#train.csv)\n- [train_eegs](#train_eegs)\n- [train_spectrograms](#train_spectrograms)\n- [test data](#test-data)\n- [sample_submission.csv](#sample_submission.csv)","metadata":{}},{"cell_type":"code","source":"import os\nimport pathlib\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-01-12T15:48:03.482378Z","iopub.execute_input":"2024-01-12T15:48:03.482935Z","iopub.status.idle":"2024-01-12T15:48:03.970442Z","shell.execute_reply.started":"2024-01-12T15:48:03.482861Z","shell.execute_reply":"2024-01-12T15:48:03.969009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_dir = pathlib.Path(\"/kaggle/input/hms-harmful-brain-activity-classification\")\nbase_dir","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:03.973055Z","iopub.execute_input":"2024-01-12T15:48:03.973551Z","iopub.status.idle":"2024-01-12T15:48:03.983988Z","shell.execute_reply.started":"2024-01-12T15:48:03.973520Z","shell.execute_reply":"2024-01-12T15:48:03.982608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir(base_dir)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:03.986207Z","iopub.execute_input":"2024-01-12T15:48:03.986735Z","iopub.status.idle":"2024-01-12T15:48:03.996243Z","shell.execute_reply.started":"2024-01-12T15:48:03.986689Z","shell.execute_reply":"2024-01-12T15:48:03.994949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# train.csv\n- Each row corresponds to a sample of an EEG recording and respective spectrogram.\n- Eeg and spec id are referring to longer recordings, for judgement the experts only used parts of these, that can be inferred by the given offsets and which are enumerated by the sub_ids.\n- For the judgement, apparently, the following is used \n    - 50 second subsample of an EEG recording\n    - 10 minute center matched spectrograms\n        - **How is this center matched, in particular for those that have offset=0?**\n    - Expert labelled the central 10 seconds.\n        - **What is labelled, EEG or spec, or both? Must look into discussions.**\n- Combination of eeg_id + eeg_sub_id is a unique identifier.\n- The 50 second intervals overlap but are not completely identical.\n- A eeg_id has at most 1 spectrogram, but a spectrogram can cover multiple eeg_id.\n    - **How can this work? Is EEG stopped in between?**\n- Labels are well balanced.\n- Number of expert votes differs.\n- Expert agreement as described on the Overview page is not given (idealized, proto, edge), but can be (partially) calculated. Uncertainty might be worth to consider!\n","metadata":{}},{"cell_type":"code","source":"path_train = base_dir / \"train.csv\"\ntrain = pd.read_csv(path_train, dtype={\"eeg_id\": \"str\", \"spectrogram_id\": \"str\"})\ntrain","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:03.998075Z","iopub.execute_input":"2024-01-12T15:48:03.998452Z","iopub.status.idle":"2024-01-12T15:48:04.435208Z","shell.execute_reply.started":"2024-01-12T15:48:03.998421Z","shell.execute_reply":"2024-01-12T15:48:04.434073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"eeg_id\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:04.439182Z","iopub.execute_input":"2024-01-12T15:48:04.440433Z","iopub.status.idle":"2024-01-12T15:48:04.484253Z","shell.execute_reply.started":"2024-01-12T15:48:04.440385Z","shell.execute_reply":"2024-01-12T15:48:04.482967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Duplicates in id + sub_id:\", train[[\"eeg_id\", \"eeg_sub_id\"]].duplicated().any())\nprint(\"Duplicates in id + offset_seconds:\", train[[\"eeg_id\", \"eeg_label_offset_seconds\"]].duplicated().any())\nprint(\"Max number of spec by eeg_id:\", train.groupby(\"eeg_id\")[\"spectrogram_id\"].nunique().max())\nprint(\"Max number of eeg_id by spec:\", train.groupby(\"spectrogram_id\")[\"eeg_id\"].nunique().max())\nprint(\"Max number of patient_id by spec:\", train.groupby(\"spectrogram_id\")[\"patient_id\"].nunique().max())\nprint(\"Max number of patient_id by eeg_id:\", train.groupby(\"eeg_id\")[\"patient_id\"].nunique().max())","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:04.486027Z","iopub.execute_input":"2024-01-12T15:48:04.486342Z","iopub.status.idle":"2024-01-12T15:48:04.723691Z","shell.execute_reply.started":"2024-01-12T15:48:04.486316Z","shell.execute_reply":"2024-01-12T15:48:04.722320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec_id = train.groupby(\"spectrogram_id\")[\"eeg_id\"].nunique().idxmax()\ntrain.query(\"spectrogram_id == @spec_id\")","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:04.728033Z","iopub.execute_input":"2024-01-12T15:48:04.728444Z","iopub.status.idle":"2024-01-12T15:48:04.805305Z","shell.execute_reply.started":"2024-01-12T15:48:04.728412Z","shell.execute_reply":"2024-01-12T15:48:04.804137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"expert_consensus\"].value_counts().plot(kind='bar')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:04.806465Z","iopub.execute_input":"2024-01-12T15:48:04.806797Z","iopub.status.idle":"2024-01-12T15:48:05.140393Z","shell.execute_reply.started":"2024-01-12T15:48:04.806769Z","shell.execute_reply":"2024-01-12T15:48:05.139093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vote_cols = [x for x in train.columns if \"vote\" in x]\ntrain[vote_cols].sum(axis=1).value_counts().sort_index()","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.141985Z","iopub.execute_input":"2024-01-12T15:48:05.142844Z","iopub.status.idle":"2024-01-12T15:48:05.185527Z","shell.execute_reply.started":"2024-01-12T15:48:05.142806Z","shell.execute_reply":"2024-01-12T15:48:05.184312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# train_eegs\n- There are more eeg signals in the train folder than in the train CSV.\n    - See [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467058).\n- Sampling frequency: 200 Hz\n- Columns are those from 10-20-system and EKG signal.\n- Not sure what is the best way to look at the data. In the example pdf they look to be differences.","metadata":{}},{"cell_type":"code","source":"sample_rate = 200","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.187736Z","iopub.execute_input":"2024-01-12T15:48:05.188714Z","iopub.status.idle":"2024-01-12T15:48:05.194995Z","shell.execute_reply.started":"2024-01-12T15:48:05.188668Z","shell.execute_reply":"2024-01-12T15:48:05.193454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eeg_dir = base_dir / \"train_eegs\"\neeg_dir","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.196603Z","iopub.execute_input":"2024-01-12T15:48:05.197017Z","iopub.status.idle":"2024-01-12T15:48:05.211274Z","shell.execute_reply.started":"2024-01-12T15:48:05.196980Z","shell.execute_reply":"2024-01-12T15:48:05.209810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eeg_ids_folder = set(x.stem for x in eeg_dir.glob(\"*.parquet\"))\neeg_ids_train = set(train[\"eeg_id\"].unique())\nprint(\"eeg_id_train == eeg_id_folder:\", eeg_ids_train == eeg_ids_folder)\nprint(\"all train eeg_id present:\", eeg_ids_train.issubset(eeg_ids_folder))\ntoo_much_eegs = sorted(eeg_ids_folder.difference(eeg_ids_train))\nprint(\"too many in folder:\", len(too_much_eegs))\nprint(\"too many in folder:\", too_much_eegs)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.212595Z","iopub.execute_input":"2024-01-12T15:48:05.213125Z","iopub.status.idle":"2024-01-12T15:48:05.590198Z","shell.execute_reply.started":"2024-01-12T15:48:05.213091Z","shell.execute_reply":"2024-01-12T15:48:05.588858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rec = train.iloc[1]\nrec","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.592497Z","iopub.execute_input":"2024-01-12T15:48:05.593721Z","iopub.status.idle":"2024-01-12T15:48:05.605290Z","shell.execute_reply.started":"2024-01-12T15:48:05.593672Z","shell.execute_reply":"2024-01-12T15:48:05.603533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_eeg = eeg_dir / f\"{rec.eeg_id}.parquet\"\neeg = pd.read_parquet(path_eeg)\neeg[\"time\"] = eeg.index / sample_rate\neeg","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.611042Z","iopub.execute_input":"2024-01-12T15:48:05.611475Z","iopub.status.idle":"2024-01-12T15:48:05.892289Z","shell.execute_reply.started":"2024-01-12T15:48:05.611440Z","shell.execute_reply":"2024-01-12T15:48:05.891074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eeg.columns","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.894180Z","iopub.execute_input":"2024-01-12T15:48:05.894952Z","iopub.status.idle":"2024-01-12T15:48:05.903845Z","shell.execute_reply.started":"2024-01-12T15:48:05.894887Z","shell.execute_reply":"2024-01-12T15:48:05.902372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"i_start = int(rec.eeg_label_offset_seconds) * sample_rate\ni_stop = i_start + 50 * sample_rate\nfig, axs = plt.subplots(nrows=20, figsize=(16, 10), tight_layout=True, sharex=True, gridspec_kw={\"hspace\": 0})\neeg.loc[i_start:i_stop].set_index(\"time\").plot(subplots=True, ax=axs)\nfor ax in axs.flat:\n    ax.legend(loc=\"upper right\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:05.906166Z","iopub.execute_input":"2024-01-12T15:48:05.907176Z","iopub.status.idle":"2024-01-12T15:48:12.423308Z","shell.execute_reply.started":"2024-01-12T15:48:05.907132Z","shell.execute_reply":"2024-01-12T15:48:12.421798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# train_spectrograms\n- All specs present, and none too much!\n- Column names indicate frequency and position/region.\n- Unsure about scale of spectrogram. Might require a separate analysis.","metadata":{}},{"cell_type":"code","source":"spec_dir = base_dir / \"train_spectrograms\"\nspec_dir","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:12.425125Z","iopub.execute_input":"2024-01-12T15:48:12.425565Z","iopub.status.idle":"2024-01-12T15:48:12.433625Z","shell.execute_reply.started":"2024-01-12T15:48:12.425528Z","shell.execute_reply":"2024-01-12T15:48:12.432638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec_ids_folder = set(x.stem for x in spec_dir.glob(\"*.parquet\"))\nspec_ids_train = set(train[\"spectrogram_id\"].unique())\nprint(\"spec_ids_train == spec_ids_folder:\", spec_ids_train == spec_ids_folder)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:12.435139Z","iopub.execute_input":"2024-01-12T15:48:12.435716Z","iopub.status.idle":"2024-01-12T15:48:12.811266Z","shell.execute_reply.started":"2024-01-12T15:48:12.435684Z","shell.execute_reply":"2024-01-12T15:48:12.809975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rec","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:12.812735Z","iopub.execute_input":"2024-01-12T15:48:12.813103Z","iopub.status.idle":"2024-01-12T15:48:12.821050Z","shell.execute_reply.started":"2024-01-12T15:48:12.813071Z","shell.execute_reply":"2024-01-12T15:48:12.819887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_spec = spec_dir / f\"{rec.spectrogram_id}.parquet\"\nspec = pd.read_parquet(path_spec)\nspec = spec.set_index(\"time\")\nspec","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:12.823087Z","iopub.execute_input":"2024-01-12T15:48:12.823600Z","iopub.status.idle":"2024-01-12T15:48:12.933286Z","shell.execute_reply.started":"2024-01-12T15:48:12.823557Z","shell.execute_reply":"2024-01-12T15:48:12.931986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# columns: region (str) x freq (float)\ncolumns = spec.columns.str.split(\"_\", expand=True)\ncolumns = pd.MultiIndex.from_tuples(\n    [(x[0], float(x[1])) for x in columns], names=[\"region\", \"freq\"]\n)\nspec.columns = columns\nspec = spec.T\nspec","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:12.935137Z","iopub.execute_input":"2024-01-12T15:48:12.935636Z","iopub.status.idle":"2024-01-12T15:48:12.996023Z","shell.execute_reply.started":"2024-01-12T15:48:12.935591Z","shell.execute_reply":"2024-01-12T15:48:12.994931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regions = list(spec.index.get_level_values(0).unique())\nregions","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:12.997531Z","iopub.execute_input":"2024-01-12T15:48:12.997861Z","iopub.status.idle":"2024-01-12T15:48:13.006923Z","shell.execute_reply.started":"2024-01-12T15:48:12.997832Z","shell.execute_reply":"2024-01-12T15:48:13.005498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.Series(spec.loc[\"LL\"].values.ravel()).describe()","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:13.008368Z","iopub.execute_input":"2024-01-12T15:48:13.008735Z","iopub.status.idle":"2024-01-12T15:48:13.028648Z","shell.execute_reply.started":"2024-01-12T15:48:13.008706Z","shell.execute_reply":"2024-01-12T15:48:13.027517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regions","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:13.030190Z","iopub.execute_input":"2024-01-12T15:48:13.030500Z","iopub.status.idle":"2024-01-12T15:48:13.037790Z","shell.execute_reply.started":"2024-01-12T15:48:13.030472Z","shell.execute_reply":"2024-01-12T15:48:13.036631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=2, ncols=2, sharex=\"all\", sharey=\"all\", tight_layout=True, figsize=(16, 10))\nfor region, ax in zip(regions, axs.flat):\n    df = spec.loc[region]\n    times = df.columns\n    freqs = df.index\n    ax.pcolormesh(times, freqs, df.values, cmap=\"viridis\")\n    ax.set_title(region)\naxs[0,0].set_ylabel(\"freq\")\naxs[1,0].set_ylabel(\"freq\")\naxs[1,0].set_xlabel(\"time\")\naxs[1,1].set_xlabel(\"time\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:48:13.039312Z","iopub.execute_input":"2024-01-12T15:48:13.039856Z","iopub.status.idle":"2024-01-12T15:48:14.624961Z","shell.execute_reply.started":"2024-01-12T15:48:13.039821Z","shell.execute_reply":"2024-01-12T15:48:14.623642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# test data\n\n## test.csv\n- Only ids of eeg, spec and patient given. No further information.\n\n## test_eegs, test_spectrograms\n- Single eeg is exactly 50 seconds long.\n- Single spectrogram is 10 minutes long.","metadata":{}},{"cell_type":"code","source":"path_test = base_dir / \"test.csv\"\ntest = pd.read_csv(path_test, dtype={\"eeg_id\": \"str\", \"spectrogram_id\": \"str\"})\ntest","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:52:14.886548Z","iopub.execute_input":"2024-01-12T15:52:14.886954Z","iopub.status.idle":"2024-01-12T15:52:14.904693Z","shell.execute_reply.started":"2024-01-12T15:52:14.886923Z","shell.execute_reply":"2024-01-12T15:52:14.903244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_eeg_dir = base_dir / \"test_eegs\"\nos.listdir(test_eeg_dir)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:51:20.763178Z","iopub.execute_input":"2024-01-12T15:51:20.763641Z","iopub.status.idle":"2024-01-12T15:51:20.779129Z","shell.execute_reply.started":"2024-01-12T15:51:20.763604Z","shell.execute_reply":"2024-01-12T15:51:20.778111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_eeg = test_eeg_dir / f\"{test.iloc[0].eeg_id}.parquet\"\neeg = pd.read_parquet(path_eeg)\neeg","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:58:41.176359Z","iopub.execute_input":"2024-01-12T15:58:41.176845Z","iopub.status.idle":"2024-01-12T15:58:41.228704Z","shell.execute_reply.started":"2024-01-12T15:58:41.176811Z","shell.execute_reply":"2024-01-12T15:58:41.227473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"EEG is 50 seconds:\", len(eeg) == 50 * sample_rate)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:58:42.175240Z","iopub.execute_input":"2024-01-12T15:58:42.175762Z","iopub.status.idle":"2024-01-12T15:58:42.183187Z","shell.execute_reply.started":"2024-01-12T15:58:42.175725Z","shell.execute_reply":"2024-01-12T15:58:42.181447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_spec_dir = base_dir / \"test_spectrograms\"\nos.listdir(test_spec_dir)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:58:42.916446Z","iopub.execute_input":"2024-01-12T15:58:42.916937Z","iopub.status.idle":"2024-01-12T15:58:42.927106Z","shell.execute_reply.started":"2024-01-12T15:58:42.916887Z","shell.execute_reply":"2024-01-12T15:58:42.925336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_spec = test_spec_dir / f\"{test.iloc[0].spectrogram_id}.parquet\"\nspec = pd.read_parquet(path_spec)\nspec = spec.set_index(\"time\")\nspec","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:58:46.154054Z","iopub.execute_input":"2024-01-12T15:58:46.154542Z","iopub.status.idle":"2024-01-12T15:58:46.265749Z","shell.execute_reply.started":"2024-01-12T15:58:46.154505Z","shell.execute_reply":"2024-01-12T15:58:46.264317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Spectrogram is 10 minutes:\", len(spec) == 10 * 30)","metadata":{"execution":{"iopub.status.busy":"2024-01-12T15:59:28.073113Z","iopub.execute_input":"2024-01-12T15:59:28.074249Z","iopub.status.idle":"2024-01-12T15:59:28.081180Z","shell.execute_reply.started":"2024-01-12T15:59:28.074191Z","shell.execute_reply":"2024-01-12T15:59:28.079433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# sample_submission.csv\n- Labelled by eeg_id.\n- Probability of individual classes.\n    - **Can we exploit that we know that 3-20 annotators have labelled?**","metadata":{}},{"cell_type":"code","source":"path_submission = base_dir / \"sample_submission.csv\"\nsubmission = pd.read_csv(path_submission)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2024-01-12T16:06:15.347046Z","iopub.execute_input":"2024-01-12T16:06:15.347624Z","iopub.status.idle":"2024-01-12T16:06:15.366980Z","shell.execute_reply.started":"2024-01-12T16:06:15.347583Z","shell.execute_reply":"2024-01-12T16:06:15.365825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}