{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data exploration: from zero to spectrograms","metadata":{}},{"cell_type":"markdown","source":"In this notebook we are gonna take a first look into the data and we will generate our first spectrogram 🧠","metadata":{}},{"cell_type":"code","source":"import os\n\nimport numpy as np\nimport pandas as pd\n\nimport matplotlib.pyplot as plt\nimport matplotlib.ticker as mticker","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:19.521991Z","iopub.execute_input":"2024-04-01T15:33:19.522280Z","iopub.status.idle":"2024-04-01T15:33:19.853859Z","shell.execute_reply.started":"2024-04-01T15:33:19.522258Z","shell.execute_reply":"2024-04-01T15:33:19.852743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kaggle_data_dirpath = '/kaggle/input/hms-harmful-brain-activity-classification' ","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:19.855286Z","iopub.execute_input":"2024-04-01T15:33:19.855641Z","iopub.status.idle":"2024-04-01T15:33:19.859813Z","shell.execute_reply.started":"2024-04-01T15:33:19.855618Z","shell.execute_reply":"2024-04-01T15:33:19.858863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ref_fname = 'train.csv'\ntrain_ref_fpath = os.path.join(kaggle_data_dirpath, train_ref_fname)\n\ntrain_ref_df = pd.read_csv(train_ref_fpath)\nTARGETS = train_ref_df.columns[-6:]\nprint('Train shape:', train_ref_df.shape )\nprint('Targets', list(TARGETS))\ntrain_ref_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:19.861269Z","iopub.execute_input":"2024-04-01T15:33:19.861579Z","iopub.status.idle":"2024-04-01T15:33:20.079548Z","shell.execute_reply.started":"2024-04-01T15:33:19.861553Z","shell.execute_reply":"2024-04-01T15:33:20.078770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the first rows of the table we can see that there are more than one eeg and spectrgoram data for each patient. In fact, from data description we know that each eeg (and the related spectrogram) were subsampled and analyzed. Moreover, the subsamples extracted can also overlap partially. The starting time of each sub-sample is written in the 'eeg_label_offset_seconds' and 'spectrogram_label_offset_seconds' columns.","metadata":{}},{"cell_type":"markdown","source":"Lets' try to open a random train spectrogram file (the first that I see in the train_sepctrograms folder):","metadata":{}},{"cell_type":"code","source":"spectrogram_id = 1000086677\n\ntrain_specs_dirpath = os.path.join(kaggle_data_dirpath, 'train_spectrograms')\n\nspec_df = pd.read_parquet(os.path.join(train_specs_dirpath, f'{spectrogram_id}.parquet'))\nspec_df","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.082779Z","iopub.execute_input":"2024-04-01T15:33:20.084019Z","iopub.status.idle":"2024-04-01T15:33:20.270885Z","shell.execute_reply.started":"2024-04-01T15:33:20.083984Z","shell.execute_reply":"2024-04-01T15:33:20.270271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The related row in the csv file is: ","metadata":{}},{"cell_type":"code","source":"train_row = train_ref_df.loc[train_ref_df['spectrogram_id'] == spectrogram_id]\ntrain_row","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.271949Z","iopub.execute_input":"2024-04-01T15:33:20.272342Z","iopub.status.idle":"2024-04-01T15:33:20.289520Z","shell.execute_reply.started":"2024-04-01T15:33:20.272320Z","shell.execute_reply":"2024-04-01T15:33:20.288821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, we know that in the spectrogram file we opened wil be only one sub-sample spectrogram. From data description we know that spectrogram are recorde in a 10 minute window (i.e. 600 secs) and in fact in the time column we see that the values go from 1 to 599.\n\nIn another case, we could have had more spectrogram sub-samples in the same spectrogram parquet file. In that case we should have utilized the offset label from the train.csv file. We can generalize this approach in the following code cell:","metadata":{}},{"cell_type":"code","source":"spec_offset = int(train_row.spectrogram_label_offset_seconds.iloc[0])\nspectrogram = spec_df.loc[(spec_df.time>=spec_offset)\n                     &(spec_df.time<spec_offset+600)]\nspectrogram","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.292738Z","iopub.execute_input":"2024-04-01T15:33:20.294413Z","iopub.status.idle":"2024-04-01T15:33:20.366941Z","shell.execute_reply.started":"2024-04-01T15:33:20.294380Z","shell.execute_reply":"2024-04-01T15:33:20.366182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's try to visualize the spectrogram","metadata":{}},{"cell_type":"markdown","source":"As data description says:\n\n> The column names indicate the frequency in hertz and the recording regions of the EEG electrodes. The latter are abbreviated as LL = left lateral; RL = right lateral; LP = left parasagittal; RP = right parasagittal.\n\nSo, there are 4 regions where the signal was recorded, and for each region different frequencies were analyzed. Since from the shape of the dataframe we know that there are 400 columns, and that presumably the same frequencies for each region were recorded, we know that a set of 100 column represents the spectrogram data for that region. We can think to the spectrogram as a classical mulitchannel image (e.g. RGB) where the channel dimension of the image doesn't represent the color but the data regarding a specific region. In this case we will have a 4-dimensional spectrogram.","metadata":{}},{"cell_type":"code","source":"## LL cols\nLL_cols_idxs = [i for i, col_name in enumerate(spec_df.columns) if 'LL' in col_name]\nprint(f'LL_cols_idxs: {LL_cols_idxs} \\n')\n## LL cols\nRL_cols_idxs = [i for i, col_name in enumerate(spec_df.columns) if 'RL' in col_name]\nprint(f'RL_cols_idxs: {RL_cols_idxs} \\n')\n## LL cols\nLP_cols_idxs = [i for i, col_name in enumerate(spec_df.columns) if 'LP' in col_name]\nprint(f'LP_cols_idxs: {LP_cols_idxs} \\n')\n## LL cols\nRP_cols_idxs = [i for i, col_name in enumerate(spec_df.columns) if 'RP' in col_name]\nprint(f'RP_cols_idxs: {RP_cols_idxs}')","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.368158Z","iopub.execute_input":"2024-04-01T15:33:20.368669Z","iopub.status.idle":"2024-04-01T15:33:20.383795Z","shell.execute_reply.started":"2024-04-01T15:33:20.368642Z","shell.execute_reply":"2024-04-01T15:33:20.382619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the results of the print statements above we are sure that each set of 100 columns represents the frequencise recorded in each region.","metadata":{}},{"cell_type":"markdown","source":"Let's analyze the spectrogram of a region for the considered sample:","metadata":{}},{"cell_type":"code","source":"spec_df.iloc[:, 1:101]","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.385083Z","iopub.execute_input":"2024-04-01T15:33:20.385370Z","iopub.status.idle":"2024-04-01T15:33:20.423819Z","shell.execute_reply.started":"2024-04-01T15:33:20.385339Z","shell.execute_reply":"2024-04-01T15:33:20.422676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First thing we can notice is that time dimension (represented by the row index going from 0 to 300) appears on the y-axis, while frequencies on the x-axis. This means that we need to transpose the data for a correct visualization","metadata":{}},{"cell_type":"code","source":"new_spec = spec_df.iloc[:, 0:101].values.T\n\nplt.imshow(new_spec)\n\nax = plt.gca()\nax.set_ylim(top=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.426496Z","iopub.execute_input":"2024-04-01T15:33:20.426752Z","iopub.status.idle":"2024-04-01T15:33:20.626835Z","shell.execute_reply.started":"2024-04-01T15:33:20.426731Z","shell.execute_reply":"2024-04-01T15:33:20.626047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can apply a log transformation to tackle the skewness of the data (they are too compressed on low frequencies):","metadata":{}},{"cell_type":"code","source":"log_spec = np.clip(new_spec, np.exp(-4), np.exp(8))\nlog_spec = np.log(log_spec)\n\nplt.imshow(log_spec)\nax = plt.gca()\nax.set_ylim(top=0, bottom=99)","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.630269Z","iopub.execute_input":"2024-04-01T15:33:20.630496Z","iopub.status.idle":"2024-04-01T15:33:20.833539Z","shell.execute_reply.started":"2024-04-01T15:33:20.630476Z","shell.execute_reply":"2024-04-01T15:33:20.832867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"On the y-axis we have the col index that goes from 0 to 100, that is not meaningful per se but we know that the higher the col idx, the higher the samplig frequency. So it may be useful, for a better interpretation, to flip the array, in order to have ascending frequency and not descending like it is now. We can salso replace the nan values with the zero frequency.","metadata":{}},{"cell_type":"code","source":"new_spec = np.nan_to_num(new_spec, nan=0.)\nlog_spec = np.clip(new_spec, np.exp(-4), np.exp(8))\nlog_spec = np.log(log_spec)\n\nplt.imshow(log_spec)\naxs = plt.gca()\nax.set_ylim(top=0, bottom=99)\n#axs.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x,_: f\"$10^{{{int(x)}}}$\"))\naxs.invert_yaxis()","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:20.834276Z","iopub.execute_input":"2024-04-01T15:33:20.834486Z","iopub.status.idle":"2024-04-01T15:33:21.035205Z","shell.execute_reply.started":"2024-04-01T15:33:20.834466Z","shell.execute_reply":"2024-04-01T15:33:21.034278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now change the labels on the y-axis to the more meaningful frequencies instead of the col_idxs","metadata":{}},{"cell_type":"code","source":"# creation of dict mapping row_idx to the relative frequency\n\ncol_idx_name_map = {i: col_name for i, col_name in enumerate(spec_df.iloc[:, 1:101].columns)}\ncol_idx_freq_map = {k: col_idx_name_map[k].split('_')[-1] \n                    for k in col_idx_name_map}\n\ncol_idx_freq_map","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:21.035995Z","iopub.execute_input":"2024-04-01T15:33:21.036226Z","iopub.status.idle":"2024-04-01T15:33:21.046872Z","shell.execute_reply.started":"2024-04-01T15:33:21.036204Z","shell.execute_reply":"2024-04-01T15:33:21.045878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(log_spec)\naxs = plt.gca()\nax.set_ylim(top=0, bottom=99)\naxs.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x,_: f\"Log_10({col_idx_freq_map[int(x)]})\" if int(x) in col_idx_freq_map else ''))\naxs.invert_yaxis()","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:21.048421Z","iopub.execute_input":"2024-04-01T15:33:21.048825Z","iopub.status.idle":"2024-04-01T15:33:21.252892Z","shell.execute_reply.started":"2024-04-01T15:33:21.048791Z","shell.execute_reply":"2024-04-01T15:33:21.251964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now combine the other regions recorded in the spectrogram by stacking them onto this first result:","metadata":{}},{"cell_type":"code","source":"idxs = [LL_cols_idxs, RL_cols_idxs, LP_cols_idxs, RP_cols_idxs]\n\nfinal_spectrogram = []\nfor idxs_lst in idxs:\n    start_idx, end_idx = idxs_lst[0], idxs_lst[-1]\n    new_spec = spec_df.iloc[:, start_idx:end_idx].values.T\n    new_spec = np.nan_to_num(new_spec, nan=0.)\n    log_spec = np.clip(new_spec,np.exp(-4),np.exp(8))\n    log_spec = np.log(log_spec)\n    final_spectrogram.append(log_spec)","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:21.254052Z","iopub.execute_input":"2024-04-01T15:33:21.254336Z","iopub.status.idle":"2024-04-01T15:33:21.262663Z","shell.execute_reply.started":"2024-04-01T15:33:21.254310Z","shell.execute_reply":"2024-04-01T15:33:21.261747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_spectrogram = np.array(final_spectrogram)\nfinal_spectrogram = np.swapaxes(final_spectrogram, 0, -1) # move region channel (ranging from 1 to 4) to the last dimension\nfinal_spectrogram = np.swapaxes(final_spectrogram, 0, 1) # tranposition of rows and columns as we saw before\nfinal_spectrogram.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:21.263854Z","iopub.execute_input":"2024-04-01T15:33:21.264204Z","iopub.status.idle":"2024-04-01T15:33:21.272534Z","shell.execute_reply.started":"2024-04-01T15:33:21.264139Z","shell.execute_reply":"2024-04-01T15:33:21.271648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"region_idx = 0\nplt.imshow(np.array(final_spectrogram)[..., region_idx]) # we can access the spectrogram of a specific recorded region using the last dimension","metadata":{"execution":{"iopub.status.busy":"2024-04-01T15:33:21.273555Z","iopub.execute_input":"2024-04-01T15:33:21.273851Z","iopub.status.idle":"2024-04-01T15:33:21.467214Z","shell.execute_reply.started":"2024-04-01T15:33:21.273826Z","shell.execute_reply":"2024-04-01T15:33:21.466586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}