{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30635,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"In this competition, EEGs and Spectrograms are supposed to be important data. Why EEG first? Because spectrograms are made from EEGs. We are going to delve down into EEG first, them look into Spectrograms in next notebook. ","metadata":{}},{"cell_type":"markdown","source":"### What is EEG?","metadata":{}},{"cell_type":"markdown","source":"EEG is the observation of the sum of the electrical activity of a population of neurons in the brain from a given direction. In the observation, electrodes are arranged according to a rule called the \"International 10-20 system\", Except for A1 and A2 are removed and EKG(ElectroCardiogram) is added.\nThe electrode arrangement in the International 10-20 system is as follows:","metadata":{}},{"cell_type":"markdown","source":"![electrogram](https://upload.wikimedia.org/wikipedia/commons/7/70/21_electrodes_of_International_10-20_system_for_EEG.svg)\n(by https://en.wikipedia.org/wiki/10%E2%80%9320_system_(EEG))","metadata":{}},{"cell_type":"markdown","source":"Based on the above, it is possible to express the positional relationship of the electrodes in a polar coordinate system. In the following plots, Hue is represented by θ in polar coordinates and r by cloma to represent the neighborhood relationship.","metadata":{}},{"cell_type":"code","source":"#location:[r, θ] , r, θ ∈ [0,1]\nelectrodes_map = {\"Cz\": [0, 0, 0],\n                  \"Fz\": [0.5, 0, 0], \"F3\": [0.5, 0.125, 0], \"C3\": [0.5, 0.25, 0], \"P3\": [0.5, 0.375, 0], \"Pz\": [0.5, 0.5, 0],\n                  \"P4\": [0.5, 0.625, 0], \"C4\": [0.5, 0.75, 0], \"F4\": [0.5, 0.875, 0],\n                  \"Fp1\": [1.0, 0.05, 0], \"F7\": [1.0, 0.15, 0], \"T3\": [1.0, 0.25, 0], \"T5\": [1.0, 0.35, 0], \"O1\": [1.0, 0.45, 0],\n                  \"O2\": [1.0, 0.55, 0], \"T6\": [1.0, 0.65, 0], \"T4\": [1.0, 0.75, 0], \"F8\": [1.0, 0.85, 0], \"Fp2\": [1.0, 0.95, 0],\n                  \"EKG\": [0,0, 1]}# might be better annotation for EKG","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:08:21.264404Z","iopub.execute_input":"2024-01-15T08:08:21.265302Z","iopub.status.idle":"2024-01-15T08:08:21.274763Z","shell.execute_reply.started":"2024-01-15T08:08:21.265256Z","shell.execute_reply":"2024-01-15T08:08:21.273562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport glob\nimport matplotlib.pyplot as plt\nimport matplotlib.cm as cm\nimport seaborn as sns\nimport matplotlib.colors as mcolors\n\n#gradient coloring\ndef collor(col_name):\n    if col_name==\"EKG\": return (0,0,0,0.1)#gray\n    val, hue, _ = electrodes_map[col_name]\n    if val==0.5: \n        val=0.4;\n    rgb = mcolors.hsv_to_rgb((hue,1, val))\n    if val==0.5: return np.append(rgb, 0.35)\n    return np.append(rgb, 0.2)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-15T08:08:21.276021Z","iopub.execute_input":"2024-01-15T08:08:21.276348Z","iopub.status.idle":"2024-01-15T08:08:21.288185Z","shell.execute_reply.started":"2024-01-15T08:08:21.276319Z","shell.execute_reply":"2024-01-15T08:08:21.286961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax1 = plt.subplots(nrows=1,ncols=1, figsize=(8, 8),tight_layout=True)#()30,30,\nfor loc, vec in electrodes_map.items(): #plot each columns\n    if loc==\"EKG\":continue\n    y = vec[0]*np.cos(vec[1]*2*np.pi)\n    x = -vec[0]*np.sin(vec[1]*2*np.pi)\n    ax1.scatter(x, y, color=collor(loc), label=loc, s=1000)\n    ax1.annotate(loc, xy=(x-0.04, y-0.02))\ntheta = np.linspace(0,2*np.pi,num=10000)\nfor ang in theta:\n    ax1.plot(1.3*np.sin(theta),1.3*np.cos(theta), color='k', lw=0.1)\nax1.set(title=f'Location and Color',ylabel='Back and Forth', xlabel='Left and Right')\nax1.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:08:21.290954Z","iopub.execute_input":"2024-01-15T08:08:21.292123Z","iopub.status.idle":"2024-01-15T08:08:52.448197Z","shell.execute_reply.started":"2024-01-15T08:08:21.292075Z","shell.execute_reply":"2024-01-15T08:08:52.447010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prep","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:08:52.449896Z","iopub.execute_input":"2024-01-15T08:08:52.450891Z","iopub.status.idle":"2024-01-15T08:08:52.740631Z","shell.execute_reply.started":"2024-01-15T08:08:52.450852Z","shell.execute_reply":"2024-01-15T08:08:52.739533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot function\ndef show_eeg(df_eeg,  title='', ylim=(-250,250), vote=False, filtering=False):\n    plt.rcParams[\"font.size\"] = 12#32\n    #set time axis\n    t_eeg = np.arange(0, df_eeg.shape[0]*0.01, 0.01) # step = 0.01s/slice, df_eeg.shape[0]: 5000 entries for 50 s EEG records, \n    \n    # Make fig and ax\n    fig, ax1 = plt.subplots(nrows=1,ncols=1, figsize=(10, 8),tight_layout=True)#()30,30,\n    \n    # EEG plotting\n    for i,col in enumerate(df_eeg): #plot each columns\n        y = df_eeg[col]\n        if filtering: \n            if (y.max()<filtering and y.min()>-filtering): continue\n        ax1.plot(t_eeg, y, color=collor(col), label=col)\n    ax1.set(title=title, xlabel='Time (s)', ylabel='intensity')\n    if ylim:ax1.set_ylim(ylim)\n    ax1.legend(loc='center left', bbox_to_anchor=(1, .5), ncols=1) # legend loc adjustment\n\n    plt.show()\n    if type(vote)!=bool:\n        print(f\"seizure_vote: {vote['seizure_vote']} | lpd_vote:  {vote['lpd_vote']} | gpd_vote:   {vote['gpd_vote']}\\nlrda_vote:    {vote['lrda_vote']} | grda_vote: {vote['grda_vote']} | other_vote: {vote['other_vote']}\")\n    plt.rcParams[\"font.size\"] = 12","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:51:49.375230Z","iopub.execute_input":"2024-01-15T08:51:49.375634Z","iopub.status.idle":"2024-01-15T08:51:49.386829Z","shell.execute_reply.started":"2024-01-15T08:51:49.375603Z","shell.execute_reply":"2024-01-15T08:51:49.385672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot","metadata":{}},{"cell_type":"markdown","source":"### Random","metadata":{}},{"cell_type":"code","source":"for i in range(2,-1,-1):\n    sample = train.sample(n=1, random_state=i).reset_index()\n    eeg_id = sample['eeg_id'][0]\n    offset_idx = int(sample['eeg_label_offset_seconds'][0]*100)\n    eeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg = eeg[offset_idx:offset_idx+5000]\n    vote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[0]\n\n    show_eeg(eeg, title=f'Random', vote=vote, ylim=False)","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:08:52.755631Z","iopub.execute_input":"2024-01-15T08:08:52.756072Z","iopub.status.idle":"2024-01-15T08:08:55.527689Z","shell.execute_reply.started":"2024-01-15T08:08:52.756030Z","shell.execute_reply":"2024-01-15T08:08:55.526869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conditions\n>Plot a few EEGs for each to get a sense of the characteristics in each condition.","metadata":{}},{"cell_type":"markdown","source":"## Seizure","metadata":{}},{"cell_type":"code","source":"sample = train.sort_values('seizure_vote', ascending=False).reset_index()\ni, count=0, 0\npatient_li=[]\nwhile count<3:\n    pat_id = sample['patient_id'][i]\n    if pat_id in patient_li:\n        i+=1\n        continue\n    patient_li.append(pat_id)\n    eeg_id = sample['eeg_id'][i]\n    offset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\n    eeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg = eeg[offset_idx:offset_idx+5000]\n    vote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n    \n    show_eeg(eeg, title=f'Random', vote=vote)\n    count+=1\n    i+=1","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:30:42.032867Z","iopub.execute_input":"2024-01-15T08:30:42.033336Z","iopub.status.idle":"2024-01-15T08:30:45.486894Z","shell.execute_reply.started":"2024-01-15T08:30:42.033289Z","shell.execute_reply":"2024-01-15T08:30:45.485895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LPD","metadata":{}},{"cell_type":"code","source":"sample = train.sort_values('lpd_vote', ascending=False).reset_index()\ni, count=0, 0\npatient_li=[]\nwhile count<3:\n    pat_id = sample['patient_id'][i]\n    if pat_id in patient_li:\n        i+=1\n        continue\n    patient_li.append(pat_id)\n    eeg_id = sample['eeg_id'][i]\n    offset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\n    eeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg = eeg[offset_idx:offset_idx+5000]\n    vote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n    \n    show_eeg(eeg, title=f'LPD patient: {pat_id}', vote=vote)\n    count+=1\n    i+=1","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:32:16.848598Z","iopub.execute_input":"2024-01-15T08:32:16.849018Z","iopub.status.idle":"2024-01-15T08:32:20.126874Z","shell.execute_reply.started":"2024-01-15T08:32:16.848987Z","shell.execute_reply":"2024-01-15T08:32:20.125599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GPD","metadata":{}},{"cell_type":"code","source":"sample = train.sort_values('gpd_vote', ascending=False).reset_index()\ni, count=0, 0\npatient_li=[]\nwhile count<3:\n    pat_id = sample['patient_id'][i]\n    if pat_id in patient_li:\n        i+=1\n        continue\n    patient_li.append(pat_id)\n    eeg_id = sample['eeg_id'][i]\n    offset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\n    eeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg = eeg[offset_idx:offset_idx+5000]\n    vote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n    \n    show_eeg(eeg, title=f'GPD patient: {pat_id}', vote=vote)\n    count+=1\n    i+=1","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:32:43.585436Z","iopub.execute_input":"2024-01-15T08:32:43.586184Z","iopub.status.idle":"2024-01-15T08:32:50.497673Z","shell.execute_reply.started":"2024-01-15T08:32:43.586146Z","shell.execute_reply":"2024-01-15T08:32:50.496687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LRDA","metadata":{}},{"cell_type":"code","source":"sample = train.sort_values('lrda_vote', ascending=False).reset_index()\ni, count=0, 0\npatient_li=[]\nwhile count<3:\n    pat_id = sample['patient_id'][i]\n    if pat_id in patient_li:\n        i+=1\n        continue\n    patient_li.append(pat_id)\n    eeg_id = sample['eeg_id'][i]\n    offset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\n    eeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg = eeg[offset_idx:offset_idx+5000]\n    vote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n    \n    show_eeg(eeg, title=f'LRDA patient: {pat_id}', vote=vote)\n    count+=1\n    i+=1","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:37:49.991584Z","iopub.execute_input":"2024-01-15T08:37:49.992029Z","iopub.status.idle":"2024-01-15T08:37:53.226105Z","shell.execute_reply.started":"2024-01-15T08:37:49.991996Z","shell.execute_reply":"2024-01-15T08:37:53.224966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GRDA","metadata":{}},{"cell_type":"code","source":"sample = train.sort_values('grda_vote', ascending=False).reset_index()\ni, count=0, 0\npatient_li=[]\nwhile count<3:\n    pat_id = sample['patient_id'][i]\n    if pat_id in patient_li:\n        i+=1\n        continue\n    patient_li.append(pat_id)\n    eeg_id = sample['eeg_id'][i]\n    offset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\n    eeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg = eeg[offset_idx:offset_idx+5000]\n    vote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n    \n    show_eeg(eeg, title=f'GRDA patient: {pat_id}', vote=vote)\n    count+=1\n    i+=1","metadata":{"execution":{"iopub.status.busy":"2024-01-15T08:48:22.062341Z","iopub.execute_input":"2024-01-15T08:48:22.062910Z","iopub.status.idle":"2024-01-15T08:48:30.359119Z","shell.execute_reply.started":"2024-01-15T08:48:22.062868Z","shell.execute_reply":"2024-01-15T08:48:30.358110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Filtering\n> Let's restrict the EEG displayed in the plot to electrodes whose maximum signal strength exceeds a certain one and see what the trend is.","metadata":{}},{"cell_type":"code","source":"def show_eeg_sub(ax1, df_eeg,  title='', ylim=(-250,250), vote=False, filtering=False):\n    plt.rcParams[\"font.size\"] = 12#32\n    #set time axis\n    t_eeg = np.arange(0, df_eeg.shape[0]*0.01, 0.01) # step = 0.01s/slice, df_eeg.shape[0]: 5000 entries for 50 s EEG records, \n    # EEG plotting\n    for i,col in enumerate(df_eeg): #plot each columns\n        y = df_eeg[col]\n        if filtering: \n            if (y.max()<filtering and y.min()>-filtering): continue\n        ax1.plot(t_eeg, y, color=collor(col), label=col)\n    ax1.set(title=title, xlabel='Time (s)', ylabel='intensity')\n    if ylim:ax1.set_ylim(ylim)\n    ax1.legend(loc='center left', bbox_to_anchor=(1, .5), ncols=1) # legend loc adjustment\n    if type(vote)!=bool:\n        print(f\"seizure_vote: {vote['seizure_vote']} | lpd_vote:  {vote['lpd_vote']} | gpd_vote:   {vote['gpd_vote']}\\nlrda_vote:    {vote['lrda_vote']} | grda_vote: {vote['grda_vote']} | other_vote: {vote['other_vote']}\")\n    plt.rcParams[\"font.size\"] = 12","metadata":{"execution":{"iopub.status.busy":"2024-01-15T09:21:50.319392Z","iopub.execute_input":"2024-01-15T09:21:50.319863Z","iopub.status.idle":"2024-01-15T09:21:50.329495Z","shell.execute_reply.started":"2024-01-15T09:21:50.319830Z","shell.execute_reply":"2024-01-15T09:21:50.328636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LPD vs GPD","metadata":{}},{"cell_type":"code","source":"fig, ((ax1, ax2), (ax3, ax4)) = plt.subplots(nrows=2,ncols=2, figsize=(15, 10),tight_layout=True)\n\nsample = train.sort_values('lpd_vote', ascending=False).reset_index()\ni=6\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax1, eeg, title=f'LPD patient: {pat_id}',  filtering=80, ylim=(-200,200))\n\nsample = train.sort_values('lpd_vote', ascending=False).reset_index()\ni=12\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax2, eeg, title=f'LPD patient: {pat_id}', filtering=160, ylim=(-200,200))\n\nsample = train.sort_values('gpd_vote', ascending=False).reset_index()\ni=2\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax3, eeg, title=f'GPD patient: {pat_id}', filtering=170, ylim=(-200,200))\n\nsample = train.sort_values('gpd_vote', ascending=False).reset_index()\ni=6\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax4, eeg, title=f'GPD patient: {pat_id}', filtering=400, ylim=(400,-400))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-15T09:30:25.922894Z","iopub.execute_input":"2024-01-15T09:30:25.923400Z","iopub.status.idle":"2024-01-15T09:30:29.785197Z","shell.execute_reply.started":"2024-01-15T09:30:25.923362Z","shell.execute_reply":"2024-01-15T09:30:29.784062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compared to LPD, GPD appears to overlap many signals in a spike. This seems to be the case given that GPD affects a wider range of brain waves.\nThe spectrograms were measured in four separate locations (LL, LR, RL, and RR), but it seems that features that account for resonance in a wider range are needed.","metadata":{}},{"cell_type":"markdown","source":"## LRDA vs GRDA","metadata":{}},{"cell_type":"code","source":"fig, ((ax1, ax2), (ax3, ax4)) = plt.subplots(nrows=2,ncols=2, figsize=(15, 10),tight_layout=True)\n\nsample = train.sort_values('lrda_vote', ascending=False).reset_index()\ni=0\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax1, eeg, title=f'LRDA patient: {pat_id}',  filtering=210, ylim=(-200,200))\n\ni=58\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax2, eeg, title=f'LRDA patient: {pat_id}', filtering=109, ylim=(-200,200))\n\nsample = train.sort_values('grda_vote', ascending=False).reset_index()\ni=1\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax3, eeg, title=f'GRDA patient: {pat_id}', filtering=210, ylim=(-200,200))\n\ni=3\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax4, eeg, title=f'GRDA patient: {pat_id}', filtering=200, ylim=(-200,200))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-15T09:42:36.395560Z","iopub.execute_input":"2024-01-15T09:42:36.395973Z","iopub.status.idle":"2024-01-15T09:42:39.375292Z","shell.execute_reply.started":"2024-01-15T09:42:36.395940Z","shell.execute_reply":"2024-01-15T09:42:39.374082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears difficult to distinguish between LRDA and GRDA, but there still appears to be more overlap of the spike areas in GRDA. Also, in GRDA, the behavior of EKG, represented in gray in the figure, might be important","metadata":{}},{"cell_type":"markdown","source":"## LPD vs LRDA","metadata":{}},{"cell_type":"code","source":"fig, ((ax1, ax2), (ax3, ax4)) = plt.subplots(nrows=2,ncols=2, figsize=(15, 10),tight_layout=True)\n\nsample = train.sort_values('lpd_vote', ascending=False).reset_index()\ni=6\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax1, eeg, title=f'LPD patient: {pat_id}',  filtering=80, ylim=(-200,200))\n\nsample = train.sort_values('lpd_vote', ascending=False).reset_index()\ni=12\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax2, eeg, title=f'LPD patient: {pat_id}', filtering=160, ylim=(-200,200))\n\nsample = train.sort_values('lrda_vote', ascending=False).reset_index()\ni=0\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax3, eeg, title=f'LRDA patient: {pat_id}',  filtering=210, ylim=(-200,200))\n\ni=58\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax4, eeg, title=f'LRDA patient: {pat_id}', filtering=109, ylim=(-200,200))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-15T10:03:04.752937Z","iopub.execute_input":"2024-01-15T10:03:04.753327Z","iopub.status.idle":"2024-01-15T10:03:07.092420Z","shell.execute_reply.started":"2024-01-15T10:03:04.753295Z","shell.execute_reply":"2024-01-15T10:03:07.091357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GPD vs GRDA","metadata":{}},{"cell_type":"code","source":"fig, ((ax1, ax2), (ax3, ax4)) = plt.subplots(nrows=2,ncols=2, figsize=(15, 10),tight_layout=True)\n\nsample = train.sort_values('gpd_vote', ascending=False).reset_index()\ni=2\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax1, eeg, title=f'GPD patient: {pat_id}', filtering=170, ylim=(-200,200))\n\nsample = train.sort_values('gpd_vote', ascending=False).reset_index()\ni=6\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax2, eeg, title=f'GPD patient: {pat_id}', filtering=400, ylim=(400,-400))\n\nsample = train.sort_values('grda_vote', ascending=False).reset_index()\ni=1\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax3, eeg, title=f'GRDA patient: {pat_id}', filtering=210, ylim=(-200,200))\n\ni=3\npat_id = sample['patient_id'][i]\neeg_id = sample['eeg_id'][i]\noffset_idx = int(sample['eeg_label_offset_seconds'][i]*100)\neeg = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\neeg = eeg[offset_idx:offset_idx+5000]\nvote = sample[['seizure_vote' ,'lpd_vote' ,'gpd_vote', 'lrda_vote', 'grda_vote' ,'other_vote']].iloc[i]\n\nshow_eeg_sub(ax4, eeg, title=f'GRDA patient: {pat_id}', filtering=200, ylim=(-200,200))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-15T10:07:08.132339Z","iopub.execute_input":"2024-01-15T10:07:08.132785Z","iopub.status.idle":"2024-01-15T10:07:11.737124Z","shell.execute_reply.started":"2024-01-15T10:07:08.132750Z","shell.execute_reply":"2024-01-15T10:07:11.735989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen in [my previous notebook](https://www.kaggle.com/code/shunsukekikuchi/fast-eda), periodic discharges (LPD and GPD) vs rhythmic delta activity (LRDA, GRDA and seizures) will be distinguished by frequency (spectrograms). Though it's a guess...","metadata":{"execution":{"iopub.status.busy":"2024-01-15T10:04:10.499994Z","iopub.execute_input":"2024-01-15T10:04:10.500454Z","iopub.status.idle":"2024-01-15T10:04:10.510083Z","shell.execute_reply.started":"2024-01-15T10:04:10.500416Z","shell.execute_reply":"2024-01-15T10:04:10.508632Z"}}},{"cell_type":"markdown","source":"# To be Continued","metadata":{}},{"cell_type":"markdown","source":"- EDA on Spectrograms, probably more important than EEGs...?\n- Baseline: GDBT with some feature discovery might be a good start\n\nI will continue to update this notebook for my EDA. If you have any suggestions, please leave a comment!","metadata":{}}]}