{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7766563,"sourceType":"datasetVersion","datasetId":4519089}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**The goal and motivation of this notebook**  \n\nWhen training DNN models using spectrogram, I tried various data augmentations, but there was only a slight improvement in both LB and CV. Therefore, as an alternative, I tried to expand it by random sampling when [creating EEG Spectrograms](https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg).  \nI thought about creating a spectrogram from all the eeg data and then doing random sampling, but I decided that the amount of data would be too large and would be difficult to read.\n\nThe eeg parquet files have various lengths, and the original EEG specs are only data for the middle 50 seconds, so I thought random sampling would be effective.I think this technique may also work for kaggle_specs. Regarding kaggle_specs, I will try reducing the sampling width from 10 mins to 50 seconds and then perform random sampling.    \nI haven't verified whether this is the right approach yet, but I'm posting this to get feedback from kagglers. Please feel free to comment if you like!   \n \nPlease also upvote if it helps \n\n**Rough steps for random sampling**\n1. Group the df by eeg_id and expert_consensus.  \nSince there is data with different conclusions even within the same eeg_id, the characteristics may change depending on the observation time even within a unique eeg parquet file. -> I decided that it should be treated as separate data.\n\n2. Create observation time for each group in seconds.   \nthe observation time for a single conclusion is not necessarily continuous because there is data in which the conclusion changes in the middle of the electroencephalogram data.  \n  \n3. From each observation data, a total of 50 seconds is acquired by random sampling using three time windows.  \nThe length of the time window is randomized and adjusted so that the total time window is 50 seconds. Each acquired element should not be duplicated.    \n\n4. Create multiple samples depending on the data length. (Increase the number of items created for long observation times)  \nOverlapping with another random sampling from the same eeg  \n\nConditional branching on data length:\n* 50 - 100 (seconds)...Create one sampling data\n* 100 - 150 (seconds)...Create 2\n* 150 - 200 (seconds)...Create 3\n* 200~ (seconds)...Create 4\n\n\n5. Create spectrograms from the acquired eeg data. The created specs are labeled with eeg_id, expert_consensus, and dup_count (duplication count).  \n\n\n**[Here](https://www.kaggle.com/datasets/suzukisatsuki/eeg-spectrogram-by-extended-data) is the dataset I created**","metadata":{}},{"cell_type":"markdown","source":"**このノートブックの動機と目的**  \n\nspectrogramを用いたDNNモデルの訓練において、様々なData Augmentationを試行しましたがLB,CVともにわずかな向上のみでした。そこで、代替案として[EEG Spectorgramsの作成](https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg)の際にrandom samplingによる拡張を試みました。  \n全ての脳波データを利用してeeg_spectrogramsを作成してからランダムサンプリングをすることも考えましたが、データ量が膨大になる可能性を考えて今回は観測時間からサンプルを取得した後にspectrogramを作成しました。    \n\nこの試みは、それぞれのeeg parquetファイルが持つデータの長さが一意でなく、かなりのばらつきがあることと、元のEEG Specsは各ファイルの真ん中50秒のみにフォーカスしたデータであることから有用な方法ではないかと判断しました。この手法はkaggle_specsに対しても有効な可能性があると考えています。kaggle_specsに関してはサンプリングの幅を10分から50秒に縮小するなどしてから試行しようと思います。  \nこれが正しいアプローチかはまだ検証できていませんが、kagglerの皆さんに意見を頂戴したいと思い投稿しました。もしよければぜひコメントしてください！ \n\n役に立った場合、良ければupvoteしていただければ励みになります！\n\n**randam samplingの大まかな手順**\n1. データをeeg_idとexpert_consensusでグループ化する。  \n※同一eeg_id内でも結論の異なるデータが存在するため、一意のeeg parquet file内でも観測時刻によって特徴が変化すると思われる。＝＞別個のデータとして切り分けるべきだと判断しました。  \n\n2. グループごとに観測時間を秒ごとに作成。  \nこれは脳波データの途中で結論が変更されるデータが内在しているため、単一の結論の観測時間が必ず連続であるとは限らないためである。  \n\n3. 上記の処理によってグループごとに長さの異なる観測時間が取得できる。それぞれのデータから、3つの時間窓を用いてランダムサンプリングによって合計50秒のデータを取得。  \n※同一eeg file内で複数のspec作成の際はサンプリングごとの重複は許すようにしました。  \n  \n4. データ長に応じて複数のサンプリングを行う。（観測時間が長いものに対して作成個数を増やす）  \n※時間窓の長さはランダムかつ時間窓の合計が50秒になるように調整。時間窓の取得要素は重複しないようにする。   \n\nデータ長の条件分岐：  \n* 50～100(秒)・・・1つのサンプリングデータ作成  \n* 100～150（秒）・・・2つ作成  \n* 150～200（秒）・・・3つ作成  \n* 200～（秒）・・・4つ作成  \n\n\n5. 取得したeeg dataより、spectrogramを作成。作成したspecsはeeg_idとexpert_consensus、dup_count(重複回数)でラベリング。  \n\n作成したdatasetは[こちら](https://www.kaggle.com/datasets/suzukisatsuki/eeg-spectrogram-by-extended-data)です。","metadata":{}},{"cell_type":"code","source":"import pandas as pd, numpy as np, os\nimport matplotlib.pyplot as plt, gc\n\nSAVE_INDIVIDUAL_EEGS = False\nSAVE_EEGS = False\n\ndf = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\nTARGETS = df.columns[-6:]\nprint('Train shape', df.shape )\ndisplay( df.head() )\n\npd.set_option('display.max_columns', None)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:42:18.583250Z","iopub.execute_input":"2024-03-05T11:42:18.583652Z","iopub.status.idle":"2024-03-05T11:42:19.922780Z","shell.execute_reply.started":"2024-03-05T11:42:18.583618Z","shell.execute_reply":"2024-03-05T11:42:19.921659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# step1 , 2\n**step1 : Group train_df by eeg_id and expert_consensus**  \n**step2 : Create observation time for each group in seconds**\n","metadata":{}},{"cell_type":"code","source":"\n# 各行に対して観測時間の連続値と重複カウントを取得\ntrain = df.copy()\ntrain['dup_count'] = train.groupby(['eeg_id', 'expert_consensus']).cumcount()\ntrain['observation_time'] = train['eeg_label_offset_seconds'].apply(lambda x: list(range(int(x),int(x)+50)))\n\n# 「expert_consensus, eeg_id」でグルーピングして、グループごとに「observation_time」の和集合を取得\nelapsed_df = train.groupby(['expert_consensus', 'eeg_id'])['observation_time'].apply(lambda x: list(set().union(*x))).reset_index()\n\n# 観測の合計時間を算出\nelapsed_df['length_t'] = elapsed_df['observation_time'].apply(lambda x: len(x))\n\n# trainに観測時間、観測時間の長さを結合\ntrain = train.iloc[:,:-1].merge(elapsed_df, on=['expert_consensus', 'eeg_id'], how='left')\n\n# targetsをグループごとに足し合わせてから正規化\ntmp = df.groupby(['eeg_id','expert_consensus'])[TARGETS].agg('sum').reset_index()\ntrain = train.drop(columns=TARGETS).merge(tmp, on=['eeg_id','expert_consensus'], how='left')\n\ny_data = train[TARGETS].values\ny_data = y_data / y_data.sum(axis=1,keepdims=True)\ntrain[TARGETS] = y_data\n\ntrain.reset_index(drop=True, inplace=True)\n\nelapsed_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:46:30.075567Z","iopub.execute_input":"2024-03-05T11:46:30.076537Z","iopub.status.idle":"2024-03-05T11:46:31.426461Z","shell.execute_reply.started":"2024-03-05T11:46:30.076488Z","shell.execute_reply":"2024-03-05T11:46:31.425299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ３つまで重複を許した時のデータ量の推移\n# Change in data amount when up to 3 duplications are allowed\n\ntrain_ = train.copy()\ntrain_ = train_[train_['dup_count']<= 3]\ntrain_['dup_count'].value_counts().cumsum().plot(kind='bar')\nplt.xlabel('dup cumsum count')","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:46:34.901435Z","iopub.execute_input":"2024-03-05T11:46:34.901849Z","iopub.status.idle":"2024-03-05T11:46:35.164455Z","shell.execute_reply.started":"2024-03-05T11:46:34.901818Z","shell.execute_reply":"2024-03-05T11:46:35.163396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* The above trends are contained all duplicate rows, regardless of observation time.  \n* Add conditional extraction regarding length_t to the above duplication.  \n(The longer the observation time, the more duplication will be adjusted.)  \n  \n※上の推移は観測時間によらず全ての重複行を足し合わせた数である。  \n　ランダムサンプリングによるデータの拡張を効果的にしたいので、上記の重複に対してlength_tについての条件抽出を加える  \n（観測時間が長いほど、重複を許容するように調整する。）","metadata":{}},{"cell_type":"code","source":"# 「length_t」が50～100（未満）の時、重複はなし\ntrain =train[~((train['length_t']>=50) & (train['length_t']<100) & (train['dup_count']>0))]\n\n# 「length_t」が100～150（未満）の時、重複を1度許容する\ntrain =train[~((train['length_t']>=100) & (train['length_t']<150) & (train['dup_count']>1))]\n\n# 「length_t」が150～200（未満）の時、重複を2度許容する\ntrain =train[~((train['length_t']>=150) & (train['length_t']<200) & (train['dup_count']>2))]\n\n# 「length_t」が200以上の時、重複を3度許容する\ntrain =train[~((train['length_t']>=200) & (train['dup_count']>3))]\n\nprint('train shape is : ', train.shape)\n\n# 重複回数ごとのデータ量を可視化\ntrain['dup_count'].value_counts().plot(kind='bar')\nfor i, value in enumerate(train['dup_count'].value_counts().sort_index()):\n    plt.annotate(str(value), xy=(i, value), ha='center', va='bottom')\ntrain.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:47:08.813868Z","iopub.execute_input":"2024-03-05T11:47:08.814231Z","iopub.status.idle":"2024-03-05T11:47:09.078632Z","shell.execute_reply.started":"2024-03-05T11:47:08.814202Z","shell.execute_reply":"2024-03-05T11:47:09.077797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # データ内にどの程度観測結果の異なるeeg_idが存在するか確認\n# data = elapsed_df.groupby(['eeg_id'])['expert_consensus'].nunique()\n# print('Multi targets eeg_ids:', len(data[data>1]))","metadata":{"execution":{"iopub.status.busy":"2024-03-04T05:33:05.169705Z","iopub.execute_input":"2024-03-04T05:33:05.170054Z","iopub.status.idle":"2024-03-04T05:33:05.182450Z","shell.execute_reply.started":"2024-03-04T05:33:05.170027Z","shell.execute_reply":"2024-03-04T05:33:05.181472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# step3 : Create a function for random_sampling( +Visualazation)","metadata":{}},{"cell_type":"code","source":"# randam_sampling用の関数作成\n# サンプリングの対象は、各グループの全観測時間の配列であり、非連続な観測時間の配列にも対応\n\nimport numpy as np\n\ndef random_sampling(time_series, total_width, num_windows):\n    sampled_windows = []\n    \n    for window in range(num_windows):\n    \n        # ランダムに時間窓の幅を選ぶ\n        if window == num_windows-1 or total_width<=15:\n            window_width = total_width\n        else:\n            window_width = np.random.randint(5, 16)\n            \n        # 開始位置をランダムに選ぶ\n        start_index = np.random.randint(0, len(time_series) - window_width + 1)\n        # 選択された時間窓を追加\n        sampled_windows.append(time_series[start_index : start_index + window_width])\n        \n        # 選択された時間窓に対応する部分を元のデータから削除\n        time_series = np.delete(time_series, range(start_index, start_index + window_width))\n        # 合計時間窓の幅を調整\n        total_width = total_width - window_width\n        \n    # 配列を結合、昇順に並べなおす\n    sampled_windows = np.sort(np.concatenate(sampled_windows, axis=0))\n    \n    return sampled_windows\n","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:47:42.872716Z","iopub.execute_input":"2024-03-05T11:47:42.873102Z","iopub.status.idle":"2024-03-05T11:47:42.881413Z","shell.execute_reply.started":"2024-03-05T11:47:42.873069Z","shell.execute_reply":"2024-03-05T11:47:42.880271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # 仮の時系列データを作成\n# time_series_data = np.arange(1000)\n\n# # 合計が50になるようなランダムな時間窓を4つサンプリング\n# sampled_windows = random_sampling(time_series_data, total_width=50, num_windows=4)\n\n# # # 結果を表示\n# # for i, window in enumerate(sampled_windows):\n# #     print(f\"Time Window {i + 1}: {window}\")\n# plt.plot(sampled_windows)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ランダムサンプリングの可視化(expert_consensusを複数持つ脳波データを可視化)\nimport matplotlib\n\n# 複数のconsensus持つ脳波データの中で、対応する適当なeeg_idを取得\neeg_ids = elapsed_df.groupby(['eeg_id'])['expert_consensus'].nunique()\nmulti_cons_id = np.random.choice(eeg_ids[eeg_ids>1].index.to_list())\n\nsample_eeg = multi_cons_id\n# ファイルをdfとして取得\neeg_df = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{multi_cons_id}.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:47:50.872286Z","iopub.execute_input":"2024-03-05T11:47:50.873282Z","iopub.status.idle":"2024-03-05T11:47:51.045277Z","shell.execute_reply.started":"2024-03-05T11:47:50.873238Z","shell.execute_reply":"2024-03-05T11:47:51.044494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # 全観測時間が100秒を超えるeegの可視化\n# eeg_ids = elapsed_df[elapsed_df['length_t']>= 100]['eeg_id'].unique()\n# sample_eeg = np.random.choice(eeg_ids)\n# eeg_df = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{sample_eeg}.parquet')\n# sample_eeg","metadata":{"execution":{"iopub.status.busy":"2024-03-05T02:14:56.792947Z","iopub.execute_input":"2024-03-05T02:14:56.793407Z","iopub.status.idle":"2024-03-05T02:14:56.853571Z","shell.execute_reply.started":"2024-03-05T02:14:56.793362Z","shell.execute_reply":"2024-03-05T02:14:56.852462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20,8))\n\ncolors = ['red', 'orange', 'yellow', 'green', 'white', 'blue', 'black']\n\n# 「eeg_id,expert_consensus」でグループ化された適当なデータからのランダムサンプリングを実行して可視化\nfor i, (_, row) in enumerate(elapsed_df[elapsed_df['eeg_id']==sample_eeg].iterrows()):\n    \n    plt.subplot(4,1,i+1)\n    eeg_df['Fp1'].plot()    #任意の特徴量を描画\n    \n    data = np.array(row['observation_time'])   # 観測時間をnumpy配列で表現\n    print('observation range: ', len(data))\n    target = row['expert_consensus']\n\n    # 最大3つの時間窓（長さはランダム）で観測時間を取得\n    window = random_sampling(data, total_width=50, num_windows=3)\n    print(window)\n    indices = np.concatenate([np.arange(ind*200, (ind+1)*200) for ind in window],axis=0)\n    print('sample_size: ', len(indices))\n    \n    for index in indices:\n        plt.axvline(x=index, color=colors[i],linestyle='-',alpha=0.02,\n                linewidth=1,)\n    plt.title(f'EEG:{sample_eeg}- target:{target}')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:47:54.960804Z","iopub.execute_input":"2024-03-05T11:47:54.961481Z","iopub.status.idle":"2024-03-05T11:48:15.799266Z","shell.execute_reply.started":"2024-03-05T11:47:54.961429Z","shell.execute_reply":"2024-03-05T11:48:15.798444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note: Although I wrote that there are a maximum of three time windows, the number of consecutive time windows after sampling may be four or more.**  \n(Because in a single time window, it is possible to obtain a list of non-consecutive observation times.)  \n\n**注意点：時間窓が最大3つと書きましたが、サンプリング後の連続する時間幅の個数は4つ以上になる可能性もあります。**  \n（単一の時間窓で、非連続な観測時間のリストを取得する可能性があるためです。）","metadata":{}},{"cell_type":"markdown","source":"# step4 : Create multiple samples depending on the data length","metadata":{}},{"cell_type":"code","source":"NAMES = ['LL','LP','RP','RR']\n\nFEATS = [['Fp1','F7','T3','T5','O1'],\n         ['Fp1','F3','C3','P3','O1'],\n         ['Fp2','F8','T4','T6','O2'],\n         ['Fp2','F4','C4','P4','O2']]","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:48:30.071236Z","iopub.execute_input":"2024-03-05T11:48:30.072033Z","iopub.status.idle":"2024-03-05T11:48:30.077856Z","shell.execute_reply.started":"2024-03-05T11:48:30.071993Z","shell.execute_reply":"2024-03-05T11:48:30.076829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pywt\nprint(\"The wavelet functions we can use:\")\nprint(pywt.wavelist())\n\nUSE_WAVELET = None #or \"db8\" or anything below","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-03-05T11:48:32.530186Z","iopub.execute_input":"2024-03-05T11:48:32.530582Z","iopub.status.idle":"2024-03-05T11:48:32.721079Z","shell.execute_reply.started":"2024-03-05T11:48:32.530545Z","shell.execute_reply":"2024-03-05T11:48:32.719874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def maddest(d, axis=None):\n    return np.mean(np.absolute(d - np.mean(d, axis)), axis)\n\ndef denoise(x, wavelet='haar', level=1):    \n    coeff = pywt.wavedec(x, wavelet, mode=\"per\")\n    sigma = (1/0.6745) * maddest(coeff[-level])\n\n    uthresh = sigma * np.sqrt(2*np.log(len(x)))\n    coeff[1:] = (pywt.threshold(i, value=uthresh, mode='hard') for i in coeff[1:])\n\n    ret=pywt.waverec(coeff, wavelet, mode='per')\n    \n    return ret","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:48:34.731234Z","iopub.execute_input":"2024-03-05T11:48:34.732348Z","iopub.status.idle":"2024-03-05T11:48:34.738944Z","shell.execute_reply.started":"2024-03-05T11:48:34.732311Z","shell.execute_reply":"2024-03-05T11:48:34.737954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" \nimport librosa\n\ndef spectrogram_from_eeg(row, display=False):\n    \n    dir_path = '/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/'\n    eeg_id = row['eeg_id']\n    parquet_path = os.path.join(dir_path, f'{eeg_id}.parquet')\n    \n    data = np.array(row['observation_time'])\n    window = random_sampling(data, total_width=50, num_windows=3)\n    indices = np.concatenate([np.arange(ind*200, (ind+1)*200) for ind in window],axis=0)\n    \n    target = row['expert_consensus']\n    dup_count = row['dup_count']\n    \n    eeg = pd.read_parquet(parquet_path)\n    eeg = eeg[eeg.index.isin(indices)]\n\n#     if offset is None:\n#         middle = (len(eeg)-10_000)//2\n#         eeg = eeg.iloc[middle:middle+10_000]\n#     else:\n#         eeg = eeg.iloc[offset:offset+10_000]\n    \n    # VARIABLE TO HOLD SPECTROGRAM\n    img = np.zeros((128,256,4),dtype='float32')\n    \n    if display: plt.figure(figsize=(10,7))\n    signals = []\n    for k in range(4):\n        COLS = FEATS[k]\n        \n        for kk in range(4):\n        \n            # COMPUTE PAIR DIFFERENCES\n            x = eeg[COLS[kk]].values - eeg[COLS[kk+1]].values\n\n            # FILL NANS\n            m = np.nanmean(x)\n            if np.isnan(x).mean() < 1: \n                x = np.nan_to_num(x,nan=m)\n            else: x[:] = 0\n\n            # DENOISE\n            if USE_WAVELET:\n                x = denoise(x, wavelet=USE_WAVELET)\n            signals.append(x)\n\n            # RAW SPECTROGRAM\n            mel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, \n                  n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)\n\n            # LOG TRANSFORM\n            width = (mel_spec.shape[1]//32)*32\n            mel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width]\n\n            # STANDARDIZE TO -1 TO 1\n            mel_spec_db = (mel_spec_db+40)/40 \n            img[:,:,k] += mel_spec_db\n                \n        # AVERAGE THE 4 MONTAGE DIFFERENCES\n        img[:,:,k] /= 4.0\n        \n        if display:\n            plt.subplot(2,2,k+1)\n            plt.imshow(img[:,:,k],aspect='auto',origin='lower')\n            plt.title(f'EEG {eeg_id}-{target}-{d_count} - Spectrogram {NAMES[k]}')\n            \n    if display: \n        plt.show()\n        plt.figure(figsize=(10,5))\n        offset = 0\n        for k in range(4):\n            if k>0: offset -= signals[3-k].min()\n            plt.plot(range(10_000),signals[k]+offset,label=NAMES[3-k])\n            offset += signals[3-k].max()\n        plt.legend()\n        plt.title(f'EEG {eeg_id}-{target}-d{dup_count} Signals')\n        plt.show()\n        print(); print('#'*25); print()\n        \n    return img\n\nprint(\"The wavelet functions we can use:\")\nprint(pywt.wavelist())","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:48:37.419481Z","iopub.execute_input":"2024-03-05T11:48:37.419883Z","iopub.status.idle":"2024-03-05T11:48:37.443322Z","shell.execute_reply.started":"2024-03-05T11:48:37.419848Z","shell.execute_reply":"2024-03-05T11:48:37.442538Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# step5 : Create spectrograms from the acquired eeg data","metadata":{}},{"cell_type":"code","source":"%%time\n\nif SAVE_INDIVIDUAL_EEGS:\n    directory_path = 'EEG_Spectrograms/'\n    if not os.path.exists(directory_path):\n        os.makedirs(directory_path)\n\nPATH = '/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/'\nDISPLAY = 4\nall_eegs = {}\n\nfor i,(idx, row) in enumerate(train.iterrows()):\n    if (i%100==0)&(i!=0): print(i,', ',end='')\n    \n    \n    eeg_id = row['eeg_id']\n    target = row['expert_consensus']\n    d_count = row['dup_count']\n    label = f'{eeg_id}-{target}-{d_count}'\n    \n    # CREATE SPECTROGRAM FROM EEG PARQUET\n    img = spectrogram_from_eeg(row, i<DISPLAY)\n    \n    # SAVE TO DISK\n    if i==DISPLAY:\n        print(f'Creating and writing {train.shape[0]} spectrograms to disk... ',end='')\n    \n    if SAVE_INDIVIDUAL_EEGS:\n        np.save(f'{directory_path}{label}',img)\n    all_eegs[label] = img\n    \n    if not SAVE_EEGS and i == 4:\n        break\n   \n# SAVE EEG SPECTROGRAM DICTIONARY\nif SAVE_EEGS:\n    np.save('eeg_specs',all_eegs)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T11:48:50.880258Z","iopub.execute_input":"2024-03-05T11:48:50.880982Z","iopub.status.idle":"2024-03-05T11:49:10.372306Z","shell.execute_reply.started":"2024-03-05T11:48:50.880949Z","shell.execute_reply":"2024-03-05T11:49:10.370989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize the created Dataset","metadata":{}},{"cell_type":"code","source":"EEG_VISUALIZE = True\nif EEG_VISUALIZE:\n    all_eegs = np.load('/kaggle/input/eeg-spectrogram-by-extended-data/eeg_specs.npy', \n                       allow_pickle=True).item()","metadata":{"execution":{"iopub.status.busy":"2024-03-05T12:01:24.607711Z","iopub.execute_input":"2024-03-05T12:01:24.608114Z","iopub.status.idle":"2024-03-05T12:03:13.161946Z","shell.execute_reply.started":"2024-03-05T12:01:24.608082Z","shell.execute_reply":"2024-03-05T12:03:13.160728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if EEG_VISUALIZE:\n    sample_row = train.iloc[np.random.randint(0,train.shape[0])]\n\n    eeg_id = sample_row['eeg_id']\n    target = sample_row['expert_consensus']\n    d_count = sample_row['dup_count']\n    label = f'{eeg_id}-{target}-{d_count}'\n\n    eeg_df = all_eegs[label]\n    \n    plt.figure(figsize=(20,8))\n    for i in range(4):\n        plt.subplot(2,2,i+1)\n        img = eeg_df[:,:,i]\n        plt.imshow(img)\n        plt.title(f'label{label} spectrogram{NAMES[i]}')\n\n    plt.tight_layout()\n    ","metadata":{"execution":{"iopub.status.busy":"2024-03-05T13:00:10.309219Z","iopub.execute_input":"2024-03-05T13:00:10.310365Z","iopub.status.idle":"2024-03-05T13:00:12.085294Z","shell.execute_reply.started":"2024-03-05T13:00:10.310315Z","shell.execute_reply":"2024-03-05T13:00:12.084220Z"},"trusted":true},"execution_count":null,"outputs":[]}]}