{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7392733,"sourceType":"datasetVersion","datasetId":4297749},{"sourceId":7392775,"sourceType":"datasetVersion","datasetId":4297782},{"sourceId":7402356,"sourceType":"datasetVersion","datasetId":4304475},{"sourceId":7403069,"sourceType":"datasetVersion","datasetId":4304949},{"sourceId":7447509,"sourceType":"datasetVersion","datasetId":4334995},{"sourceId":7450712,"sourceType":"datasetVersion","datasetId":4336944},{"sourceId":158958765,"sourceType":"kernelVersion"}],"dockerImageVersionId":30646,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Inspiration from [Chris Deotte](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010), [Muhammad Ahmed](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467000), ","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## Terminologies\n\n* Imagine your brain is a bustling city, full of activity. Every thought, feeling, and movement is like a car zipping through the streets. But instead of using gasoline, your brain uses electricity! EEG is like a traffic camera that captures snapshots of this electrical activity.\n\n\n* **what is EEG?**\n1. Imagine a bunch of wavy lines, like mountains on a graph. Each line represents the electrical activity from a different part of your brain.\n2. The height of the waves (amplitude) shows how strong the signal is.\n3. The speed of the waves (frequency) tells you how fast the brain cells are firing.\n4. Doctors use EEG to diagnose conditions like epilepsy, where the brain's electrical activity is abnormal. They can also use it to study sleep disorders, coma, and even brain injuries.\n5. Think of it like a detective using clues to solve a mystery about what's going on inside your brain.\n6. **Spectograms**: magine taking the EEG data and turning it into a rainbow. That's what a spectrogram does!\n\n\n\n* **LPD/GPD/LRDA/GRDA**\n* These are different patterns that can show up on an EEG, like traffic jams or parades in this busy city. Understanding these patterns helps doctors diagnose problems in the brain.\n\n* **LPD (Lateralized Periodic Discharges)**:\n1. Like: A one-sided parade marching in circles around a specific part of the city.\n2. Means: One side of the brain is sending out repetitive bursts of strong signals, often due to injury or disease.\n3. Think of it as: A flashing alarm on one side of the brain, indicating something's not right.\n\n* **GPD (Generalized Periodic Discharges)**:\n\n1. Like: A city-wide traffic jam where all the streets are clogged with moving cars.\n2. Means: Both sides of the brain are sending out synchronized bursts of strong signals, indicating widespread problems.\n3. Think of it as: A city-wide blackout, suggesting the entire brain is struggling.\n\n* **LRDA (Lateralized Rhythmic Delta Activity)**:\n \n1. Like: Slow, rhythmic music playing only in one part of the city.\n2. Means: One side of the brain is sending out slow, regular waves, suggesting a localized issue.\n3. Think of it as: A broken record playing on repeat in a specific district, hinting at a problem there.\n\n*  **GRDA (Generalized Rhythmic Delta Activity)**:\n \n1. Like: A city-wide lullaby where everything slows down and hums at the same low rhythm.\n2. Means: Both sides of the brain are sending out slow, regular waves, indicating widespread dysfunction.\n3. Think of it as: All the clocks in the city running slow and in sync, suggesting something is affecting the entire brain's rhythm.","metadata":{}},{"cell_type":"markdown","source":"## Understanding Competition Data\n\nThe data in this competition is confusing. \n**Question**: We are given train.csv with 106,800 rows but there are only 17089 unique eeg_ids, 11138 unique spectrogram_ids, and 1950 unique patients. What is going on?? \n**Answer**: Each row of train is a window of time from one specific patient. And the corresponding eeg and spectrogram can be found in the corresponding Kaggle parquet files.\n\n\n### Data Explained\n\nSince each row of train.csv is a specific window in time (for a specific patient_id), each row has a specific middle timestamp in seconds. For example, maybe row 235 has center timestamp T = May 3 2023 19:30:06 exactly in the middle of both its EEG time window and Spectrogram time window. (Note that the dataframe does not give us the middle timestamp).\n\nThe EEG time window is length 50 seconds and the Spectrogram time window is length 600 seconds. And both have the same center timestamp. In this competition, we are asked to predict the event occurring in the middle 10 seconds of both these time windows:\n*   Center is T\n*   EEG is [T-25:T+25]\n*   Spectrogram is [T-300:T+300]\n*   We predict event in [T-5:T+5]\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain8.png)\n\n#### EEG Parquet Files\n\nThe EEG parquet files are longer than 50 seconds. One EEG parquet file has multiple rows of time windows inside it. Similarily, the Spectrogram parquet files are longer than 600 seconds. One Spectrogram parquet file has multiple rows of time windows inside it.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain9.png)\n\n#### Spectrogram Parquet Files\n\nThere are fewer Spectrogram parquet files than EEG parquet files. This is because two rows may have the same Spectrogram parquet file but different EEG parquet files.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain11.png)","metadata":{}},{"cell_type":"markdown","source":"### Code to Retrive EEG & Spectrogram\n\nFor a specific row of train.csv, here is code to retrieve the corresponding EEG and Spectrogram. Note that train.csv does not give us the middle timestamp. Instead it gives us the beginning timestamp for each timewindow. The two beginnings are determined from eeg_label_offset_seconds and spectrogram_label_offset_seconds. These are offsets from the start of the parquet file which tells us where the eeg and spectrogram time windows begin respectively.","metadata":{}},{"cell_type":"code","source":"import os\n#os.environ[\"KERAS_BACKEND\"] = \"tensorflow\" # you can also use tensorflow or torch\n\nimport keras_cv\nimport keras\n#from keras import ops\nimport tensorflow as tf\n\nimport cv2\nimport pandas as pd\nimport numpy as np\nfrom glob import glob\nfrom tqdm.notebook import tqdm\nimport joblib\n\nimport matplotlib.pyplot as plt ","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:32.241097Z","iopub.execute_input":"2024-03-09T21:24:32.241396Z","iopub.status.idle":"2024-03-09T21:24:53.518594Z","shell.execute_reply.started":"2024-03-09T21:24:32.241370Z","shell.execute_reply":"2024-03-09T21:24:53.517526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''class CFG:\n    verbose = 1  # Verbosity\n    seed = 42  # Random seed\n    preset = \"efficientnetv2_b2_imagenet\"  # Name of pretrained classifier\n    image_size = [400, 300]  # Input image size\n    epochs = 13 # Training epochs\n    batch_size = 64  # Batch size\n    lr_mode = \"cos\" # LR scheduler mode from one of \"cos\", \"step\", \"exp\"\n    drop_remainder = True  # Drop incomplete batches\n    num_classes = 6 # Number of classes in the dataset\n    fold = 0 # Which fold to set as validation data\n    class_names = ['Seizure', 'LPD', 'GPD', 'LRDA','GRDA', 'Other']\n    label2name = dict(enumerate(class_names))\n    name2label = {v:k for k, v in label2name.items()}\n    '''","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:53.520605Z","iopub.execute_input":"2024-03-09T21:24:53.521592Z","iopub.status.idle":"2024-03-09T21:24:53.529603Z","shell.execute_reply.started":"2024-03-09T21:24:53.521555Z","shell.execute_reply":"2024-03-09T21:24:53.528612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''BASE_PATH = \"/kaggle/input/hms-harmful-brain-activity-classification\"\n\nSPEC_DIR = \"/tmp/dataset/hms-hbac\"\nos.makedirs(SPEC_DIR+'/train_spectrograms', exist_ok=True)\nos.makedirs(SPEC_DIR+'/test_spectrograms', exist_ok=True)\n'''","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:53.531142Z","iopub.execute_input":"2024-03-09T21:24:53.532349Z","iopub.status.idle":"2024-03-09T21:24:53.548393Z","shell.execute_reply.started":"2024-03-09T21:24:53.532314Z","shell.execute_reply":"2024-03-09T21:24:53.547422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# Train + Valid\ndf = pd.read_csv(f'{BASE_PATH}/train.csv')\ndf['eeg_path'] = f'{BASE_PATH}/train_eegs/'+df['eeg_id'].astype(str)+'.parquet'\ndf['spec_path'] = f'{BASE_PATH}/train_spectrograms/'+df['spectrogram_id'].astype(str)+'.parquet'\ndf['spec2_path'] = f'{SPEC_DIR}/train_spectrograms/'+df['spectrogram_id'].astype(str)+'.npy'\ndf['class_name'] = df.expert_consensus.copy()\ndf['class_label'] = df.expert_consensus.map(CFG.name2label)\ndisplay(df.head(2))\n\n# Test\ntest_df = pd.read_csv(f'{BASE_PATH}/test.csv')\ntest_df['eeg_path'] = f'{BASE_PATH}/test_eegs/'+test_df['eeg_id'].astype(str)+'.parquet'\ntest_df['spec_path'] = f'{BASE_PATH}/test_spectrograms/'+test_df['spectrogram_id'].astype(str)+'.parquet'\ntest_df['spec2_path'] = f'{SPEC_DIR}/test_spectrograms/'+test_df['spectrogram_id'].astype(str)+'.npy'\ndisplay(test_df.head(2))\n'''","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:53.550698Z","iopub.execute_input":"2024-03-09T21:24:53.551245Z","iopub.status.idle":"2024-03-09T21:24:53.559237Z","shell.execute_reply.started":"2024-03-09T21:24:53.551221Z","shell.execute_reply":"2024-03-09T21:24:53.558473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, gc\nos.environ[\"CUDA_VISIBLE_DEVICES\"]=\"0,1\"\nimport tensorflow as tf\nimport pandas as pd, numpy as np\nimport matplotlib.pyplot as plt\nprint('TensorFlow version =',tf.__version__)\n\n# USE MULTIPLE GPUS\ngpus = tf.config.list_physical_devices('GPU')\nif len(gpus)<=1: \n    strategy = tf.distribute.OneDeviceStrategy(device=\"/gpu:0\")\n    print(f'Using {len(gpus)} GPU')\nelse: \n    strategy = tf.distribute.MirroredStrategy()\n    print(f'Using {len(gpus)} GPUs')\n\nVER = 5\n\n# IF THIS EQUALS NONE, THEN WE TRAIN NEW MODELS\n# IF THIS EQUALS DISK PATH, THEN WE LOAD PREVIOUSLY TRAINED MODELS\nLOAD_MODELS_FROM = '/kaggle/input/brain-efficientnet-models-v3-v4-v5/'\n\nUSE_KAGGLE_SPECTROGRAMS = True\nUSE_EEG_SPECTROGRAMS = True","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:53.560187Z","iopub.execute_input":"2024-03-09T21:24:53.560420Z","iopub.status.idle":"2024-03-09T21:24:54.496597Z","shell.execute_reply.started":"2024-03-09T21:24:53.560400Z","shell.execute_reply":"2024-03-09T21:24:54.495572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# USE MIXED PRECISION\nMIX = True\nif MIX:\n    tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n    print('Mixed precision enabled')\nelse:\n    print('Using full precision')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:54.497810Z","iopub.execute_input":"2024-03-09T21:24:54.498115Z","iopub.status.idle":"2024-03-09T21:24:54.503577Z","shell.execute_reply.started":"2024-03-09T21:24:54.498089Z","shell.execute_reply":"2024-03-09T21:24:54.502679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load Train Data","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\nTARGETS = df.columns[-6:]\nprint('Train shape:', df.shape )\nprint('Targets', list(TARGETS))\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:54.504868Z","iopub.execute_input":"2024-03-09T21:24:54.505211Z","iopub.status.idle":"2024-03-09T21:24:54.780810Z","shell.execute_reply.started":"2024-03-09T21:24:54.505187Z","shell.execute_reply":"2024-03-09T21:24:54.779246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 106_800 rows in our training data, which represents labeled eeg subsample\n\nEach row has an id corresponding to the EEG and spectrogram data, which is the data we'll use to make our predictions.\n\nlpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote are our target columns, which must be probabilities\n\nEach row has sub_id and offset_seconds for EEG and spectrogram data. During inference, we'll obtain 50-second-long subsample of EEG and 10-minute-long subsample of the spectrogram to make a prediction. Therefore, target values are applicable for this subsamples, but not for the entire EEG and spectrograms\n\nThe patient_id column would be benefitial to make a train/val/test split\n\nexpert_consensus will be the useful indicator of contraversial samples. When consensus is low, the row is, probably, the row is probably going to be dropped.\n","metadata":{}},{"cell_type":"markdown","source":"#### inspect classes distribution","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nplt.figure(figsize=(10, 6))\nsns.countplot(data=df, x='expert_consensus')\nplt.title('Distribution of Expert Consensus')\nplt.xlabel('Expert Consensus')\nplt.ylabel('Count')\nplt.xticks(rotation=45)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:54.782801Z","iopub.execute_input":"2024-03-09T21:24:54.783191Z","iopub.status.idle":"2024-03-09T21:24:55.565800Z","shell.execute_reply.started":"2024-03-09T21:24:54.783161Z","shell.execute_reply":"2024-03-09T21:24:55.564862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Inspect number of patients","metadata":{}},{"cell_type":"code","source":"print(f\"Number of patients: {df['patient_id'].nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:55.567016Z","iopub.execute_input":"2024-03-09T21:24:55.567732Z","iopub.status.idle":"2024-03-09T21:24:55.575230Z","shell.execute_reply.started":"2024-03-09T21:24:55.567696Z","shell.execute_reply":"2024-03-09T21:24:55.574236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Inspect offset seconds number","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,6))\nplt.hist(df['eeg_label_offset_seconds'], bins='auto', log=True)\nplt.xlabel('Offset Seconds')\nplt.ylabel('Frequency (Log Scale)')\nplt.title('Histogram of Offset Seconds (Log Scale)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:55.578828Z","iopub.execute_input":"2024-03-09T21:24:55.579150Z","iopub.status.idle":"2024-03-09T21:24:57.963890Z","shell.execute_reply.started":"2024-03-09T21:24:55.579119Z","shell.execute_reply":"2024-03-09T21:24:57.962932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### correlation ","metadata":{}},{"cell_type":"code","source":"correlation_targets = df[TARGETS].corr()\nplt.figure(figsize=(12, 8))\nsns.heatmap(correlation_targets, annot=True, cmap='coolwarm', fmt=\".2f\")\nplt.title('Correlation Matrix of Vote Columns')\nplt.show()\n\nplt.figure(figsize=(12, 10))\nfor i, column in enumerate(TARGETS, 1):\n    plt.subplot(3, 2, i)\n    sns.violinplot(data=df, x='expert_consensus', y=column)\n    plt.title(f'Distribution of {column} by Expert Consensus')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:24:57.965107Z","iopub.execute_input":"2024-03-09T21:24:57.965396Z","iopub.status.idle":"2024-03-09T21:25:04.423502Z","shell.execute_reply.started":"2024-03-09T21:24:57.965372Z","shell.execute_reply":"2024-03-09T21:25:04.422524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Non-Overlapping Eeg Id Train Data","metadata":{}},{"cell_type":"code","source":"train = df.groupby('eeg_id')[['spectrogram_id','spectrogram_label_offset_seconds']].agg(\n    {'spectrogram_id':'first','spectrogram_label_offset_seconds':'min'})\ntrain.columns = ['spec_id','min']\n\ntmp = df.groupby('eeg_id')[['spectrogram_id','spectrogram_label_offset_seconds']].agg(\n    {'spectrogram_label_offset_seconds':'max'})\ntrain['max'] = tmp\n\ntmp = df.groupby('eeg_id')[['patient_id']].agg('first')\ntrain['patient_id'] = tmp\n\ntmp = df.groupby('eeg_id')[TARGETS].agg('sum')\nfor t in TARGETS:\n    train[t] = tmp[t].values\n    \ny_data = train[TARGETS].values\ny_data = y_data / y_data.sum(axis=1,keepdims=True)\ntrain[TARGETS] = y_data\n\ntmp = df.groupby('eeg_id')[['expert_consensus']].agg('first')\ntrain['target'] = tmp\n\ntrain = train.reset_index()\nprint('Train non-overlapp eeg_id shape:', train.shape )\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:25:04.424839Z","iopub.execute_input":"2024-03-09T21:25:04.425234Z","iopub.status.idle":"2024-03-09T21:25:04.511190Z","shell.execute_reply.started":"2024-03-09T21:25:04.425202Z","shell.execute_reply":"2024-03-09T21:25:04.510264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Read Train Spectrograms\n\n","metadata":{}},{"cell_type":"code","source":"%%time\nREAD_SPEC_FILES = False\n\n# READ ALL SPECTROGRAMS\nPATH = '/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/'\nfiles = os.listdir(PATH)\nprint(f'There are {len(files)} spectrogram parquets')\n\nif READ_SPEC_FILES:    \n    spectrograms = {}\n    for i,f in enumerate(files):\n        if i%100==0: print(i,', ',end='')\n        tmp = pd.read_parquet(f'{PATH}{f}')\n        name = int(f.split('.')[0])\n        spectrograms[name] = tmp.iloc[:,1:].values\nelse:\n    spectrograms = np.load('/kaggle/input/brain-spectrograms/specs.npy',allow_pickle=True).item()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:25:04.512565Z","iopub.execute_input":"2024-03-09T21:25:04.512953Z","iopub.status.idle":"2024-03-09T21:25:56.847694Z","shell.execute_reply.started":"2024-03-09T21:25:04.512920Z","shell.execute_reply":"2024-03-09T21:25:56.846678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Read EEG Spectrograms","metadata":{}},{"cell_type":"code","source":"%%time\nREAD_EEG_SPEC_FILES = False\n\nif READ_EEG_SPEC_FILES:\n    all_eegs = {}\n    for i,e in enumerate(train.eeg_id.values):\n        if i%100==0: print(i,', ',end='')\n        x = np.load(f'/kaggle/input/brain-eeg-spectrograms/EEG_Spectrograms/{e}.npy')\n        all_eegs[e] = x\nelse:\n    all_eegs = np.load('/kaggle/input/brain-eeg-spectrograms/eeg_specs.npy',allow_pickle=True).item()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:25:56.849047Z","iopub.execute_input":"2024-03-09T21:25:56.849756Z","iopub.status.idle":"2024-03-09T21:27:00.802444Z","shell.execute_reply.started":"2024-03-09T21:25:56.849700Z","shell.execute_reply":"2024-03-09T21:27:00.801217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train DataLoader","metadata":{}},{"cell_type":"code","source":"import albumentations as albu\nTARS = {'Seizure':0, 'LPD':1, 'GPD':2, 'LRDA':3, 'GRDA':4, 'Other':5}\nTARS2 = {x:y for y,x in TARS.items()}\n\nclass DataGenerator(tf.keras.utils.Sequence):\n    'Generates data for Keras'\n    def __init__(self, data, batch_size=32, shuffle=False, augment=False, mode='train',\n                 specs = spectrograms, eeg_specs = all_eegs): \n\n        self.data = data\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.augment = augment\n        self.mode = mode\n        self.specs = specs\n        self.eeg_specs = eeg_specs\n        self.on_epoch_end()\n        \n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        ct = int( np.ceil( len(self.data) / self.batch_size ) )\n        return ct\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        X, y = self.__data_generation(indexes)\n        if self.augment: X = self.__augment_batch(X) \n        return X, y\n\n    def on_epoch_end(self):\n        'Updates indexes after each epoch'\n        self.indexes = np.arange( len(self.data) )\n        if self.shuffle: np.random.shuffle(self.indexes)\n                        \n    def __data_generation(self, indexes):\n        'Generates data containing batch_size samples' \n        \n        X = np.zeros((len(indexes),128,256,8),dtype='float32')\n        y = np.zeros((len(indexes),6),dtype='float32')\n        img = np.ones((128,256),dtype='float32')\n        \n        for j,i in enumerate(indexes):\n            row = self.data.iloc[i]\n            if self.mode=='test': \n                r = 0\n            else: \n                r = int( (row['min'] + row['max'])//4 )\n\n            for k in range(4):\n                # EXTRACT 300 ROWS OF SPECTROGRAM\n                img = self.specs[row.spec_id][r:r+300,k*100:(k+1)*100].T\n                \n                # LOG TRANSFORM SPECTROGRAM\n                img = np.clip(img,np.exp(-4),np.exp(8))\n                img = np.log(img)\n                \n                # STANDARDIZE PER IMAGE\n                ep = 1e-6\n                m = np.nanmean(img.flatten())\n                s = np.nanstd(img.flatten())\n                img = (img-m)/(s+ep)\n                img = np.nan_to_num(img, nan=0.0)\n                \n                # CROP TO 256 TIME STEPS\n                X[j,14:-14,:,k] = img[:,22:-22] / 2.0\n        \n            # EEG SPECTROGRAMS\n            img = self.eeg_specs[row.eeg_id]\n            X[j,:,:,4:] = img\n                \n            if self.mode!='test':\n                y[j,] = row[TARGETS]\n            \n        return X,y\n    \n    def __random_transform(self, img):\n        composition = albu.Compose([\n            albu.HorizontalFlip(p=0.5),\n            #albu.CoarseDropout(max_holes=8,max_height=32,max_width=32,fill_value=0,p=0.5),\n        ])\n        return composition(image=img)['image']\n            \n    def __augment_batch(self, img_batch):\n        for i in range(img_batch.shape[0]):\n            img_batch[i, ] = self.__random_transform(img_batch[i, ])\n        return img_batch","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:00.807295Z","iopub.execute_input":"2024-03-09T21:27:00.807689Z","iopub.status.idle":"2024-03-09T21:27:01.574554Z","shell.execute_reply.started":"2024-03-09T21:27:00.807651Z","shell.execute_reply":"2024-03-09T21:27:01.573752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Display DataLoader","metadata":{}},{"cell_type":"code","source":"gen = DataGenerator(train, batch_size=32, shuffle=False)\nROWS=2; COLS=3; BATCHES=2\n\nfor i,(x,y) in enumerate(gen):\n    plt.figure(figsize=(20,8))\n    for j in range(ROWS):\n        for k in range(COLS):\n            plt.subplot(ROWS,COLS,j*COLS+k+1)\n            t = y[j*COLS+k]\n            img = x[j*COLS+k,:,:,0][::-1,]\n            mn = img.flatten().min()\n            mx = img.flatten().max()\n            img = (img-mn)/(mx-mn)\n            plt.imshow(img)\n            tars = f'[{t[0]:0.2f}'\n            for s in t[1:]: tars += f', {s:0.2f}'\n            eeg = train.eeg_id.values[i*32+j*COLS+k]\n            plt.title(f'EEG = {eeg}\\nTarget = {tars}',size=12)\n            plt.yticks([])\n            plt.ylabel('Frequencies (Hz)',size=14)\n            plt.xlabel('Time (sec)',size=16)\n    plt.show()\n    if i==BATCHES-1: break","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:01.576012Z","iopub.execute_input":"2024-03-09T21:27:01.577283Z","iopub.status.idle":"2024-03-09T21:27:04.049749Z","shell.execute_reply.started":"2024-03-09T21:27:01.577246Z","shell.execute_reply":"2024-03-09T21:27:04.048863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Scheduler","metadata":{}},{"cell_type":"code","source":"import math\nLR_START = 1e-6\nLR_MAX = 1e-3\nLR_MIN = 1e-6\nLR_RAMPUP_EPOCHS = 0\nLR_SUSTAIN_EPOCHS = 0\nEPOCHS2 = 10\n\ndef lrfn(epoch):\n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        decay_total_epochs = EPOCHS2 - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS - 1\n        decay_epoch_index = epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS\n        phase = math.pi * decay_epoch_index / decay_total_epochs\n        cosine_decay = 0.5 * (1 + math.cos(phase))\n        lr = (LR_MAX - LR_MIN) * cosine_decay + LR_MIN\n    return lr\n\nrng = [i for i in range(EPOCHS2)]\nlr_y = [lrfn(x) for x in rng]\nplt.figure(figsize=(10, 4))\nplt.plot(rng, lr_y, '-o')\nplt.xlabel('epoch',size=14); plt.ylabel('learning rate',size=14)\nplt.title('Cosine Training Schedule',size=16); plt.show()\n\nLR2 = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:04.051191Z","iopub.execute_input":"2024-03-09T21:27:04.051532Z","iopub.status.idle":"2024-03-09T21:27:04.253038Z","shell.execute_reply.started":"2024-03-09T21:27:04.051503Z","shell.execute_reply":"2024-03-09T21:27:04.252132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LR_START = 1e-4\nLR_MAX = 1e-3\nLR_RAMPUP_EPOCHS = 0\nLR_SUSTAIN_EPOCHS = 1\nLR_STEP_DECAY = 0.1\nEVERY = 1\nEPOCHS = 4\n\ndef lrfn(epoch):\n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = LR_MAX * LR_STEP_DECAY**((epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS)//EVERY)\n    return lr\n\nrng = [i for i in range(EPOCHS)]\ny = [lrfn(x) for x in rng]\nplt.figure(figsize=(10, 4))\nplt.plot(rng, y, 'o-'); \nplt.xlabel('epoch',size=14); plt.ylabel('learning rate',size=14)\nplt.title('Step Training Schedule',size=16); plt.show()\n\nLR = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:04.254450Z","iopub.execute_input":"2024-03-09T21:27:04.254792Z","iopub.status.idle":"2024-03-09T21:27:04.462667Z","shell.execute_reply.started":"2024-03-09T21:27:04.254762Z","shell.execute_reply":"2024-03-09T21:27:04.461782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Build EfficientNet Model","metadata":{}},{"cell_type":"code","source":"!pip install --no-index --find-links=/kaggle/input/tf-efficientnet-whl-files /kaggle/input/tf-efficientnet-whl-files/efficientnet-1.1.1-py3-none-any.whl","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:04.463821Z","iopub.execute_input":"2024-03-09T21:27:04.464110Z","iopub.status.idle":"2024-03-09T21:27:18.650173Z","shell.execute_reply.started":"2024-03-09T21:27:04.464084Z","shell.execute_reply":"2024-03-09T21:27:18.649025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import efficientnet.tfkeras as efn\n\ndef build_model():\n    \n    inp = tf.keras.Input(shape=(128,256,8))\n    base_model = efn.EfficientNetB0(include_top=False, weights=None, input_shape=None)\n    base_model.load_weights('/kaggle/input/tf-efficientnet-imagenet-weights/efficientnet-b0_weights_tf_dim_ordering_tf_kernels_autoaugment_notop.h5')\n    \n    # RESHAPE INPUT 128x256x8 => 512x512x3 MONOTONE IMAGE\n    # KAGGLE SPECTROGRAMS\n    x1 = [inp[:,:,:,i:i+1] for i in range(4)]\n    x1 = tf.keras.layers.Concatenate(axis=1)(x1)\n    # EEG SPECTROGRAMS\n    x2 = [inp[:,:,:,i+4:i+5] for i in range(4)]\n    x2 = tf.keras.layers.Concatenate(axis=1)(x2)\n    # MAKE 512X512X3\n    if USE_KAGGLE_SPECTROGRAMS & USE_EEG_SPECTROGRAMS:\n        x = tf.keras.layers.Concatenate(axis=2)([x1,x2])\n    elif USE_EEG_SPECTROGRAMS: x = x2\n    else: x = x1\n    x = tf.keras.layers.Concatenate(axis=3)([x,x,x])\n    \n    # OUTPUT\n    x = base_model(x)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(6,activation='softmax', dtype='float32')(x)\n        \n    # COMPILE MODEL\n    model = tf.keras.Model(inputs=inp, outputs=x)\n    opt = tf.keras.optimizers.Adam(learning_rate = 1e-3)\n    loss = tf.keras.losses.KLDivergence()\n\n    model.compile(loss=loss, optimizer = opt) \n        \n    return model","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:18.651819Z","iopub.execute_input":"2024-03-09T21:27:18.652208Z","iopub.status.idle":"2024-03-09T21:27:18.691718Z","shell.execute_reply.started":"2024-03-09T21:27:18.652169Z","shell.execute_reply":"2024-03-09T21:27:18.691012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Model","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import KFold, GroupKFold\nimport tensorflow.keras.backend as K, gc\n\nall_oof = []\nall_true = []\n\ngkf = GroupKFold(n_splits=5)\nfor i, (train_index, valid_index) in enumerate(gkf.split(train, train.target, train.patient_id)):  \n    \n    print('#'*25)\n    print(f'### Fold {i+1}')\n    \n    train_gen = DataGenerator(train.iloc[train_index], shuffle=True, batch_size=32, augment=False)\n    valid_gen = DataGenerator(train.iloc[valid_index], shuffle=False, batch_size=64, mode='valid')\n    \n    print(f'### train size {len(train_index)}, valid size {len(valid_index)}')\n    print('#'*25)\n    \n    K.clear_session()\n    with strategy.scope():\n        model = build_model()\n    if LOAD_MODELS_FROM is None:\n        model.fit(train_gen, verbose=1,\n              validation_data = valid_gen,\n              epochs=EPOCHS, callbacks = [LR])\n        model.save_weights(f'EffNet_v{VER}_f{i}.h5')\n    else:\n        model.load_weights(f'{LOAD_MODELS_FROM}EffNet_v{VER}_f{i}.h5')\n        \n    oof = model.predict(valid_gen, verbose=1)\n    all_oof.append(oof)\n    all_true.append(train.iloc[valid_index][TARGETS].values)\n    \n    del model, oof\n    gc.collect()\n    \nall_oof = np.concatenate(all_oof)\nall_true = np.concatenate(all_true)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:27:18.692680Z","iopub.execute_input":"2024-03-09T21:27:18.693718Z","iopub.status.idle":"2024-03-09T21:30:14.464232Z","shell.execute_reply.started":"2024-03-09T21:27:18.693694Z","shell.execute_reply":"2024-03-09T21:30:14.463144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\n\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:30:14.466519Z","iopub.execute_input":"2024-03-09T21:30:14.467139Z","iopub.status.idle":"2024-03-09T21:30:14.471733Z","shell.execute_reply.started":"2024-03-09T21:30:14.467111Z","shell.execute_reply":"2024-03-09T21:30:14.470651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## CV Score for EfficientNet","metadata":{}},{"cell_type":"code","source":"import sys\nsys.path.append('/kaggle/input/kaggle-kl-div')\nfrom kaggle_kl_div import score\n\noof = pd.DataFrame(all_oof.copy())\noof['id'] = np.arange(len(oof))\n\ntrue = pd.DataFrame(all_true.copy())\ntrue['id'] = np.arange(len(true))\n\ncv = score(solution=true, submission=oof, row_id_column_name='id')\nprint('CV Score KL-Div for EfficientNetB2 =',cv)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:30:14.475361Z","iopub.execute_input":"2024-03-09T21:30:14.476123Z","iopub.status.idle":"2024-03-09T21:30:14.561034Z","shell.execute_reply.started":"2024-03-09T21:30:14.476090Z","shell.execute_reply":"2024-03-09T21:30:14.560020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Infer Test and Create Submission CSVv","metadata":{}},{"cell_type":"code","source":"del all_eegs, spectrograms; gc.collect()\ntest = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\nprint('Test shape',test.shape)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:30:26.490418Z","iopub.execute_input":"2024-03-09T21:30:26.490826Z","iopub.status.idle":"2024-03-09T21:30:26.796570Z","shell.execute_reply.started":"2024-03-09T21:30:26.490796Z","shell.execute_reply":"2024-03-09T21:30:26.795638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# READ ALL SPECTROGRAMS\nPATH2 = '/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms/'\nfiles2 = os.listdir(PATH2)\nprint(f'There are {len(files2)} test spectrogram parquets')\n    \nspectrograms2 = {}\nfor i,f in enumerate(files2):\n    if i%100==0: print(i,', ',end='')\n    tmp = pd.read_parquet(f'{PATH2}{f}')\n    name = int(f.split('.')[0])\n    spectrograms2[name] = tmp.iloc[:,1:].values\n    \n# RENAME FOR DATALOADER\ntest = test.rename({'spectrogram_id':'spec_id'},axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:30:50.089489Z","iopub.execute_input":"2024-03-09T21:30:50.090230Z","iopub.status.idle":"2024-03-09T21:30:50.397040Z","shell.execute_reply.started":"2024-03-09T21:30:50.090199Z","shell.execute_reply":"2024-03-09T21:30:50.396001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pywt, librosa\n\nUSE_WAVELET = None \n\nNAMES = ['LL','LP','RP','RR']\n\nFEATS = [['Fp1','F7','T3','T5','O1'],\n         ['Fp1','F3','C3','P3','O1'],\n         ['Fp2','F8','T4','T6','O2'],\n         ['Fp2','F4','C4','P4','O2']]\n\n# DENOISE FUNCTION\ndef maddest(d, axis=None):\n    return np.mean(np.absolute(d - np.mean(d, axis)), axis)\n\ndef denoise(x, wavelet='haar', level=1):    \n    coeff = pywt.wavedec(x, wavelet, mode=\"per\")\n    sigma = (1/0.6745) * maddest(coeff[-level])\n\n    uthresh = sigma * np.sqrt(2*np.log(len(x)))\n    coeff[1:] = (pywt.threshold(i, value=uthresh, mode='hard') for i in coeff[1:])\n\n    ret=pywt.waverec(coeff, wavelet, mode='per')\n    \n    return ret\n\ndef spectrogram_from_eeg(parquet_path, display=False):\n    \n    # LOAD MIDDLE 50 SECONDS OF EEG SERIES\n    eeg = pd.read_parquet(parquet_path)\n    middle = (len(eeg)-10_000)//2\n    eeg = eeg.iloc[middle:middle+10_000]\n    \n    # VARIABLE TO HOLD SPECTROGRAM\n    img = np.zeros((128,256,4),dtype='float32')\n    \n    if display: plt.figure(figsize=(10,7))\n    signals = []\n    for k in range(4):\n        COLS = FEATS[k]\n        \n        for kk in range(4):\n        \n            # COMPUTE PAIR DIFFERENCES\n            x = eeg[COLS[kk]].values - eeg[COLS[kk+1]].values\n\n            # FILL NANS\n            m = np.nanmean(x)\n            if np.isnan(x).mean()<1: x = np.nan_to_num(x,nan=m)\n            else: x[:] = 0\n\n            # DENOISE\n            if USE_WAVELET:\n                x = denoise(x, wavelet=USE_WAVELET)\n            signals.append(x)\n\n            # RAW SPECTROGRAM\n            mel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, \n                  n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)\n\n            # LOG TRANSFORM\n            width = (mel_spec.shape[1]//32)*32\n            mel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width]\n\n            # STANDARDIZE TO -1 TO 1\n            mel_spec_db = (mel_spec_db+40)/40 \n            img[:,:,k] += mel_spec_db\n                \n        # AVERAGE THE 4 MONTAGE DIFFERENCES\n        img[:,:,k] /= 4.0\n        \n        if display:\n            plt.subplot(2,2,k+1)\n            plt.imshow(img[:,:,k],aspect='auto',origin='lower')\n            plt.title(f'EEG {eeg_id} - Spectrogram {NAMES[k]}')\n            \n    if display: \n        plt.show()\n        plt.figure(figsize=(10,5))\n        offset = 0\n        for k in range(4):\n            if k>0: offset -= signals[3-k].min()\n            plt.plot(range(10_000),signals[k]+offset,label=NAMES[3-k])\n            offset += signals[3-k].max()\n        plt.legend()\n        plt.title(f'EEG {eeg_id} Signals')\n        plt.show()\n        print(); print('#'*25); print()\n        \n    return img","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:31:50.910835Z","iopub.execute_input":"2024-03-09T21:31:50.911177Z","iopub.status.idle":"2024-03-09T21:31:51.010858Z","shell.execute_reply.started":"2024-03-09T21:31:50.911151Z","shell.execute_reply":"2024-03-09T21:31:51.009869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# READ ALL EEG SPECTROGRAMS\nPATH2 = '/kaggle/input/hms-harmful-brain-activity-classification/test_eegs/'\nDISPLAY = 1\nEEG_IDS2 = test.eeg_id.unique()\nall_eegs2 = {}\n\nprint('Converting Test EEG to Spectrograms...'); print()\nfor i,eeg_id in enumerate(EEG_IDS2):\n        \n    # CREATE SPECTROGRAM FROM EEG PARQUET\n    img = spectrogram_from_eeg(f'{PATH2}{eeg_id}.parquet', i<DISPLAY)\n    all_eegs2[eeg_id] = img","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:32:13.214643Z","iopub.execute_input":"2024-03-09T21:32:13.215971Z","iopub.status.idle":"2024-03-09T21:32:25.315894Z","shell.execute_reply.started":"2024-03-09T21:32:13.215937Z","shell.execute_reply":"2024-03-09T21:32:25.314662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# INFER EFFICIENTNET ON TEST\npreds = []\nmodel = build_model()\ntest_gen = DataGenerator(test, shuffle=False, batch_size=64, mode='test',\n                         specs = spectrograms2, eeg_specs = all_eegs2)\n\nfor i in range(5):\n    print(f'Fold {i+1}')\n    if LOAD_MODELS_FROM:\n        model.load_weights(f'{LOAD_MODELS_FROM}EffNet_v{VER}_f{i}.h5')\n    else:\n        model.load_weights(f'EffNet_v{VER}_f{i}.h5')\n    pred = model.predict(test_gen, verbose=1)\n    preds.append(pred)\npred = np.mean(preds,axis=0)\nprint()\nprint('Test preds shape',pred.shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:32:35.582298Z","iopub.execute_input":"2024-03-09T21:32:35.583620Z","iopub.status.idle":"2024-03-09T21:32:43.617030Z","shell.execute_reply.started":"2024-03-09T21:32:35.583584Z","shell.execute_reply":"2024-03-09T21:32:43.616013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.DataFrame({'eeg_id':test.eeg_id.values})\nsub[TARGETS] = pred\nsub.to_csv('submission.csv',index=False)\nprint('Submissionn shape',sub.shape)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:32:53.595199Z","iopub.execute_input":"2024-03-09T21:32:53.595578Z","iopub.status.idle":"2024-03-09T21:32:53.617578Z","shell.execute_reply.started":"2024-03-09T21:32:53.595548Z","shell.execute_reply":"2024-03-09T21:32:53.616616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SANITY CHECK TO CONFIRM PREDICTIONS SUM TO ONE\nsub.iloc[:,-6:].sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T21:33:05.191646Z","iopub.execute_input":"2024-03-09T21:33:05.192003Z","iopub.status.idle":"2024-03-09T21:33:05.201009Z","shell.execute_reply.started":"2024-03-09T21:33:05.191978Z","shell.execute_reply":"2024-03-09T21:33:05.200024Z"},"trusted":true},"execution_count":null,"outputs":[]}]}