{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30636,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import pandas as pd\nimport pathlib\n\nimport matplotlib.pyplot as plt\nimport numpy as np\n\nimport keras\n\ndset_path = pathlib.Path(\"/kaggle/input/hms-harmful-brain-activity-classification/\")\ntrain_eegs = dset_path/\"train_eegs\"\ntrain_specs = dset_path/\"train_spectrograms\"\ntest_eegs = dset_path/\"test_eegs\"\ntest_specs = dset_path/\"test_eegs\"","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-16T22:18:41.088218Z","iopub.execute_input":"2024-01-16T22:18:41.088671Z","iopub.status.idle":"2024-01-16T22:18:41.094965Z","shell.execute_reply.started":"2024-01-16T22:18:41.088633Z","shell.execute_reply":"2024-01-16T22:18:41.093996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Models\n\nThese models are from https://github.com/vlawhern/arl-eegmodels, the papers these models are based on can also be found in this repository.","metadata":{}},{"cell_type":"code","source":"\"\"\"\n ARL_EEGModels - A collection of Convolutional Neural Network models for EEG\n Signal Processing and Classification, using Keras and Tensorflow\n\n Requirements:\n    (1) tensorflow == 2.X (as of this writing, 2.0 - 2.3 have been verified\n        as working)\n \n To run the EEG/MEG ERP classification sample script, you will also need\n\n    (4) mne >= 0.17.1\n    (5) PyRiemann >= 0.2.5\n    (6) scikit-learn >= 0.20.1\n    (7) matplotlib >= 2.2.3\n    \n To use:\n    \n    (1) Place this file in the PYTHONPATH variable in your IDE (i.e.: Spyder)\n    (2) Import the model as\n        \n        from EEGModels import EEGNet    \n        \n        model = EEGNet(nb_classes = ..., Chans = ..., Samples = ...)\n        \n    (3) Then compile and fit the model\n    \n        model.compile(loss = ..., optimizer = ..., metrics = ...)\n        fitted    = model.fit(...)\n        predicted = model.predict(...)\n\n Portions of this project are works of the United States Government and are not\n subject to domestic copyright protection under 17 USC Sec. 105.  Those \n portions are released world-wide under the terms of the Creative Commons Zero \n 1.0 (CC0) license.  \n \n Other portions of this project are subject to domestic copyright protection \n under 17 USC Sec. 105.  Those portions are licensed under the Apache 2.0 \n license.  The complete text of the license governing this material is in \n the file labeled LICENSE.TXT that is a part of this project's official \n distribution. \n\"\"\"\n\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.layers import Dense, Activation, Permute, Dropout\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, AveragePooling2D\nfrom tensorflow.keras.layers import SeparableConv2D, DepthwiseConv2D\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.layers import SpatialDropout2D\nfrom tensorflow.keras.regularizers import l1_l2\nfrom tensorflow.keras.layers import Input, Flatten\nfrom tensorflow.keras.constraints import max_norm\nfrom tensorflow.keras import backend as K\n\n\ndef EEGNet(nb_classes, Chans = 64, Samples = 128, \n             dropoutRate = 0.5, kernLength = 64, F1 = 8, \n             D = 2, F2 = 16, norm_rate = 0.25, dropoutType = 'Dropout'):\n    \"\"\" Keras Implementation of EEGNet\n    http://iopscience.iop.org/article/10.1088/1741-2552/aace8c/meta\n\n    Note that this implements the newest version of EEGNet and NOT the earlier\n    version (version v1 and v2 on arxiv). We strongly recommend using this\n    architecture as it performs much better and has nicer properties than\n    our earlier version. For example:\n        \n        1. Depthwise Convolutions to learn spatial filters within a \n        temporal convolution. The use of the depth_multiplier option maps \n        exactly to the number of spatial filters learned within a temporal\n        filter. This matches the setup of algorithms like FBCSP which learn \n        spatial filters within each filter in a filter-bank. This also limits \n        the number of free parameters to fit when compared to a fully-connected\n        convolution. \n        \n        2. Separable Convolutions to learn how to optimally combine spatial\n        filters across temporal bands. Separable Convolutions are Depthwise\n        Convolutions followed by (1x1) Pointwise Convolutions. \n        \n    \n    While the original paper used Dropout, we found that SpatialDropout2D \n    sometimes produced slightly better results for classification of ERP \n    signals. However, SpatialDropout2D significantly reduced performance \n    on the Oscillatory dataset (SMR, BCI-IV Dataset 2A). We recommend using\n    the default Dropout in most cases.\n        \n    Assumes the input signal is sampled at 128Hz. If you want to use this model\n    for any other sampling rate you will need to modify the lengths of temporal\n    kernels and average pooling size in blocks 1 and 2 as needed (double the \n    kernel lengths for double the sampling rate, etc). Note that we haven't \n    tested the model performance with this rule so this may not work well. \n    \n    The model with default parameters gives the EEGNet-8,2 model as discussed\n    in the paper. This model should do pretty well in general, although it is\n\tadvised to do some model searching to get optimal performance on your\n\tparticular dataset.\n\n    We set F2 = F1 * D (number of input filters = number of output filters) for\n    the SeparableConv2D layer. We haven't extensively tested other values of this\n    parameter (say, F2 < F1 * D for compressed learning, and F2 > F1 * D for\n    overcomplete). We believe the main parameters to focus on are F1 and D. \n\n    Inputs:\n        \n      nb_classes      : int, number of classes to classify\n      Chans, Samples  : number of channels and time points in the EEG data\n      dropoutRate     : dropout fraction\n      kernLength      : length of temporal convolution in first layer. We found\n                        that setting this to be half the sampling rate worked\n                        well in practice. For the SMR dataset in particular\n                        since the data was high-passed at 4Hz we used a kernel\n                        length of 32.     \n      F1, F2          : number of temporal filters (F1) and number of pointwise\n                        filters (F2) to learn. Default: F1 = 8, F2 = F1 * D. \n      D               : number of spatial filters to learn within each temporal\n                        convolution. Default: D = 2\n      dropoutType     : Either SpatialDropout2D or Dropout, passed as a string.\n\n    \"\"\"\n    \n    if dropoutType == 'SpatialDropout2D':\n        dropoutType = SpatialDropout2D\n    elif dropoutType == 'Dropout':\n        dropoutType = Dropout\n    else:\n        raise ValueError('dropoutType must be one of SpatialDropout2D '\n                         'or Dropout, passed as a string.')\n    \n    input1   = Input(shape = (Chans, Samples, 1))\n\n    ##################################################################\n    block1       = Conv2D(F1, (1, kernLength), padding = 'same',\n                                   input_shape = (Chans, Samples, 1),\n                                   use_bias = False)(input1)\n    block1       = BatchNormalization()(block1)\n    block1       = DepthwiseConv2D((Chans, 1), use_bias = False, \n                                   depth_multiplier = D,\n                                   depthwise_constraint = max_norm(1.))(block1)\n    block1       = BatchNormalization()(block1)\n    block1       = Activation('elu')(block1)\n#     block1       = AveragePooling2D((1, 6))(block1)\n    block1       = AveragePooling2D((1, 4))(block1)\n    block1       = dropoutType(dropoutRate)(block1)\n    \n    block2       = SeparableConv2D(F2, (1, 16),\n                                   use_bias = False, padding = 'same')(block1)\n    block2       = BatchNormalization()(block2)\n    block2       = Activation('elu')(block2)\n#     block2       = AveragePooling2D((1, 12))(block2)\n    block2       = AveragePooling2D((1, 8))(block2)\n    block2       = dropoutType(dropoutRate)(block2)\n        \n    flatten      = Flatten(name = 'flatten')(block2)\n    \n    dense        = Dense(nb_classes, name = 'dense', \n                         kernel_constraint = max_norm(norm_rate))(flatten)\n    softmax      = Activation('softmax', name = 'softmax')(dense)\n    \n    return Model(inputs=input1, outputs=softmax)\n\n\n\n\ndef EEGNet_SSVEP(nb_classes = 12, Chans = 8, Samples = 256, \n             dropoutRate = 0.5, kernLength = 256, F1 = 96, \n             D = 1, F2 = 96, dropoutType = 'Dropout'):\n    \"\"\" SSVEP Variant of EEGNet, as used in [1]. \n\n    Inputs:\n        \n      nb_classes      : int, number of classes to classify\n      Chans, Samples  : number of channels and time points in the EEG data\n      dropoutRate     : dropout fraction\n      kernLength      : length of temporal convolution in first layer\n      F1, F2          : number of temporal filters (F1) and number of pointwise\n                        filters (F2) to learn. \n      D               : number of spatial filters to learn within each temporal\n                        convolution.\n      dropoutType     : Either SpatialDropout2D or Dropout, passed as a string.\n      \n      \n    [1]. Waytowich, N. et. al. (2018). Compact Convolutional Neural Networks\n    for Classification of Asynchronous Steady-State Visual Evoked Potentials.\n    Journal of Neural Engineering vol. 15(6). \n    http://iopscience.iop.org/article/10.1088/1741-2552/aae5d8\n\n    \"\"\"\n    \n    if dropoutType == 'SpatialDropout2D':\n        dropoutType = SpatialDropout2D\n    elif dropoutType == 'Dropout':\n        dropoutType = Dropout\n    else:\n        raise ValueError('dropoutType must be one of SpatialDropout2D '\n                         'or Dropout, passed as a string.')\n    \n    input1   = Input(shape = (Chans, Samples, 1))\n\n    ##################################################################\n    block1       = Conv2D(F1, (1, kernLength), padding = 'same',\n                                   input_shape = (Chans, Samples, 1),\n                                   use_bias = False)(input1)\n    block1       = BatchNormalization()(block1)\n    block1       = DepthwiseConv2D((Chans, 1), use_bias = False, \n                                   depth_multiplier = D,\n                                   depthwise_constraint = max_norm(1.))(block1)\n    block1       = BatchNormalization()(block1)\n    block1       = Activation('elu')(block1)\n    block1       = AveragePooling2D((1, 4))(block1)\n    block1       = dropoutType(dropoutRate)(block1)\n    \n    block2       = SeparableConv2D(F2, (1, 16),\n                                   use_bias = False, padding = 'same')(block1)\n    block2       = BatchNormalization()(block2)\n    block2       = Activation('elu')(block2)\n    block2       = AveragePooling2D((1, 8))(block2)\n    block2       = dropoutType(dropoutRate)(block2)\n        \n    flatten      = Flatten(name = 'flatten')(block2)\n    \n    dense        = Dense(nb_classes, name = 'dense')(flatten)\n    softmax      = Activation('softmax', name = 'softmax')(dense)\n    \n    return Model(inputs=input1, outputs=softmax)\n\n\n\ndef EEGNet_old(nb_classes, Chans = 64, Samples = 128, regRate = 0.0001,\n           dropoutRate = 0.25, kernels = [(2, 32), (8, 4)], strides = (2, 4)):\n    \"\"\" Keras Implementation of EEGNet_v1 (https://arxiv.org/abs/1611.08024v2)\n\n    This model is the original EEGNet model proposed on arxiv\n            https://arxiv.org/abs/1611.08024v2\n    \n    with a few modifications: we use striding instead of max-pooling as this \n    helped slightly in classification performance while also providing a \n    computational speed-up. \n    \n    Note that we no longer recommend the use of this architecture, as the new\n    version of EEGNet performs much better overall and has nicer properties.\n    \n    Inputs:\n        \n        nb_classes     : total number of final categories\n        Chans, Samples : number of EEG channels and samples, respectively\n        regRate        : regularization rate for L1 and L2 regularizations\n        dropoutRate    : dropout fraction\n        kernels        : the 2nd and 3rd layer kernel dimensions (default is \n                         the [2, 32] x [8, 4] configuration)\n        strides        : the stride size (note that this replaces the max-pool\n                         used in the original paper)\n    \n    \"\"\"\n\n    # start the model\n    input_main   = Input((Chans, Samples))\n    layer1       = Conv2D(16, (Chans, 1), input_shape=(Chans, Samples, 1),\n                                 kernel_regularizer = l1_l2(l1=regRate, l2=regRate))(input_main)\n    layer1       = BatchNormalization()(layer1)\n    layer1       = Activation('elu')(layer1)\n    layer1       = Dropout(dropoutRate)(layer1)\n    \n    permute_dims = 2, 1, 3\n    permute1     = Permute(permute_dims)(layer1)\n    \n    layer2       = Conv2D(4, kernels[0], padding = 'same', \n                            kernel_regularizer=l1_l2(l1=0.0, l2=regRate),\n                            strides = strides)(permute1)\n    layer2       = BatchNormalization()(layer2)\n    layer2       = Activation('elu')(layer2)\n    layer2       = Dropout(dropoutRate)(layer2)\n    \n    layer3       = Conv2D(4, kernels[1], padding = 'same',\n                            kernel_regularizer=l1_l2(l1=0.0, l2=regRate),\n                            strides = strides)(layer2)\n    layer3       = BatchNormalization()(layer3)\n    layer3       = Activation('elu')(layer3)\n    layer3       = Dropout(dropoutRate)(layer3)\n    \n    flatten      = Flatten(name = 'flatten')(layer3)\n    \n    dense        = Dense(nb_classes, name = 'dense')(flatten)\n    softmax      = Activation('softmax', name = 'softmax')(dense)\n    \n    return Model(inputs=input_main, outputs=softmax)\n\n\n\ndef DeepConvNet(nb_classes, Chans = 64, Samples = 256,\n                dropoutRate = 0.5):\n    \"\"\" Keras implementation of the Deep Convolutional Network as described in\n    Schirrmeister et. al. (2017), Human Brain Mapping.\n    \n    This implementation assumes the input is a 2-second EEG signal sampled at \n    128Hz, as opposed to signals sampled at 250Hz as described in the original\n    paper. We also perform temporal convolutions of length (1, 5) as opposed\n    to (1, 10) due to this sampling rate difference. \n    \n    Note that we use the max_norm constraint on all convolutional layers, as \n    well as the classification layer. We also change the defaults for the\n    BatchNormalization layer. We used this based on a personal communication \n    with the original authors.\n    \n                      ours        original paper\n    pool_size        1, 2        1, 3\n    strides          1, 2        1, 3\n    conv filters     1, 5        1, 10\n    \n    Note that this implementation has not been verified by the original \n    authors. \n    \n    \"\"\"\n\n    # start the model\n    input_main   = Input((Chans, Samples, 1))\n    block1       = Conv2D(25, (1, 5), \n                                 input_shape=(Chans, Samples, 1),\n                                 kernel_constraint = max_norm(2., axis=(0,1,2)))(input_main)\n    block1       = Conv2D(25, (Chans, 1),\n                                 kernel_constraint = max_norm(2., axis=(0,1,2)))(block1)\n    block1       = BatchNormalization(epsilon=1e-05, momentum=0.9)(block1)\n    block1       = Activation('elu')(block1)\n    block1       = MaxPooling2D(pool_size=(1, 2), strides=(1, 2))(block1)\n    block1       = Dropout(dropoutRate)(block1)\n  \n    block2       = Conv2D(50, (1, 5),\n                                 kernel_constraint = max_norm(2., axis=(0,1,2)))(block1)\n    block2       = BatchNormalization(epsilon=1e-05, momentum=0.9)(block2)\n    block2       = Activation('elu')(block2)\n    block2       = MaxPooling2D(pool_size=(1, 2), strides=(1, 2))(block2)\n    block2       = Dropout(dropoutRate)(block2)\n    \n    block3       = Conv2D(100, (1, 5),\n                                 kernel_constraint = max_norm(2., axis=(0,1,2)))(block2)\n    block3       = BatchNormalization(epsilon=1e-05, momentum=0.9)(block3)\n    block3       = Activation('elu')(block3)\n    block3       = MaxPooling2D(pool_size=(1, 2), strides=(1, 2))(block3)\n    block3       = Dropout(dropoutRate)(block3)\n    \n    block4       = Conv2D(200, (1, 5),\n                                 kernel_constraint = max_norm(2., axis=(0,1,2)))(block3)\n    block4       = BatchNormalization(epsilon=1e-05, momentum=0.9)(block4)\n    block4       = Activation('elu')(block4)\n    block4       = MaxPooling2D(pool_size=(1, 2), strides=(1, 2))(block4)\n    block4       = Dropout(dropoutRate)(block4)\n    \n    flatten      = Flatten()(block4)\n    \n    dense        = Dense(nb_classes, kernel_constraint = max_norm(0.5))(flatten)\n    softmax      = Activation('softmax')(dense)\n    \n    return Model(inputs=input_main, outputs=softmax)\n\n\n# need these for ShallowConvNet\ndef square(x):\n    return K.square(x)\n\ndef log(x):\n    return K.log(K.clip(x, min_value = 1e-7, max_value = 10000))   \n\n\ndef ShallowConvNet(nb_classes, Chans = 64, Samples = 128, dropoutRate = 0.5):\n    \"\"\" Keras implementation of the Shallow Convolutional Network as described\n    in Schirrmeister et. al. (2017), Human Brain Mapping.\n    \n    Assumes the input is a 2-second EEG signal sampled at 128Hz. Note that in \n    the original paper, they do temporal convolutions of length 25 for EEG\n    data sampled at 250Hz. We instead use length 13 since the sampling rate is \n    roughly half of the 250Hz which the paper used. The pool_size and stride\n    in later layers is also approximately half of what is used in the paper.\n    \n    Note that we use the max_norm constraint on all convolutional layers, as \n    well as the classification layer. We also change the defaults for the\n    BatchNormalization layer. We used this based on a personal communication \n    with the original authors.\n    \n                     ours        original paper\n    pool_size        1, 35       1, 75\n    strides          1, 7        1, 15\n    conv filters     1, 13       1, 25    \n    \n    Note that this implementation has not been verified by the original \n    authors. We do note that this implementation reproduces the results in the\n    original paper with minor deviations. \n    \"\"\"\n\n    # start the model\n    input_main   = Input((Chans, Samples, 1))\n    block1       = Conv2D(40, (1, 13), \n                                 input_shape=(Chans, Samples, 1),\n                                 kernel_constraint = max_norm(2., axis=(0,1,2)))(input_main)\n    block1       = Conv2D(40, (Chans, 1), use_bias=False, \n                          kernel_constraint = max_norm(2., axis=(0,1,2)))(block1)\n    block1       = BatchNormalization(epsilon=1e-05, momentum=0.9)(block1)\n    block1       = Activation(square)(block1)\n    block1       = AveragePooling2D(pool_size=(1, 35), strides=(1, 7))(block1)\n    block1       = Activation(log)(block1)\n    block1       = Dropout(dropoutRate)(block1)\n    flatten      = Flatten()(block1)\n    dense        = Dense(nb_classes, kernel_constraint = max_norm(0.5))(flatten)\n    softmax      = Activation('softmax')(dense)\n    \n    return Model(inputs=input_main, outputs=softmax)\n\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-16T22:18:41.104675Z","iopub.execute_input":"2024-01-16T22:18:41.105605Z","iopub.status.idle":"2024-01-16T22:18:41.169708Z","shell.execute_reply.started":"2024-01-16T22:18:41.105547Z","shell.execute_reply":"2024-01-16T22:18:41.168337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training","metadata":{}},{"cell_type":"code","source":"def paired_eeg(eeg_id, offset):\n    consolidated_eeg = pd.read_parquet(train_eegs/f\"{eeg_id}.parquet\")\n    # 200 rows = 1 second\n    # Want 50 seconds starting at the offset\n    start = offset * 200\n    end = start + (200 * 50)\n    return consolidated_eeg.iloc[start:end,]\n\ndef paired_spectrogram(spec_id, offset):\n    consolidated_spec = pd.read_parquet(train_specs/f\"{spec_id}.parquet\")\n    start = offset\n    end = offset + 600\n    return consolidated_spec[(consolidated_spec[\"time\"] <= end) & (consolidated_spec[\"time\"] >= start)]","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:18:41.172311Z","iopub.execute_input":"2024-01-16T22:18:41.172755Z","iopub.status.idle":"2024-01-16T22:18:41.188166Z","shell.execute_reply.started":"2024-01-16T22:18:41.172716Z","shell.execute_reply":"2024-01-16T22:18:41.186992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_meta = pd.read_csv(dset_path/\"train.csv\", dtype={\"eeg_label_offset_seconds\": \"Int64\",\n                                                       \"spectrogram_label_offset_seconds\": \"Int64\",\n                                                       \"expert_consensus\": \"category\"})\n\nrand = train_meta.sample(n=1).iloc[0]\neeg = paired_eeg(rand[\"eeg_id\"], rand[\"eeg_label_offset_seconds\"])\neeg","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:18:41.189808Z","iopub.execute_input":"2024-01-16T22:18:41.190431Z","iopub.status.idle":"2024-01-16T22:18:41.665878Z","shell.execute_reply.started":"2024-01-16T22:18:41.190392Z","shell.execute_reply":"2024-01-16T22:18:41.664552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = EEGNet(nb_classes=6, Chans=20, Samples=2000, kernLength=96, F1=12, D=4, F2=48, dropoutRate=0.75)\n\nlr_schedule = keras.optimizers.schedules.ExponentialDecay(\n    initial_learning_rate=0.1,\n    decay_steps=10000,\n    decay_rate=0.95\n)\nopt = keras.optimizers.Nadam(learning_rate=lr_schedule)\nmodel.compile(loss=\"kl_divergence\", optimizer=opt)\n\n# model.predict(np.expand_dims(eeg.to_numpy().T, (0, 3)))","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:18:41.668948Z","iopub.execute_input":"2024-01-16T22:18:41.669399Z","iopub.status.idle":"2024-01-16T22:18:41.818305Z","shell.execute_reply.started":"2024-01-16T22:18:41.669360Z","shell.execute_reply":"2024-01-16T22:18:41.816959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_meta.sample(frac=0.8)\nval_df = train_meta.drop(train_df.index)","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:18:41.820032Z","iopub.execute_input":"2024-01-16T22:18:41.820817Z","iopub.status.idle":"2024-01-16T22:18:41.850353Z","shell.execute_reply.started":"2024-01-16T22:18:41.820753Z","shell.execute_reply":"2024-01-16T22:18:41.848818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We want to first preprocess the eeg files by filtering and saving them as a numpy array to speed up training\nUnfortunately dataset is slightly too big to store in /kaggle/working so wil have to be done for each session (20~ mins).\nCould be worth looking into alternative storage methods.\n\nLimiting the signals to be between 10 and 40 Hz seems to be a common approach so I have used that here.","metadata":{}},{"cell_type":"code","source":"!mkdir -p /kaggle/temp/train_eeg\nprocessed_eeg_train = pathlib.Path(\"/kaggle/temp/train_eeg\")\n\nfrom scipy import signal\nb, a = signal.butter(4, [10., 40.], btype=\"band\", fs=200)\n\nfrom tqdm import tqdm\n\nfor f in tqdm(train_eegs.iterdir(), total=17300): # more or less 17300 anyway\n    arr = pd.read_parquet(f).fillna(0).to_numpy()\n    filtered = np.empty_like(arr)\n    for i in range(20):\n        filtered[:, i] = signal.filtfilt(b, a, arr[:, i])\n    filtered = filtered.T\n    np.save(processed_eeg_train / (f.stem + \".npy\"), arr)","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:18:41.853696Z","iopub.execute_input":"2024-01-16T22:18:41.854247Z","iopub.status.idle":"2024-01-16T22:18:41.860481Z","shell.execute_reply.started":"2024-01-16T22:18:41.854197Z","shell.execute_reply":"2024-01-16T22:18:41.859184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Creating a dataset generator for keras, includes batching (drops incomplete batch if batchsize doesn't fit into #files) and softmaxing the votes to get probabilities.","metadata":{}},{"cell_type":"code","source":"def softmax(votes):\n    exps = np.exp(votes)\n    return exps / np.sum(exps)\n\ndef make_gen(df, batchsize=16):\n    def data_gen():\n        batch_eeg = np.empty((batchsize, 20, 2000, 1))\n        batch_votes = np.empty((batchsize, 6))\n        batch_indx = 0\n        for i in df.eeg_id.unique():\n            subset = df[df.eeg_id == i]\n#             full_eeg = pd.read_parquet(train_eegs/f\"{i}.parquet\").fillna(0)\n            full_eeg = np.load(processed_eeg_train / (f\"{i}.npy\"))\n            for ind, data in subset.iterrows():\n                offset = data.eeg_label_offset_seconds\n                # start = offset * 200\n                start = (offset+20) * 200\n#                 end = start + (200 * 50)\n                end = start + (200 * 10)\n#                 eeg = full_eeg.iloc[start:end,].to_numpy().T\n                eeg = full_eeg[start:end,].T\n                eeg = np.expand_dims(eeg, 2)\n                votes = softmax(data.loc[\"seizure_vote\":\"other_vote\"].to_numpy(dtype=np.float32).squeeze())\n                batch_eeg[batch_indx] = eeg\n                batch_votes[batch_indx] = votes\n                batch_indx += 1\n                \n                if batch_indx == batchsize:\n                    batch_indx = 0\n                    yield batch_eeg, batch_votes\n#         if batch_indx != 0:\n#             batch_eeg = np.delete(batch_eeg, np.arange(batchsize-batch_indx, batchsize), axis=0)\n#             batch_votes = np.delete(batch_votes, np.arange(batchsize-batch_indx, batchsize), axis=0)\n#             yield batch_eeg, batch_votes\n        \n    return data_gen\n        \nimport tensorflow as tf\nbatch_size = 32\ntrain_dset = tf.data.Dataset.from_generator(make_gen(train_df, batch_size),\n                                           output_signature=(\n                                               tf.TensorSpec(shape=(batch_size, 20, 2000, 1), dtype=tf.float32),\n                                               tf.TensorSpec(shape=(batch_size, 6), dtype=tf.float32)\n                                           ))\nval_dset = tf.data.Dataset.from_generator(make_gen(val_df, batch_size),\n                                           output_signature=(\n                                               tf.TensorSpec(shape=(batch_size, 20, 2000, 1), dtype=tf.float32),\n                                               tf.TensorSpec(shape=(batch_size, 6), dtype=tf.float32)\n                                           ))","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:19:06.766021Z","iopub.execute_input":"2024-01-16T22:19:06.766469Z","iopub.status.idle":"2024-01-16T22:19:06.853153Z","shell.execute_reply.started":"2024-01-16T22:19:06.766435Z","shell.execute_reply":"2024-01-16T22:19:06.851504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(train_dset, validation_data=val_dset, epochs=10)","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:19:08.870576Z","iopub.execute_input":"2024-01-16T22:19:08.871036Z","iopub.status.idle":"2024-01-16T22:45:10.227480Z","shell.execute_reply.started":"2024-01-16T22:19:08.870998Z","shell.execute_reply":"2024-01-16T22:45:10.224938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I have tried various different hyperparameters/model configurations and am unable to get the training/val loss under 1.0, which seems to be much inferior to the models on the leaderboard learning from the spectrograms.\n\nPerhaps these models are not complex enough to learn this dataset? It also seems to be overfitting.","metadata":{}},{"cell_type":"markdown","source":"# Inference","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/test.csv\")\ntest_df","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:45:12.981794Z","iopub.execute_input":"2024-01-16T22:45:12.982220Z","iopub.status.idle":"2024-01-16T22:45:13.001514Z","shell.execute_reply.started":"2024-01-16T22:45:12.982187Z","shell.execute_reply":"2024-01-16T22:45:13.000594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def make_gen_test(df):\n    def gen():\n        for i, data in df.iterrows():\n            eeg_id = data[\"eeg_id\"]\n            eeg = pd.read_parquet(test_eegs/f\"{eeg_id}.parquet\").fillna(0).to_numpy().T\n            filtered = np.empty((20, 2000))\n            for i in range(20):\n                start = 20 * 200\n                end = start + (200 * 10)\n                filtered[i, :] = signal.filtfilt(b, a, eeg[i, start:end])\n            eeg = np.expand_dims(filtered, (0, 3))\n            yield eeg\n        \n    return gen\n    \ntest_dset = tf.data.Dataset.from_generator(make_gen_test(test_df),\n                                           output_signature=(\n                                               tf.TensorSpec(shape=(1, 20, 2000, 1), dtype=tf.float32)\n                                           ))\npreds = model.predict(test_dset)","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:46:01.051869Z","iopub.execute_input":"2024-01-16T22:46:01.052370Z","iopub.status.idle":"2024-01-16T22:46:01.243496Z","shell.execute_reply.started":"2024-01-16T22:46:01.052325Z","shell.execute_reply":"2024-01-16T22:46:01.242492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame(columns=[\"eeg_id\", \"seizure_vote\", \"lpd_vote\", \"gpd_vote\", \"lrda_vote\", \"grda_vote\", \"other_vote\"])\nfor i, row in test_df.iterrows():\n    submission.loc[i] = [row[\"eeg_id\"], *preds[i]]","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:46:20.631482Z","iopub.execute_input":"2024-01-16T22:46:20.631984Z","iopub.status.idle":"2024-01-16T22:46:20.643175Z","shell.execute_reply.started":"2024-01-16T22:46:20.631946Z","shell.execute_reply":"2024-01-16T22:46:20.641484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2024-01-16T22:46:21.567287Z","iopub.execute_input":"2024-01-16T22:46:21.568318Z","iopub.status.idle":"2024-01-16T22:46:21.590236Z","shell.execute_reply.started":"2024-01-16T22:46:21.568272Z","shell.execute_reply":"2024-01-16T22:46:21.589039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\nThough this has not been very successful, it is still useful to show how we can apply a model to the EEG signals themselves.\n\nMy intuition to why these models do not perform well is that I believe that the original datasets used had a much shorter window than the 50 seconds of signals present in this dataset. It would be worth looking into larger models or doing more testing on the hyperparameters for EEGNet, specifically kernlength as well as the average pooling kernel size and F1 + D, they could probably be changed to work better on the large 50 seconds of 200Hz samples.","metadata":{}}]}