{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7694235,"sourceType":"datasetVersion","datasetId":4490607}],"dockerImageVersionId":30646,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"vscode":{"interpreter":{"hash":"0649532bc6f99b8c98fa427cee141766b0dad28d95dfd84af62041cd5d6a3f21"}}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# HMS - Harmful Brain Activity Classification \n* EEG monitoring is an important tool used by clinicians to detect seizures and other types of brain activity that might cause brain damage in severely ill patients.\n* The existing EEG analysis procedure is highly reliant on manual review by specialized neurologists, which is laborious, time-consuming, and error-prone.\n* Manual evaluation of EEG recordings is also expensive and has reliability concerns among reviewers, even if they are specialists in the subject.\n* Automating EEG analysis can increase the efficiency and accuracy of detecting seizures and other aberrant brain activity.\n* By automating EEG analysis, clinicians and brain researchers can offer medicines to patients more quickly and precisely, potentially lowering the risk of brain injury.\n* Our goal is to categorize EEG data into one of the following categories: seizure (SZ), generalized periodic discharges (GPD), lateralized periodic discharges (LPD), lateralized rhythmic delta activity (LRDA), generalized rhythmic delta activity (GRDA), or \"other.\"\n","metadata":{"execution":{"iopub.execute_input":"2024-02-22T15:12:23.614889Z","iopub.status.busy":"2024-02-22T15:12:23.614534Z","iopub.status.idle":"2024-02-22T15:12:23.707500Z","shell.execute_reply":"2024-02-22T15:12:23.706683Z","shell.execute_reply.started":"2024-02-22T15:12:23.614864Z"}}},{"cell_type":"markdown","source":"## Table of Content\n\n- [1 - Packages](#1)\n    - [1.1 Create the Dataset](#1-1)\n    - [1.2 Initilize the DataLoader Class and Load the DataSets](#1-2)\n    - [1.3 Split the DataSets for Training and Validation](#1-3)\n- [2 - Build HMS Model using MobileNetV2 and Transfer Learning](#2)\n    - [2.1 - Layer Freezing with the Functional API](#2-1)\n    - [2.2 - Fine-tuning the Model](#2-2)\n    - [2.3 - Create Model](#2-3)\n    - [2.4 - Train Model](#2-4)\n    - [2.5 - Test Model](#2-5)\n- [3 - Submission](#2)\n\n     ","metadata":{}},{"cell_type":"markdown","source":"<a name='1'></a>\n## 1 - Packages","metadata":{"execution":{"iopub.execute_input":"2024-02-22T16:02:47.712685Z","iopub.status.busy":"2024-02-22T16:02:47.712384Z","iopub.status.idle":"2024-02-22T16:02:47.719281Z","shell.execute_reply":"2024-02-22T16:02:47.717831Z","shell.execute_reply.started":"2024-02-22T16:02:47.712663Z"}}},{"cell_type":"code","source":"# import required libraries\nimport pandas as pd \nimport numpy as np\nimport os \nfrom concurrent.futures import ThreadPoolExecutor\nimport timeit\nfrom sklearn.model_selection import train_test_split\nimport pickle\nimport tensorflow as tf\nimport tensorflow.keras.layers as tfl\nimport cv2\nimport matplotlib.pyplot as plt\nfrom tensorflow.keras.layers import Dense\nfrom sklearn.utils import shuffle\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.applications import MobileNetV2","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:49:01.917222Z","iopub.execute_input":"2024-02-24T20:49:01.917923Z","iopub.status.idle":"2024-02-24T20:49:08.940428Z","shell.execute_reply.started":"2024-02-24T20:49:01.917887Z","shell.execute_reply":"2024-02-24T20:49:08.938917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='1-1'></a>\n### 1.1 Load Dataset \n- Load the metadata from the training dataset.\n- Load all spectrogram data.\n- Resize spectrogram images to the input size of (224, 224, 3), which is compatible with the MobileNetV2 model.\n- MobileNetV2 is a type of Efficient Convolutional Neural Network designed for mobile vision applications.\n- We will create a DataLoader class that reads spectrogram image data, resizes it to  fit the input format of MobileNetV2, and calculates the probabilities of the following categories: seizure (SZ), generalized periodic discharges (GPD), lateralized periodic discharges (LPD), lateralized rhythmic delta activity (LRDA), generalized rhythmic delta activity (GRDA), or \"other.\"\n\n","metadata":{}},{"cell_type":"code","source":"## This is the class for loading data sets \n\nclass DataLoader:\n    def __init__(self, base_floder_path):\n        \"\"\"\n        Initializes the DataLoader class with the file path.\n        \n        Parameters:\n        - file_path (str): The path to the data file.\n        \"\"\"\n        self.base_folder = base_floder_path\n        self.train_meta_file = os.path.join(self.base_folder, 'train.csv')\n        self.cutoff=1e-10 ### this is for removing zeros\n        self.desired_size = (224, 224)\n        self.data = None\n        \n    def resize_image_to_square(self,image):\n        \"\"\"\n        Resize a rectangular image to a square image by adding padding.\n\n        Parameters:\n        - image: Input image as a NumPy array.\n        - desired_size: Desired size of the square image (e.g., (224, 224)).\n\n        Returns:\n        - square_image: Square image as a NumPy array.\n        \"\"\"\n        desired_size=self.desired_size\n        # Get the dimensions of the input image\n        height, width = image.shape[:2]\n\n        # Compute the dimensions of the padding\n        delta_width = max(0, desired_size[1] - width)\n        delta_height = max(0, desired_size[0] - height)\n\n        # Compute the padding amounts for the top, bottom, left, and right sides\n        top = delta_height // 2\n        bottom = delta_height - top\n        left = delta_width // 2\n        right = delta_width - left\n\n        # Add padding to the image\n        padded_image = cv2.copyMakeBorder(image, top, bottom, left, right, cv2.BORDER_CONSTANT, value=(0, 0, 0))\n\n        # Resize the padded image to the desired size\n        square_image = cv2.resize(padded_image, desired_size, interpolation=cv2.INTER_AREA)\n\n        return square_image\n    \n\n    def load_data(self):\n        \"\"\"\n        Loads the data from the specified file path.\n        This method should be implemented according to the data source.\n        \"\"\"\n        ## load the meta data from training set \n        self.df=pd.read_csv(self.train_meta_file)\n        res_df=self.get_voting_probability()\n#         print(res_df.head())\n        self.spec_df=res_df\n        # Assuming filenames is a list of filenames to read\n        filenames = res_df['filename'].to_list()\n\n        # Read files in parallel and store the returned data\n        data = self.read_files_parallel(filenames)\n        # Print the returned data    \n        print(len(data))\n        \n        return data\n        \n        \n    def read_file(self,filename):\n        \"\"\"\n        Read the spectrogram file. \n\n        Input Parameters:\n        filename: *.parquet\n\n        Output:\n        - res_df: DataFrame with probabilities for each spectrogram\n        \"\"\"\n        spectrogram_df = pd.read_parquet(filename)\n        time_points = spectrogram_df['time']\n        array=spectrogram_df.to_numpy()\n        array[np.isnan(array)] = 0\n        array[array<self.cutoff]=self.cutoff\n        \n        array =np.log10(array)\n        ## resize it \n        \n        # Resize the images to square images\n        image = self.resize_image_to_square(array.T)\n        image = (image - np.min(image)) / (np.max(image) - np.min(image)) * 255\n\n        # Convert the 2D spectrogram image to a 3D RGB image\n        rgb_image = np.stack((image,) * 3, axis=-1).astype(np.uint8)\n    \n        return rgb_image,time_points\n        \n    def read_files_parallel(self, filenames):\n        \"\"\"\n        Read files in parallel using ThreadPoolExecutor.\n\n        Input Parameters:\n        filenames: List of filenames to read.\n\n        Output:\n        - results: List of tuples containing returned data from read_file function.\n        \"\"\"\n        results = []\n        # Define the number of threads (workers) to use\n        num_threads = 10  # You can adjust this number based on your system and workload\n        # Create a ThreadPoolExecutor with the desired number of threads\n        with ThreadPoolExecutor(max_workers=num_threads) as executor:\n            # Use map to submit tasks and maintain order\n            for result in executor.map(self.read_file, filenames):\n                results.append(result)\n        ##\n        return results\n    \n    # set augmentation \n    def data_augmenter():\n        '''\n        Create a Sequential model composed of 2 layers\n        Returns:\n            tf.keras.Sequential\n        '''\n\n        data_augmentation = tf.keras.Sequential()\n        data_augmentation.add(tf.keras.layers.experimental.preprocessing.RandomFlip('horizontal'))\n        data_augmentation.add(tf.keras.layers.experimental.preprocessing.RandomRotation(0.2))\n\n\n        return data_augmentation\n        \n    def get_voting_probability(self):\n        \"\"\"\n        Returns the probabilities for each spectrogram. \n\n        Input Parameters:\n        None\n\n        Output:\n        - res_df: DataFrame with probabilities for each spectrogram\n        \"\"\"\n        # Create a copy of the DataFrame to avoid modifying the original\n        df = self.df.copy()\n        dfs=[]\n        # Create an empty DataFrame with columns to store results\n        res_df = pd.DataFrame(columns=['spectrogram_id', 'filename', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote'])\n\n        # Get unique spectrogram IDs\n        all_spectrogram_ids = df['spectrogram_id'].unique()\n\n        # Iterate over each unique spectrogram ID\n        for id in all_spectrogram_ids:\n            # Filter DataFrame to include only rows with the current spectrogram ID\n            tmp = df[df['spectrogram_id'] == id]\n\n            # Generate filename for the spectrogram\n            filenames = os.path.join(self.base_folder, 'train_spectrograms', str(id) + '.parquet')\n\n            # Extract relevant voting columns and calculate voting probabilities\n            tmp1 = tmp[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].copy()\n            voting = tmp1.sum(axis=0) / tmp1.sum(axis=0).sum()\n            # Append results to the list of DataFrames\n            tmp_res_df = pd.DataFrame({'spectrogram_id': [id], 'filename': [filenames], **voting})\n            dfs.append(tmp_res_df)\n            # Check if the sum of voting probabilities is less than 0.99\n            if voting.sum() < 0.99:\n                print('Check:', id)\n        res_df = pd.concat(dfs, ignore_index=True)\n        \n        return res_df\n\n            \n        ","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:49:08.947382Z","iopub.execute_input":"2024-02-24T20:49:08.947735Z","iopub.status.idle":"2024-02-24T20:49:08.970824Z","shell.execute_reply.started":"2024-02-24T20:49:08.947706Z","shell.execute_reply":"2024-02-24T20:49:08.969724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='1-2'></a>\n### 1.2 Initilize the DataLoader Class and Load the DataSets \nTo initialize the DataLoader class and load the datasets, follow these steps:\n\n- Initialize a data_loader object using the DataLoader class.\n- Execute the load_data function.","metadata":{}},{"cell_type":"code","source":"base_floder_path=\"/kaggle/input/hms-harmful-brain-activity-classification\"\ndata_loader = DataLoader(base_floder_path)\n\n# Your code here\n\nstart_time = timeit.default_timer()\n\ndata=data_loader.load_data()\n\nelapsed_time = timeit.default_timer() - start_time\nprint(\"Elapsed time in loading the data\", elapsed_time, \"seconds\")","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:49:08.972178Z","iopub.execute_input":"2024-02-24T20:49:08.973732Z","iopub.status.idle":"2024-02-24T20:53:22.239991Z","shell.execute_reply.started":"2024-02-24T20:49:08.973665Z","shell.execute_reply":"2024-02-24T20:53:22.237385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='1-3'></a>\n### 1.3 Split the DataSets for Training and Validation \nWe will split the entire dataset into 80% for training and 20% for validation sets","metadata":{}},{"cell_type":"markdown","source":"Retrieve a list containing spectrogram image data along with their corresponding probabilities for the six classes.","metadata":{}},{"cell_type":"code","source":"X=[ d[0] for d in data]\nY=data_loader.spec_df[['seizure_vote', 'lpd_vote', 'gpd_vote',\n       'lrda_vote', 'grda_vote', 'other_vote']].copy()\ndel data","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:22.243541Z","iopub.execute_input":"2024-02-24T20:53:22.244206Z","iopub.status.idle":"2024-02-24T20:53:22.350209Z","shell.execute_reply.started":"2024-02-24T20:53:22.244172Z","shell.execute_reply":"2024-02-24T20:53:22.348491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This function prepares input data (X) and corresponding labels (y) for TensorFlow training by shuffling the data and labels, converting them into TensorFlow datasets, and setting a specified batch size for efficient training. It takes parameters X (input data), y (labels), and batch_size (size of training batches) and returns a TensorFlow dataset containing shuffled and batched data and labels.\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"def data_generator(X,Y,batch_size = 32):\n    # This function prepares the input data (X) and corresponding labels (y) for TensorFlow training.\n    # It performs the following steps:\n    # 1. Shuffles the input data and labels to ensure randomness during training.\n    # 2. Converts the shuffled data and labels into TensorFlow datasets, which offer efficient data handling.\n    # 3. Sets a specified batch size for the datasets, allowing the model to process data in smaller subsets (batches)\n    #    during training, enhancing computational efficiency and memory usage.\n    #\n    # Parameters:\n    # - X: Input data, typically a NumPy array or a list containing features.\n    # - y: Labels associated with the input data, usually a NumPy array or a list.\n    # - batch_size: The size of batches used during training.\n    #\n    # Returns:\n    # - TensorFlow dataset containing the shuffled and batched data and labels.\n    #\n\n    # Shuffle X_train and Y_train together\n    X_shuffled, Y_shuffled = shuffle(X, Y)\n    # Zip X_train and Y_train together\n    data = zip(X_shuffled, Y_shuffled.values)  # Convert Y_train to values to extract arrays\n   # Create TensorFlow dataset from zipped data\n    dataset = tf.data.Dataset.from_generator(\n    lambda: data,\n    output_signature=(\n        tf.TensorSpec(shape=(224, 224, 3), dtype=tf.float64),  # Shape of X_train\n        tf.TensorSpec(shape=(6,), dtype=tf.float64)  # Assuming Y_train has 6 classes\n        )\n    )\n\n    # Batch the dataset\n    dataset = dataset.batch(batch_size)\n    return dataset\n","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:22.352133Z","iopub.execute_input":"2024-02-24T20:53:22.353595Z","iopub.status.idle":"2024-02-24T20:53:22.362251Z","shell.execute_reply.started":"2024-02-24T20:53:22.353520Z","shell.execute_reply":"2024-02-24T20:53:22.360782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we split the data into training and validation sets and then invoke the data_generator function to prepare the training and validation datasets.","metadata":{}},{"cell_type":"code","source":"X_train, X_valid, Y_train, Y_valid = train_test_split(X, Y, test_size=0.2, random_state=42)\n# X_train, X_valid, Y_train, Y_valid\nprint(len(X_train),len(X_valid))\ntrain_dataset=data_generator(X_train,Y_train)\ndel X_train\ndel Y_train\nvalid_dataset=data_generator(X_valid,Y_valid)\ndel X_valid\ndel Y_valid\ntrain_dataset","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:22.363763Z","iopub.execute_input":"2024-02-24T20:53:22.364169Z","iopub.status.idle":"2024-02-24T20:53:22.524266Z","shell.execute_reply.started":"2024-02-24T20:53:22.364132Z","shell.execute_reply":"2024-02-24T20:53:22.523133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='2'></a>\n## 2 - Build HMS Model using MobileNetV2 and Transfer Learning\n\nMobileNetV2, originally trained on ImageNet, is designed for mobile and low-power applications, featuring 155 layers. It's highly efficient for tasks such as object detection, image segmentation, and classification, including the current one. The architecture boasts three key features: \n*   Depthwise separable convolutions\n*   Thin input and output bottlenecks between layers\n*   Shortcut connections between bottleneck layers\n\nThese characteristics contribute to its efficiency and effectiveness across various tasks and platforms.\n\nWe aim to construct a model for Harmful Brain Activity Classification by excluding the top layer from MobileNetV2 and adapting it to this specific problem using transfer learning techniques.","metadata":{}},{"cell_type":"markdown","source":"<a name='2-1'></a>\n### 2.1 - Layer Freezing with the Functional API\n\nIn the following sections, I'll guide you through the process of utilizing a pre-trained model to modify the classification task for identifying Harmful Brain Activity effectively. This process involves three simple steps:\n\n- Remove the `top layer`, which serves as the classification layer, by setting `include_top` in base_model to False.\n- Introduce a new classifier layer where you train only one layer while keeping the remainder of the network frozen. \n- Freeze the base model and exclusively train the newly-added classifier layer. This is achieved by setting `base_model.trainable=False` to prevent weight alteration.\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"<a name='2-2'></a>\n### 2.2 - Fine-tuning the Model\nConsider fine-tuning the model to improve accuracy by lowering the optimizer in the last layers to a lower learning rate. In transfer learning, unfreezing the end layers and re-training them with a very low learning rate enables finer modifications to align the model with new input, potentially boosting accuracy. This approach is facilitated by unfreezing the final layers and changing the optimizer to a lower learning rate while keeping other levels frozen. ","metadata":{}},{"cell_type":"code","source":"# change the pretrained models \n\ndef hms_model(input_shape):\n    ''' Define a tf.keras model for multi class classification out of the MobileNetV2 model\n    Arguments:\n        input_shape -- Image width, height, and channels\n    Returns:\n        tf.keras.model\n    '''\n    \n    # Load the MobileNetV2 model with pre-trained weights, excluding the top (classification) layer\n#     base_model = tf.keras.applications.MobileNetV2(input_shape=input_shape,\n#                                                include_top=False,\n#                                                weights='imagenet')\n    base_model = MobileNetV2(weights='/kaggle/input/mobilenetv2weight/mobilenet_v2_weights_tf_dim_ordering_tf_kernels_1.0_224_no_top.h5', include_top=False, input_shape=input_shape)\n    \n    # Freeze the parameters of the base model\n#     base_model.trainable = False\n    \n     # Freeze all layers except the last two\n    for layer in base_model.layers[:-2]:\n        layer.trainable = False\n\n    # Preprocess the input images to the range of [-1, 1]\n    inputs = tf.keras.layers.Input(shape=input_shape)\n    x = tf.keras.applications.mobilenet_v2.preprocess_input(inputs)\n\n    # Pass the inputs through the base model\n    x = base_model(x)\n\n    # Add a Global Average Pooling layer to reduce spatial dimensions\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n\n    # Add a custom output layer for 6 classes on top of the base model\n    output_layer = tf.keras.layers.Dense(6, activation='softmax')(x)\n\n    # Create a new model by defining inputs and outputs\n    model = tf.keras.Model(inputs=inputs, outputs=output_layer)\n    \n    return model\n","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:22.525679Z","iopub.execute_input":"2024-02-24T20:53:22.526468Z","iopub.status.idle":"2024-02-24T20:53:22.534748Z","shell.execute_reply.started":"2024-02-24T20:53:22.526436Z","shell.execute_reply":"2024-02-24T20:53:22.532747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='2-3'></a>\n### 2.3 - Create HMS model","metadata":{}},{"cell_type":"code","source":"model=hms_model(X[0].shape)\nprint(model.summary())","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:22.536479Z","iopub.execute_input":"2024-02-24T20:53:22.536945Z","iopub.status.idle":"2024-02-24T20:53:25.368879Z","shell.execute_reply.started":"2024-02-24T20:53:22.536901Z","shell.execute_reply":"2024-02-24T20:53:25.367186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='2-4'></a>\n### 2.4 - Train Model","metadata":{}},{"cell_type":"code","source":"\n\n# Define your learning rate\nlearning_rate = 0.0001  # You can adjust this value based on your needs\n\n# Create an instance of the Adam optimizer with the specified learning rate\noptimizer = Adam(learning_rate=learning_rate)\n\n# model.compile(optimizer='adam', loss='kullback_leibler_divergence')\n\nmodel.compile(optimizer=optimizer, loss='kullback_leibler_divergence')","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:25.370536Z","iopub.execute_input":"2024-02-24T20:53:25.370902Z","iopub.status.idle":"2024-02-24T20:53:25.398667Z","shell.execute_reply.started":"2024-02-24T20:53:25.370871Z","shell.execute_reply":"2024-02-24T20:53:25.396658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='2-5'></a>\n### 2.4 - Test Model","metadata":{}},{"cell_type":"code","source":"initial_epochs = 5\nhistory = model.fit(train_dataset, validation_data=valid_dataset, epochs=initial_epochs)","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:53:25.400383Z","iopub.execute_input":"2024-02-24T20:53:25.400810Z","iopub.status.idle":"2024-02-24T20:56:29.072935Z","shell.execute_reply.started":"2024-02-24T20:53:25.400779Z","shell.execute_reply":"2024-02-24T20:56:29.071986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#predict the probabilities for test case \nfilename = os.path.join(base_floder_path, 'test_spectrograms', '853520.parquet')\ntest_arr,test_time_points=data_loader.read_file(filename)\nprint(test_arr.shape,test_time_points.shape)\ntest_square_image=data_loader.resize_image_to_square(test_arr)\nprint(test_square_image.shape)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:56:29.074961Z","iopub.execute_input":"2024-02-24T20:56:29.076245Z","iopub.status.idle":"2024-02-24T20:56:29.236735Z","shell.execute_reply.started":"2024-02-24T20:56:29.076183Z","shell.execute_reply":"2024-02-24T20:56:29.235270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## test the probabilities with prediction \ninput_data_expanded = np.expand_dims(test_square_image, axis=0)\npredictions = model.predict(input_data_expanded, batch_size=1)\nprint(predictions)","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:56:29.238892Z","iopub.execute_input":"2024-02-24T20:56:29.239300Z","iopub.status.idle":"2024-02-24T20:56:30.111252Z","shell.execute_reply.started":"2024-02-24T20:56:29.239267Z","shell.execute_reply":"2024-02-24T20:56:30.110077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name='3'></a>\n## 3 - Submission","metadata":{}},{"cell_type":"code","source":"submission_df = pd.read_csv(os.path.join(base_floder_path, 'sample_submission.csv'))\n# sub_df = sub_df[[\"eeg_id\"]].copy()\n# sub_df = sub_df.merge(pred_df, on=\"eeg_id\", how=\"left\")\n# sub_df.to_csv(\"submission.csv\", index=False)\nsubmission_df.head()\nsubmission_df.iloc[0,1:]=predictions\nsubmission_df.to_csv(\"submission.csv\", index=False)\nsubmission_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-02-24T20:56:30.115952Z","iopub.execute_input":"2024-02-24T20:56:30.116423Z","iopub.status.idle":"2024-02-24T20:56:30.149232Z","shell.execute_reply.started":"2024-02-24T20:56:30.116389Z","shell.execute_reply":"2024-02-24T20:56:30.147426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}