{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30635,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **I will develop this Model using EfficientNetB2 **\n\nEfficientNetB2 is not a library but rather a specific variant of the EfficientNet architecture, which is a family of convolutional neural networks (CNNs) designed for efficient and effective use of computational resources. The EfficientNet model was proposed by Mingxing Tan and Quoc V. Le in the paper titled \"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,\" published in 2019.\n\nEfficientNetB2 is part of the EfficientNet family and represents a specific configuration or scale within the broader architecture. The \"B2\" indicates the model's depth, width, and resolution configuration, where higher numbers generally represent larger and more computationally intensive models.\n\nThe EfficientNet architecture introduces a compound scaling method that uniformly scales the network's depth, width, and resolution to find a good balance between model size and performance. This approach allows EfficientNet models to achieve state-of-the-art performance with fewer parameters compared to other architectures.\n\nTo use EfficientNetB2 or other variants, you would typically leverage deep learning frameworks like TensorFlow or PyTorch, which provide pre-trained models and APIs for working with these architectures. These frameworks offer implementations of EfficientNet models, including pre-trained weights that can be fine-tuned on specific tasks or used for feature extraction in various computer vision applications.","metadata":{}},{"cell_type":"markdown","source":"**Refernces : **\n\n[EfficientNetB0 Starter - [LB 0.43]](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43)\n\n\nhttps://www.kaggle.com/code/andreasbis/hms-inference-lb-0-42","metadata":{}},{"cell_type":"markdown","source":"# **Overview**\nThe goal of this competition is to detect and classify seizures and other types of harmful brain activity. You will develop a model trained on electroencephalography (EEG) signals recorded from critically ill hospital patients.\n\nYour work may help rapidly improve electroencephalography pattern classification accuracy, unlocking transformative benefits for neurocritical care, epilepsy, and drug development. Advancement in this area may allow doctors and brain researchers to detect seizures or other brain damage to provide faster and more accurate treatments.","metadata":{}},{"cell_type":"markdown","source":"# **Description of the probelm :**\n    Doctors use various tools, from stethoscopes to tongue depressors, in patient treatment. Electroencephalography (EEG) aids in detecting seizures and other brain activities causing potential damage. Currently, manual EEG analysis by specialized neurologists is time-consuming, expensive, prone to errors, and lacks reliability. The Sunstella Foundation, established in 2021, supports minority graduate students in technology through mentorship and competitions. Partnering with Persyst, Jazz Pharmaceuticals, and CDAC, the foundation aims to preserve and enhance brain health. Your role in automating EEG analysis will expedite accurate detection of seizures, aiding prompt treatment. This contest's algorithms may also assist drug development for seizure treatment. The competition focuses on six patterns: SZ, GPD, LPD, LRDA, GRDA, or \"other,\" with annotated EEG segments categorized as idealized, proto patterns, or edge cases based on expert agreement.","metadata":{}},{"cell_type":"markdown","source":"# **Few details about the problem :**\nRefer to the Data tab for full-screen PDFs of EEG patterns in rows: seizure, LPDs, GPDs, LRDA, and GRDA. Columns show idealized patterns (A), partially formed patterns (B), and edge cases (C, D) with disagreement between raters. For instance, B-1 suggests rhythmic delta activity, possibly a seizure. B-2 displays frontal transients, split between LPD and \"Other.\" B-3 features semi-rhythmic delta as a proto-GPD, while B-4 shows proto-LRDA. B-5 exhibits proto-GRDA. Edge cases, like C-1 and D-1, show evolving patterns, providing insight into EEG electrode regions (LL, RL, LP, RP).\n\n# **in Arabic**\nرجاءً راجع علامة \"البيانات\" للحصول على صفحات PDF بحجم كامل لأمثلة محددة لأنماط EEG في صفوف: النوبة الصرعية، والنوم الجزئي المتكرر، والتصريف العام، ونشاط الدلتا الجانبي الناجم، ونشاط الدلتا الجانبي العام. الأعمدة تظهر أمثلة مثالية (A)، وأمثلة جزئية (B)، وحالات حافة (C، D) مع عدم اتفاق بين الحكام. على سبيل المثال، B-1 تظهر نشاطًا دلتا رائعًا، ربما يكون نوبة صرعية. B-2 تُظهر فترات حادة جزئية في الجبهة، مُقسمة بين التصريف الجزئي و\"آخر\". B-3 تعرض دلتا نصف رائعة مع خصائص التصريف العام الجزئي. B-4 تظهر نشاط دلتا نصف رائع كنمط تصريف جانبي جانبي جديد. B-5 تُظهر نمطًا جانبيًا نصف رائعًا. حالات الحافة، مثل C-1 و D-1، تُظهر أنماطًا تتطور، مما يوفر رؤية في مناطق الأقطاب EEG (LL، RL، LP، RP).","metadata":{}},{"cell_type":"markdown","source":"# **how to select the best Machine Learning Model when creating a model for solving some problem ?**\n\nRef.\nhttps://medium.com/@osama.ghandour/how-to-select-the-best-machine-learning-model-when-creating-a-model-for-solving-some-problem-08407e59ac21\n\nSelecting the best machine learning model for a given problem involves a combination of understanding the nature of the problem, exploring different algorithms, and evaluating their performance. Here is a step-by-step guide to help you choose the best machine learning model:\n\n1. Define the Problem:\n   - Clearly define the problem you are trying to solve, whether it's classification, regression, clustering, or another type of task.\n\n2. Understand the Data:\n   - Analyze the characteristics of your data, including its size, structure, and distribution.\n   - Handle missing values, outliers, and other data preprocessing tasks.\n\n3. Determine the Input Features:\n   - Identify the relevant features that will be used to train the model.\n   - Consider feature engineering to create new features or transform existing ones.\n\n4. Select Evaluation Metrics:\n   - Choose appropriate metrics for evaluating model performance based on the nature of the problem (accuracy, precision, recall, F1 score, etc.).\n\n5. Explore Different Algorithms:\n   - Start with simple models and gradually move to more complex ones.\n   - Consider algorithms suitable for your problem (e.g., decision trees, support vector machines, neural networks).\n\n6. Split the Data:\n   - Split your dataset into training, validation, and test sets.\n   - Use the training set to train models, the validation set to tune hyperparameters, and the test set to evaluate final performance.\n\n7. Train Multiple Models:\n   - Train several different algorithms on your training data.\n   - Experiment with different hyperparameter settings for each algorithm.\n\n8. Cross-Validation:\n   - Use techniques like k-fold cross-validation to assess the model's performance across different subsets of the data.\n\n9. Compare Performance:\n   - Evaluate the models using the chosen metrics on the validation set.\n   - Compare performance and identify models that generalize well to unseen data.\n\n10. Tune Hyperparameters:\n    - Fine-tune the hyperparameters of the chosen models to optimize performance.\n    - Consider techniques like grid search or random search for hyperparameter tuning.\n\n11. Feature Importance:\n    - Investigate feature importance to understand which features contribute most to the model's predictions.\n    - This is particularly useful for decision tree-based models.\n\n12. Ensemble Methods:\n    - Explore ensemble methods like random forests or gradient boosting to combine the strengths of multiple models.\n\n13. Evaluate on Test Set:\n    - Evaluate the final models on the test set to ensure unbiased performance assessment.\n    - Compare the models and select the one with the best overall performance.\n\n14. Consider Model Interpretability:\n    - Depending on the application, consider the interpretability of the model. Some models are more interpretable than others (e.g., decision trees).\n\n15. Monitor for Overfitting:\n    - Be cautious of overfitting, especially with complex models. Regularization techniques can help mitigate this.\n\n16. Documentation:\n    - Document the chosen model, hyperparameters, and any relevant information for future reference.\n\nRemember that the best model can vary depending on the specific characteristics of your data and problem, so it's essential to experiment and iterate based on your findings.","metadata":{}},{"cell_type":"markdown","source":"#  Using PyTorch to create a model using EfficientNetB2 for detecting and classifying EEG patterns:","metadata":{}},{"cell_type":"markdown","source":"# **Importing Liberaries **","metadata":{}},{"cell_type":"code","source":"# Importing essential libraries\nimport gc\nimport os\nimport random\nimport warnings\nimport numpy as np\nimport pandas as pd\nfrom IPython.display import display\n\n# PyTorch for deep learning\nimport timm\nimport torch\nimport torch.nn as nn  \nimport torch.optim as optim\nimport torch.nn.functional as F\n\nfrom torch.utils.data import DataLoader, Dataset\nfrom torchvision import transforms\nfrom efficientnet_pytorch import EfficientNet\n\n\n# torchvision for image processing and augmentation\nimport torchvision.transforms as transforms\n\n# Suppressing minor warnings to keep the output clean\nwarnings.filterwarnings('ignore', category=Warning)\n\n# Reclaim memory no longer in use.\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-09T02:42:51.056761Z","iopub.execute_input":"2024-02-09T02:42:51.057263Z","iopub.status.idle":"2024-02-09T02:42:56.887872Z","shell.execute_reply.started":"2024-02-09T02:42:51.057219Z","shell.execute_reply":"2024-02-09T02:42:56.886461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"from ChatGPT","metadata":{}},{"cell_type":"code","source":"\n\n# Define paths to your training and testing data CSV files\ntrain_csv_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\ntest_csv_path = '/kaggle/input/hms-harmful-brain-activity-classification/test.csv'\n\n# Assuming we have a custom EEGDataset class for loading data\nclass EEGDataset(Dataset):\n    def __init__(self, csv_path, transform=None):\n        # Implement your dataset loading logic using the CSV file\n        # Make sure to preprocess the data appropriately\n        pass\n\n    def __len__(self):\n        # Return the number of samples in your dataset\n        pass\n\n    def __getitem__(self, idx):\n        # Return a sample from your dataset\n        pass\n\n# Define transformations for data augmentation\ntransform = transforms.Compose([\n    transforms.Resize((224, 224)),\n    transforms.ToTensor(),\n    # Add more transformations if needed\n])\n\n# Create instances of your custom dataset\ntrain_dataset = EEGDataset(train_csv_path, transform=transform)\ntest_dataset = EEGDataset(test_csv_path, transform=transform)\n\n# Create data loaders\ntrain_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)\ntest_loader = DataLoader(test_dataset, batch_size=32, shuffle=False)\n\n# Build EfficientNetB2 model\nclass EEGModel(nn.Module):\n    def __init__(self, num_classes=5):  # Assuming 5 classes for the EEG patterns\n        super(EEGModel, self).__init__()\n        self.efficient_net = EfficientNet.from_pretrained('efficientnet-b2')\n        self.fc = nn.Linear(1408, num_classes)\n\n    def forward(self, x):\n        x = self.efficient_net(x)\n        x = x.view(x.size(0), -1)\n        x = self.fc(x)\n        return x\n\n# Instantiate the model\nmodel = EEGModel()\n\n# Define loss function and optimizer\ncriterion = nn.CrossEntropyLoss()\noptimizer = optim.Adam(model.parameters(), lr=0.001)\n\n# Training loop\nnum_epochs = 10\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\nmodel.to(device)\n\nfor epoch in range(num_epochs):\n    model.train()\n    for inputs, labels in train_loader:\n        inputs, labels = inputs.to(device), labels.to(device)\n        optimizer.zero_grad()\n        outputs = model(inputs)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n\n# Evaluate the model on the test set\nmodel.eval()\ncorrect = 0\ntotal = 0\n\nwith torch.no_grad():\n    for inputs, labels in test_loader:\n        inputs, labels = inputs.to(device), labels.to(device)\n        outputs = model(inputs)\n        _, predicted = torch.max(outputs.data, 1)\n        total += labels.size(0)\n        correct += (predicted == labels).sum().item()\n\naccuracy = correct / total\nprint(f'Test accuracy: {accuracy}')\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****SetUP / Configuration ****","metadata":{}},{"cell_type":"code","source":"class Config:\n    seed=42\n    image_transform=transforms.Resize((512, 512))\n    num_folds=5\n    \n# Set the seed for reproducibility across multiple libraries\ndef set_seed(seed):\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = True\n    torch.manual_seed(seed)\n    np.random.seed(seed)\n    random.seed(seed)\n    \nset_seed(Config.seed)","metadata":{"execution":{"iopub.status.busy":"2024-02-09T02:43:04.598226Z","iopub.execute_input":"2024-02-09T02:43:04.598607Z","iopub.status.idle":"2024-02-09T02:43:04.605667Z","shell.execute_reply.started":"2024-02-09T02:43:04.598571Z","shell.execute_reply":"2024-02-09T02:43:04.604327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Get Data**","metadata":{}},{"cell_type":"code","source":"# Load and store the trained models for each fold into a list\nmodels = []\n\n# Load ResNet34d\nfor i in range(Config.num_folds):\n    # Create the same model architecture as during training\n    model_resnet = timm.create_model('resnet34d', pretrained=False, num_classes=6, in_chans=1)\n    \n    # Load the trained weights from the corresponding file\n    model_resnet.load_state_dict(torch.load(f'/kaggle/input/hms-train-resnet34d/resnet34d_fold{i}.pth', map_location=torch.device('cpu')))\n    \n    # Append the loaded model to the models list\n    models.append(model_resnet)\n\n# Reclaim memory no longer in use.\ngc.collect()\n\n# Load EfficientNetB0\nfor j in range(Config.num_folds):\n    # Create the same model architecture as during training\n    model_effnet = timm.create_model('efficientnet_b0', pretrained=False, num_classes=6, in_chans=1)\n    \n    # Load the trained weights from the corresponding file\n    model_effnet.load_state_dict(torch.load(f'/kaggle/input/hms-train-efficientnetb0/efficientnet_b0_fold{j}.pth', map_location=torch.device('cpu')))\n    \n    # Append the loaded model to the models list\n    models.append(model_effnet)\n\n# Reclaim memory no longer in use.\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-02-09T02:43:29.18409Z","iopub.execute_input":"2024-02-09T02:43:29.184529Z","iopub.status.idle":"2024-02-09T02:43:29.583383Z","shell.execute_reply.started":"2024-02-09T02:43:29.184496Z","shell.execute_reply":"2024-02-09T02:43:29.58165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load test data and sample submission dataframe\ntrain_df=pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/test.csv\")\nsubmission = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")\n\n# Merge the submission dataframe with the test data on EEG IDs\nsubmission = submission.merge(test_df, on='eeg_id', how='left')\n\n# Generate file paths for each spectrogram based on the EEG data in the submission dataframe\nsubmission['path'] = submission['spectrogram_id'].apply(lambda x: f\"/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms/{x}.parquet\")\n\n# Display the first few rows of the submission dataframe\ndisplay(submission.head())\n\n# Reclaim memory no longer in use\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-09T02:43:52.08271Z","iopub.execute_input":"2024-02-09T02:43:52.083054Z","iopub.status.idle":"2024-02-09T02:43:52.525761Z","shell.execute_reply.started":"2024-02-09T02:43:52.08303Z","shell.execute_reply":"2024-02-09T02:43:52.524903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Predictions**","metadata":{}},{"cell_type":"code","source":"# Get file paths for test spectrograms\npaths = submission['path'].values\ntest_preds = []\n\n# Generate predictions for each spectrogram using all models\nfor path in paths:\n    eps = 1e-6\n    # Read and preprocess spectrogram data\n    data = pd.read_parquet(path)\n    data = data.fillna(-1).values[:, 1:].T\n    data = np.clip(data, np.exp(-6), np.exp(10))\n    data = np.log(data)\n    \n    # Normalize the data\n    data_mean = data.mean(axis=(0, 1))\n    data_std = data.std(axis=(0, 1))\n    data = (data - data_mean) / (data_std + eps)\n    data_tensor = torch.unsqueeze(torch.Tensor(data), dim=0)\n    data = Config.image_transform(data_tensor)\n\n    test_pred = []\n    \n    # Generate predictions using all models\n    for model in models:\n        model.eval()\n        with torch.no_grad():\n            pred = F.softmax(model(data.unsqueeze(0)))[0]\n            pred = pred.detach().cpu().numpy()\n        test_pred.append(pred)\n        \n    # Combine predictions from all models using soft voting\n    test_pred = np.mean(test_pred, axis=0)  # Soft Voting (equal weights)\n    test_preds.append(test_pred)\n\n# Convert the list of predictions to a NumPy array for further processing\ntest_preds = np.array(test_preds)\n\n# Reclaim memory no longer in use.\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-02-09T02:44:00.062564Z","iopub.execute_input":"2024-02-09T02:44:00.062944Z","iopub.status.idle":"2024-02-09T02:44:00.443972Z","shell.execute_reply.started":"2024-02-09T02:44:00.062916Z","shell.execute_reply":"2024-02-09T02:44:00.443215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  **Evaluation code ** \nwas copied from this link: https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook in Kaggel  between the predicted probability and the observed target.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pandas.api.types\n\nimport kaggle_metric_utilities\n\nfrom typing import Optional\n\n\nclass ParticipantVisibleError(Exception):\n    pass\n\n\ndef kl_divergence(solution: pd.DataFrame, submission: pd.DataFrame, epsilon: float, micro_average: bool, sample_weights: Optional[pd.Series]):\n    # Overwrite solution for convenience\n    for col in solution.columns:\n        # Prevent issue with populating int columns with floats\n        if not pandas.api.types.is_float_dtype(solution[col]):\n            solution[col] = solution[col].astype(float)\n\n        # Clip both the min and max following Kaggle conventions for related metrics like log loss\n        # Clipping the max avoids cases where the loss would be infinite or undefined, clipping the min\n        # prevents users from playing games with the 20th decimal place of predictions.\n        submission[col] = np.clip(submission[col], epsilon, 1 - epsilon)\n\n        y_nonzero_indices = solution[col] != 0\n        solution[col] = solution[col].astype(float)\n        solution.loc[y_nonzero_indices, col] = solution.loc[y_nonzero_indices, col] * np.log(solution.loc[y_nonzero_indices, col] / submission.loc[y_nonzero_indices, col])\n        # Set the loss equal to zero where y_true equals zero following the scipy convention:\n        # https://docs.scipy.org/doc/scipy/reference/generated/scipy.special.rel_entr.html#scipy.special.rel_entr\n        solution.loc[~y_nonzero_indices, col] = 0\n\n    if micro_average:\n        return np.average(solution.sum(axis=1), weights=sample_weights)\n    else:\n        return np.average(solution.mean())\n\n\ndef score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        epsilon: float=10**-15,\n        micro_average: bool=True,\n        sample_weights_column_name: Optional[str]=None\n    ) -> float:\n    ''' The Kullback–Leibler divergence.\n    The KL divergence is technically undefined/infinite where the target equals zero.\n\n    This implementation always assigns those cases a score of zero; effectively removing them from consideration.\n    The predictions in each row must add to one so any probability assigned to a case where y == 0 reduces\n    another prediction where y > 0, so crucially there is an important indirect effect.\n\n    https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence\n\n    solution: pd.DataFrame\n    submission: pd.DataFrame\n    epsilon: KL divergence is undefined for p=0 or p=1. If epsilon is not null, solution and submission probabilities are clipped to max(eps, min(1 - eps, p).\n    row_id_column_name: str\n    micro_average: bool. Row-wise average if True, column-wise average if False.\n\n    Examples\n    --------\n    >>> import pandas as pd\n    >>> row_id_column_name = \"id\"\n    >>> score(pd.DataFrame({'id': range(4), 'ham': [0, 1, 1, 0], 'spam': [1, 0, 0, 1]}), pd.DataFrame({'id': range(4), 'ham': [.1, .9, .8, .35], 'spam': [.9, .1, .2, .65]}), row_id_column_name=row_id_column_name)\n    0.216161...\n    >>> solution = pd.DataFrame({'id': range(3), 'ham': [0, 0.5, 0.5], 'spam': [0.1, 0.5, 0.5], 'other': [0.9, 0, 0]})\n    >>> submission = pd.DataFrame({'id': range(3), 'ham': [0, 0.5, 0.5], 'spam': [0.1, 0.5, 0.5], 'other': [0.9, 0, 0]})\n    >>> score(solution, submission, 'id')\n    0.0\n    >>> solution = pd.DataFrame({'id': range(3), 'ham': [0, 0.5, 0.5], 'spam': [0.1, 0.5, 0.5], 'other': [0.9, 0, 0]})\n    >>> submission = pd.DataFrame({'id': range(3), 'ham': [0.2, 0.3, 0.5], 'spam': [0.1, 0.5, 0.5], 'other': [0.7, 0.2, 0]})\n    >>> score(solution, submission, 'id')\n    0.160531...\n    '''\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n\n    sample_weights = None\n    if sample_weights_column_name:\n        if sample_weights_column_name not in solution.columns:\n            raise ParticipantVisibleError(f'{sample_weights_column_name} not found in solution columns')\n        sample_weights = solution.pop(sample_weights_column_name)\n\n    if sample_weights_column_name and not micro_average:\n        raise ParticipantVisibleError('Sample weights are only valid if `micro_average` is `True`')\n\n    for col in solution.columns:\n        if col not in submission.columns:\n            raise ParticipantVisibleError(f'Missing submission column {col}')\n\n    kaggle_metric_utilities.verify_valid_probabilities(solution, 'solution')\n    kaggle_metric_utilities.verify_valid_probabilities(submission, 'submission')\n\n\n    return kaggle_metric_utilities.safe_call_score(kl_divergence, solution, submission, epsilon=epsilon, micro_average=micro_average, sample_weights=sample_weights)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **submision**\n\nFor each eeg_id in the test set, you must predict a probability for each of the vote columns. The file should contain a header and have the following format:\n\neeg_id,seizure_vote,lpd_vote,gpd_vote,lrda_vote,grda_vote,other_vote\n0,0.166,0.166,0.167,0.167,0.167,0.167\n1,0.166,0.166,0.167,0.167,0.167,0.167\netc.\n\nYour total predicted probabilities for each row must sum to one or your submission will fail.","metadata":{}},{"cell_type":"code","source":"# Load the sample submission file and update it with model predictions for each label\nsubmission = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")\nlabels = ['seizure', 'lpd', 'gpd', 'lrda', 'grda', 'other']\n\n# Assign model predictions to respective columns in the submission DataFrame\nfor i in range(len(labels)):\n    submission[f'{labels[i]}_vote'] = test_preds[:, i]\n#test_preds[:, i]\n# Save the updated DataFrame as the final submission file\nsubmission.to_csv(\"submission.csv\", index=None)\n\n# Display the first few rows of the submission file\ndisplay(submission.head())\n\n# Reclaim memory no longer in use.\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-02-09T02:44:08.962805Z","iopub.execute_input":"2024-02-09T02:44:08.963152Z","iopub.status.idle":"2024-02-09T02:44:09.000003Z","shell.execute_reply.started":"2024-02-09T02:44:08.963126Z","shell.execute_reply":"2024-02-09T02:44:08.998189Z"},"trusted":true},"execution_count":null,"outputs":[]}]}