{"metadata":{"kernelspec":{"display_name":"ml_env","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.9.18"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Harmful Brain Activity Classification - Data Augmentation with Pandas and NumPy**\n\n## **Project by:** [Aarish Asif Khan]()\n\n## **Date:** 4th April 2024\n\n## **Dataset:** [HMC - Harmful Brain Activity Dataset]()","metadata":{}},{"cell_type":"markdown","source":"## **Data Augmentation:**\n\n`Data augmentation` is a technique `commonly used in machine learning and deep learning` to artificially increase the size of a dataset by creating modified versions of the existing data. The `goal of data augmentation is to improve the performance and generalization of machine learning models` by exposing them to a wider variety of training examples.","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will try to perform data augmentation on the HMC Dataset by using these two famous libraries in Python, NumPy and Pandas. Without wasting any time, let's get started!","metadata":{}},{"cell_type":"code","source":"# !pip install pandas \n# !pip install numpy","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import the basic libraries\nimport pandas as pd \nimport numpy as np \n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the number of augmented samples\nnum_augmented_samples = 5  # You can adjust this number as needed","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create an empty list to store augmented metadata\naugmented_data = []\n\n# Perform data augmentation for each original sample\nfor index, row in train_data.iterrows():\n    original_data = row  # Original metadata for the current sample\n\n    # Generate augmented samples by perturbing the original metadata\n    for i in range(num_augmented_samples):\n        # Modify some of the metadata values to create variation\n        # For example, you can add noise, randomize values, or apply other transformations\n        augmented_data_point = original_data.copy()\n\n        # Example: adding noise to the label_id\n        noise = np.random.normal(0, 1, size=len(augmented_data_point))\n        augmented_data_point['label_id'] += noise\n\n        # Append the augmented data point to the list\n        augmented_data.append(augmented_data_point)\n\n# Create a DataFrame from the list of augmented data points\naugmented_data = pd.DataFrame(augmented_data)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check shape between original dataset and augmented dataset\nprint(\"Original dataset shape: \", train_data.shape)\nprint(\"Augmented dataset shape: \", augmented_data.shape)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(augmented_data)\n# -------------------------------------------\nprint(\"Data augmentation has been completed!\")","metadata":{},"execution_count":null,"outputs":[]}]}