{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1>Identify Bird Species in the Western Ghats</h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-14T15:17:45.263712Z","iopub.execute_input":"2024-04-14T15:17:45.264479Z","iopub.status.idle":"2024-04-14T15:17:45.399834Z","shell.execute_reply.started":"2024-04-14T15:17:45.264440Z","shell.execute_reply":"2024-04-14T15:17:45.398804Z"}}},{"cell_type":"markdown","source":"<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><b>Identify Bird Species in the Western Ghats</b> is a specific application of audio classification, where the goal is to classify bird species based on their calls captured in audio recordings. Deep learning techniques, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), can be implemented to effectively classify bird species from audio data. Therefore, the project involves implementing deep learning models for audio classification to achieve the objective of identifying bird species in the Western Ghats region.</p>\n\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">Introduces the concept of audio and sound classification using machine learning and deep learning techniques. It highlights the significance of natural language processing (NLP) in enabling machines to understand human language, whether in the form of speech or text.</p>\n\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">This concept aligns well with the project \"Identify Bird Species in the Western Ghats,\" as it involves utilizing machine learning and potentially deep learning methods to classify bird species based on audio recordings, which are essentially representations of sound. By applying NLP principles, machines can analyze and interpret the audio data, extracting meaningful features to classify different bird species based on their distinct calls.</p>\n","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"font-family: 'Times New Roman'; font-size: 24pt;\">Table of Contents</h1>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Introduction to Audio Classification</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Project Overview</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Dataset Overview</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Hands-on Implementing Audio Classification project</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>EDA On Audio Data</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Data Preprocessing</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Building ANN for Audio Classification</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Testing some unknown Audio</li></ul>\n<ul style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>End Notes</li></ul>\n","metadata":{}},{"cell_type":"markdown","source":"<ol style=\"font-family: 'Times New Roman'; font-size: 12pt;\">\n    <li><b>Introduction to Audio Classification:</b> This section introduces the concept of audio classification and its relevance in machine learning and deep learning applications. It sets the stage for understanding how audio data can be processed and classified to solve real-world problems, such as identifying bird species based on their calls in the Western Ghats project.</li>\n    <li><b>Project Overview:</b> In this section, the project \"Identify Bird Species in the Western Ghats\" is introduced. It provides an overview of the project's objectives, scope, and significance, highlighting the goal of classifying bird species from audio recordings in the Western Ghats region.</li>\n    <li><b>Dataset Overview:</b> Here, the dataset used in the project is described. It includes details about the audio recordings of bird calls collected from the Western Ghats, along with metadata such as species labels, file paths, and other relevant information.</li>\n    <li><b>Hands-on Implementing Audio Classification project:</b> This section delves into the practical aspect of implementing audio classification using machine learning and deep learning techniques. It discusses the steps involved in preprocessing the audio data, extracting features, building and training models, and evaluating their performance.</li>\n    <li><b>EDA On Audio Data:</b> Exploratory Data Analysis (EDA) is conducted on the audio data to gain insights into its characteristics, distributions, and patterns. This helps in understanding the nature of the data and informs subsequent preprocessing and modeling steps.</li>\n    <li><b>Data Preprocessing:</b> This section focuses on preparing the audio data for model training by performing preprocessing tasks such as feature extraction, normalization, and data transformation. It ensures that the data is in a suitable format for input to machine learning and deep learning models.</li>\n    <li><b>Building ANN for Audio Classification:</b> Here, artificial neural networks (ANNs) are constructed and trained for audio classification. Various architectures, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), may be explored and implemented to classify bird species from audio recordings effectively.</li>\n    <li><b>Testing some unknown Audio:</b> The trained models are tested on unseen or unknown audio recordings to evaluate their performance and generalization capabilities. This section assesses the models' ability to accurately classify bird species in real-world scenarios.</li>\n    <li><b>End Notes:</b> Finally, this section provides concluding remarks, summarizing the key findings, challenges, and potential areas for further exploration in the project. It reflects on the project's outcomes and implications for future research and applications in audio classification.</li>\n</ol>\n","metadata":{}},{"cell_type":"markdown","source":"<ol style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><h1>Overview</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">The competition aims to address the challenges of monitoring biodiversity, particularly bird species, in the Western Ghats region using advanced technology and machine learning. By leveraging passive acoustic monitoring (PAM) and machine learning techniques, participants are tasked with automating the detection and classification of bird species based on soundscapes.</p>\n\n<h1>Goal of the Competition</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">The primary goal is to develop computational solutions capable of identifying under-studied Indian bird species by their calls. Participants are expected to process continuous audio data and train reliable classifiers with limited training data. Successful solutions will contribute to ongoing efforts to protect avian biodiversity in the Western Ghats, India.</p>\n\n<h1>Context</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">The Western Ghats, a Global Biodiversity Hotspot, is facing significant anthropogenic pressures such as habitat modification and climate change. This necessitates the use of advanced conservation tools and technologies to monitor biodiversity effectively. The competition aims to automate the detection and classification of bird species in this region, given its diverse ecosystems and high levels of bird diversity.</p>\n\n<h1>Broad Goals</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\"><li>Identify endemic bird species of the sky-islands of the Western Ghats in soundscape data.</li>\n<li>Detect/classify endangered bird species featuring limited training data.</li>\n    <li>Detect/classify nocturnal bird species which are poorly understood.</li></p>\n<h1>Timeline</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">The competition timeline spans from April 3, 2024, to June 10, 2024, with various deadlines for entry, team merger, and final submissions. All deadlines are at 11:59 PM UTC on the corresponding day unless otherwise noted.</p>\n\n<h1>Evaluation</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">Evaluation is based on a version of macro-averaged ROC-AUC, with classes that have no true positive labels skipped. Participants are required to submit predictions for each bird species present in the audio recordings, with one column per species.</p>\n\n<h1>Working Note Award Criteria</h1>\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">Participants have the option to submit working notes, which will be evaluated based on criteria such as originality, quality, contribution, and presentation. Each working note will be reviewed by two reviewers, and scores will be averaged.</p>\n\n<p style=\"font-family: 'Times New Roman'; font-size: 12pt;\">This structured overview provides key insights into the competition's objectives, context, timeline, evaluation criteria, and optional working note submission process.</p></ol> ","metadata":{}},{"cell_type":"markdown","source":"<h1>Best Working Note Award (Optional)</h1>\n<ol>\n  <li><h2>Participants Submission:</h2></li>\n  <p style=\"font-family: Arial, sans-serif;\">Participants in the competition are invited to submit working notes to the LifeCLEF 2024 conference. A special BirdCLEF working note competition will be conducted during the conference, where participants can showcase their findings, methodologies, and insights gained from the competition.</p>\n  <li><h2>Award Details:</h2></li>\n  <p style=\"font-family: Arial, sans-serif;\">The top two best working note award winners will each receive a cash prize of $2,500.</p>\n  <li><h2>Submission Process:</h2></li>\n  <p style=\"font-family: Arial, sans-serif;\">Participants should submit their working notes in accordance with the guidelines provided by the conference organizers.</p>\n  <li><h2>Evaluation Criteria:</h2></li>\n  <p style=\"font-family: Arial, sans-serif;\">The evaluation of working notes will be based on various criteria, including originality, quality, contribution, and presentation. For detailed judging criteria, refer to the Evaluation page of the competition.</p>\n</ol>\nThe Best Working Note Award offers participants an opportunity to showcase their innovative approaches, novel findings, and significant contributions to the field of audio classification and biodiversity monitoring.","metadata":{}},{"cell_type":"markdown","source":"\n<h1>This is a Code Competition</h1>\n<p style=\"font-family: Arial, sans-serif;\">Submissions to this competition must be made through Notebooks. For the \"Submit\" button to be active after a commit, the following conditions must be met:</p>\n<ol>\n  <li>CPU Notebook <= 120 minutes run-time</li>\n  <li>GPU Notebook submissions are disabled. You can technically submit but will only have 1 minute of runtime.</li>\n  <li>Internet access disabled</li>\n  <li>Freely & publicly available external data is allowed, including pre-trained models</li>\n  <li>Submission file must be named submission.csv</li>\n</ol>\n<p style=\"font-family: Arial, sans-serif;\">Please see the Code Competition FAQ for more information on how to submit. And review the code debugging doc if you encounter submission errors.</p>","metadata":{}},{"cell_type":"markdown","source":"<h1>Dataset Description</h1>\n<ol>\n  <li>\n    <p><b>train_audio/</b>: The training data consists of short recordings of individual bird calls generously uploaded by users of xenocanto.org. These files have been downsampled to 32 kHz where applicable to match the test set audio and converted to the ogg format. The training data should have nearly all relevant files; we expect there is no benefit to looking for more on xenocanto.org and appreciate your cooperation in limiting the burden on their servers.</p>\n  </li>\n  <li>\n    <p><b>test_soundscapes/</b>: When you submit a notebook, the test_soundscapes directory will be populated with approximately 1,100 recordings to be used for scoring. They are 4 minutes long and in ogg audio format. The file names are randomized but have the general form of soundscape_xxxxxx.ogg. It should take your submission notebook approximately five minutes to load all of the test soundscapes.</p>\n  </li>\n  <li>\n    <p><b>unlabeled_soundscapes/</b>: Unlabeled audio data from the same recording locations as the test soundscapes.</p>\n  </li>\n  <li>\n    <p><b>train_metadata.csv</b>: A wide range of metadata is provided for the training data. The most directly relevant fields are:\n      <ul>\n        <li>primary_label: A code for the bird species. You can review detailed information about the bird codes by appending the code to <a href=\"https://ebird.org/species/\">https://ebird.org/species/</a>, such as <a href=\"https://ebird.org/species/amecro\">https://ebird.org/species/amecro</a> for the American Crow. Not all species have their own pages; some links will fail.</li>\n        <li>latitude & longitude: Coordinates for where the recording was taken. Some bird species may have local call 'dialects,' so you may want to seek geographic diversity in your training data.</li>\n        <li>author: The user who provided the recording.</li>\n        <li>filename: The name of the associated audio file.</li>\n      </ul>\n    </p>\n  </li>\n  <li>\n    <p><b>sample_submission.csv</b>: A valid sample submission.\n      <ul>\n        <li>row_id: A slug of soundscape_[soundscape_id]_[end_time] for the prediction.</li>\n        <li>[bird_id]: There are 182 bird ID columns. You will need to predict the probability of the presence of each bird for each row.</li>\n      </ul>\n    </p>\n  </li>\n  <li>\n    <p><b>eBird_Taxonomy_v2021.csv</b>: Data on the relationships between different species.</p>\n  </li>\n</ol>\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"<h1>Define the monitoring function</h1>","metadata":{}},{"cell_type":"markdown","source":"<ol><li>At the beginning of each cell, call the function with the current time</li>\n    <b>start_time = time.time()</b>\n\n<li>At the end of each cell, call the function again with the start time</li>\n    <b>monitor_runtime(start_time)</b>   \n</ol>","metadata":{}},{"cell_type":"code","source":"import time\n# Start counting notebook running time\ntime_start = time.time()\nimport os\nimport IPython.display as ipd\n\nfrom datetime import datetime\n","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:03:59.933686Z","iopub.execute_input":"2024-05-23T22:03:59.934082Z","iopub.status.idle":"2024-05-23T22:03:59.970647Z","shell.execute_reply.started":"2024-05-23T22:03:59.934050Z","shell.execute_reply":"2024-05-23T22:03:59.969567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# File path to the directory containing subdirectories with audio files\ndirectory_path = \"/kaggle/input/birdclef-2024/train_audio/\"\n\n# Get the list of subdirectories\nsubdirectories = [d for d in os.listdir(directory_path) if os.path.isdir(os.path.join(directory_path, d))]\n\n# Find the first subdirectory with audio files\naudio_files = []\nfor subdirectory in subdirectories:\n    subdirectory_path = os.path.join(directory_path, subdirectory)\n    files = os.listdir(subdirectory_path)\n    audio_files.extend([os.path.join(subdirectory_path, f) for f in files if f.endswith(\".ogg\")])\n\n# Display the first audio file found\nif audio_files:\n    first_audio_file = audio_files[0]\n    display(ipd.Audio(first_audio_file))  # Display the audio\nelse:\n    print(\"No audio files found in the subdirectories.\")\n    # Define the monitor_runtime function\ndef monitor_runtime(start_time):\n    end_time = datetime.now()\n    print(f\"Runtime: {end_time - start_time}\")\n    \n    \n   # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:03:59.972409Z","iopub.execute_input":"2024-05-23T22:03:59.972914Z","iopub.status.idle":"2024-05-23T22:04:07.480344Z","shell.execute_reply.started":"2024-05-23T22:03:59.972884Z","shell.execute_reply":"2024-05-23T22:04:07.479148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport librosa\nimport librosa.display\ndata, sample_rate = librosa.load(first_audio_file)\nplt.figure(figsize=(12, 5))\nlibrosa.display.waveshow(data, sr=sample_rate)\n\n# Define the monitor_runtime function\ndef monitor_runtime(start_time):\n    end_time = datetime.now()\n    print(f\"Runtime: {end_time - start_time}\")\n\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:07.482190Z","iopub.execute_input":"2024-05-23T22:04:07.482600Z","iopub.status.idle":"2024-05-23T22:04:20.810926Z","shell.execute_reply.started":"2024-05-23T22:04:07.482564Z","shell.execute_reply":"2024-05-23T22:04:20.808756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the audio file\nwaveform, sample_rate = librosa.load(first_audio_file, sr=None)\n\n# Plot the waveform\nplt.figure(figsize=(12, 4))\nplt.plot(waveform)\nplt.xlabel('Sample')\nplt.ylabel('Amplitude')\nplt.title('Waveform')\nplt.show()\n# Define the monitor_runtime function\ndef monitor_runtime(start_time):\n    end_time = datetime.now()\n    print(f\"Runtime: {end_time - start_time}\")\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:20.812917Z","iopub.execute_input":"2024-05-23T22:04:20.814332Z","iopub.status.idle":"2024-05-23T22:04:21.303108Z","shell.execute_reply.started":"2024-05-23T22:04:20.814291Z","shell.execute_reply":"2024-05-23T22:04:21.302029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nmetadata = pd.read_csv('/kaggle/input/birdclef-2024/train_metadata.csv')\nmetadata.head(10)\n\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:21.305770Z","iopub.execute_input":"2024-05-23T22:04:21.306865Z","iopub.status.idle":"2024-05-23T22:04:21.890258Z","shell.execute_reply.started":"2024-05-23T22:04:21.306823Z","shell.execute_reply":"2024-05-23T22:04:21.888890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n# Count occurrences of each class label\nclass_counts = metadata['primary_label'].value_counts()\n\nplt.figure(figsize=(10, 6))\nsns.barplot(x=class_counts.index, y=class_counts.values)\nplt.title(\"Count of records in each class\")\nplt.xticks(rotation=\"vertical\")\nplt.show()\n# Define the monitor_runtime function\ndef monitor_runtime(start_time):\n    end_time = datetime.now()\n    print(f\"Runtime: {end_time - start_time}\")\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:21.891335Z","iopub.execute_input":"2024-05-23T22:04:21.891648Z","iopub.status.idle":"2024-05-23T22:04:24.109411Z","shell.execute_reply.started":"2024-05-23T22:04:21.891624Z","shell.execute_reply":"2024-05-23T22:04:24.108274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save submission CSV file\n#submission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.111087Z","iopub.execute_input":"2024-05-23T22:04:24.111438Z","iopub.status.idle":"2024-05-23T22:04:24.116555Z","shell.execute_reply.started":"2024-05-23T22:04:24.111408Z","shell.execute_reply":"2024-05-23T22:04:24.115143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def extract_features(audio_file_path, sr=22050, n_mfcc=20):\n    \"\"\"\n    Extracts Mel-frequency cepstral coefficients (MFCCs) from an audio file.\n    \n    Parameters:\n        audio_file_path (str): Path to the audio file.\n        sr (int): Sampling rate. Default is 22050 Hz.\n        n_mfcc (int): Number of MFCCs to extract. Default is 20.\n    \n    Returns:\n        numpy.ndarray: Extracted MFCC features.\n    \"\"\"\n    # Load audio file\n    y, sr = librosa.load(audio_file_path, sr=sr)\n    \n    # Extract MFCC features\n    mfccs = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=n_mfcc)\n    \n    return mfccs\n\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.118379Z","iopub.execute_input":"2024-05-23T22:04:24.118689Z","iopub.status.idle":"2024-05-23T22:04:24.134875Z","shell.execute_reply.started":"2024-05-23T22:04:24.118664Z","shell.execute_reply":"2024-05-23T22:04:24.133543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def extract_features_from_directory(directory_path, chunk_size=1000, sr=22050, n_mfcc=20):\n    \"\"\"\n    Extracts MFCC features from audio files in a directory and its subdirectories.\n    \n    Parameters:\n        directory_path (str): Path to the directory containing audio files.\n        chunk_size (int): Size of each chunk. Default is 1000.\n        sr (int): Sampling rate. Default is 22050 Hz.\n        n_mfcc (int): Number of Mel-frequency cepstral coefficients (MFCCs) to extract. Default is 20.\n    \n    Yields:\n        list: List of tuples containing (filename, features) for each audio file in a chunk.\n    \"\"\"\n    # Initialize an empty list to store features in the current chunk\n    chunk_features = []\n    # Initialize a counter to keep track of the number of files processed\n    count = 0\n    \n    # Iterate through each directory and subdirectory\n    for root, dirs, files in os.walk(directory_path):\n        for filename in files:\n            if filename.endswith(\".ogg\"):\n                # Construct full path to audio file\n                file_path = os.path.join(root, filename)\n                \n                # Extract features from audio file\n                features = extract_features(file_path, sr=sr, n_mfcc=n_mfcc)\n                \n                # Append filename and features to the current chunk\n                chunk_features.append((filename, features))\n                # Increment the counter\n                count += 1\n                \n                # If the chunk size is reached, yield the current chunk and reset it\n                if count % chunk_size == 0:\n                    yield chunk_features\n                    chunk_features = []\n    \n    # If there are remaining features not yielded yet (i.e., the last chunk),\n    # yield them before exiting the function\n    if chunk_features:\n        yield chunk_features\n\n         # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.138702Z","iopub.execute_input":"2024-05-23T22:04:24.139093Z","iopub.status.idle":"2024-05-23T22:04:24.149782Z","shell.execute_reply.started":"2024-05-23T22:04:24.139054Z","shell.execute_reply":"2024-05-23T22:04:24.148483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n\ndef monitor_runtime(start_time):\n    end_time = datetime.now()\n    duration = end_time - start_time\n    print(f\"Runtime: {duration}\")\n\n# Define train_dir with the path to your training directory\ntrain_dir = '/kaggle/input/birdclef-2024/train_audio'\n\n# Step 1: Collect class labels (assumes subdirectories in train_dir are class labels)\nclass_labels = [dir_name for dir_name in os.listdir(train_dir) if os.path.isdir(os.path.join(train_dir, dir_name))]\n\n# Step 2: Perform one-hot encoding on the class labels\none_hot_encoded_labels = pd.get_dummies(class_labels)\n\n\n# Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.151193Z","iopub.execute_input":"2024-05-23T22:04:24.151542Z","iopub.status.idle":"2024-05-23T22:04:24.173875Z","shell.execute_reply.started":"2024-05-23T22:04:24.151512Z","shell.execute_reply":"2024-05-23T22:04:24.172133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Step 3: Perform one-hot encoding on the class labels\none_hot_encoded_labels = pd.get_dummies(class_labels)\n\n# Step 4: Create a DataFrame with filenames and one-hot encoded labels\ndata = []\n\nfor label in class_labels:\n    label_dir = os.path.join(train_dir, label)\n    filenames = [os.path.join(label, file) for file in os.listdir(label_dir) if file.endswith('.ogg')]\n    data.extend([(filename, label) for filename in filenames])\n\ndf = pd.DataFrame(data, columns=['filename', 'label'])\ndf_encoded = pd.concat([df['filename'], pd.get_dummies(df['label'])], axis=1)\n\n# Generate row IDs\nrow_ids = list(range(len(df_encoded)))\ndf_encoded['row_id'] = row_ids\n\n# Arrange the data such that each row corresponds to a soundscape and each column represents a label\nsoundscape_data = df_encoded.groupby('row_id').sum()\n\n# print(\"\\nSoundscape data after arranging:\")\n# print(soundscape_data)\n\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.179880Z","iopub.execute_input":"2024-05-23T22:04:24.180322Z","iopub.status.idle":"2024-05-23T22:04:24.762202Z","shell.execute_reply.started":"2024-05-23T22:04:24.180290Z","shell.execute_reply":"2024-05-23T22:04:24.760063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.764630Z","iopub.execute_input":"2024-05-23T22:04:24.765527Z","iopub.status.idle":"2024-05-23T22:04:24.771074Z","shell.execute_reply.started":"2024-05-23T22:04:24.765484Z","shell.execute_reply":"2024-05-23T22:04:24.769931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# One-hot encode the 'filename' column\nsoundscape_data_encoded = pd.get_dummies(soundscape_data['filename'])","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.773883Z","iopub.execute_input":"2024-05-23T22:04:24.775174Z","iopub.status.idle":"2024-05-23T22:04:24.913832Z","shell.execute_reply.started":"2024-05-23T22:04:24.775118Z","shell.execute_reply":"2024-05-23T22:04:24.912602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# Assuming you have already loaded the soundscape_data DataFrame\n\n# Extract class label from filename\nsoundscape_data['class_label'] = soundscape_data['filename'].str.split('/').str[0]\n\n# Reorder the columns to have 'class_label' as the last column\nsoundscape_data = soundscape_data.reindex(columns=[col for col in soundscape_data.columns if col != 'class_label'] + ['class_label'])\n\n# Display the DataFrame\nsoundscape_data.head()\n # Display the runtime\nstart_time = datetime.now()\nmonitor_runtime(start_time)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.915466Z","iopub.execute_input":"2024-05-23T22:04:24.916085Z","iopub.status.idle":"2024-05-23T22:04:24.991875Z","shell.execute_reply.started":"2024-05-23T22:04:24.916023Z","shell.execute_reply":"2024-05-23T22:04:24.990681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"soundscape_data","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:24.993956Z","iopub.execute_input":"2024-05-23T22:04:24.994344Z","iopub.status.idle":"2024-05-23T22:04:25.043306Z","shell.execute_reply.started":"2024-05-23T22:04:24.994315Z","shell.execute_reply":"2024-05-23T22:04:25.042033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"soundscape_data.columns","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.044674Z","iopub.execute_input":"2024-05-23T22:04:25.045084Z","iopub.status.idle":"2024-05-23T22:04:25.053689Z","shell.execute_reply.started":"2024-05-23T22:04:25.045055Z","shell.execute_reply":"2024-05-23T22:04:25.052662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.impute import SimpleImputer\nimport numpy as np\nimport time\n\n# Start timing the process\nstart_time = time.time()\n\n# Load your soundscape_data DataFrame (assuming it is already loaded)\nsoundscape_data \n\n# Step 1: Aggregate features by species\n# Group by class_label and sum the binary presence indicators for each species\naggregated_data = soundscape_data.groupby('class_label').sum()\n\n# Reset the index to make class_label a column again\naggregated_data = aggregated_data.reset_index()\n\n# Step 2: Apply imputation to handle any missing values\n# Initialize SimpleImputer with the mean strategy\nimputer = SimpleImputer(strategy='mean')\n\n# Perform imputation on the aggregated data\ncolumns_for_imputation = aggregated_data.columns.difference(['class_label', 'filename'])\n\naggregated_data[columns_for_imputation] = imputer.fit_transform(aggregated_data[columns_for_imputation])\n\n# Stop timing the process\nend_time = time.time()\n\n# Print the DataFrame to verify its structure\naggregated_data.head()\n\n# Print the time taken to execute the process\nprint(f\"Time taken to execute the process: {end_time - start_time:.2f} seconds\")\n\n# Save the new DataFrame to a CSV file\n#aggregated_data.to_csv('aggregated_features_imputed.csv', index=False)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.055240Z","iopub.execute_input":"2024-05-23T22:04:25.055885Z","iopub.status.idle":"2024-05-23T22:04:25.538354Z","shell.execute_reply.started":"2024-05-23T22:04:25.055849Z","shell.execute_reply":"2024-05-23T22:04:25.536762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aggregated_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.540394Z","iopub.execute_input":"2024-05-23T22:04:25.540987Z","iopub.status.idle":"2024-05-23T22:04:25.587302Z","shell.execute_reply.started":"2024-05-23T22:04:25.540916Z","shell.execute_reply":"2024-05-23T22:04:25.586375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.impute import SimpleImputer\nimport numpy as np\nimport time\n\n# Start timing the process\nstart_time = time.time()\n\n# Load your soundscape_data DataFrame (assuming it is already loaded)\n# soundscape_data = pd.read_csv('path_to_your_soundscape_data.csv')\n\n# Step 1: Aggregate features by species, ignoring the filename during aggregation\n# Group by class_label and sum the binary presence indicators for each species\n# We exclude the filename from the aggregation by dropping it before grouping\naggregated_data = soundscape_data.drop(columns=['filename']).groupby('class_label').sum().reset_index()\n\n# Step 2: Apply imputation to handle any missing values\n# Initialize SimpleImputer with the mean strategy\nimputer = SimpleImputer(strategy='mean')\n\n# Perform imputation on the aggregated data\ncolumns_for_imputation = aggregated_data.columns.difference(['class_label'])\n\naggregated_data[columns_for_imputation] = imputer.fit_transform(aggregated_data[columns_for_imputation])\n\n# Stop timing the process\nend_time = time.time()\n\n# Print the DataFrame to verify its structure\nprint(aggregated_data.head())\n\n# Print the time taken to execute the process\nprint(f\"Time taken to execute the process: {end_time - start_time:.2f} seconds\")\n\n# Save the new DataFrame to a CSV file\n#aggregated_data.to_csv('aggregated_features_imputed.csv', index=False)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.588526Z","iopub.execute_input":"2024-05-23T22:04:25.589509Z","iopub.status.idle":"2024-05-23T22:04:25.697738Z","shell.execute_reply.started":"2024-05-23T22:04:25.589474Z","shell.execute_reply":"2024-05-23T22:04:25.695508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.impute import SimpleImputer\nimport os\nfrom datetime import datetime\n\ndef monitor_runtime(start_time):\n    end_time = datetime.now()\n    duration = end_time - start_time\n    print(f\"Time taken to execute the process: {duration.total_seconds()} seconds\")\n\n# Assuming `soundscape_data` DataFrame is already created as shown in your example\n# and it contains the 'class_label' and one-hot encoded columns\n\n# Start timing\nstart_time = datetime.now()\n\n# Initialize SimpleImputer with the mean strategy\nimputer = SimpleImputer(strategy='mean')\n\n# Select columns for imputation (excluding 'class_label' and 'filename')\ncolumns_for_imputation = soundscape_data.columns.difference(['class_label', 'filename'])\n\n# Perform imputation on the selected columns\nsoundscape_data[columns_for_imputation] = imputer.fit_transform(soundscape_data[columns_for_imputation])\n\n# Print the DataFrame to verify the imputation\nprint(\"\\nSoundscape data after imputation:\")\nprint(soundscape_data.head())\n\n# Monitor runtime\nmonitor_runtime(start_time)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.701113Z","iopub.execute_input":"2024-05-23T22:04:25.703299Z","iopub.status.idle":"2024-05-23T22:04:25.853544Z","shell.execute_reply.started":"2024-05-23T22:04:25.703243Z","shell.execute_reply":"2024-05-23T22:04:25.851996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"soundscape_data","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.856779Z","iopub.execute_input":"2024-05-23T22:04:25.857423Z","iopub.status.idle":"2024-05-23T22:04:25.911208Z","shell.execute_reply.started":"2024-05-23T22:04:25.857392Z","shell.execute_reply":"2024-05-23T22:04:25.909980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"soundscape_data.to_csv('submission', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-05-23T22:04:25.912389Z","iopub.execute_input":"2024-05-23T22:04:25.912705Z","iopub.status.idle":"2024-05-23T22:04:29.290610Z","shell.execute_reply.started":"2024-05-23T22:04:25.912678Z","shell.execute_reply":"2024-05-23T22:04:29.289323Z"},"trusted":true},"execution_count":null,"outputs":[]}]}