{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"},{"sourceId":8649454,"sourceType":"datasetVersion","datasetId":5180902}],"dockerImageVersionId":30732,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Bird Species Classification(BirdCLEF 2024)**","metadata":{}},{"cell_type":"markdown","source":"Monitoring and assessing avian biodiversity over large areas is critical for evaluating ecological ‎restoration projects. Traditional observer-based surveys are costly and logistically challenging, ‎especially in regions like the Western Ghats of India. This biodiversity hotspot faces significant ‎threats from landscape alterations and climate changes, endangering its unique ecosystems and ‎bird species. Birds, being excellent indicators of biodiversity changes, necessitate effective ‎monitoring methods to gauge the success of conservation efforts.‎\nTraditional surveys struggle with:‎\n\n**High Costs:** Extensive surveys require significant financial and human resources.‎\n\n**Logistical Challenges:** Accessing remote or vast areas regularly is difficult.‎\n\n**Temporal and Spatial Coverage:** Limited ability to provide continuous monitoring over large ‎regions.‎\n\n\nBy combining **Passive Acoustic Monitoring (PAM) with sophisticated machine learning ‎techniques****, this solution offers an efficient and scalable method for monitoring bird species in ‎the Western Ghats, addressing the limitations of traditional surveys and bolstering conservation ‎efforts to protect and restore avian biodiversity in this critical region. PAM utilizes automated ‎recording devices that are cost-effective to deploy and maintain, significantly reducing expenses ‎compared to human observers. The integration of machine learning enables the analysis of vast ‎amounts of audio data, facilitating extensive area monitoring without requiring constant human ‎presence. Additionally, PAM provides continuous, real-time monitoring, delivering a more ‎comprehensive understanding of avian biodiversity and its fluctuations. This approach holds ‎significant potential for broader application in global biodiversity conservation initiatives.‎","metadata":{}},{"cell_type":"markdown","source":"# **Import Libraries**","metadata":{}},{"cell_type":"code","source":"!pip install noisereduce\nimport os\nimport glob\nimport shutil\nimport zipfile\nimport matplotlib.pyplot as plt\nplt.style.use('dark_background')\nimport seaborn as sns\nimport numpy as np\n!pip install plotly\nimport plotly.express as px\nimport librosa\nfrom IPython.display import Audio\nimport pandas as pd\nimport pickle\nfrom joblib import dump, load\nfrom pathlib import Path\nimport noisereduce as nr\n!pip install -U imbalanced-learn\nfrom imblearn.over_sampling import RandomOverSampler\nimport sklearn\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.ensemble import RandomForestClassifier\nimport xgboost as xgb\nfrom sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix\nfrom sklearn.metrics import roc_auc_score, roc_curve, auc\nfrom sklearn.multiclass import OneVsRestClassifier\nfrom sklearn.preprocessing import label_binarize\nfrom itertools import cycle","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:17:23.76304Z","iopub.execute_input":"2024-06-10T20:17:23.763501Z","iopub.status.idle":"2024-06-10T20:18:23.910182Z","shell.execute_reply.started":"2024-06-10T20:17:23.763465Z","shell.execute_reply":"2024-06-10T20:18:23.908716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Collection**","metadata":{}},{"cell_type":"code","source":"meta_data = pd.read_csv('/kaggle/input/birdclef-2024/train_metadata.csv')\nmeta_data.head(4)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:23:27.574758Z","iopub.execute_input":"2024-06-10T20:23:27.575217Z","iopub.status.idle":"2024-06-10T20:23:27.827555Z","shell.execute_reply.started":"2024-06-10T20:23:27.575182Z","shell.execute_reply":"2024-06-10T20:23:27.826297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Exploratory Data Analysis (EDA)**","metadata":{}},{"cell_type":"code","source":"# Check for missing values\nprint(meta_data.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-06-10T15:29:49.360614Z","iopub.execute_input":"2024-06-10T15:29:49.361042Z","iopub.status.idle":"2024-06-10T15:29:49.394242Z","shell.execute_reply.started":"2024-06-10T15:29:49.361006Z","shell.execute_reply":"2024-06-10T15:29:49.393172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Basic statistics\nprint(meta_data.describe())","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:41:00.747949Z","iopub.execute_input":"2024-06-09T17:41:00.74842Z","iopub.status.idle":"2024-06-09T17:41:00.774031Z","shell.execute_reply.started":"2024-06-09T17:41:00.748386Z","shell.execute_reply":"2024-06-09T17:41:00.772624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of birds in each species\nmeta_data.primary_label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:41:05.715193Z","iopub.execute_input":"2024-06-09T17:41:05.71565Z","iopub.status.idle":"2024-06-09T17:41:05.734836Z","shell.execute_reply.started":"2024-06-09T17:41:05.715615Z","shell.execute_reply":"2024-06-09T17:41:05.733479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualize the relationship between primary and secondary bird labels \ntmp = meta_data[meta_data['secondary_labels'] != '[]'].iloc[:40]\nfig = px.sunburst(tmp, path=['primary_label', 'secondary_labels'] )\nfig.update_layout(\n         title=\"Primary & Secondary Labels \",\n    )\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:41:09.975999Z","iopub.execute_input":"2024-06-09T17:41:09.97643Z","iopub.status.idle":"2024-06-09T17:41:12.034899Z","shell.execute_reply.started":"2024-06-09T17:41:09.976397Z","shell.execute_reply":"2024-06-09T17:41:12.033747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a copy of the df_train DataFrame\ntmp = meta_data.copy()\n\n# Create a scatter mapbox plot\nfig = px.scatter_mapbox(\n    tmp,  # Use the tmp DataFrame as the data source\n    lat=\"latitude\",  # Latitude column for plotting points\n    lon=\"longitude\",  # Longitude column for plotting points\n    color=\"primary_label\",  # Color points based on the primary_label column\n    zoom=0.1,  # Initial zoom level of the map\n    title='Bird Recordings Loaction'  # Title of the plot\n)\n\n# Update the layout of the plot to use the \"open-street-map\" style for the map background\nfig.update_layout(mapbox_style=\"open-street-map\")\n\n# Update the layout of the plot to set the margin around the map\nfig.update_layout(margin={\"r\":0,\"t\":30,\"l\":0,\"b\":0})\n\n# Display the scatter mapbox plot\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:41:18.081648Z","iopub.execute_input":"2024-06-09T17:41:18.083039Z","iopub.status.idle":"2024-06-09T17:41:18.793117Z","shell.execute_reply.started":"2024-06-09T17:41:18.082985Z","shell.execute_reply":"2024-06-09T17:41:18.791729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter_mapbox(meta_data, lat='latitude', lon='longitude', color='common_name', \n                        hover_name='common_name', hover_data=['latitude', 'longitude'], \n                        title='Origin of Bird Species',\n                        zoom=1, height=600, template='plotly_dark')\nfig.update_layout(\n    mapbox_style=\"white-bg\",\n    mapbox_layers=[\n        {\n            \"below\": 'traces',\n            \"sourcetype\": \"raster\",\n            \"sourceattribution\": \"United States Geological Survey\",\n            \"source\": [\n                \"https://basemap.nationalmap.gov/arcgis/rest/services/USGSImageryOnly/MapServer/tile/{z}/{y}/{x}\"\n            ]\n        }\n      ])\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:42:28.02881Z","iopub.execute_input":"2024-06-09T17:42:28.029973Z","iopub.status.idle":"2024-06-09T17:42:29.385179Z","shell.execute_reply.started":"2024-06-09T17:42:28.029933Z","shell.execute_reply":"2024-06-09T17:42:29.38278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It visualizes the geographical distribution of bird species represented in the 'meta_data' DataFrame. Each point on the map represents a bird species, colored by its common name. Hovering over a point displays additional information such as latitude and longitude. ","metadata":{}},{"cell_type":"code","source":"sns.displot(data=meta_data,x='latitude',bins=30,kde=True)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:41:34.32836Z","iopub.execute_input":"2024-06-09T17:41:34.328818Z","iopub.status.idle":"2024-06-09T17:41:35.061391Z","shell.execute_reply.started":"2024-06-09T17:41:34.328785Z","shell.execute_reply":"2024-06-09T17:41:35.060187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.displot(data=meta_data,x='longitude',bins=30,kde=True)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:41:39.796998Z","iopub.execute_input":"2024-06-09T17:41:39.797525Z","iopub.status.idle":"2024-06-09T17:41:40.444269Z","shell.execute_reply.started":"2024-06-09T17:41:39.797485Z","shell.execute_reply":"2024-06-09T17:41:40.443046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For latitude and longitude in the train_metadata dataset, it might be more appropriate to fill missing values with the median to avoid potential distortions from outliers.","metadata":{}},{"cell_type":"code","source":"median_latitude = meta_data['latitude'].median()\nmedian_longitude = meta_data['longitude'].median()\n#Fill missing latitude and longitude with the mean values\nmeta_data['latitude'] = meta_data['latitude'].fillna(median_latitude)\nmeta_data['longitude'] = meta_data['longitude'].fillna(median_longitude)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:24:07.950377Z","iopub.execute_input":"2024-06-10T20:24:07.95077Z","iopub.status.idle":"2024-06-10T20:24:07.96587Z","shell.execute_reply.started":"2024-06-10T20:24:07.950739Z","shell.execute_reply":"2024-06-10T20:24:07.964289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if there are any remaining missing values\nprint(meta_data.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:42:07.576501Z","iopub.execute_input":"2024-06-09T17:42:07.576965Z","iopub.status.idle":"2024-06-09T17:42:07.60971Z","shell.execute_reply.started":"2024-06-09T17:42:07.576935Z","shell.execute_reply":"2024-06-09T17:42:07.608309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def audio_waveframe(file_path):\n    # Load the audio file\n    audio_data, sampling_rate = librosa.load(file_path)\n    # Calculate the duration of the audio file\n    duration = len(audio_data) / sampling_rate\n    # Create a time array for plotting\n    time = np.arange(0, duration, 1/sampling_rate)\n    # Plot the waveform\n    plt.figure(figsize=(30, 4))\n    plt.plot(time, audio_data, color='blue')\n    plt.title('Audio Waveform')\n    plt.xlabel('Time (s)')\n    plt.ylabel('Amplitude')\n    plot = plt.show()\n    return plot","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:46:42.879404Z","iopub.execute_input":"2024-06-10T20:46:42.879839Z","iopub.status.idle":"2024-06-10T20:46:42.888178Z","shell.execute_reply.started":"2024-06-10T20:46:42.879806Z","shell.execute_reply":"2024-06-10T20:46:42.886726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**waveform** visualizes the amplitude (volume) of an audio signal over time. This allows you to see how the sound changes, including the presence of different frequencies, the amplitude of the signal, and any patterns or noise within the audio.","metadata":{}},{"cell_type":"code","source":"def spectrogram(file_path):\n    # Compute the short-time Fourier transform (STFT)\n    n_fft = 500  # Number of FFT points 2048\n    hop_length = 50  # Hop length for STFT 512\n    audio_data, sampling_rate = librosa.load(file_path)\n    stft = librosa.stft(audio_data, n_fft=n_fft, hop_length=hop_length)\n    # Convert the magnitude spectrogram to decibels (log scale)\n    spectrogram = librosa.amplitude_to_db(np.abs(stft))\n    # Plot the spectrogram\n    plt.figure(figsize=(30, 6))\n    librosa.display.specshow(spectrogram, sr=sampling_rate, hop_length=hop_length, x_axis='time', y_axis='linear')\n    plt.colorbar(format='%+2.0f dB')\n    plt.title('Spectrogram')\n    plt.xlabel('Time (s)')\n    plt.ylabel('Frequency (Hz)')\n    plt.tight_layout()\n    plot = plt.show()\n    return plot\n\n","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:46:47.454798Z","iopub.execute_input":"2024-06-10T20:46:47.455214Z","iopub.status.idle":"2024-06-10T20:46:47.464629Z","shell.execute_reply.started":"2024-06-10T20:46:47.455181Z","shell.execute_reply":"2024-06-10T20:46:47.463371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**A spectrogram** is a visual representation of the spectrum of frequencies in a signal as it varies with time. It is commonly used in the fields of audio processing, signal processing, and speech analysis to analyze the frequency content of signals, especially non-stationary signals.","metadata":{}},{"cell_type":"code","source":"def audio_analysis(file_path):\n    aw = audio_waveframe(file_path)\n    spg = spectrogram(file_path)\n    return aw, spg","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:46:51.675554Z","iopub.execute_input":"2024-06-10T20:46:51.675979Z","iopub.status.idle":"2024-06-10T20:46:51.68233Z","shell.execute_reply.started":"2024-06-10T20:46:51.675947Z","shell.execute_reply":"2024-06-10T20:46:51.680949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_analysis('/kaggle/input/birdclef-2024/train_audio/asbfly/XC134896.ogg')\nAudio('/kaggle/input/birdclef-2024/train_audio/asbfly/XC134896.ogg')","metadata":{"execution":{"iopub.status.busy":"2024-06-09T17:45:08.573521Z","iopub.execute_input":"2024-06-09T17:45:08.573993Z","iopub.status.idle":"2024-06-09T17:45:24.9047Z","shell.execute_reply.started":"2024-06-09T17:45:08.573959Z","shell.execute_reply":"2024-06-09T17:45:24.903292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Feature Engineering**","metadata":{}},{"cell_type":"markdown","source":"# 1. Feature Extraction","metadata":{}},{"cell_type":"code","source":"def normalize_and_denoise(audio,sample_rate):\n#     # Normalize audio data\n    audio = audio / np.max(np.abs(audio))\n    \n    # Apply Spectral Subtraction Filter\n    filtered_audio = nr.reduce_noise(y=audio, sr=sample_rate)\n    \n    return filtered_audio","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:46:58.121574Z","iopub.execute_input":"2024-06-10T20:46:58.122022Z","iopub.status.idle":"2024-06-10T20:46:58.128968Z","shell.execute_reply.started":"2024-06-10T20:46:58.121987Z","shell.execute_reply":"2024-06-10T20:46:58.127555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NoiseReduce** is a Python library designed for noise reduction in audio signals. It provides simple and effective ‎methods to reduce background noise from audio recordings, enhancing the clarity of the audio ‎signal, which is especially useful in tasks such as speech recognition, audio processing, and, as in ‎your case, bird song identification.‎\nThe library typically applies spectral subtraction techniques to reduce noise. This involves ‎estimating the noise spectrum from the audio signal and subtracting it from the original signal ‎to produce a cleaner version.‎\n","metadata":{}},{"cell_type":"code","source":"# Function to extract features from audio file\ndef extract_features(file_path):\n    # Load audio file\n    audio, sample_rate = librosa.load(file_path)\n      # Normalize and denoise audio\n    audio = normalize_and_denoise(audio, sample_rate)\n    # Extract features using Mel-Frequency Cepstral Coefficients (MFCC)\n    mfccs = librosa.feature.mfcc(y=audio, sr=sample_rate, n_mfcc=40)\n    # Flatten the features into a 1D array\n    flattened_features = np.mean(mfccs.T, axis=0)\n    return flattened_features","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:47:01.731295Z","iopub.execute_input":"2024-06-10T20:47:01.731735Z","iopub.status.idle":"2024-06-10T20:47:01.741028Z","shell.execute_reply.started":"2024-06-10T20:47:01.7317Z","shell.execute_reply":"2024-06-10T20:47:01.739565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Mel-Frequency Cepstral Coefficients (MFCCs)**  are a widely used feature extraction technique ‎in audio signal processing. They provide a compact representation of the spectral properties of ‎an audio signal, making them suitable for tasks such as speech and bird call recognition.","metadata":{}},{"cell_type":"code","source":"# Function to load dataset and extract features\ndef load_data_and_extract_features(data_dir):\n    labels = []\n    features = []\n    # Loop through each audio file in the dataset directory\n    for filename in os.listdir(data_dir):\n        if filename.endswith('.ogg'):\n            file_path = os.path.join(data_dir, filename)\n            # Extract label from filename\n            label = filename.split('-')[0]\n            labels.append(label)\n            # Extract features from audio file\n            feature = extract_features(file_path)\n            features.append(feature)\n    return np.array(features), np.array(labels)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:47:10.913346Z","iopub.execute_input":"2024-06-10T20:47:10.913731Z","iopub.status.idle":"2024-06-10T20:47:10.921644Z","shell.execute_reply.started":"2024-06-10T20:47:10.9137Z","shell.execute_reply":"2024-06-10T20:47:10.920259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\n\nextracted_features = []\n\nfor i in tqdm(annotated_data['audio_file_path']):\n\n    features = extract_features(file_path=i)\n    # print(features)\n    extracted_features.append(features)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open(\"extracted_features\", \"wb\") as file:   #Pickling\n\tpickle.dump(extracted_features, file)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open(\"/kaggle/input/extracted-features/extracted_featuress\", \"rb\") as file:   # Unpickling\n\tpickled_extracted_features = pickle.load(file)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:24:21.001261Z","iopub.execute_input":"2024-06-10T20:24:21.002706Z","iopub.status.idle":"2024-06-10T20:24:21.146017Z","shell.execute_reply.started":"2024-06-10T20:24:21.002642Z","shell.execute_reply":"2024-06-10T20:24:21.144641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Label Encoding","metadata":{}},{"cell_type":"code","source":"# Path to the directory containing your audio dataset\ndataset_dir = '/kaggle/input/birdclef-2024/train_audio'\n# Initialize an empty dictionary to store the mapping between audio files and labels\nlabel_mapping = {}\n# Iterate over subdirectories (classes) in the dataset directory\nfor label in os.listdir(dataset_dir):\n    label_dir = os.path.join(dataset_dir, label)\n    # Check if the item in the dataset directory is a directory\n    if os.path.isdir(label_dir):\n        # Iterate over audio files in the subdirectory (class)\n        for audio_file in os.listdir(label_dir):\n            # Add the mapping between audio file path and label to the dictionary\n            audio_file_path = os.path.join(label_dir, audio_file)\n            label_mapping[audio_file_path] = label","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:24:27.169768Z","iopub.execute_input":"2024-06-10T20:24:27.170256Z","iopub.status.idle":"2024-06-10T20:24:33.291017Z","shell.execute_reply.started":"2024-06-10T20:24:27.170218Z","shell.execute_reply":"2024-06-10T20:24:33.289911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a list of tuples containing the audio file paths and labels\ndata = [(audio_file_path, label) for audio_file_path, label in label_mapping.items()]\n# Create a Pandas DataFrame from the list of tuples\nannotated_data = pd.DataFrame(data, columns=['audio_file_path', 'label'])\nannotated_data","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:25:03.526516Z","iopub.execute_input":"2024-06-10T20:25:03.526975Z","iopub.status.idle":"2024-06-10T20:25:03.555086Z","shell.execute_reply.started":"2024-06-10T20:25:03.526943Z","shell.execute_reply":"2024-06-10T20:25:03.553882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_encoder = LabelEncoder()\nannotated_data['encoded_label'] = label_encoder.fit_transform(annotated_data['label'])\nannotated_data","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:25:07.218629Z","iopub.execute_input":"2024-06-10T20:25:07.219085Z","iopub.status.idle":"2024-06-10T20:25:07.246418Z","shell.execute_reply.started":"2024-06-10T20:25:07.219049Z","shell.execute_reply":"2024-06-10T20:25:07.244938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Label encoding** is an essential step when you need to convert categorical labels into a numerical ‎format that can be used by machine learning models. You can incorporate label encoding after ‎loading the dataset and extracting features but before feeding the data into your model.‎","metadata":{}},{"cell_type":"markdown","source":"# 3. Random Sampling Data","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(24, 12))\nsns.countplot(x='primary_label', data=meta_data, order=meta_data['primary_label'].value_counts().index)\nplt.xticks(rotation=45)\nplt.rc('font', size=6)\nplt.title('Count of Bird Species Classes')\nplt.xlabel('Bird Species')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-09T18:04:40.793146Z","iopub.execute_input":"2024-06-09T18:04:40.793697Z","iopub.status.idle":"2024-06-09T18:04:42.6811Z","shell.execute_reply.started":"2024-06-09T18:04:40.793658Z","shell.execute_reply":"2024-06-09T18:04:42.679938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Imbalanced data** refers to a situation in classification problems where the distribution of class ‎labels in the dataset is skewed or disproportionate. In other words, one or more classes have ‎significantly fewer samples compared to other classes‎","metadata":{}},{"cell_type":"code","source":"x = np.vstack(pickled_extracted_features)\ny = annotated_data['encoded_label']\n\nprint(x.shape)\nprint(y.shape)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:25:14.06581Z","iopub.execute_input":"2024-06-10T20:25:14.066365Z","iopub.status.idle":"2024-06-10T20:25:14.132243Z","shell.execute_reply.started":"2024-06-10T20:25:14.066321Z","shell.execute_reply":"2024-06-10T20:25:14.130818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ros = RandomOverSampler(random_state=42)\nfeatures_resampled, labels_reshampled = ros.fit_resample(x, y)\n\nprint(features_resampled.shape)\nprint(labels_reshampled.shape)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:25:17.855094Z","iopub.execute_input":"2024-06-10T20:25:17.856126Z","iopub.status.idle":"2024-06-10T20:25:17.925982Z","shell.execute_reply.started":"2024-06-10T20:25:17.856085Z","shell.execute_reply":"2024-06-10T20:25:17.924654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The RandomOverSampler** is a technique used in imbalanced classification problems to balance ‎the class distribution by randomly duplicating samples from the minority class(es) until they are ‎represented in a proportion similar to the majority class. This technique helps address the issue ‎of imbalanced datasets, where one or more classes are significantly underrepresented ‎compared to others.‎","metadata":{}},{"cell_type":"markdown","source":"# **Model Training**","metadata":{}},{"cell_type":"code","source":"# Split data into training and testing sets\nx_train, x_test, y_train, y_test = train_test_split(features_resampled, labels_reshampled, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:25:21.401298Z","iopub.execute_input":"2024-06-10T20:25:21.402203Z","iopub.status.idle":"2024-06-10T20:25:21.429837Z","shell.execute_reply.started":"2024-06-10T20:25:21.402161Z","shell.execute_reply":"2024-06-10T20:25:21.428638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize the XGBoost classifier\nxgb_model = xgb.XGBClassifier(n_estimators=100, random_state=42)\nmodel = xgb_model.fit(x_train, y_train)\ny_predict = model.predict(x_test)\naccuracy = accuracy_score(y_test, y_predict)\nprint(\"Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T18:09:00.787459Z","iopub.execute_input":"2024-06-09T18:09:00.787989Z","iopub.status.idle":"2024-06-09T18:15:34.903328Z","shell.execute_reply.started":"2024-06-09T18:09:00.787954Z","shell.execute_reply":"2024-06-09T18:15:34.901963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_forest_classifier = RandomForestClassifier(n_estimators=100, random_state=42)\nrandom_forest_model = random_forest_classifier.fit(x_train, y_train)\ny_predict = random_forest_model.predict(x_test)\naccuracy = accuracy_score(y_test, y_predict)\nprint(\"Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:41:34.221255Z","iopub.execute_input":"2024-06-10T20:41:34.222176Z","iopub.status.idle":"2024-06-10T20:44:57.153799Z","shell.execute_reply.started":"2024-06-10T20:41:34.222106Z","shell.execute_reply":"2024-06-10T20:44:57.152027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Hyperparameter tuning**","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\nimport xgboost as xgb\nfrom sklearn.metrics import accuracy_score\n\n# Define the parameter grid for XGBoost\nxgb_param_grid = {\n    'n_estimators': [50, 100],\n    'max_depth': [3, 4],\n\n}\n\n# Initialize the XGBoost classifier\nxgb_model = xgb.XGBClassifier(random_state=42)\n\n# Set up GridSearchCV\nxgb_grid_search = GridSearchCV(estimator=xgb_model, param_grid=xgb_param_grid, \n                               scoring='accuracy', cv=5, verbose=1, n_jobs=1)\n\n# Fit GridSearchCV\nxgb_grid_search.fit(x_train, y_train)\n\n# Get the best parameters and best score\nxgb_best_params = xgb_grid_search.best_params_\nxgb_best_score = xgb_grid_search.best_score_\n\nprint(\"Best XGBoost Parameters:\", xgb_best_params)\nprint(\"Best XGBoost CV Accuracy:\", xgb_best_score)\n\n# Train the model with the best parameters\nxgb_best_model = xgb_grid_search.best_estimator_\ny_predict = xgb_best_model.predict(x_test)\naccuracy = accuracy_score(y_test, y_predict)\nprint(\"XGBoost Test Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T21:27:23.87139Z","iopub.execute_input":"2024-06-09T21:27:23.87262Z","iopub.status.idle":"2024-06-09T22:38:26.665892Z","shell.execute_reply.started":"2024-06-09T21:27:23.872528Z","shell.execute_reply":"2024-06-09T22:38:26.664162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the parameter grid for Random Forest\nrf_param_grid = {\n    'n_estimators': [50, 100],\n    'max_depth': [20, 30],\n\n}\n\n# Initialize the Random Forest classifier\nrf_model = RandomForestClassifier(random_state=42)\n\n# Set up GridSearchCV with limited parallel jobs\nrf_grid_search = GridSearchCV(estimator=rf_model, param_grid=rf_param_grid, \n                              scoring='accuracy', cv=5, verbose=1, n_jobs=1)  # Limited to 2 jobs\n\n# Fit GridSearchCV\nrf_grid_search.fit(x_train, y_train)\n\n# Get the best parameters and best score\nrf_best_params = rf_grid_search.best_params_\nrf_best_score = rf_grid_search.best_score_\n\nprint(\"Best Random Forest Parameters:\", rf_best_params)\nprint(\"Best Random Forest CV Accuracy:\", rf_best_score)\n\n# Train the model with the best parameters\nrf_best_model = rf_grid_search.best_estimator_\ny_predict = rf_best_model.predict(x_test)\naccuracy = accuracy_score(y_test, y_predict)\nprint(\"Random Forest Test Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T22:47:48.869903Z","iopub.execute_input":"2024-06-09T22:47:48.870945Z","iopub.status.idle":"2024-06-09T23:21:34.271166Z","shell.execute_reply.started":"2024-06-09T22:47:48.870907Z","shell.execute_reply":"2024-06-09T23:21:34.269771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Model Evaluation**","metadata":{}},{"cell_type":"code","source":"def evaluate_model(y_true, y_pred):\n    # Calculate accuracy\n    accuracy = accuracy_score(y_true, y_pred)\n    # Calculate precision\n    precision = precision_score(y_true, y_pred, average='weighted')\n    # Calculate recall\n    recall = recall_score(y_true, y_pred, average='weighted')\n    # Calculate F1 score\n    f1 = f1_score(y_true, y_pred, average='weighted')\n    \n    return accuracy, precision, recall, f1\n\n# Evaluate the model\naccuracy, precision, recall, f1 = evaluate_model(y_test, y_predict)\n# Print evaluation metrics\nprint(\"Accuracy:\", accuracy)\nprint(\"Precision:\", precision)\nprint(\"Recall:\", recall)\nprint(\"F1 Score:\", f1)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T17:42:20.817742Z","iopub.execute_input":"2024-06-10T17:42:20.818144Z","iopub.status.idle":"2024-06-10T17:42:20.880045Z","shell.execute_reply.started":"2024-06-10T17:42:20.818112Z","shell.execute_reply":"2024-06-10T17:42:20.879026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_proba = random_forest_model.predict_proba(x_test)\ny_test_bin = label_binarize(y_test, classes=np.unique(y_train))\nn_classes = y_test_bin.shape[1]\n# Compute ROC curve and ROC area for each class\nfpr = dict()\ntpr = dict()\nroc_auc = dict()\nfor i in range(n_classes):\n    fpr[i], tpr[i], _ = roc_curve(y_test_bin[:, i], y_proba[:, i])\n    roc_auc[i] = auc(fpr[i], tpr[i])\n\n# Compute micro-average ROC curve and ROC area\nfpr[\"micro\"], tpr[\"micro\"], _ = roc_curve(y_test_bin.ravel(), y_proba.ravel())\nroc_auc[\"micro\"] = auc(fpr[\"micro\"], tpr[\"micro\"])\n\n# Plot ROC curve for each class\nplt.figure()\ncolors = cycle(['aqua', 'darkorange', 'cornflowerblue'])\nfor i, color in zip(range(n_classes), colors):\n    plt.plot(fpr[i], tpr[i], color=color, lw=2,\n             label='ROC curve of class {0} (area = {1:0.2f})'\n             ''.format(i, roc_auc[i]))\n\nplt.plot([0, 1], [0, 1], 'k--', lw=2)\nplt.xlim([0.0, 1.0])\nplt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Receiver Operating Characteristic (ROC) Curve')\nplt.legend(loc=\"lower right\")\nplt.show()\n\n# Calculate the average AUC score\naverage_auc = np.mean(list(roc_auc.values()))\nprint(\"Average AUC Score:\", average_auc)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T17:42:36.010839Z","iopub.execute_input":"2024-06-10T17:42:36.011693Z","iopub.status.idle":"2024-06-10T17:42:41.381364Z","shell.execute_reply.started":"2024-06-10T17:42:36.011652Z","shell.execute_reply":"2024-06-10T17:42:41.379975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Model Testing and Deployment**","metadata":{}},{"cell_type":"code","source":"dump(random_forest_model, 'random_forest.joblib')","metadata":{"execution":{"iopub.status.busy":"2024-06-10T21:59:39.942867Z","iopub.execute_input":"2024-06-10T21:59:39.943486Z","iopub.status.idle":"2024-06-10T22:00:07.576759Z","shell.execute_reply.started":"2024-06-10T21:59:39.943415Z","shell.execute_reply":"2024-06-10T22:00:07.574209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = load('/kaggle/working/random_forest.joblib')","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:00:39.970533Z","iopub.execute_input":"2024-06-10T22:00:39.970991Z","iopub.status.idle":"2024-06-10T22:01:02.313013Z","shell.execute_reply.started":"2024-06-10T22:00:39.970957Z","shell.execute_reply":"2024-06-10T22:01:02.311582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def audio_classification(file_path):\n    audio = file_path\n    print(audio)\n    extracted_features = extract_features(audio).reshape(1, -1)\n    # extracted_features = x_test[112].reshape(1, -1)\n    y_predict = model.predict(extracted_features)\n    labels_list = annotated_data['label'].unique()\n    encoded_label = annotated_data['encoded_label'].unique()\n\n    labels = {}\n    for label, prediction in zip(encoded_label, labels_list):\n        labels[label] = prediction\n    if y_predict[0] in labels.keys():\n        predicted = ('Predicted Class:', labels[y_predict[0]])\n    return predicted","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:01:25.800534Z","iopub.execute_input":"2024-06-10T22:01:25.80096Z","iopub.status.idle":"2024-06-10T22:01:25.810975Z","shell.execute_reply.started":"2024-06-10T22:01:25.800928Z","shell.execute_reply":"2024-06-10T22:01:25.809452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file_path = '/kaggle/input/birdclef-2024/unlabeled_soundscapes/1001358022.ogg'\naudio_analysis(file_path)\nAudio(file_path)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:47:30.33319Z","iopub.execute_input":"2024-06-10T20:47:30.334498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_classification(file_path)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T20:48:23.78188Z","iopub.execute_input":"2024-06-10T20:48:23.78235Z","iopub.status.idle":"2024-06-10T20:48:30.008935Z","shell.execute_reply.started":"2024-06-10T20:48:23.782307Z","shell.execute_reply":"2024-06-10T20:48:30.007316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Project Submission**","metadata":{}},{"cell_type":"code","source":"test_soundscapes = '/kaggle/input/birdclef-2024/test_soundscapes'\n\nfor path in Path(test_soundscapes).glob(\"*.ogg\"):\n    print(path)\n    print(path.stem)\n    print(path.stem.split(\"_\"))","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:01:39.321114Z","iopub.execute_input":"2024-06-10T22:01:39.322221Z","iopub.status.idle":"2024-06-10T22:01:39.343494Z","shell.execute_reply.started":"2024-06-10T22:01:39.322174Z","shell.execute_reply":"2024-06-10T22:01:39.342284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.DataFrame(\n     [(path.stem, *path.stem.split(\"_\"), path) for path in Path(test_soundscapes).glob(\"*.ogg\")],\n    columns = [\"filename\", \"name\" ,\"id\", \"path\"]\n)\nprint(test.shape)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:01:48.839918Z","iopub.execute_input":"2024-06-10T22:01:48.84038Z","iopub.status.idle":"2024-06-10T22:01:48.867095Z","shell.execute_reply.started":"2024-06-10T22:01:48.840338Z","shell.execute_reply":"2024-06-10T22:01:48.865629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filenames = test.filename.values.tolist()\nbird_cols = list(pd.get_dummies(meta_data['primary_label']).columns)\nsubmission_df = pd.DataFrame(columns=['row_id']+bird_cols)\nsubmission_df","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:01:52.1157Z","iopub.execute_input":"2024-06-10T22:01:52.116122Z","iopub.status.idle":"2024-06-10T22:01:52.15519Z","shell.execute_reply.started":"2024-06-10T22:01:52.116088Z","shell.execute_reply":"2024-06-10T22:01:52.153969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, file in enumerate(filenames):\n    predicted = model.predict[i]\n    num_rows = len(predicted)\n    row_ids = [f'{file}_{(i+1)*5}' for i in range(num_rows)]\n    df = pd.DataFrame(columns=['row_id']+bird_cols)\n    \n    df['row_id'] = row_ids\n    df[bird_cols] = predicted\n    \n    submission_df = pd.concat([submission_df,df]).reset_index(drop=True)\n ","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:01:56.546055Z","iopub.execute_input":"2024-06-10T22:01:56.546497Z","iopub.status.idle":"2024-06-10T22:01:56.555671Z","shell.execute_reply.started":"2024-06-10T22:01:56.546462Z","shell.execute_reply":"2024-06-10T22:01:56.55418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"   submission_df","metadata":{"execution":{"iopub.status.busy":"2024-06-10T21:03:31.560544Z","iopub.execute_input":"2024-06-10T21:03:31.560982Z","iopub.status.idle":"2024-06-10T21:03:31.578169Z","shell.execute_reply.started":"2024-06-10T21:03:31.560949Z","shell.execute_reply":"2024-06-10T21:03:31.576698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-06-10T22:02:01.155668Z","iopub.execute_input":"2024-06-10T22:02:01.156206Z","iopub.status.idle":"2024-06-10T22:02:01.164028Z","shell.execute_reply.started":"2024-06-10T22:02:01.156086Z","shell.execute_reply":"2024-06-10T22:02:01.162691Z"},"trusted":true},"execution_count":null,"outputs":[]}]}