{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"The purpose of this notebook is twofold; to serve as a baseline entry for me, and to serve as an educational guide for those who may be undertaking a competition of this complexity for the first time. You are most welcome to use this notebook as a guide to build upon. Feel free to comment on things you find incorrect or could be improved.","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"# Standard data manipulation libraries\nimport numpy as np\nimport pandas as pd\nimport os\nimport joblib # Used for exporting and importing trained models\nimport librosa # Used for processing audio\n\n# Visualization libraries\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nfrom IPython.display import Audio # Used for displaying an interactive audio player\n\n# Machine learning libraries\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nfrom imblearn.over_sampling import RandomOverSampler\nfrom xgboost import XGBClassifier\n\n%matplotlib inline\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:42.885606Z","iopub.execute_input":"2024-04-24T07:15:42.886157Z","iopub.status.idle":"2024-04-24T07:15:46.037609Z","shell.execute_reply.started":"2024-04-24T07:15:42.886121Z","shell.execute_reply":"2024-04-24T07:15:46.036414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Data","metadata":{}},{"cell_type":"code","source":"df_metadata = pd.read_csv('/kaggle/input/birdclef-2024/train_metadata.csv')\ndf_ebird_taxonomy = pd.read_csv('/kaggle/input/birdclef-2024/eBird_Taxonomy_v2021.csv')\nsample_submission = pd.read_csv('/kaggle/input/birdclef-2024/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:46.039498Z","iopub.execute_input":"2024-04-24T07:15:46.040909Z","iopub.status.idle":"2024-04-24T07:15:46.389860Z","shell.execute_reply.started":"2024-04-24T07:15:46.040853Z","shell.execute_reply":"2024-04-24T07:15:46.388570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Metadata summary\ndisplay(df_metadata.head(2))\ndisplay(df_metadata.info())","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:46.392232Z","iopub.execute_input":"2024-04-24T07:15:46.393091Z","iopub.status.idle":"2024-04-24T07:15:46.478996Z","shell.execute_reply.started":"2024-04-24T07:15:46.393037Z","shell.execute_reply":"2024-04-24T07:15:46.477277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This tells us the training data has **24459 entries**","metadata":{}},{"cell_type":"code","source":"# Taxonomy data summary\ndisplay(df_ebird_taxonomy.head(2))\ndisplay(df_ebird_taxonomy.info())","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:46.482122Z","iopub.execute_input":"2024-04-24T07:15:46.483405Z","iopub.status.idle":"2024-04-24T07:15:46.530243Z","shell.execute_reply.started":"2024-04-24T07:15:46.483360Z","shell.execute_reply":"2024-04-24T07:15:46.528752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sample submission file\ndisplay(sample_submission.head(2))","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:46.532020Z","iopub.execute_input":"2024-04-24T07:15:46.532448Z","iopub.status.idle":"2024-04-24T07:15:46.566666Z","shell.execute_reply.started":"2024-04-24T07:15:46.532415Z","shell.execute_reply":"2024-04-24T07:15:46.565144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the sample submission file, we see that for each test soundscape, we need to predict the **probability** of the bird species in the audio being each unique class. Moreover, the test soundscapes are 4 minutes each with each soundscape recording possibly containing multiple bird species. We therefore need to split each soundscape into 5-second intervals and make a prediction for each interval. This means that for each 4-minute soundscape, there will be 48 predictions ($4mins / 5secs = 48$). The row_id for each prediction will be in the format; \"soundscape_[soundscape_id]_[end_time]\" with [end time] referring to the interval of the soundscape.","metadata":{}},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"# Checking for null values\nprint('Metadata')\ndisplay(df_metadata.isnull().sum())\nprint('-' * 25)\nprint('Taxonomy Data')\ndisplay(df_ebird_taxonomy.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:46.568978Z","iopub.execute_input":"2024-04-24T07:15:46.569525Z","iopub.status.idle":"2024-04-24T07:15:46.628226Z","shell.execute_reply.started":"2024-04-24T07:15:46.569479Z","shell.execute_reply":"2024-04-24T07:15:46.626901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that there are several null values but seeing as we will using the audio files only for training, we will ignore these features","metadata":{}},{"cell_type":"code","source":"bird_freq = df_metadata['primary_label'].value_counts()\n\nfig, (ax1, ax2) = plt.subplots(2, 1, figsize = (20, 6), sharey = True)\nfig.tight_layout(pad = 6)\n\nsns.barplot(x = bird_freq[0:91].index, y = bird_freq[0:91].values, ax = ax1)\nax1.set_xticklabels(bird_freq[0:91].index, rotation = 90)\nax1.set_ylim(0, 500)\nax1.set_xlabel('')\nax1.set_ylabel('Count')\n\nsns.barplot(x = bird_freq[91:].index, y = bird_freq[91:].values, ax = ax2)\nax2.set_xticklabels(bird_freq[91:].index, rotation = 90)\nax2.set_ylim(0, 500)\nax2.set_xlabel('Primary Label')\nax2.set_ylabel('Count')\n\nfig.suptitle('Distribution of Bird Species')\n\nfig.show()\n\n# Prints the total number of unique bird species\nprint('Total unique bird species: ', len(df_metadata['primary_label'].unique()), '\\n')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:15:48.775576Z","iopub.execute_input":"2024-04-24T07:15:48.776071Z","iopub.status.idle":"2024-04-24T07:15:51.791537Z","shell.execute_reply.started":"2024-04-24T07:15:48.776034Z","shell.execute_reply":"2024-04-24T07:15:51.790101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the figure above, we observe a major class imbalance where 12% (22 out of 182) of the unique species make up 45% (11000 out of 24459) of the total population. To mitigate the potentially skewed prediction this will bring, we can use an **oversampler** to equalize the number of samples between each class.","metadata":{}},{"cell_type":"code","source":"# Plots the bird species according to their geographical location (latitude, longitude)\nfig = px.scatter_mapbox(df_metadata, lat = 'latitude', lon = 'longitude', color = 'primary_label', title = 'Geographical Distribution of Bird Species', zoom = 0, height = 600, width = 1000, mapbox_style = 'open-street-map')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-24T07:18:00.427147Z","iopub.execute_input":"2024-04-24T07:18:00.427683Z","iopub.status.idle":"2024-04-24T07:18:01.074475Z","shell.execute_reply.started":"2024-04-24T07:18:00.427643Z","shell.execute_reply":"2024-04-24T07:18:01.073377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Sample audio file","metadata":{}},{"cell_type":"code","source":"# Function for displaying a waveform\ndef plot_waveform(input_path, title = 'Waveform'):\n    waveform, sample_rate = librosa.load(input_path)\n    \n    fig, ax = plt.subplots(1,1, figsize = (8, 3))\n    librosa.display.waveshow(waveform, sr=sample_rate, ax=ax)\n    plt.title(title)\n    plt.show()\n\n# Audio player\ndisplay(Audio('/kaggle/input/birdclef-2024/train_audio/asbfly/XC134896.ogg'))\n\n# Plots the waveform of one of the training audio files\nplot_waveform('/kaggle/input/birdclef-2024/train_audio/asbfly/XC134896.ogg', 'Muscicapa dauurica')","metadata":{"execution":{"iopub.status.busy":"2024-04-23T12:39:30.928736Z","iopub.execute_input":"2024-04-23T12:39:30.929120Z","iopub.status.idle":"2024-04-23T12:39:39.802449Z","shell.execute_reply.started":"2024-04-23T12:39:30.929092Z","shell.execute_reply":"2024-04-23T12:39:39.801586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each training audio file is a unique clean recording of a single bird species.","metadata":{}},{"cell_type":"markdown","source":"# Data Wrangling","metadata":{}},{"cell_type":"code","source":"# Adding the full file path to the metadata\naudio_dir_path = '/kaggle/input/birdclef-2024/train_audio/'\ndf_metadata['file_path'] = audio_dir_path + df_metadata['filename']\ndf_metadata.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T03:30:10.591764Z","iopub.execute_input":"2024-04-24T03:30:10.592121Z","iopub.status.idle":"2024-04-24T03:30:10.610916Z","shell.execute_reply.started":"2024-04-24T03:30:10.592091Z","shell.execute_reply":"2024-04-24T03:30:10.609905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"To be able to use a classifier, we must extract or create **numerical** features from the audio files. One way to do that is by extracting the Mel-Frequency Cepstral Coefficients (MFFCs) of an audio file. MFFCs are a way to numerically represent the spectral characteristics of an audio file.","metadata":{}},{"cell_type":"code","source":"def get_coefficients(input_path):\n    # Returns the Mel-Frequency Cepstral Coefficients (MFCCs) from an audio file\n    \n    # loading the audio file into librosa which returns its waveform and sample_rate\n    waveform, sample_rate = librosa.load(input_path)\n    \n    # Extracting the MFCCs from the waveform\n    mfccs = librosa.feature.mfcc(y = waveform, sr = sample_rate, n_mfcc = 40)\n    scaled_mfccs = np.mean(mfccs.T, axis = 0)\n    \n    return scaled_mfccs\n\ndef feature_engineer(df):\n    # This function returns a dataframe containing the MFFCs of each audio file in the dataframe\n    \n    # Empty array where we will store the new features (MFCCs)\n    new_features = []\n    \n    # From earlier, we added the full file path of each training audio file to the metadata.\n    # The purpose of that was so we can iterate over the metadata instead of having to perform a crawl on the training directory\n    for i, file in df.iterrows():\n        input_path = file['file_path']\n        \n        # Using the get_coefficients function to extract the MFCCs\n        coefficients = get_coefficients(input_path)\n        new_features.append(coefficients)\n    \n    # hard coding the feature names for the MFCCs\n    train_columns = np.arange(0, 40)\n    train_columns_str = [str(name) for name in train_columns]\n    \n    df_train = pd.DataFrame(new_features, columns = train_columns_str)\n    df_train = pd.concat([df_train, df['primary_label']], axis = 1)\n    \n    return df_train","metadata":{"execution":{"iopub.status.busy":"2024-04-24T03:55:01.480411Z","iopub.execute_input":"2024-04-24T03:55:01.481436Z","iopub.status.idle":"2024-04-24T03:55:01.489421Z","shell.execute_reply.started":"2024-04-24T03:55:01.481399Z","shell.execute_reply":"2024-04-24T03:55:01.488440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Running the feature engineering functions\ndf_train = feature_engineer(df_metadata)","metadata":{"execution":{"iopub.status.busy":"2024-04-24T03:55:02.630200Z","iopub.execute_input":"2024-04-24T03:55:02.630803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Summary of the training dataset with the new features\ndisplay(df_train.head(2))\ndisplay(df_train['primary_label'].value_counts())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Label Encoding","metadata":{}},{"cell_type":"markdown","source":"The classifier will expect numerical classes. We therefore need to encode the class names into numerical classes by using a **label encoder**.","metadata":{}},{"cell_type":"code","source":"# Extract the unique labels from the metadata\noriginal_labels = df_metadata['primary_label'].unique()\n\n# Fitting the label encoder to the list of labels from the metadata\nlabel_encoder = LabelEncoder()\nlabel_encoder.fit(original_labels)\n\n# encoding the primary_labels\nencoded_labels = label_encoder.transform(df_train['primary_label'])\ndf_train['encoded_label'] = encoded_labels\n\ndisplay(df_train['encoded_label'].value_counts())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Splitting the training dataset into X(independent features) and y(dependent/target feature)\nX, y = df_train.drop(columns = ['primary_label', 'encoded_label']), df_train['encoded_label']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Oversampling\nAs mentioned in the EDA section, oversampling is required to address the class imbalance.","metadata":{}},{"cell_type":"code","source":"# initializing the oversampler\nsampler = RandomOverSampler(random_state = 0)\n\n# fit oversampler and resample\nX_resampled, y_resampled = sampler.fit_resample(X, y)\n\ndf_train_resampled = pd.DataFrame(X_resampled)\ndf_train_resampled['encoded_label'] = y_resampled\ndisplay(df_train_resampled.head(2))\ndisplay(df_train_resampled.info())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Splitting the data into a training set and a testing set\nX_train, X_test, y_train, y_test = train_test_split(X_resampled, y_resampled, test_size = 0.2, random_state = 12)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training","metadata":{}},{"cell_type":"markdown","source":"We will use XGBClassifier as our classifier using only the default parameters and hyperparameters. You may utilize a grid search or random search algorithm to tune the hyperparameters.","metadata":{}},{"cell_type":"code","source":"# Initializing XGBClassifier\n# NOTE: IF you will be running this notebook without a GPU accelerator, remove the tree_method = 'gpu_hist' parameter below. This argument ensures that XGBoost runs the notebook using a GPU.\nclf = XGBClassifier(tree_method = 'gpu_hist')\n\n# Training the model\nclf.fit(X_train, y_train)\n\n# Saving/exporting the model\njoblib.dump(clf, 'clf.joblib')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Simple Model Evaluation","metadata":{}},{"cell_type":"code","source":"# making predictions on the test set\ny_predictions = clf.predict(X_test)\n\n# accuracy score of the model\naccuracy = accuracy_score(y_predictions, y_test)\nprint('Accuracy score: ', accuracy)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# What's next?\nAfter saving the model, we can use it to predict and submit in a separate notebook: [Inferring Guide](https://www.kaggle.com/code/lorenzojayd/birdclef-2024-beginner-guide-inferring/notebook)\n\nAs for improvements, you may look at different feature engineering techniques such as turning the audio chunks into spectograms and using computer vision libraries to train neural networks. You may also experiment with tuning the hyperparameters of the classifier, or even use different classfying algorithms such as random forests, or use different libraries such as LGBM or Scikit-learn.","metadata":{}}]}