{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BirdCLEF EDA 🐦 Quickstart: Sound 🔊ON\n\nBirdCLEF is a long running annual competition on Kaggle. Below are the links to the previous editions. Keep in mind, everytime it's a little different.\n\n[**2020 Edition**](https://www.kaggle.com/competitions/birdsong-recognition/) | [**2021 Edition**](https://www.kaggle.com/competitions/birdclef-2021) | [**2022 Edition**](https://www.kaggle.com/competitions/birdclef-2022)\n\nThis year, people will be identifying (classifying) bird calls recorded in Kenya. The biggest change of this year seems to be the [**constrained runtimes**](https://www.kaggle.com/competitions/birdclef-2023/overview/code-requirements). Basically, \n\n* No GPU during submission\n* CPU submission runtime capped at 120 mins!\n\nEvaluation is based on a modified [average precision](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html#sklearn.metrics.average_precision_score). More on this when we explore the data below.\n\n**Please upvote if you find the notebook helpful**\n\n# Contents\n\n[**1. Explore Train Data**](#The-Train-Data)\\\n[**2. Listen to Bird Calls 🎵**](#Let's-hear-the-birdcalls!🎵)\\\n[**3. Similarity b/w calls of same bird 🦜**](#Similarity-b/w-calls-of-same-bird-🦜)\\\n[**4. Visualizing the Audios 🔔**](#Visualizing-the-Audios-🔔)\\\n[**5. Spectrograms ✨**](#Spectrograms)\\\n[**6. Evaluation**](#Evaluation)\\\n[**7. Submission**](#Demo-Submission)","metadata":{}},{"cell_type":"markdown","source":"![](https://images.unsplash.com/photo-1612213158843-5e27ccaa59e5?ixlib=rb-4.0.3&ixid=MnwxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8&auto=format&fit=crop&w=1665&q=80)","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nimport random\nimport numpy as np\nimport pandas as pd\n\nimport plotly.express as px\nimport librosa\nimport librosa.display\nimport IPython.display as ipd\nimport sklearn\nimport warnings\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nwarnings.filterwarnings('ignore')\nsns.set()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-09T05:51:55.763262Z","iopub.execute_input":"2023-03-09T05:51:55.764000Z","iopub.status.idle":"2023-03-09T05:51:58.817031Z","shell.execute_reply.started":"2023-03-09T05:51:55.763949Z","shell.execute_reply":"2023-03-09T05:51:58.815686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TRAIN_METADATA_FILE = \"/kaggle/input/birdclef-2023/train_metadata.csv\"\nTAXONOMY_FILE = \"/kaggle/input/birdclef-2023/eBird_Taxonomy_v2021.csv\"\nAUDIO_DIR = \"/kaggle/input/birdclef-2023/train_audio\"","metadata":{"execution":{"iopub.status.busy":"2023-03-09T05:52:03.991327Z","iopub.execute_input":"2023-03-09T05:52:03.991753Z","iopub.status.idle":"2023-03-09T05:52:03.996864Z","shell.execute_reply.started":"2023-03-09T05:52:03.991716Z","shell.execute_reply":"2023-03-09T05:52:03.995596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The Train Data\n\nLet's first have a preliminary understanding of the main train dataset. ","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv(TRAIN_METADATA_FILE)\ndata.sample(5)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-09T05:52:04.518813Z","iopub.execute_input":"2023-03-09T05:52:04.519225Z","iopub.status.idle":"2023-03-09T05:52:04.683485Z","shell.execute_reply.started":"2023-03-09T05:52:04.519175Z","shell.execute_reply":"2023-03-09T05:52:04.682238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our task is to predict the **PRIMARY_LABEL** for each sample. There are a lot of additional metadata too. However, they seem little of importance for our main objective which is identifying audio of bird calls. There are 264 such primary_labels or bird codes. So let's do some analysis to check how balanced the dataset is in terms of class distribution.","metadata":{}},{"cell_type":"code","source":"label_counts = data.groupby(['primary_label']).agg(count=('filename', 'count')).reset_index().sort_values(['count'], ascending=False)\nprint(f\"Number of birds with over 100 samples: {np.sum(label_counts['count'] >= 100)}\")\nprint(f\"Number of birds with less than 20 samples: {np.sum(label_counts['count'] < 20)}\")\n\n\nfig = px.bar(label_counts, x='primary_label', y='count')\nfig.update_traces(marker_color='lightsalmon')\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-09T05:52:05.200579Z","iopub.execute_input":"2023-03-09T05:52:05.201030Z","iopub.status.idle":"2023-03-09T05:52:06.828536Z","shell.execute_reply.started":"2023-03-09T05:52:05.200989Z","shell.execute_reply":"2023-03-09T05:52:06.827240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like it varies wildly! From 5 birds with 500 samples each to quite a few with only 1 sample. Also our data has a long tail with nearly 2/3 of the birds have less than 50 samples. Sure looks like a disaster waiting during inference, right? We'll explore this in the evaluation section.","metadata":{}},{"cell_type":"markdown","source":"# Let's hear the birdcalls!🎵\n\nOkay, enough about the data! Now let's hear some birds chirping! Below we listen to 3 randomly selected birds to begin with.","metadata":{}},{"cell_type":"code","source":"sample_data = data.sample(3)\nmedia = []\nfor _, row in sample_data.iterrows():\n    path = os.path.join(AUDIO_DIR, row['filename'])\n    media.append((f\"Common_Name: {row['common_name']} ({row['primary_label']})\", path))","metadata":{"execution":{"iopub.status.busy":"2023-03-09T05:52:09.740714Z","iopub.execute_input":"2023-03-09T05:52:09.741108Z","iopub.status.idle":"2023-03-09T05:52:09.749536Z","shell.execute_reply.started":"2023-03-09T05:52:09.741074Z","shell.execute_reply":"2023-03-09T05:52:09.748424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio = media[0]\nprint(audio[0])\nipd.Audio(audio[1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-08T02:05:34.551423Z","iopub.execute_input":"2023-03-08T02:05:34.552693Z","iopub.status.idle":"2023-03-08T02:05:34.600898Z","shell.execute_reply.started":"2023-03-08T02:05:34.552633Z","shell.execute_reply":"2023-03-08T02:05:34.599999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio = media[1]\nprint(audio[0])\nipd.Audio(audio[1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-08T02:05:34.847779Z","iopub.execute_input":"2023-03-08T02:05:34.848590Z","iopub.status.idle":"2023-03-08T02:05:34.861295Z","shell.execute_reply.started":"2023-03-08T02:05:34.848545Z","shell.execute_reply":"2023-03-08T02:05:34.859918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio = media[2]\nprint(audio[0])\nipd.Audio(audio[1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-08T02:05:35.052499Z","iopub.execute_input":"2023-03-08T02:05:35.053109Z","iopub.status.idle":"2023-03-08T02:05:35.070197Z","shell.execute_reply.started":"2023-03-08T02:05:35.053071Z","shell.execute_reply":"2023-03-08T02:05:35.069189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Similarity b/w calls of same bird 🦜\n\nIf an ML model is to recognise a bird, first we might check we as humans are able to distinguish it or not. We pick the top 3 most frequent bird from our dataset. You'll find there are definite traits of similarities between the different calls of a bird.","metadata":{}},{"cell_type":"code","source":"call_samples = []\nbirds = label_counts['primary_label'][:3]\n\nfor i,bird in enumerate(birds):\n    samples = data.loc[data['primary_label'] == bird, 'filename'][:3].reset_index(drop=True)\n    call_samples.append((bird, [os.path.join(AUDIO_DIR, samples[p]) for p in range(3)]))","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:18:27.382992Z","iopub.execute_input":"2023-03-08T02:18:27.383716Z","iopub.status.idle":"2023-03-08T02:18:27.397200Z","shell.execute_reply.started":"2023-03-08T02:18:27.383671Z","shell.execute_reply":"2023-03-08T02:18:27.395893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### -------------------------------- barswa --------------------------------","metadata":{}},{"cell_type":"code","source":"sample = call_samples[0]\nipd.Audio(sample[1][0])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-08T02:19:06.218817Z","iopub.execute_input":"2023-03-08T02:19:06.220067Z","iopub.status.idle":"2023-03-08T02:19:06.251869Z","shell.execute_reply.started":"2023-03-08T02:19:06.219997Z","shell.execute_reply":"2023-03-08T02:19:06.250875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = call_samples[0]\nipd.Audio(sample[1][1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-08T02:19:07.006617Z","iopub.execute_input":"2023-03-08T02:19:07.007930Z","iopub.status.idle":"2023-03-08T02:19:07.048516Z","shell.execute_reply.started":"2023-03-08T02:19:07.007869Z","shell.execute_reply":"2023-03-08T02:19:07.047523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = call_samples[0]\nipd.Audio(sample[1][2])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-08T02:19:07.696236Z","iopub.execute_input":"2023-03-08T02:19:07.697584Z","iopub.status.idle":"2023-03-08T02:19:07.713274Z","shell.execute_reply.started":"2023-03-08T02:19:07.697421Z","shell.execute_reply":"2023-03-08T02:19:07.711839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### -------------------------------- wlwwar --------------------------------","metadata":{}},{"cell_type":"code","source":"sample = call_samples[1]\nipd.Audio(sample[1][2])","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = call_samples[1]\nipd.Audio(sample[1][2])","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = call_samples[1]\nipd.Audio(sample[1][2])","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### -------------------------------- thrnig1 --------------------------------","metadata":{}},{"cell_type":"code","source":"sample = call_samples[2]\nipd.Audio(sample[1][2])","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:20:24.096736Z","iopub.execute_input":"2023-03-08T02:20:24.097771Z","iopub.status.idle":"2023-03-08T02:20:24.145494Z","shell.execute_reply.started":"2023-03-08T02:20:24.097722Z","shell.execute_reply":"2023-03-08T02:20:24.144526Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = call_samples[2]\nipd.Audio(sample[1][2])","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:20:24.378307Z","iopub.execute_input":"2023-03-08T02:20:24.379265Z","iopub.status.idle":"2023-03-08T02:20:24.396890Z","shell.execute_reply.started":"2023-03-08T02:20:24.379211Z","shell.execute_reply":"2023-03-08T02:20:24.395651Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = call_samples[2]\nipd.Audio(sample[1][2])","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:20:24.886683Z","iopub.execute_input":"2023-03-08T02:20:24.887356Z","iopub.status.idle":"2023-03-08T02:20:24.905379Z","shell.execute_reply.started":"2023-03-08T02:20:24.887311Z","shell.execute_reply":"2023-03-08T02:20:24.904468Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing the Audios 🔔\n\nDuring the modeling phase, we'll be converting the audios into a visual medium, which are melspectograms. But before moving on to that, let's visualize our audio in a more simpler way! Let's take the 3 samples that we heard above.\n\nAs you can see below, the 3 examples in our case are pretty distinct even visually since they come from 3 separate classes. ","metadata":{}},{"cell_type":"code","source":"audio = media[0]\nprint(audio[0])\ny, sr = librosa.load(audio[1])\nfig, ax = plt.subplots(figsize=(10, 5))\nsns.lineplot(x=np.arange(len(y)), y=y, ax=ax);","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:38:01.787737Z","iopub.execute_input":"2023-03-08T02:38:01.788857Z","iopub.status.idle":"2023-03-08T02:38:03.878257Z","shell.execute_reply.started":"2023-03-08T02:38:01.788811Z","shell.execute_reply":"2023-03-08T02:38:03.876514Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio = media[1]\nprint(audio[0])\ny, sr = librosa.load(audio[1])\nfig, ax = plt.subplots(figsize=(10, 5))\nsns.lineplot(x=np.arange(len(y)), y=y, ax=ax);","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:38:16.800184Z","iopub.execute_input":"2023-03-08T02:38:16.801024Z","iopub.status.idle":"2023-03-08T02:38:18.029356Z","shell.execute_reply.started":"2023-03-08T02:38:16.800971Z","shell.execute_reply":"2023-03-08T02:38:18.028027Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio = media[2]\nprint(audio[0])\ny, sr = librosa.load(audio[1])\nfig, ax = plt.subplots(figsize=(10, 5))\nsns.lineplot(x=np.arange(len(y)), y=y, ax=ax);","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:38:22.493666Z","iopub.execute_input":"2023-03-08T02:38:22.494080Z","iopub.status.idle":"2023-03-08T02:38:23.810989Z","shell.execute_reply.started":"2023-03-08T02:38:22.494043Z","shell.execute_reply":"2023-03-08T02:38:23.809507Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Spectrograms\n\nSpectrograms help us visualize an audio. It is usually obtained by converting the data from the time domain to the frequency domain at which point we can plot the amplitudes at different frequencies which ultimately represent an image. This is very important since almost all of the top solutions that we will witness in this competition will use this method to convert an audio to spectrogram as preprocessing, virtually converting this to a computer vision challenge.","metadata":{}},{"cell_type":"code","source":"n_fft=2048\nhop_length=512\n\nfig, ax = plt.subplots(3, 1, figsize=(15, 12))\n\nfor i in range(3):\n    audio = media[i]\n    y, sr = librosa.load(audio[1])\n    D = np.abs(librosa.stft(y, n_fft = n_fft, hop_length = hop_length))\n    DB = librosa.amplitude_to_db(D, ref = np.max)\n    \n    fig.suptitle('Log Frequency Spectrogram      ', fontsize=16)\n    img=librosa.display.specshow(DB, sr = sr, hop_length = hop_length, y_axis = 'log', ax=ax[i])\n    ax[i].set_title(audio[0], fontsize=12) \nplt.colorbar(img,ax=ax)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-09T06:03:07.471288Z","iopub.execute_input":"2023-03-09T06:03:07.471685Z","iopub.status.idle":"2023-03-09T06:03:16.295815Z","shell.execute_reply.started":"2023-03-09T06:03:07.471652Z","shell.execute_reply":"2023-03-09T06:03:16.294715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can already see some distinguishing factors, can't you! \n\nSimilar to log frequency spectograms, there are other slight variations like Melspectrograms etc. Surely explore those and see which you find better.","metadata":{}},{"cell_type":"markdown","source":"# Evaluation\n\nWe need to talk a little about evaluation and the evaluation metrics. The official metric for this competition is a variation of average precision. We saw in the train data the classes are very unbalanced. So a simple average precision can have not so ideal effects. Well organizers realized this and hence:\n\n> The evaluation metric for this contest is padded cmAP, a derivative of the macro-averaged average precision score as implemented by scikit-learn. In order to support accepting predictions for species with zero true positive labels and to reduce the impact of species with very few positive labels, prior to scoring we pad each submission and the solution with five rows of true positives. This means that even a baseline submission will get a relatively strong score.\n\n-- [From the evaluation page](https://www.kaggle.com/competitions/birdclef-2023/overview/evaluation)\n\nWhat does this mean? First you need to look at the given code below:","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn import metrics\n\ndef padded_cmap(solution, submission, padding_factor=5):\n    solution = solution.drop(['row_id'], axis=1, errors='ignore')\n    submission = submission.drop(['row_id'], axis=1, errors='ignore')\n    new_rows = []\n    for i in range(padding_factor):\n        new_rows.append([1 for i in range(len(solution.columns))])\n    new_rows = pd.DataFrame(new_rows)\n    new_rows.columns = solution.columns\n    padded_solution = pd.concat([solution, new_rows]).reset_index(drop=True).copy()\n    padded_submission = pd.concat([submission, new_rows]).reset_index(drop=True).copy()\n    score = metrics.average_precision_score(\n        padded_solution.values,\n        padded_submission.values,\n        average='macro',\n    )\n    return score","metadata":{"execution":{"iopub.status.busy":"2023-03-09T06:10:29.413398Z","iopub.execute_input":"2023-03-09T06:10:29.414361Z","iopub.status.idle":"2023-03-09T06:10:29.496931Z","shell.execute_reply.started":"2023-03-09T06:10:29.414316Z","shell.execute_reply":"2023-03-09T06:10:29.495897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **macro** parameter at the end means that the average precision is calculated for each 264 labels separately and then averaged on an unweighted manner. Hence, the effect of a very rare class and a very common class is the same. The way the counteracted this is by introducing 5 rows of true positives. For a very rare class, we will have only very few positives. Hence, even if you do a bad job at identifying those, your score will not take that bad of a hit. Also, to note here is that your scores are not to be thresholded. It is imperative that they stay in the raw form in the submission.","metadata":{}},{"cell_type":"markdown","source":"# Demo Submission\n\nBelow we do a submission for demonstration purposes. During submission time, the model should output logits for each 5 second segment of the 100 odd test files available. Expect to see around 100 * 12 * 10 = 12,000 rows. We will output zeros for all the labels. Note that if we output a constant value for each label, that is equivalent to the baseline (0.71). The metric above rewards better discriminative quality **between different rows for a particular class**, not between different classes of a certain row.","metadata":{}},{"cell_type":"code","source":"from copy import deepcopy\n\nlabels = {key:0 for key in data['primary_label'].unique()}\nrow_ids = pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")['row_id']\n\npreds = []\nfor row_id in row_ids:\n    pred_i = deepcopy(labels)\n    pred_i['row_id'] = row_id\n    preds.append(pred_i)\n    \nsubmission = pd.DataFrame(preds)\nsubmission.to_csv(\"submission.csv\", index=False)\n\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-09T06:27:53.176501Z","iopub.execute_input":"2023-03-09T06:27:53.176940Z","iopub.status.idle":"2023-03-09T06:27:53.232474Z","shell.execute_reply.started":"2023-03-09T06:27:53.176902Z","shell.execute_reply":"2023-03-09T06:27:53.231618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### That's all for now. **Please upvote if you find this notebook helpful.** And have a great competition ahead!","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}