{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport torchaudio\nimport torch\nfrom IPython.display import Audio\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-05-13T16:52:17.779133Z","iopub.execute_input":"2022-05-13T16:52:17.780498Z","iopub.status.idle":"2022-05-13T16:52:19.762934Z","shell.execute_reply.started":"2022-05-13T16:52:17.780165Z","shell.execute_reply":"2022-05-13T16:52:19.761629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's first open the training metadata CSV to see what's available.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/birdclef-2022/train_metadata.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T16:52:19.764854Z","iopub.execute_input":"2022-05-13T16:52:19.765125Z","iopub.status.idle":"2022-05-13T16:52:19.920029Z","shell.execute_reply.started":"2022-05-13T16:52:19.765094Z","shell.execute_reply":"2022-05-13T16:52:19.919274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And let's also take a look at the labels that will be used in the submission.","metadata":{}},{"cell_type":"code","source":"scored_birds = set(pd.read_json('/kaggle/input/birdclef-2022/scored_birds.json')[0])\nscored_birds","metadata":{"execution":{"iopub.status.busy":"2022-05-13T16:52:19.928106Z","iopub.execute_input":"2022-05-13T16:52:19.928347Z","iopub.status.idle":"2022-05-13T16:52:19.941722Z","shell.execute_reply.started":"2022-05-13T16:52:19.928316Z","shell.execute_reply":"2022-05-13T16:52:19.940731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's also grab the lengths of each audio file (in seconds) and add that as a column. This might take a few minutes, since there's about ~15,000 files!","metadata":{}},{"cell_type":"code","source":"lengths = []\nfor filename in tqdm(df['filename']):\n    metadata = torchaudio.info('/kaggle/input/birdclef-2022/train_audio/' + filename)\n    lengths.append(metadata.num_frames / metadata.sample_rate)\ndf['audio_length'] = lengths","metadata":{"execution":{"iopub.status.busy":"2022-05-13T16:52:19.943180Z","iopub.execute_input":"2022-05-13T16:52:19.944266Z","iopub.status.idle":"2022-05-13T16:55:38.163377Z","shell.execute_reply.started":"2022-05-13T16:52:19.944196Z","shell.execute_reply":"2022-05-13T16:55:38.162652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In total, the training data has 21,952,580,839 audio samples and is ~190.6 hours long.","metadata":{}},{"cell_type":"code","source":"total_length = sum(df['audio_length'])\ntotal_length, total_length * 32000, total_length / 60 / 60","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:09:08.859131Z","iopub.execute_input":"2022-05-13T17:09:08.859712Z","iopub.status.idle":"2022-05-13T17:09:08.868767Z","shell.execute_reply.started":"2022-05-13T17:09:08.859675Z","shell.execute_reply":"2022-05-13T17:09:08.867618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Play a sample\n\nLet's listen a single sample from the training set. Each sample is rated based on its quality, so let's just listen to one that has a 5 star rating.","metadata":{}},{"cell_type":"code","source":"datum = df[df['rating'] == 5].iloc[0] # get the first 5 star rated recording\naudio, rate = torchaudio.load('/kaggle/input/birdclef-2022/train_audio/' + datum['filename'])\nnum_samples = audio[0].numel()\nprint(f'{num_samples} samples at {rate / 1000} kHz (~{round(num_samples / rate)} seconds)')\ndisplay(Audio(audio, rate=rate))\ndatum","metadata":{"execution":{"iopub.status.busy":"2022-05-13T16:59:29.023063Z","iopub.execute_input":"2022-05-13T16:59:29.023455Z","iopub.status.idle":"2022-05-13T16:59:29.346273Z","shell.execute_reply.started":"2022-05-13T16:59:29.023411Z","shell.execute_reply":"2022-05-13T16:59:29.345606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look at the first 5 seconds of audio from this sample to see what the first bird vocalization in that sample actually looks like. You can do this a number of ways, but the simplest is to average the audio's channels and then plot the samples over time.\n\nSpectrograms are also useful visualization technqiue. Similar to convolution, they slide a window over the audio and extract frequencies at each step. Taking the log of the spectrogram makes it a bit easier to see its structure.","metadata":{}},{"cell_type":"code","source":"averaged_audio_clip = audio.mean(0)[:rate*5]\n\nplt.title('first 5 seconds of audio as pressure over time')\nplt.plot(averaged_audio_clip)\nplt.show()\n\nplt.title('first 5 seconds of audio as a log magnitude spectrogram over time')\nplt.imshow(torch.log10(torchaudio.transforms.Spectrogram(n_fft=512)(averaged_audio_clip)));","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:26:50.216348Z","iopub.execute_input":"2022-05-13T17:26:50.216880Z","iopub.status.idle":"2022-05-13T17:26:52.326751Z","shell.execute_reply.started":"2022-05-13T17:26:50.216840Z","shell.execute_reply":"2022-05-13T17:26:52.325903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Columns\n\nNow let's go back to the CSV, and investigate each column.","metadata":{}},{"cell_type":"code","source":"for key in df:\n    print(key)\n    display(df[key].describe())\n    print()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:10:13.397053Z","iopub.execute_input":"2022-05-13T17:10:13.398022Z","iopub.status.idle":"2022-05-13T17:10:13.563321Z","shell.execute_reply.started":"2022-05-13T17:10:13.397978Z","shell.execute_reply":"2022-05-13T17:10:13.562574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Labels\n\nHere's the competition's description of the \"primary_label\" and \"secondary_labels\" columns:\n\n> primary_label: a code for the bird species. You can review detailed information about the bird codes by appending the code to https://ebird.org/species/, such as https://ebird.org/species/amecro for the American Crow.\n\n> secondary_labels: Background species as annotated by the recordist. An empty list does not mean that no background birds are audible.\n\nNote that the distribution of individual secondary labels seems to be quite a bit different than that of primary labels. Also, it looks like there's some primary labels that never appear as secondary labels, but all secondary labels appear as a primary label.","metadata":{}},{"cell_type":"code","source":"df.primary_label.value_counts()[:30].plot.bar(width=0.9, title=\"Top 30 primary labels\", ylabel=\"occurrences\")\ndf.primary_label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:11:19.252484Z","iopub.execute_input":"2022-05-13T17:11:19.252796Z","iopub.status.idle":"2022-05-13T17:11:19.672553Z","shell.execute_reply.started":"2022-05-13T17:11:19.252755Z","shell.execute_reply":"2022-05-13T17:11:19.671463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"secondary_labels_flattened = pd.Series(np.concatenate(df.secondary_labels.apply(lambda v: v[2:-2].split(\"', '\"))))\nsecondary_labels_flattened = secondary_labels_flattened[secondary_labels_flattened != '']\nsecondary_labels_flattened.value_counts()[:30].plot.bar(width=0.9, title=\"Top 30 secondary labels\", ylabel=\"occurrences\")\nsecondary_labels_flattened.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:11:26.833505Z","iopub.execute_input":"2022-05-13T17:11:26.834177Z","iopub.status.idle":"2022-05-13T17:11:27.357694Z","shell.execute_reply.started":"2022-05-13T17:11:26.834129Z","shell.execute_reply":"2022-05-13T17:11:27.356745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_labels = df.primary_label.append(secondary_labels_flattened)\nall_labels.value_counts()[:30].plot.bar(width=0.9, title=\"Top 30 primary+secondary labels\", ylabel=\"occurrences\")\nall_labels.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:11:32.314462Z","iopub.execute_input":"2022-05-13T17:11:32.315295Z","iopub.status.idle":"2022-05-13T17:11:32.898698Z","shell.execute_reply.started":"2022-05-13T17:11:32.315247Z","shell.execute_reply":"2022-05-13T17:11:32.898044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Filter labels by scored birds\n\nOnly about 10% of training samples have a primary or secondary labels that is a scored bird.","metadata":{}},{"cell_type":"code","source":"primary_labels_filter = df.primary_label.isin(scored_birds)\nsecondary_labels = df.secondary_labels.apply(lambda v: v[2:-2].split(\"', '\"))\nsecondary_label_filter = np.array([len(scored_birds.intersection(x)) > 0 for x in secondary_labels])\n\npct_scored_bird = 100 * len(df[primary_labels_filter | secondary_label_filter]) / len(df)\nprint(f'% of samples with a scored bird in the primary or secondary labels: {pct_scored_bird}%')","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:13:08.930397Z","iopub.execute_input":"2022-05-13T17:13:08.931356Z","iopub.status.idle":"2022-05-13T17:13:08.961512Z","shell.execute_reply.started":"2022-05-13T17:13:08.931315Z","shell.execute_reply":"2022-05-13T17:13:08.960599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.primary_label[df.primary_label.isin(scored_birds)].value_counts().plot.bar(\n    width=0.9, title=\"Primary labels - scored\", ylabel=\"occurrences\"\n)\ndf.primary_label[df.primary_label.isin(scored_birds)].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:16:44.599205Z","iopub.execute_input":"2022-05-13T17:16:44.599546Z","iopub.status.idle":"2022-05-13T17:16:44.954547Z","shell.execute_reply.started":"2022-05-13T17:16:44.599512Z","shell.execute_reply":"2022-05-13T17:16:44.953624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Audio lengths\n\nLet's take a look at the distribution of recordings based on the length of audio.","metadata":{}},{"cell_type":"code","source":"(df[df.primary_label.isin(scored_birds)].groupby(['primary_label']).audio_length.agg(sum) / 60 / 60).sort_values(ascending=False).plot.bar(\n    width=0.9, title=\"Primary labels - scored\", ylabel=\"minutes\"\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:17:59.537400Z","iopub.execute_input":"2022-05-13T17:17:59.537949Z","iopub.status.idle":"2022-05-13T17:17:59.837231Z","shell.execute_reply.started":"2022-05-13T17:17:59.537903Z","shell.execute_reply":"2022-05-13T17:17:59.836468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.yscale('log')\nplt.ylabel('# of samples')\nplt.xlabel('length of sample (minutes)')\nplt.hist([l/60 for l in df.audio_length], bins=np.arange(0, 81, 0.5));","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:18:00.979897Z","iopub.execute_input":"2022-05-13T17:18:00.980625Z","iopub.status.idle":"2022-05-13T17:18:01.966158Z","shell.execute_reply.started":"2022-05-13T17:18:00.980587Z","shell.execute_reply":"2022-05-13T17:18:01.965225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.yscale('log')\nplt.ylabel('# of samples')\nplt.xlabel('length of sample (minutes)')\nplt.hist([l/60 for l in df[df.primary_label.isin(scored_birds)].audio_length], bins=np.arange(0, 81, 0.5));","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:18:01.968203Z","iopub.execute_input":"2022-05-13T17:18:01.968745Z","iopub.status.idle":"2022-05-13T17:18:02.761364Z","shell.execute_reply.started":"2022-05-13T17:18:01.968698Z","shell.execute_reply":"2022-05-13T17:18:02.760249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(df[df.primary_label.isin(scored_birds)].audio_length/60).describe()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:18:02.762607Z","iopub.execute_input":"2022-05-13T17:18:02.762870Z","iopub.status.idle":"2022-05-13T17:18:02.778161Z","shell.execute_reply.started":"2022-05-13T17:18:02.762839Z","shell.execute_reply":"2022-05-13T17:18:02.777240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum(df[df.primary_label.isin(scored_birds)].audio_length/60)","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:18:05.014305Z","iopub.execute_input":"2022-05-13T17:18:05.014648Z","iopub.status.idle":"2022-05-13T17:18:05.025767Z","shell.execute_reply.started":"2022-05-13T17:18:05.014614Z","shell.execute_reply":"2022-05-13T17:18:05.024759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_labels[all_labels.isin(scored_birds)].value_counts().plot.bar(\n    width=0.9, title=\"All labels - scored\", ylabel=\"occurrences\"\n)\nall_labels[all_labels.isin(scored_birds)].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:18:05.155772Z","iopub.execute_input":"2022-05-13T17:18:05.156201Z","iopub.status.idle":"2022-05-13T17:18:05.484219Z","shell.execute_reply.started":"2022-05-13T17:18:05.156171Z","shell.execute_reply":"2022-05-13T17:18:05.483456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calls\n\nCalls appears to be a list of tag-like descriptions of the recording. The top two are \"call\" and \"song\". \n\nBy the way, apparently there's a big difference between a \"birdcall\" and a \"birdsong\", as explained by the [wikipedia page on Bird Vocalization](https://en.wikipedia.org/wiki/Bird_vocalization): \n\n> In ornithology and birding, songs (relatively complex vocalizations) are distinguished by function from calls (relatively simple vocalizations).","metadata":{}},{"cell_type":"code","source":"types = pd.Series(np.concatenate(df.type.apply(lambda v: v[2:-2].split(\"', '\"))))\ntypes.value_counts()[:30].plot.bar(width=0.9)\ntypes.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-04-24T22:19:35.300095Z","iopub.execute_input":"2022-04-24T22:19:35.300428Z","iopub.status.idle":"2022-04-24T22:19:36.005889Z","shell.execute_reply.started":"2022-04-24T22:19:35.300395Z","shell.execute_reply":"2022-04-24T22:19:36.005138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Author\n\nAccording to the competition description, \"Author\" is\n\n> the eBird user who provided the recording\n\nPaul Marvin is the top contributor by far--what a hero! From googling I found a note about him [on Cornel Lab's facebook page](https://www.facebook.com/macaulaylibrary/posts/10156633504705424):\n\n> Paul lives in Cocoa, Florida, where he has taken advantage of the warm weather and abundant birdlife to make numerous recordings this winter in conjunction with complete eBird checklists. When he is not out in the field making new recordings, he has also been creating historical checklists to upload some of the many recordings he has made since he started recording in earnest in 2011.","metadata":{}},{"cell_type":"code","source":"df.author.value_counts()[:30].plot.bar(width=0.9, title=\"Author\")\ndf.author.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:18:51.709540Z","iopub.execute_input":"2022-05-13T17:18:51.709999Z","iopub.status.idle":"2022-05-13T17:18:52.153341Z","shell.execute_reply.started":"2022-05-13T17:18:51.709966Z","shell.execute_reply":"2022-05-13T17:18:52.151980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Rating\n\nAccording to the competition description, \"Rating\" is:\n\n> Float value between 0.0 and 5.0 as an indicator of the quality rating on Xeno-canto and the number of background species, where 5.0 is the highest and 1.0 is the lowest. 0.0 means that this recording has no user rating yet.","metadata":{}},{"cell_type":"code","source":"df.rating.plot.hist(bins=10, title=\"Rating\")\ndf.rating.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:19:16.418766Z","iopub.execute_input":"2022-05-13T17:19:16.419106Z","iopub.status.idle":"2022-05-13T17:19:16.660249Z","shell.execute_reply.started":"2022-05-13T17:19:16.419058Z","shell.execute_reply":"2022-05-13T17:19:16.659141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Time\n\nI'm guessing the \"Time\" column is is the time of day when the recording took place. This column is pretty messy, so I wrote some code to clean it up and convert it to a proper time of day. You can see an interesting peak at 8 AM--turns out, we have a name for that: the [dawn chorus](https://en.wikipedia.org/wiki/Dawn_chorus_(birds))!","metadata":{}},{"cell_type":"code","source":"def clean_time(t):\n    add_12 = False\n    if t.endswith('am'):\n        t = t[:-2]\n    elif t.endswith('pm'):\n        t = t[:-2]\n        add_12 = True\n    parts = t.split(':')\n    if parts[0].isnumeric():\n        parts[0] = str(int(parts[0]) + 12) if add_12 else parts[0]\n    t = ':'.join(parts)\n    if len(parts) == 2:\n        t = t + ':00'\n    return t\n\ncoerce_times = pd.to_timedelta(df.time.apply(clean_time), errors=\"coerce\")\ncoerce_times.astype('timedelta64[h]').plot.hist(bins=24, title=\"Time of day (hour)\")\nplt.ylabel(\"# of recordings\")\nplt.grid()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:23:07.760829Z","iopub.execute_input":"2022-05-13T17:23:07.761320Z","iopub.status.idle":"2022-05-13T17:23:08.170900Z","shell.execute_reply.started":"2022-05-13T17:23:07.761267Z","shell.execute_reply":"2022-05-13T17:23:08.169944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here's all the times that I wasn't able to parse. You can see that some of them are just question marks.","metadata":{}},{"cell_type":"code","source":"# Cases that couldn't be converted\ndf.time[coerce_times.isnull()].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-13T17:19:20.856443Z","iopub.execute_input":"2022-05-13T17:19:20.856922Z","iopub.status.idle":"2022-05-13T17:19:20.868259Z","shell.execute_reply.started":"2022-05-13T17:19:20.856887Z","shell.execute_reply":"2022-05-13T17:19:20.867238Z"},"trusted":true},"execution_count":null,"outputs":[]}]}