{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🐦 Audio 101. 2- Detailed EDA\n\n## [BirdCLEF 2022](https://www.kaggle.com/c/birdclef-2022)\n### Identify bird calls in soundscapes\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/33246/logos/header.png)\n\n\n## Hi and welcome! This is the second kernel of the series `Audio 101`, the documentation of my learning process in the amazing world of audio processing.\n\n**In this short kernel we will perform a detailed EDA of the input data [BirdCLEF 2022](https://www.kaggle.com/c/birdclef-2022/) competition, bringing data from the `ebirds.org` site with `requests` and creating a concise \"bird card\" for each species.**\n\n\nThis series aims to get a good understanding of the specific topic from zero.\n\nThe ideal reader is a Data Scientist noob with some general knowledge about Deep Learning, but no technical expertise in Audio Processing. \n\n---\n\nThe full series consists of the following notebooks:\n1. [🐦 Audio 101. 1-Audio manipulation & musical notes](https://www.kaggle.com/julian3833/audio-101-1-audio-manipulation-musical-notes/)\n2. _[🐦 Audio 101. 2- Detailed EDA](https://www.kaggle.com/julian3833/audio-101-2-detailed-eda/) (This notebook)_\n\n\n\nThis is an ongoing project, so expect more notebooks to be added to the series soon. Actually, we are currently working on the following ones:\n* **Plot Fourier Transforms and Spectrograms**\n* **Build a simple CNN classifier model over image features**\n* **Study the previous competition [BirdCLEF 2021 - Birdcall Identification](https://www.kaggle.com/c/birdclef-2021) and migrate some good models**\n\n---\n\n\n\n#  Please _DO_ upvote if you found this useful or interesting!\n\n\nEnough chitchat, let's code!","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport json\nimport requests\nimport numpy as np\nimport pandas as pd\nimport matplotlib\nimport matplotlib.pyplot as plt\n\nfrom bs4 import BeautifulSoup\n\nimport torchaudio\nfrom IPython.display import Audio, display, HTML, Image","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-02-18T15:25:35.625349Z","iopub.execute_input":"2022-02-18T15:25:35.625839Z","iopub.status.idle":"2022-02-18T15:25:37.104676Z","shell.execute_reply.started":"2022-02-18T15:25:35.625747Z","shell.execute_reply":"2022-02-18T15:25:37.104015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore the input data","metadata":{}},{"cell_type":"code","source":"BASE_PATH = \"../input/birdclef-2022/\"\nos.listdir(BASE_PATH)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.106113Z","iopub.execute_input":"2022-02-18T15:25:37.106705Z","iopub.status.idle":"2022-02-18T15:25:37.117372Z","shell.execute_reply.started":"2022-02-18T15:25:37.106668Z","shell.execute_reply":"2022-02-18T15:25:37.116707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# train_metadata.csv\n\nFrom the [Data tab](https://www.kaggle.com/c/birdclef-2022/data):\n- A wide range of metadata is provided for the training data. The most directly relevant fields are:\n\n* **primary_label** - a code for the bird species. You can review detailed information about the bird codes by appending the code to https://ebird.org/species/, such as https://ebird.org/species/amecro for the American Crow.\n* **secondary_labels**: Background species as annotated by the recordist. An empty list does not mean that no background birds are audible.\n* **author** - the eBird user who provided the recording.\n* **filename**: the associated audio file.\n* **rating**: Float value between 0.0 and 5.0 as an indicator of the quality rating on Xeno-canto and the number of background species, where 5.0 is the highest and 1.0 is the lowest. 0.0 means that this recording has no user rating yet.","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(f\"{BASE_PATH}train_metadata.csv\")\ndf_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.118542Z","iopub.execute_input":"2022-02-18T15:25:37.119048Z","iopub.status.idle":"2022-02-18T15:25:37.246789Z","shell.execute_reply.started":"2022-02-18T15:25:37.119012Z","shell.execute_reply":"2022-02-18T15:25:37.246086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.248760Z","iopub.execute_input":"2022-02-18T15:25:37.249229Z","iopub.status.idle":"2022-02-18T15:25:37.276378Z","shell.execute_reply.started":"2022-02-18T15:25:37.249192Z","shell.execute_reply":"2022-02-18T15:25:37.275517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.277610Z","iopub.execute_input":"2022-02-18T15:25:37.277902Z","iopub.status.idle":"2022-02-18T15:25:37.297307Z","shell.execute_reply.started":"2022-02-18T15:25:37.277863Z","shell.execute_reply":"2022-02-18T15:25:37.296704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's drop a few columns to grasp the core of the problem","metadata":{}},{"cell_type":"code","source":"df_train = df_train[['primary_label', 'time', 'filename']]\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.298550Z","iopub.execute_input":"2022-02-18T15:25:37.298936Z","iopub.status.idle":"2022-02-18T15:25:37.313833Z","shell.execute_reply.started":"2022-02-18T15:25:37.298899Z","shell.execute_reply":"2022-02-18T15:25:37.313058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is time an hour of the day or a pointer to a position in the audio?","metadata":{}},{"cell_type":"code","source":"df_train['time'].value_counts().head(20)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.314961Z","iopub.execute_input":"2022-02-18T15:25:37.315313Z","iopub.status.idle":"2022-02-18T15:25:37.327857Z","shell.execute_reply.started":"2022-02-18T15:25:37.315285Z","shell.execute_reply":"2022-02-18T15:25:37.326881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Each filename is unique\ndf_train['filename'].nunique() == len(df_train)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.328823Z","iopub.execute_input":"2022-02-18T15:25:37.329296Z","iopub.status.idle":"2022-02-18T15:25:37.340585Z","shell.execute_reply.started":"2022-02-18T15:25:37.329267Z","shell.execute_reply":"2022-02-18T15:25:37.339843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is a time of the day. Each `filename` is unique. Let's drop `time` as well:","metadata":{}},{"cell_type":"code","source":"df_train = df_train[['primary_label', 'filename']]\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.341714Z","iopub.execute_input":"2022-02-18T15:25:37.342287Z","iopub.status.idle":"2022-02-18T15:25:37.353670Z","shell.execute_reply.started":"2022-02-18T15:25:37.342256Z","shell.execute_reply":"2022-02-18T15:25:37.353168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Total number of species\n\nThere are 152 different bird species:","metadata":{}},{"cell_type":"code","source":"df_train['primary_label'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.356834Z","iopub.execute_input":"2022-02-18T15:25:37.357179Z","iopub.status.idle":"2022-02-18T15:25:37.363862Z","shell.execute_reply.started":"2022-02-18T15:25:37.357151Z","shell.execute_reply":"2022-02-18T15:25:37.363215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['primary_label'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:37.364908Z","iopub.execute_input":"2022-02-18T15:25:37.365198Z","iopub.status.idle":"2022-02-18T15:25:37.379189Z","shell.execute_reply.started":"2022-02-18T15:25:37.365172Z","shell.execute_reply":"2022-02-18T15:25:37.378567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Audio samples per species","metadata":{}},{"cell_type":"code","source":"audio_samples_per_species = df_train.groupby(\"primary_label\")['filename'].count().sort_values(ascending=False)\naudio_samples_per_species.sort_values(ascending=True).plot.barh(figsize=(25, 40), alpha=0.5, title=\"Samples per species\");","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:28:13.077267Z","iopub.execute_input":"2022-02-18T15:28:13.077607Z","iopub.status.idle":"2022-02-18T15:28:14.752451Z","shell.execute_reply.started":"2022-02-18T15:28:13.077567Z","shell.execute_reply":"2022-02-18T15:28:14.751541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"audio_samples_per_species[:50].plot.bar(figsize=(25, 5), rot=45, alpha=0.5, title=\"Samples per species. 50 species with more samples\")\nplt.show()\naudio_samples_per_species[50:100].plot.bar(figsize=(25, 5), rot=45, alpha=0.5, title=\"Samples per species. Species 50-100\")\nplt.show()\naudio_samples_per_species[100:].plot.bar(figsize=(25, 5), rot=45, alpha=0.5, title=\"Samples per species. Species with less samples\");","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:39.451844Z","iopub.execute_input":"2022-02-18T15:25:39.452025Z","iopub.status.idle":"2022-02-18T15:25:41.459206Z","shell.execute_reply.started":"2022-02-18T15:25:39.452000Z","shell.execute_reply":"2022-02-18T15:25:41.458398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 20 species with less samples\naudio_samples_per_species[-20:]","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.460443Z","iopub.execute_input":"2022-02-18T15:25:41.460896Z","iopub.status.idle":"2022-02-18T15:25:41.467117Z","shell.execute_reply.started":"2022-02-18T15:25:41.460863Z","shell.execute_reply":"2022-02-18T15:25:41.466390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# scored_birds.json\n\nFrom the [Data tab](https://www.kaggle.com/c/birdclef-2022/data): **The subset of the species in the dataset that are scored.**\n\nFrom the [Evaluation tab](https://www.kaggle.com/c/birdclef-2022/overview/evaluation): **Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file.**\n\nIt seems we will only need to detect the bird species present in this file. There are  21.","metadata":{}},{"cell_type":"code","source":"with open(f\"{BASE_PATH}scored_birds.json\") as fp:\n    scored = json.load(fp)\n    \nprint(f\"Scored species: {scored}\")\nprint(f\"Total scored species: {len(scored)}\")","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.468350Z","iopub.execute_input":"2022-02-18T15:25:41.468721Z","iopub.status.idle":"2022-02-18T15:25:41.481604Z","shell.execute_reply.started":"2022-02-18T15:25:41.468692Z","shell.execute_reply":"2022-02-18T15:25:41.481104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scored_species_samples = df_train[df_train['primary_label'].isin(scored)]\\\n                            .groupby('primary_label').count()\\\n                            .rename(columns={'filename': 'total'})\\\n                            .sort_values('total', ascending=False)\nscored_species_samples","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.482423Z","iopub.execute_input":"2022-02-18T15:25:41.482984Z","iopub.status.idle":"2022-02-18T15:25:41.501688Z","shell.execute_reply.started":"2022-02-18T15:25:41.482937Z","shell.execute_reply":"2022-02-18T15:25:41.501187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scored_species_samples.plot.bar(figsize=(25, 5), rot=0, alpha=0.5, title=\"Samples per species for scored species\");","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.503914Z","iopub.execute_input":"2022-02-18T15:25:41.504464Z","iopub.status.idle":"2022-02-18T15:25:41.756824Z","shell.execute_reply.started":"2022-02-18T15:25:41.504425Z","shell.execute_reply":"2022-02-18T15:25:41.756428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are various species for which we have very little data. That will definitely be a challenge!","metadata":{}},{"cell_type":"markdown","source":"# train_audio path\n","metadata":{}},{"cell_type":"code","source":"TRAIN_AUDIO_PATH = f\"{BASE_PATH}train_audio/\"\ntrain_subfolders = os.listdir(TRAIN_AUDIO_PATH)\nprint(\"Train subfolders: \", train_subfolders)\nprint(\"Total subfolders: \", len(train_subfolders))","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.757548Z","iopub.execute_input":"2022-02-18T15:25:41.757806Z","iopub.status.idle":"2022-02-18T15:25:41.785177Z","shell.execute_reply.started":"2022-02-18T15:25:41.757784Z","shell.execute_reply":"2022-02-18T15:25:41.784806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Folders are the same as df_train labels\nset(train_subfolders) == set(df_train['primary_label'].tolist())","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.785992Z","iopub.execute_input":"2022-02-18T15:25:41.786576Z","iopub.status.idle":"2022-02-18T15:25:41.791049Z","shell.execute_reply.started":"2022-02-18T15:25:41.786552Z","shell.execute_reply":"2022-02-18T15:25:41.790400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total_files = sum([len(files) for r, d, files in os.walk(TRAIN_AUDIO_PATH)])\ntotal_files","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:41.791985Z","iopub.execute_input":"2022-02-18T15:25:41.792197Z","iopub.status.idle":"2022-02-18T15:25:45.577425Z","shell.execute_reply.started":"2022-02-18T15:25:41.792169Z","shell.execute_reply":"2022-02-18T15:25:45.576699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The total number of files is the same as the total number of rows in df train\ntotal_files == len(df_train)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:25:45.578421Z","iopub.execute_input":"2022-02-18T15:25:45.578617Z","iopub.status.idle":"2022-02-18T15:25:45.583972Z","shell.execute_reply.started":"2022-02-18T15:25:45.578590Z","shell.execute_reply":"2022-02-18T15:25:45.583225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Play some files\n\nSee [🐦Audio 101- 1) Audio manipulation & musical notes](https://www.kaggle.com/julian3833/audio-101-1-audio-manipulation-musical-notes)","metadata":{}},{"cell_type":"code","source":"def plot_waveform(waveform, sample_rate, color=None):\n    \"\"\"\n    Arguments:\n        waveform: 1D array\n        sample_rate: int\n    \"\"\"\n\n    if color is None:\n        color = np.random.choice(list(matplotlib.colors.TABLEAU_COLORS.keys()))\n    ax = pd.Series(waveform).plot(figsize=(20, 5), alpha=0.6, color=color)\n    duration = len(waveform) / sample_rate\n    \n   \n    ticks = list(range(0, sample_rate*int(duration+1), sample_rate))\n    labels = [label for label, _ in enumerate(ticks, 0)]\n    \n    if duration > 90:\n        ticks = list(range(0, sample_rate*int(duration+1), sample_rate*10))\n        labels = [10*label for label, _ in enumerate(ticks, 0)]\n    \n\n    ax.set_xticks(ticks)\n    ax.set_xticklabels(labels)\n    ax.set_xlabel(\"Time (secs)\")\n    ax.set_ylabel(\"Amplitude (dB)\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:26:57.900668Z","iopub.execute_input":"2022-02-18T15:26:57.901336Z","iopub.status.idle":"2022-02-18T15:26:57.910483Z","shell.execute_reply.started":"2022-02-18T15:26:57.901302Z","shell.execute_reply":"2022-02-18T15:26:57.909657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_file(filename):\n    waveform, sample_rate = torchaudio.load(f\"{TRAIN_AUDIO_PATH}{filename}\")\n    waveform = waveform.numpy()\n    return waveform, sample_rate\n\ndef play_file(filename, max_duration=None, show_waveform=True):\n    waveform, sample_rate = load_file(filename)\n    \n    waveform = waveform[0, : (max_duration * sample_rate) if max_duration is not None else None]\n    \n    display(Audio(waveform, rate=sample_rate))\n    \n    if show_waveform:\n        plot_waveform(waveform, sample_rate)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:26:58.562197Z","iopub.execute_input":"2022-02-18T15:26:58.562449Z","iopub.status.idle":"2022-02-18T15:26:58.568755Z","shell.execute_reply.started":"2022-02-18T15:26:58.562416Z","shell.execute_reply":"2022-02-18T15:26:58.568237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a_file = df_train.iloc[0]['filename']\nplay_file(a_file)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:26:58.921394Z","iopub.execute_input":"2022-02-18T15:26:58.921785Z","iopub.status.idle":"2022-02-18T15:26:59.229129Z","shell.execute_reply.started":"2022-02-18T15:26:58.921745Z","shell.execute_reply":"2022-02-18T15:26:59.228389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def play_species(species, df_train):\n    filename = df_train[df_train['primary_label'] == species].sample(1)['filename'].iloc[0]\n    display(HTML(f\"<h2 style='color:green'>{species.capitalize()}</h2>{filename.split('/')[1]}\"))\n    play_file(filename, max_duration=10)\n\nfor species in scored[:2]:\n    play_species(species, df_train)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:26:59.511025Z","iopub.execute_input":"2022-02-18T15:26:59.511283Z","iopub.status.idle":"2022-02-18T15:27:00.140329Z","shell.execute_reply.started":"2022-02-18T15:26:59.511255Z","shell.execute_reply":"2022-02-18T15:27:00.139682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# A more detailed overview of each species\n\nFrom the [Data tab](https://www.kaggle.com/c/birdclef-2022/data):\n**You can review detailed information about the bird codes by appending the code to https://ebird.org/species/, such as https://ebird.org/species/amecro for the American Crow.**\n","metadata":{}},{"cell_type":"code","source":"DETAIL_BASE_URL = \"https://ebird.org/species/\"\n\n\ndef get_html(species):\n    try:\n        html = requests.get(f\"{DETAIL_BASE_URL}/{species}\").text\n        return html\n    except Exception as e:\n        print(f\"Exception trying to get html for {species}: {e}\")\n        return \"\"\n    \n    \ndef get_img_data(html):\n    try:\n        img_idx = html.find(\"https://cdn.download.ams.birds.cornell.edu/api/v1/\")\n        img_url = html[img_idx:img_idx+len(\"https://cdn.download.ams.birds.cornell.edu/api/v1/asset/\") + 30]\n        img_url = img_url[:img_url.find('\"')]\n        img = requests.get(img_url).content\n        return img\n    except Exception as e:\n        print(f\"Exception trying to get img: {e}\")\n        return \"\"\n    \n    \ndef get_and_display_image(html):\n    display(Image(get_img_data(html), width=400, height=400))\n    \n\ndef get_description(html):\n    try:\n        soup = BeautifulSoup(html, 'html.parser')\n        description = soup.find_all('p', attrs={'class':'u-stack-sm'})[0].text.strip()    \n    except:\n        description = \"\"\n    return description\n\n    \ndef display_sound_types(df_species):\n    df_species['type_list'] = df_species['type'].apply(eval)\n    sounds = df_species.explode('type_list').groupby(\"type_list\")['filename'].apply(list).to_dict()\n    for sound_type, files_list in sounds.items():\n        if sound_type in ['call', 'song']:\n            a_file = files_list[0]\n            waveform, sample_rate = load_file(a_file)\n            duration = int(waveform.shape[1] / sample_rate)\n            display(HTML(f\"<h2> {sound_type}</h2>\"))\n            play_file(a_file, max_duration=None)\n  \n        \ndef show_species_details(species, df_train_full):\n\n    df_species = df_train_full[df_train_full['primary_label'] == species].copy()\n    name = df_species.iloc[0]['common_name']\n    html = get_html(species)\n    description = get_description(html)\n    \n    display(HTML(f\"<h1 style='color:green'> {name} </h1>\"))\n    display(HTML(f\"\"\"<ul> <li>Label: <b>{species}</b></li> \n    <li> Scientific name: <b>{df_species.iloc[0]['scientific_name']}</b></li> \n    <li> Total training samples: <b>{len(df_species)}</b></li> \n    <li> Description: <b>{description}</b></li> \n\n    </ul>\"\"\"))\n    \n    get_and_display_image(html)\n    display_sound_types(df_species)\n        \ndf_train_full = pd.read_csv(f\"{BASE_PATH}train_metadata.csv\")\ndf_train_full.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:27:01.595472Z","iopub.execute_input":"2022-02-18T15:27:01.595724Z","iopub.status.idle":"2022-02-18T15:27:01.660712Z","shell.execute_reply.started":"2022-02-18T15:27:01.595697Z","shell.execute_reply":"2022-02-18T15:27:01.660059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#for species in scored:\n#    show_species_details(species, df_train_full)\n#    print(\"========================\")\n\n\nfor species in ['akiapo', 'apapan', 'maupar', 'crehon']:\n    show_species_details(species, df_train_full)\n    print(\"========================\")","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:27:02.662758Z","iopub.execute_input":"2022-02-18T15:27:02.662970Z","iopub.status.idle":"2022-02-18T15:27:22.603840Z","shell.execute_reply.started":"2022-02-18T15:27:02.662947Z","shell.execute_reply":"2022-02-18T15:27:22.603151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a lot of different sounds and noises in there! In the last audio I think I can even hear flies, frogs, and various types of birds. I cannot distinguish a highlighted song over all that noise to be honest... it seems like a very hard task!","metadata":{}},{"cell_type":"markdown","source":"# test_soundscapes/ \n\nFrom the [Data tab](https://www.kaggle.com/c/birdclef-2022/data):\n**When you submit a notebook, the test_soundscapes directory will be populated with approximately 5,500 recordings to be used for scoring. These are each within a few milliseconds of 1 minute long and in the ogg audio format. Only one soundscape is available for download.**\n\nThe submission files last about 1 minute.\nThere will be about `5500` when submitting, while there is only one when saving. The increase in the prediction runtime depends on the prediction code, so it will not necessarily be a `5500` factor.\n","metadata":{}},{"cell_type":"code","source":"!ls -l {BASE_PATH}test_soundscapes/","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:34:24.558183Z","iopub.execute_input":"2022-02-18T15:34:24.558470Z","iopub.status.idle":"2022-02-18T15:34:24.863870Z","shell.execute_reply.started":"2022-02-18T15:34:24.558446Z","shell.execute_reply":"2022-02-18T15:34:24.862957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"waveform, sample_rate = torchaudio.load(f\"{BASE_PATH}test_soundscapes/soundscape_453028782.ogg\")\nplot_waveform(waveform[0], sample_rate)\nAudio(waveform, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:37:53.503288Z","iopub.execute_input":"2022-02-18T15:37:53.503797Z","iopub.status.idle":"2022-02-18T15:37:54.324201Z","shell.execute_reply.started":"2022-02-18T15:37:53.503769Z","shell.execute_reply":"2022-02-18T15:37:54.322064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I cannot detect a thing. The quality is very bad and the bird songs volumes are very low compared to the noise of the manipulation of the recorder.","metadata":{}},{"cell_type":"markdown","source":"\n# test.csv\n\nFrom the [Data tab](https://www.kaggle.com/c/birdclef-2022/data):\n\n>Metadata for the test set. Only the first three rows are available for download; the full test.csv is provided in the hidden test set.\n> * **row_id** - A unique identifier for the row.\n> * **file_id** - A unique identifier for the audio file.\n> * **bird** - The ebird code for the row. There is one row for each of the scored species per 5 second window per audio file.\n> * **end_time** - The last second of the 5 second time window (5, 10, 15, etc).\n\n","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_csv(f\"{BASE_PATH}test.csv\")\ndf_test","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:31:24.860057Z","iopub.execute_input":"2022-02-18T15:31:24.860627Z","iopub.status.idle":"2022-02-18T15:31:24.873197Z","shell.execute_reply.started":"2022-02-18T15:31:24.860597Z","shell.execute_reply":"2022-02-18T15:31:24.872409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:31:30.238615Z","iopub.execute_input":"2022-02-18T15:31:30.238924Z","iopub.status.idle":"2022-02-18T15:31:30.244915Z","shell.execute_reply.started":"2022-02-18T15:31:30.238896Z","shell.execute_reply":"2022-02-18T15:31:30.244084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# sample_submission.csv\n\nFrom the [Data tab](https://www.kaggle.com/c/birdclef-2022/data):\n\n> sample_submission.csv - A valid sample submission. Only the first three rows are available for download; the full submission.csv is provided in the hidden test set.\n\n> * **row_id** - A unique identifier for the row.\n> * **target** - True/False for whether or not the bird in question called during the 5 second window.","metadata":{}},{"cell_type":"code","source":"df_sub = pd.read_csv(f\"{BASE_PATH}sample_submission.csv\")\ndf_sub","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:47:11.577105Z","iopub.execute_input":"2022-02-18T15:47:11.577720Z","iopub.status.idle":"2022-02-18T15:47:11.588878Z","shell.execute_reply.started":"2022-02-18T15:47:11.577685Z","shell.execute_reply":"2022-02-18T15:47:11.588250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.shape","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:47:15.005477Z","iopub.execute_input":"2022-02-18T15:47:15.005723Z","iopub.status.idle":"2022-02-18T15:47:15.010263Z","shell.execute_reply.started":"2022-02-18T15:47:15.005697Z","shell.execute_reply":"2022-02-18T15:47:15.009874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, in the end, we need to have a function with the following signature:\n\n$(\\text{file} \\times \\text{scored bird} \\times \\text{5 seconds frame}) \\rightarrow \\text{presence / abscence}$\n\nAnd the submission has the tuple $(\\text{file} \\times \\text{scored bird} \\times \\text{5 seconds frame})$ encoded as `row_id`.\n\nThis is a very interesting problem. It is not a classification definitely. It might be considered a multilabel classification.\n\nA few thoughts:\n* The `secondary_labels` column might be very important in this scenario.\n* Mixing-up various audio tracks to make sure various birds are present in a given audio.\n* This fact is bugging me: the training audios don't have a 5-seconds window. The full track is labeled and the presence or absence of the bird doesn't have a 5-second resolution. This is something to address, although I don't know how right now.","metadata":{}},{"cell_type":"code","source":"df_train_full[df_train_full['secondary_labels'] != \"[]\"].head()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:53:40.676343Z","iopub.execute_input":"2022-02-18T15:53:40.676591Z","iopub.status.idle":"2022-02-18T15:53:40.700972Z","shell.execute_reply.started":"2022-02-18T15:53:40.676563Z","shell.execute_reply":"2022-02-18T15:53:40.700231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 10% of the dataset has a non-empty \"secondary label\"\n(df_train_full['secondary_labels'] != \"[]\").mean()","metadata":{"execution":{"iopub.status.busy":"2022-02-18T15:54:13.903485Z","iopub.execute_input":"2022-02-18T15:54:13.903756Z","iopub.status.idle":"2022-02-18T15:54:13.911847Z","shell.execute_reply.started":"2022-02-18T15:54:13.903726Z","shell.execute_reply":"2022-02-18T15:54:13.911179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since this notebook has Internet enabled, we will move to another one to play around with the submissions. We will start with https://www.kaggle.com/stefankahl/how-to-submit-to-birdclef-2022 and move on from there.","metadata":{}},{"cell_type":"markdown","source":"#  Please _DO_ upvote if you found this useful or interesting!","metadata":{}}]}