{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BirdCLEF 2023 EDA\n\nHello,\n\nthis notebook contains an Exploratory Data Analysis (EDA) for the BirdCLEF 2023 competition.\n\nBeside some standard EDA processes I have also added an interactive widget that allows to listen to and visualize selected audio samples.\n\nThere are a few comments here and there, including in the code. It should all be self-explanatory, in any case feel free to contact me if you have any questions.","metadata":{}},{"cell_type":"code","source":"# Importing libraries\nimport os\nimport pandas as pd\nimport librosa","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:15.917274Z","iopub.execute_input":"2023-03-21T18:26:15.918373Z","iopub.status.idle":"2023-03-21T18:26:15.955845Z","shell.execute_reply.started":"2023-03-21T18:26:15.918332Z","shell.execute_reply":"2023-03-21T18:26:15.954982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting the metadata file\ndf = pd.read_csv('/kaggle/input/birdclef-2023/train_metadata.csv')\n\n# Adding a path column\nroot_dir = '/kaggle/input/birdclef-2023/train_audio/'\n\ndf['path'] = df['filename'].apply(lambda x: os.path.join(root_dir, x))\n\ndf.sample(5)","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:15.957829Z","iopub.execute_input":"2023-03-21T18:26:15.958457Z","iopub.status.idle":"2023-03-21T18:26:16.147668Z","shell.execute_reply.started":"2023-03-21T18:26:15.958423Z","shell.execute_reply":"2023-03-21T18:26:16.146557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the number of rows and columns\nprint(\"The DataFrame has \" + str(df.shape[0]) + \" samples and \" + str(df.shape[1]) + \" columns\")","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:16.149242Z","iopub.execute_input":"2023-03-21T18:26:16.149690Z","iopub.status.idle":"2023-03-21T18:26:16.157348Z","shell.execute_reply.started":"2023-03-21T18:26:16.149626Z","shell.execute_reply":"2023-03-21T18:26:16.156228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check duplicates\nprint(\"Duplicate entries in the dataset: \" + str(df.duplicated().sum()))","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:16.158813Z","iopub.execute_input":"2023-03-21T18:26:16.159113Z","iopub.status.idle":"2023-03-21T18:26:16.212495Z","shell.execute_reply.started":"2023-03-21T18:26:16.159084Z","shell.execute_reply":"2023-03-21T18:26:16.211231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get dataframe info\ndf.info()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:16.216176Z","iopub.execute_input":"2023-03-21T18:26:16.216801Z","iopub.status.idle":"2023-03-21T18:26:16.242622Z","shell.execute_reply.started":"2023-03-21T18:26:16.216765Z","shell.execute_reply":"2023-03-21T18:26:16.241607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\n%config InlineBackend.figure_format = 'retina'\n\n# Sort the values in descending order\nsorted_counts = df['common_name'].value_counts().sort_values(ascending=True)\n\n# Create a horizontal bar chart with sorted values\nplt.figure(figsize=(10, 50))\nplt.barh(sorted_counts.index, sorted_counts.values)\n\nplt.title('Distribution of audio samples per bird', fontweight='bold') # Set title\nplt.xlabel('Number of samples') # Set axis labels X\nplt.ylabel('Bird common name')  # Set axis labels Y\nplt.xticks(rotation=90) # Rotate x-axis ticks by 90 degrees\nplt.margins(y=0)  # Adjust subplot parameters to remove top and bottom padding\n\n# Display the plot\nplt.show()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:16.243815Z","iopub.execute_input":"2023-03-21T18:26:16.244205Z","iopub.status.idle":"2023-03-21T18:26:22.564226Z","shell.execute_reply.started":"2023-03-21T18:26:16.244174Z","shell.execute_reply":"2023-03-21T18:26:22.563122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this chart it's clear that there is a significant imbalance in the number of audio samples for each bird species. While some birds have 500 samples, others have less than 10 or even only 1.\n\nLet's see now the rating distribution:","metadata":{}},{"cell_type":"code","source":"# Check target balance\ndf['rating'].value_counts().sort_index().plot.bar(figsize=(10,5))\nplt.title('Rating distribution of the audio recordings', fontweight=\"bold\")\nplt.xlabel('Target')\nplt.ylabel('Number of samples')\nplt.show()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:22.565489Z","iopub.execute_input":"2023-03-21T18:26:22.566447Z","iopub.status.idle":"2023-03-21T18:26:23.010562Z","shell.execute_reply.started":"2023-03-21T18:26:22.566409Z","shell.execute_reply":"2023-03-21T18:26:23.009414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have votes ranging from 0 to 5, with a peak rating of 4. It is important to keep in mind that audio quality across different file is not the same.","metadata":{}},{"cell_type":"markdown","source":"### Interactive widget\n\nAs we are dealing with almost 17000 audio samples, I thought it would be useful to further explore the audio data with an interactive tool that would display the waveform, spectrogram and audio player to listen to the selected file. For this task I defined a dedicated function and used `ipywidgets`.\n\nNote that the notebook needs to be running to interact properly. Simply follow the filters in the drop-down menu and click the `Display audio info` button. Also, you may need to wait a few seconds for the file to be loaded.","metadata":{}},{"cell_type":"code","source":"from IPython.display import HTML\nfrom IPython.display import display\nimport IPython.display as ipd\nimport numpy as np\n\ndef display_audio_info(file_path):\n    # Load audio file\n    y, sr = librosa.load(file_path, sr=32000)\n\n    # Compute the spectrogram of the audio\n    D = np.abs(librosa.stft(y))**2\n    S = librosa.feature.melspectrogram(S=D)\n\n    # Create subplots\n    fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(20, 6))\n\n    # Plot the waveform\n    librosa.display.waveshow(y, sr=sr, ax=axs[0])\n    axs[0].set_xlabel('Time (s)')\n    axs[0].set_ylabel('Amplitude')\n\n    # Plot the Mel spectrogram\n    librosa.display.specshow(librosa.power_to_db(S), sr=sr, x_axis='time', y_axis='mel', ax=axs[1])\n    axs[1].set_xlabel('Time (s)')\n    axs[1].set_ylabel('Frequency (Hz)')\n\n    # Add a colorbar to the Mel spectrogram plot if it has images\n    try:\n        plt.colorbar(mappable=axs[1].images[0], ax=axs[1])\n    except IndexError:\n        pass\n\n    # Get bird name, rating, and filename from the DataFrame\n    bird_info = df.loc[df['path'] == file_path, ['common_name', 'type',  'rating', 'filename']].iloc[0]\n    bird_name, bird_type, rating, filename = bird_info.values\n\n    # Set the suptitle with the bird name, rating, and filename\n    plt.suptitle(f\"Name:  {bird_name}  |  Type: {bird_type}  |  Rating: {rating}  |  Filename: {filename}\", fontsize=16)\n    \n    # Define the audio player\n    audio_player = ipd.Audio(file_path)\n\n    # Set the audio player width to 100%\n    audio_player_html = audio_player._repr_html_().replace('controls', 'controls style=\"width:100%;\"')\n\n    # Set the layout and display the plot and audio player\n    plt.tight_layout()\n    plt.show()\n    \n    # Display the audio player using HTML\n    display(HTML(audio_player_html))","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:23.011895Z","iopub.execute_input":"2023-03-21T18:26:23.012213Z","iopub.status.idle":"2023-03-21T18:26:23.024403Z","shell.execute_reply.started":"2023-03-21T18:26:23.012183Z","shell.execute_reply":"2023-03-21T18:26:23.023381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import ipywidgets as widgets\n\n# Create a dropdown widget with the list of bird names\nbird_dropdown = widgets.Dropdown(\n    options=['All'] + list(df['common_name'].unique()),\n    description='Select a bird:'\n)\n\n# Create a dropdown widget with the list of rating values\nrating_dropdown = widgets.Dropdown(\n    options=['All'],\n    description='Rating value:'\n)\n\n# Create a dropdown widget with the list of file names\nfile_dropdown = widgets.Dropdown(\n    description='Select a file:'\n)\n\n# Create a button widget\nbutton = widgets.Button(description='Display audio info')\n\n# Create an output widget for displaying the spectrogram\noutput = widgets.Output()\n\n# Define a function to update the file dropdown options based on the selected bird and rating value\ndef update_file_dropdown_options(change):\n    bird_value = bird_dropdown.value\n\n    if bird_value == 'All':\n        rating_values = df['rating'].dropna().unique()\n    else:\n        rating_values = df.loc[df['common_name'] == bird_value, 'rating'].dropna().unique()\n\n    if len(rating_values) == 0:\n        rating_values = ['All']\n\n    rating_values = sorted([rv for rv in rating_values if not pd.isna(rv)])\n    \n    rating_dropdown.options = ['All'] + rating_values\n    \n    rating_value = rating_dropdown.value\n    \n    if rating_value == 'All':\n        if bird_value == 'All':\n            file_dropdown.options = df['filename']\n        else:\n            file_dropdown.options = df.loc[df['common_name'] == bird_value, 'filename']\n    else:\n        if bird_value == 'All':\n            file_dropdown.options = df.loc[df['rating'] == rating_value, 'filename']\n        else:\n            file_dropdown.options = df.loc[(df['common_name'] == bird_value) & (df['rating'] == rating_value), 'filename']\n    \n    if rating_value not in rating_values:\n        rating_dropdown.value = 'All'\n\n# Attach the update_file_dropdown_options function to both the bird and rating dropdown widgets\nbird_dropdown.observe(update_file_dropdown_options, names='value')\nrating_dropdown.observe(update_file_dropdown_options, names='value')\n\n# Define a function to update the output when the button is clicked\ndef on_button_click(b):\n    with output:\n        output.clear_output()  # Clear the previous output\n        try:\n            file_path = df.loc[df['filename'] == file_dropdown.value, 'path'].iloc[0]\n            display_audio_info(file_path)\n        except Exception as e:\n            print('An error occurred:', e)\n\n# Attach the on_button_click function to the button widget\nbutton.on_click(on_button_click)\n\n# Display the widgets\ndisplay(bird_dropdown)\ndisplay(rating_dropdown)\ndisplay(file_dropdown)\ndisplay(button)\ndisplay(output)","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:23.025830Z","iopub.execute_input":"2023-03-21T18:26:23.026274Z","iopub.status.idle":"2023-03-21T18:26:23.150182Z","shell.execute_reply.started":"2023-03-21T18:26:23.026239Z","shell.execute_reply":"2023-03-21T18:26:23.149067Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When playing with this instrument, it can be seen that some samples are quite long, on the order of minutes. You may notice that the audio data are raw, often containing several seconds without bird sounds and background noise.\n\nWhat is useful to get from the spectogram part, is that there are patterns of frequencies with the characteristics of the selected species. Hence, let's compare more spectograms together:","metadata":{}},{"cell_type":"code","source":"import random\n\n# Define the size of the subplot grid\nnum_rows = 3\nnum_cols = 3\n\n# Get a random subset of audio file paths from the DataFrame\naudio_files = random.sample(df['path'].tolist(), num_rows*num_cols)\n\n# Create the figure and subplots\nfig, axs = plt.subplots(num_rows, num_cols, figsize=(20, 10))\n\n# Iterate over the subplots and audio files\nfor i, ax in enumerate(axs.flat):\n    # Load the audio file\n    file_path = audio_files[i]\n    y, sr = librosa.load(file_path, sr=32000)\n\n    # Compute the spectrogram of the audio\n    D = np.abs(librosa.stft(y))**2\n    S = librosa.feature.melspectrogram(S=D)\n\n    # Plot the Mel spectrogram on the current axis\n    librosa.display.specshow(librosa.power_to_db(S), sr=sr, x_axis='time', y_axis='mel', ax=ax)\n    ax.set_xlabel('Time (s)')\n    ax.set_ylabel('Frequency (Hz)')\n\n    # Add a colorbar to the Mel spectrogram plot if it has images\n    try:\n        plt.colorbar(mappable=ax.images[0], ax=ax)\n    except IndexError:\n        pass\n\n    # Set the title of the current axis to the filename\n    filename = df.loc[df['path'] == file_path, 'filename'].iloc[0]\n    ax.set_title(filename)\n\n# Set the layout and display the plot\nplt.suptitle('Random spectograms of birds audio files', fontsize=16)\nplt.tight_layout()\nplt.show()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:23.151859Z","iopub.execute_input":"2023-03-21T18:26:23.152281Z","iopub.status.idle":"2023-03-21T18:26:41.161167Z","shell.execute_reply.started":"2023-03-21T18:26:23.152238Z","shell.execute_reply":"2023-03-21T18:26:41.159506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From these random spectrograms we can see different patterns which identify our bird species.\n\nNext, I've used `folium` to display the localization of the recordings as we have latitude and longitude values for most of the samples:","metadata":{}},{"cell_type":"code","source":"import folium\nfrom folium.plugins import HeatMap\n\n# Create a new map centered on the first latitude and longitude in the DataFrame\nmap = folium.Map(location=[df['latitude'].iloc[0], df['longitude'].iloc[0]], zoom_start=2, min_zoom=2)\n\n# Create a HeatMap layer from the latitude and longitude data in the DataFrame\nheat_data = [[row['latitude'], row['longitude']] for index, row in df.dropna(subset=['latitude', 'longitude']).iterrows()]\nheatmap = folium.plugins.HeatMap(heat_data,\n                                 opacity=0.8,\n                                 radius=6,\n                                 blur=5,\n                                 gradient={0.2: 'blue', 0.4: 'cyan', 0.6: 'lime', 0.8: 'yellow', 1: 'red'})\nheatmap.add_to(map)\n\n# Display the map\nmap","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:41.163706Z","iopub.execute_input":"2023-03-21T18:26:41.164239Z","iopub.status.idle":"2023-03-21T18:26:42.670744Z","shell.execute_reply.started":"2023-03-21T18:26:41.164202Z","shell.execute_reply":"2023-03-21T18:26:42.669516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sounds recordings come from all over the world, with a dominance in the Kenya region and Europe.\n\nLast, let's check the test audio file:","metadata":{}},{"cell_type":"code","source":"# Sligthly different function than before, adapted for the test sample\n\ndef display_test_audio_info(file_path):\n    # Load audio file\n    y, sr = librosa.load(file_path, sr=32000)\n\n    # Compute the spectrogram of the audio\n    D = np.abs(librosa.stft(y)) ** 2\n    S = librosa.feature.melspectrogram(S=D)\n\n    # Create subplots\n    fig, axs = plt.subplots(nrows=2, ncols=1, figsize=(20, 10))\n\n    # Plot the waveform\n    librosa.display.waveshow(y, sr=sr, ax=axs[0])\n    axs[0].set_xlabel('Time (s)')\n    axs[0].set_ylabel('Amplitude')\n\n    # Plot the Mel spectrogram\n    librosa.display.specshow(librosa.power_to_db(S), sr=sr, x_axis='time', y_axis='mel', ax=axs[1])\n    axs[1].set_xlabel('Time (s)')\n    axs[1].set_ylabel('Frequency (Hz)')\n\n    # Add a colorbar to the Mel spectrogram plot if it has images\n    try:\n        plt.colorbar(mappable=axs[1].images[0], ax=axs[1])\n    except IndexError:\n        pass\n\n    # Define the audio player\n    audio_player = ipd.Audio(file_path)\n\n    # Set the width to 100%\n    audio_player_html = audio_player._repr_html_().replace('controls', 'controls style=\"width:100%;\"')\n\n    # Set the layout and display the plot and audio player\n    plt.suptitle('Waveform and spectrogram of the test audio recording', fontsize=16)\n    plt.tight_layout()\n    plt.show()\n\n    # Display the audio player using HTML\n    display(HTML(audio_player_html))","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:42.672519Z","iopub.execute_input":"2023-03-21T18:26:42.673189Z","iopub.status.idle":"2023-03-21T18:26:42.685138Z","shell.execute_reply.started":"2023-03-21T18:26:42.673147Z","shell.execute_reply":"2023-03-21T18:26:42.684308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_audio = \"/kaggle/input/birdclef-2023/test_soundscapes/soundscape_29201.ogg\"\n\ndisplay_test_audio_info(test_audio)","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-03-21T18:26:42.686358Z","iopub.execute_input":"2023-03-21T18:26:42.686859Z","iopub.status.idle":"2023-03-21T18:26:51.014402Z","shell.execute_reply.started":"2023-03-21T18:26:42.686828Z","shell.execute_reply":"2023-03-21T18:26:51.012491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here is the 10 minutes recording to predict. As can be seen, multiple patterns are visible across the audio file. Some background noise is also present, more evidently around 3 and 5/6 minutes.","metadata":{}}]}