{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <br>\n    <img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/44224/logos/header.png?\" style=\"display:block; margin:auto; width:95%; height:250px;\">\n</div><br><br> \n\n<div style=\"letter-spacing:normal; opacity:1.;\">\n<!--   https://xkcd.com/color/rgb/   -->\n  <p style=\"text-align:center; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; \n            line-height:1.4; font-size:32px; font-weight:bold; text-transform: uppercase; padding: 9px;\">\n            <strong>BirdCLEF 2023: Identify Bird Calls in Soundscapes</strong></p>  \n  \n  <p style=\"text-align:center; background-color:romance; color: Jaguar; border-radius:10px; font-family:monospace; \n            line-height:1.4; font-size:28px; font-weight:normal; text-transform: capitalize; padding: 10px;\"\n     >Deep Learning Module: Bird Calls Part 1 Exploratory Data Analysis <br>(Convolutional Neural Network (CNN))</p>   \n</div>","metadata":{}},{"cell_type":"markdown","source":"<h4>Dataset Info</h4>\n\n- <h5>Dataset Description</h5>\n\nYour challenge in this competition is to identify which birds are calling in long recordings made in Kenya. This is an important task for scientists who monitor bird populations for conservation purposes. More accurate solutions could enable more comprehensive monitoring. This year, your notebook must also complete inference in a more constrained time frame. This will make it easier to deploy winning models for on the ground conservation efforts where efficiency is at a premium.\n\n`This competition uses a hidden test. When your submitted notebook is scored, the actual test data (including a sample submission) will be made available to your notebook.`\n\n- <h5>Files</h5>\n\n**train_audio/** The training data consists of short recordings of individual bird calls generously uploaded by users of [xenocanto.org](https://xeno-canto.org/). These files have been downsampled to 32 kHz where applicable to match the test set audio and converted to the ogg format. The training data should have nearly all relevant files; we expect there is no benefit to looking for more on [xenocanto.org](https://xeno-canto.org/).\n\n**test_soundscapes/** When you submit a notebook, the **test_soundscapes** directory will be populated with approximately 200 recordings to be used for scoring. They are 10 minutes long and in ogg audio format. The file names are randomized. It should take your submission notebook approximately five minutes to load all of the test soundscapes.\n\n\n**train_metadata.csv** A wide range of metadata is provided for the training data. The most directly relevant fields are:\n\n- **`primary_label`** - a code for the bird species. You can review detailed information about the bird codes by appending the code to https://ebird.org/species/, such as https://ebird.org/species/amecro for the American Crow.\n- **`latitude`** & **`longitude`**: coordinates for where the recording was taken. Some bird species may have local call 'dialects,' so you may want to seek geographic diversity in your training data.\n- **`author`** - The user who provided the recording.\n- **`filename`**: the name of the associated audio file.\n\n\n**`sample_submission.csv`** A valid sample submission.\n\n- **`row_id`**: A slug of [**`soundscape_id`**]_[**`end_time`**] for the prediction.\n- [**`bird_id`**]: There are 264 bird ID columns. You will need to predict the probability of the presence of each bird for each row.\n\n\n**eBird_Taxonomy_v2021.csv** - Data on the relationships between different species.\n\n<h4>TASK</h4>\n\n<h5>The goal of the competition is to use machine learning to identify bird species in Eastern Africa based on their calls. This can help conservationists monitor and protect bird biodiversity more efficiently and over larger areas. The challenge is to develop algorithms that can accurately recognize bird species with limited training data.</h5>\n<br>\n\n<h4>Table of Contents</h4>\n\n1. Import Libraries & Ingest Data\n\n2. EDA Recognizing and Understanding Data\n    - Overview train_metadata.csv file\n        - Check for missing data\n        - Consider how many classes are present in the training set\n        - Object Columns\n            - Consider the column secondary labels\n            - Consider the column type\n            - Consider the column scientific name\n            - Consider the column common name\n            - Consider the columns author\n        - Numeric Columns\n            - Consider the columns latitude & longitude\n            - Consider the column rating\n    - Overview eBird_Taxonomy_v2021.csv file\n        - Check for missing data\n        - Histogram of the taxonomic order counts\n        - Category Taxonomic Order Distribution\n        - Family Taxonomic Order Distribution\n    - Overview directories\n        - Overview train_audio/ directory\n        - Overview test_soundscapes/ directory\n    - Audio Exploration - Most Popular Birds\n        - Black-and-white Mannikin\n        - African Black-headed Oriole\n        - African Bare-eyed Thrush\n        - African Gray Flycatcher\n        - African Goshawk<br><br>\n        \n\n> Competition Dataset: https://www.kaggle.com/competitions/birdclef-2023\n\n> Thanks :  \n    - https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models    \n    - https://www.kaggle.com/code/shreydan/audio-visualization-eda-listen<br> \n    - https://www.kaggle.com/code/leonidkulyk/eda-birdclef-plotly-vis-map-audio<br> \n    - https://www.kaggle.com/code/burhanuddinlatsaheb/eda-visualizations-audio-exploration","metadata":{}},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h1 style=\"text-align:center; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:20px;\n            line-height:1.4; font-size:32px; font-weight:bold; text-transform: uppercase; padding: 9px;\">\n            <strong>1. Import Libraries & Ingest Data</strong></h1>   \n</div>\n\n<h4>pip freeze</h4>","metadata":{}},{"cell_type":"code","source":"%%writefile requirements.txt\n# scikit-plot==0.3.7\n# tensorflow==2.12.0\n# tensorflow-text==2.12.0      # https://github.com/tensorflow/text\n# transformers==4.27.4         # Hugging Face Library \n# yellowbrick==1.5\n# matplotlib==3.5.3            # 3.7.1\n# plotly==5.10.0\nipywidgets==7.7.1\n\nmatplotlib-dashboard==0.0.4\ngsutil==5.23","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:05.248930Z","iopub.execute_input":"2023-05-03T23:36:05.249324Z","iopub.status.idle":"2023-05-03T23:36:05.279885Z","shell.execute_reply.started":"2023-05-03T23:36:05.249288Z","shell.execute_reply":"2023-05-03T23:36:05.278590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, sys, platform\nprint(\"Python  :\", sys.version)\nprint(\"Platform:\", platform.platform())\n\n!{sys.executable} -m pip install -Uq -r requirements.txt","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:07.304505Z","iopub.execute_input":"2023-05-03T23:36:07.305471Z","iopub.status.idle":"2023-05-03T23:36:35.013384Z","shell.execute_reply.started":"2023-05-03T23:36:07.305431Z","shell.execute_reply":"2023-05-03T23:36:35.012163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Print Available Devices - 'TPU', 'GPU', 'CPU'.","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport tensorflow_hub as hub\nimport tensorflow_io as tfio\nprint(\"Tensorflow version \\t\\t:\" + tf.__version__)\n\n# print(\"Available devices:\")\n# for i, device in enumerate(tf.config.list_logical_devices()):\n#     print(\"%d) %s\" % (i, device))","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:35.019039Z","iopub.execute_input":"2023-05-03T23:36:35.019465Z","iopub.status.idle":"2023-05-03T23:36:43.205035Z","shell.execute_reply.started":"2023-05-03T23:36:35.019430Z","shell.execute_reply":"2023-05-03T23:36:43.204061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Importing Related Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\n# import seaborn as sns\n# import scipy.stats as stats\n# !pip install scikit-plot -Uq\nimport scikitplot as skplt\n\n# data visualization\n# import plotly\nimport plotly.offline as po\n# Set notebook mode to work in offline\npo.init_notebook_mode()\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.figure_factory as ff\nfrom plotly.offline import plot\nfrom plotly.subplots import make_subplots\nfrom matplotlib_dashboard import MatplotlibDashboard\nimport missingno as msno\nimport geopandas as gpd\n\nimport io\nimport re\nimport ast\nimport librosa\nimport random\nimport time\nimport gc\n# gc.collect()\n# import tempfile\n\nfrom glob import glob\nfrom tqdm.notebook import tqdm \nfrom collections import Counter\n\nimport torch\nimport torchaudio\nfrom IPython.display import Audio\n# from kaggle_datasets import KaggleDatasets","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:43.206254Z","iopub.execute_input":"2023-05-03T23:36:43.206952Z","iopub.status.idle":"2023-05-03T23:36:49.369388Z","shell.execute_reply.started":"2023-05-03T23:36:43.206919Z","shell.execute_reply":"2023-05-03T23:36:49.368040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Parameters","metadata":{}},{"cell_type":"code","source":"# For old TPU's the dataset needs to be stored in Google Cloud\n# Retrieve the Google Cloud location of the dataset TRY on CPU-GPU\n# new TPU's: get_gcs_path is not required on TPU VMs which can directly use Kaggle datasets\nfrom kaggle_datasets import KaggleDatasets\n# GCS_PATH = KaggleDatasets().get_gcs_path('birdclef-2023')\nGCS_PATH = \"gs://kds-15d8808aa7b690c316f4472b24093b498ca881408e832afd16da59d9\"\n!gsutil ls $GCS_PATH  |  wc -l  | echo -e \"GCS_PATH Count: $(xargs)\\n\"\n\nDATA_DIR = \"/kaggle/input/birdclef-2023\"\n\nprint('GCS_PATH:', GCS_PATH)\nprint('DATA_DIR:', DATA_DIR)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:49.372555Z","iopub.execute_input":"2023-05-03T23:36:49.373690Z","iopub.status.idle":"2023-05-03T23:36:52.468052Z","shell.execute_reply.started":"2023-05-03T23:36:49.373621Z","shell.execute_reply":"2023-05-03T23:36:52.466605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h1 style=\"text-align:center; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:20px;\n            line-height:1.4; font-size:32px; font-weight:bold; text-transform: uppercase; padding: 9px;\">\n            <strong>2. EDA Recognizing and Understanding Data</strong></h1>   \n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h2 style=\"text-align:left; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:15px;\n            line-height:1.4; font-size:28px; font-weight:bold; text-transform: title; padding: 9px;\">\n            <strong>2.1. Overview train_metadata.csv file</strong></h2>   \n</div>\n\nA wide range of metadata is provided for the training data. The most directly relevant fields are:\n\n**primary_label** - a code for the bird species. You can review detailed information about the bird codes by appending the code, such as American Crow.\n\n**latitude**  & **longitude**: coordinates for where the recording was taken. Some bird species may have local call 'dialects,' so you may want to seek geographic diversity in your training data.\n\n**author** - The user who provided the recording.\n\n**filename**: the name of the associated audio file.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(f'{DATA_DIR}/train_metadata.csv')\n\nprint(train_df.shape)\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:52.469603Z","iopub.execute_input":"2023-05-03T23:36:52.469988Z","iopub.status.idle":"2023-05-03T23:36:52.672790Z","shell.execute_reply.started":"2023-05-03T23:36:52.469954Z","shell.execute_reply":"2023-05-03T23:36:52.671729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display (train_df.info(), train_df.describe().T)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:52.674116Z","iopub.execute_input":"2023-05-03T23:36:52.674439Z","iopub.status.idle":"2023-05-03T23:36:52.755924Z","shell.execute_reply.started":"2023-05-03T23:36:52.674411Z","shell.execute_reply":"2023-05-03T23:36:52.754885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check for missing data","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:52.757257Z","iopub.execute_input":"2023-05-03T23:36:52.758108Z","iopub.status.idle":"2023-05-03T23:36:52.794759Z","shell.execute_reply.started":"2023-05-03T23:36:52.758074Z","shell.execute_reply":"2023-05-03T23:36:52.793698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider how many classes are present in the training set","metadata":{}},{"cell_type":"code","source":"primary_label_counts = train_df['primary_label'].value_counts().sort_values()\n\nfig = px.bar(\n    x = primary_label_counts.index,\n    y = primary_label_counts.values, \n    hover_name=primary_label_counts.index,\n    color_discrete_sequence=['cornflowerblue'],\n    orientation='v', width=900, height=600, \n    title=f'Found {len(primary_label_counts)} Class in Primary Label',\n)\nfig.update_layout(xaxis_title=\"Count\", yaxis_title=\"Primary Label\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-04T00:03:07.970205Z","iopub.execute_input":"2023-05-04T00:03:07.970589Z","iopub.status.idle":"2023-05-04T00:03:08.041453Z","shell.execute_reply.started":"2023-05-04T00:03:07.970560Z","shell.execute_reply":"2023-05-04T00:03:08.040387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the column secondary labels","metadata":{}},{"cell_type":"code","source":"# all_secondary_labels1     = sum([eval(x) for x in train_df.secondary_labels], [])\nall_secondary_labels        = np.array(' '.join(train_df['secondary_labels'].astype(str).apply(lambda x: re.sub(r'[\\'\\[\\],]', '', x)).loc[lambda x : x!=''].values).split())\nall_secondary_labels_counts = pd.value_counts(all_secondary_labels).sort_values()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:54.539271Z","iopub.execute_input":"2023-05-03T23:36:54.539706Z","iopub.status.idle":"2023-05-03T23:36:54.580065Z","shell.execute_reply.started":"2023-05-03T23:36:54.539675Z","shell.execute_reply":"2023-05-03T23:36:54.578850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar( \n    x = all_secondary_labels_counts.index,\n    y = all_secondary_labels_counts.values,\n    hover_name=all_secondary_labels_counts.index,\n    color_discrete_sequence=['darkslateblue'],\n    orientation='v', width=900, height=600,\n    title=f'Found {len(all_secondary_labels)} Class in Secondary Label',\n)\nfig.update_layout(xaxis_title=\"Count\", yaxis_title=\"Secondary Label\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-04T00:02:20.573774Z","iopub.execute_input":"2023-05-04T00:02:20.574160Z","iopub.status.idle":"2023-05-04T00:02:20.642440Z","shell.execute_reply.started":"2023-05-04T00:02:20.574132Z","shell.execute_reply":"2023-05-04T00:02:20.641414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the column type","metadata":{}},{"cell_type":"code","source":"types = list(train_df['type'].map(ast.literal_eval))\ntypes = [element for sub_list in types for element in sub_list if element!='']\ntypes = Counter(types)\ntypes['Other'] = sum(types.values()) - sum(dict(types.most_common(10)).values())\ntop_types = types.most_common(11)\n\nfig = px.pie(\n    names  = [x[0] for x in top_types],\n    values = [x[1] for x in top_types],\n    title  = 'Types of Bird Sounds',\n    color_discrete_sequence=px.colors.qualitative.Pastel\n).update_traces(textinfo='label+value+percent')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:54.652808Z","iopub.execute_input":"2023-05-03T23:36:54.653279Z","iopub.status.idle":"2023-05-03T23:36:54.890299Z","shell.execute_reply.started":"2023-05-03T23:36:54.653251Z","shell.execute_reply":"2023-05-03T23:36:54.889516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# type_labels = sum([eval(x) for x in train_df.type], [])\ntype_labels = [label.strip(\"[]'\") for sublist in train_df['type'].apply(ast.literal_eval) for label in sublist]\ntype_labels = list(filter(lambda x: x not in ['$', ''], type_labels))\ntype_counts = pd.value_counts(type_labels).sort_values()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:54.891529Z","iopub.execute_input":"2023-05-03T23:36:54.892012Z","iopub.status.idle":"2023-05-03T23:36:55.068911Z","shell.execute_reply.started":"2023-05-03T23:36:54.891983Z","shell.execute_reply":"2023-05-03T23:36:55.067666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(\n    x = type_counts.index,\n    y = type_counts.values, \n    color_discrete_sequence=['darkgoldenrod'],\n    orientation='v', width=900, height=700,\n    title=f'Found {len(type_labels)} Class in Type Label',\n)\nfig.update_layout(xaxis_title=\"Count\", yaxis_title=\"Audio Type Label\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:55.070242Z","iopub.execute_input":"2023-05-03T23:36:55.070572Z","iopub.status.idle":"2023-05-03T23:36:55.140134Z","shell.execute_reply.started":"2023-05-03T23:36:55.070543Z","shell.execute_reply":"2023-05-03T23:36:55.139280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the column scientific name","metadata":{}},{"cell_type":"code","source":"scientific_name_counts = train_df.scientific_name.value_counts().sort_values()\n\nfig = px.bar(\n    x = scientific_name_counts.index,\n    y = scientific_name_counts.values, \n    color_discrete_sequence=['crimson'],\n    orientation='v', width=900, height=700,\n    title=f'Found {len(type_labels)} Class in Scientific Name Label',\n)\nfig.update_layout(xaxis_title=\"Count\", yaxis_title=\"Scientific Name Label\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:55.141224Z","iopub.execute_input":"2023-05-03T23:36:55.141717Z","iopub.status.idle":"2023-05-03T23:36:55.209305Z","shell.execute_reply.started":"2023-05-03T23:36:55.141681Z","shell.execute_reply":"2023-05-03T23:36:55.208105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the column common name","metadata":{}},{"cell_type":"code","source":"common_names = Counter(train_df['common_name'])\ncommon_names['Other'] = sum(common_names.values()) - sum(dict(common_names.most_common(15)).values())\ntop_common = common_names.most_common(16)\n\nfig = px.pie(\n    names  = [x[0] for x in top_common],\n    values = [x[1] for x in top_common],\n    title  = 'Common Names of top 15 samples'\n).update_traces(textinfo='label')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:55.210988Z","iopub.execute_input":"2023-05-03T23:36:55.211329Z","iopub.status.idle":"2023-05-03T23:36:55.269589Z","shell.execute_reply.started":"2023-05-03T23:36:55.211300Z","shell.execute_reply":"2023-05-03T23:36:55.268462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the columns author","metadata":{}},{"cell_type":"code","source":"# thanks to all the people who spent time collecting, and for sharing the samples\nprint('total authors:',len(set(Counter(train_df['author']).keys())))\nCounter(train_df['author']).most_common(10)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:55.271075Z","iopub.execute_input":"2023-05-03T23:36:55.272146Z","iopub.status.idle":"2023-05-03T23:36:55.286437Z","shell.execute_reply.started":"2023-05-03T23:36:55.272108Z","shell.execute_reply":"2023-05-03T23:36:55.285281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the columns latitude & longitude\n\n- Each record has data about the place of its creation (its latitude and longitude). Let's visualize all this data on a map using folio.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(14,14))\n\ncountries = gpd.read_file(gpd.datasets.get_path(\"naturalearth_lowres\"))\ncountries.plot(color=\"lightgreen\",ax=ax)\ntrain_df.plot(\n    x=\"longitude\", y=\"latitude\", \n    kind=\"scatter\", alpha=0.2, color=\"white\", \n    title=\"Global Distribution of Samples\", ax=ax\n)\nax.set_facecolor('darkblue')\nax.grid(visible=True, alpha=0.5)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:55.288221Z","iopub.execute_input":"2023-05-03T23:36:55.288954Z","iopub.status.idle":"2023-05-03T23:36:56.273102Z","shell.execute_reply.started":"2023-05-03T23:36:55.288912Z","shell.execute_reply":"2023-05-03T23:36:56.271975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.density_mapbox(\n    train_df, lat='latitude', lon='longitude', \n    hover_name=\"primary_label\", hover_data=[\"common_name\", \"scientific_name\"],\n    color_continuous_scale=px.colors.cyclical.IceFire, zoom=1.6, radius=7,\n    mapbox_style=\"stamen-terrain\", width=900, height=700,\n    center=dict(\n        lat=train_df['latitude'].mean(), \n        lon=train_df['longitude'].mean()\n    )\n)\nfig.update_layout(margin={\"r\":0,\"t\":0,\"l\":0,\"b\":0})\nfig.update_layout(mapbox_bounds={\"west\":0, \"east\":0, \"south\":0, \"north\":0})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:56.274440Z","iopub.execute_input":"2023-05-03T23:36:56.274814Z","iopub.status.idle":"2023-05-03T23:36:56.750229Z","shell.execute_reply.started":"2023-05-03T23:36:56.274781Z","shell.execute_reply":"2023-05-03T23:36:56.748933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.scatter_mapbox(\n    train_df, lat=\"latitude\", lon=\"longitude\", color=\"rating\", size=\"rating\",\n    hover_name=\"primary_label\", hover_data=[\"common_name\", \"scientific_name\"],\n    color_continuous_scale=px.colors.cyclical.IceFire, size_max=6, zoom=1.6,\n    mapbox_style=\"open-street-map\", width=900, height=700,\n    center=dict(\n        lat=train_df['latitude'].mean(), \n        lon=train_df['longitude'].mean()\n    )\n)\nfig.update_layout(\n    mapbox_style=\"white-bg\",\n    mapbox_layers=[\n        {\n            \"below\": 'traces',\n            \"sourcetype\": \"raster\",\n            \"sourceattribution\": \"United States Geological Survey\",\n            \"source\": [\n                \"https://basemap.nationalmap.gov/arcgis/rest/services/USGSImageryOnly/MapServer/tile/{z}/{y}/{x}\"\n            ]\n        }\n      ])\nfig.update_layout(margin={\"r\":0,\"t\":0,\"l\":0,\"b\":0})\nfig.update_layout(mapbox_bounds={\"west\":0, \"east\":0, \"south\":0, \"north\":0})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:56.751818Z","iopub.execute_input":"2023-05-03T23:36:56.752200Z","iopub.status.idle":"2023-05-03T23:36:57.572748Z","shell.execute_reply.started":"2023-05-03T23:36:56.752166Z","shell.execute_reply":"2023-05-03T23:36:57.571375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Let's combine latitute and longitude with a class label. When plotting the entire dataframe, the visualization lags a lot, so I sample 10% of the dataframe.\n\n* Each of the labels has its own unique color, and if you want to know the label on the map, you can simply click on the icon you are interested in and annotation will be shown.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_mapbox(\n    train_df, lat=\"latitude\", lon=\"longitude\", color=\"common_name\", size=\"rating\",\n    hover_name=\"filename\", hover_data=[\"common_name\", \"scientific_name\", \"author\", \"rating\"],\n    color_continuous_scale=px.colors.cyclical.IceFire, size_max=6, zoom=1.6,\n    mapbox_style=\"open-street-map\", width=900, height=800,\n    center=dict(\n        lat=train_df['latitude'].mean(), \n        lon=train_df['longitude'].mean()\n    ),\n)\nfig.update_layout(margin={\"r\":0,\"t\":0,\"l\":0,\"b\":0})\nfig.update_layout(mapbox_bounds={\"west\":0, \"east\":0, \"south\":0, \"north\":0})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:57.574379Z","iopub.execute_input":"2023-05-03T23:36:57.574757Z","iopub.status.idle":"2023-05-03T23:36:59.471032Z","shell.execute_reply.started":"2023-05-03T23:36:57.574724Z","shell.execute_reply":"2023-05-03T23:36:59.469451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Consider the column Rating","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"rating\", nbins=len(train_df[\"rating\"].unique()), \n                   color_discrete_sequence=['red'], text_auto=True, width=900, height=500)\nfig.update_layout(title_text=\"Distribution of Ratings\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:59.472490Z","iopub.execute_input":"2023-05-03T23:36:59.472848Z","iopub.status.idle":"2023-05-03T23:36:59.586513Z","shell.execute_reply.started":"2023-05-03T23:36:59.472816Z","shell.execute_reply":"2023-05-03T23:36:59.585406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop columns from correlation matrix\ncorr = train_df.select_dtypes('number').corr()\n\n# create correlation heatmap\nfig = px.imshow(\n    corr, x = corr.columns, y = corr.columns,\n    labels=dict(x=\"Columns\", y=\"Columns\", color=\"Correlation\"),\n    color_continuous_scale='Viridis', zmin=-1, zmax=1,\n    title=\"Correlation Heatmap\"\n)                \n# add text annotations\nannotations = []\nfor i, row in enumerate(corr.values):\n    for j, value in enumerate(row):\n        text = '{:.2f}'.format(value)\n        annotations.append(dict(x=corr.columns[j], y=corr.columns[i], text=text, showarrow=False))\n\nfig.update_layout(width=700, height=700)\nfig.update_traces(showscale=True, colorbar_thickness=25, colorbar_len=0.75)\nfig.update_layout(margin=dict(l=50, r=50, b=100, t=100, pad=4))\nfig.update_layout(annotations=annotations)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:59.588216Z","iopub.execute_input":"2023-05-03T23:36:59.588555Z","iopub.status.idle":"2023-05-03T23:36:59.686686Z","shell.execute_reply.started":"2023-05-03T23:36:59.588527Z","shell.execute_reply":"2023-05-03T23:36:59.685917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h2 style=\"text-align:left; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:15px;\n            line-height:1.4; font-size:28px; font-weight:bold; text-transform: title; padding: 9px;\">\n            <strong>2.2. Overview eBird_Taxonomy_v2021.csv file</strong></h2>   \n</div> \n\n In this .csv file represented the data on the relationships between different species. This data may be used to identify relationships between different species of birds based on their taxonomic classification.\n\nDescription of the columns:\n\n- `TAXON_ORDER`: The taxonomic order of the species.\n\n- `CATEGORY`: The taxonomic category of the species (e.g., species, subspecies, genus, family, etc.)\n\n- `SPECIES_CODE`: A unique code assigned to each species.\n\n- `PRIMARY_COM_NAME`: The common name of the species.\n\n- `SCI_NAME`: The scientific name of the species.\n\n- `ORDER1`: The taxonomic order of the species.\n\n- `FAMILY`: The taxonomic family of the species.\n\n- `SPECIES_GROUP`: The taxonomic group that the species belongs to.\n\n- `REPORT_AS`: A code indicating how the species should be reported.","metadata":{}},{"cell_type":"code","source":"ebt_df = pd.read_csv(f'{DATA_DIR}/eBird_Taxonomy_v2021.csv')\n\nprint(ebt_df.shape)\nebt_df","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:59.688017Z","iopub.execute_input":"2023-05-03T23:36:59.688567Z","iopub.status.idle":"2023-05-03T23:36:59.776784Z","shell.execute_reply.started":"2023-05-03T23:36:59.688539Z","shell.execute_reply":"2023-05-03T23:36:59.776022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display (ebt_df.info(), ebt_df.describe().T)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:59.777926Z","iopub.execute_input":"2023-05-03T23:36:59.778463Z","iopub.status.idle":"2023-05-03T23:36:59.826722Z","shell.execute_reply.started":"2023-05-03T23:36:59.778433Z","shell.execute_reply":"2023-05-03T23:36:59.825747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check for missing data","metadata":{}},{"cell_type":"code","source":"ebt_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:36:59.828393Z","iopub.execute_input":"2023-05-03T23:36:59.828745Z","iopub.status.idle":"2023-05-03T23:36:59.858156Z","shell.execute_reply.started":"2023-05-03T23:36:59.828715Z","shell.execute_reply":"2023-05-03T23:36:59.857000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Histogram of the taxonomic order counts","metadata":{}},{"cell_type":"code","source":"# Histogram of the taxonomic order counts\nfig = px.histogram(\n    ebt_df, x=\"TAXON_ORDER\", color_discrete_sequence=['goldenrod'],\n    nbins=300, width=900, height=600, \n    title= \"Distribution of Taxonomic Orders\",\n)\nfig.update_layout(title={'y':0.93, 'x':0.5, 'xanchor': 'center', 'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-04T00:01:16.698050Z","iopub.execute_input":"2023-05-04T00:01:16.698595Z","iopub.status.idle":"2023-05-04T00:01:16.769511Z","shell.execute_reply.started":"2023-05-04T00:01:16.698558Z","shell.execute_reply":"2023-05-04T00:01:16.768696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Category Taxonomic Order Distribution","metadata":{}},{"cell_type":"code","source":"categories = Counter(ebt_df['CATEGORY'])\nfig = px.pie(\n    ebt_df,\n    names=categories.keys(),\n    values=categories.values(),\n    color_discrete_sequence=px.colors.qualitative.Safe,\n    title='Categories',\n).update_traces(textinfo='label')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:59:55.723905Z","iopub.execute_input":"2023-05-03T23:59:55.724305Z","iopub.status.idle":"2023-05-03T23:59:55.798701Z","shell.execute_reply.started":"2023-05-03T23:59:55.724271Z","shell.execute_reply":"2023-05-03T23:59:55.797691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Box plot of the taxonomic order counts by category\nfig = px.box(\n    ebt_df, x=\"TAXON_ORDER\", y=\"CATEGORY\", \n    color_discrete_sequence=['red'],\n    orientation='h', width=900, height=800,\n    title=\"Taxonomic Order Distribution by Category\",\n)\nfig.update_layout(title={'y':0.93, 'x':0.5, 'xanchor': 'center', 'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-04T00:00:59.899154Z","iopub.execute_input":"2023-05-04T00:00:59.899522Z","iopub.status.idle":"2023-05-04T00:01:00.042164Z","shell.execute_reply.started":"2023-05-04T00:00:59.899495Z","shell.execute_reply":"2023-05-04T00:01:00.041042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Family Taxonomic Order Distribution","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(\n    ebt_df, x=\"FAMILY\",y=\"TAXON_ORDER\",  \n    color_discrete_sequence=['brown'], \n    orientation='v', width=900, height=800,\n    title=\"Taxonomic Order Distribution by Family\",\n)\nfig.update_layout(title={'y':0.93, 'x':0.5, 'xanchor': 'center', 'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-04T00:00:44.537723Z","iopub.execute_input":"2023-05-04T00:00:44.538136Z","iopub.status.idle":"2023-05-04T00:00:44.690078Z","shell.execute_reply.started":"2023-05-04T00:00:44.538101Z","shell.execute_reply":"2023-05-04T00:00:44.689176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h2 style=\"text-align:left; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:15px;\n            line-height:1.4; font-size:28px; font-weight:bold; text-transform: title; padding: 9px;\">\n            <strong>2.3. Overview directories</strong></h2>   \n</div>","metadata":{}},{"cell_type":"code","source":"# TFRecord file paths use DATA_DIR or GCS_PATH\nFILE_PATHS_TRAIN = np.sort(tf.io.gfile.glob(f'{DATA_DIR}/train_audio/*/*'))\nFILE_PATHS_TEST  = np.sort(tf.io.gfile.glob(f'{DATA_DIR}/test_soundscapes'))\n\nFILE_PATHS_TRAIN.shape, FILE_PATHS_TEST.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:37:00.279716Z","iopub.execute_input":"2023-05-03T23:37:00.280704Z","iopub.status.idle":"2023-05-03T23:37:01.533207Z","shell.execute_reply.started":"2023-05-03T23:37:00.280645Z","shell.execute_reply":"2023-05-03T23:37:01.532267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load a sample audio files from two different species\naudio_abe, sr_abe = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/abethr1/XC128013.ogg\")\naudio_abh, sr_abh = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/abhori1/XC127317.ogg\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:37:01.534420Z","iopub.execute_input":"2023-05-03T23:37:01.534738Z","iopub.status.idle":"2023-05-03T23:37:13.284091Z","shell.execute_reply.started":"2023-05-03T23:37:01.534712Z","shell.execute_reply":"2023-05-03T23:37:13.283048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Play the audio\ndisplay(Audio(data=audio_abe, rate=sr_abe), Audio(data=audio_abh, rate=sr_abh))","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:37:13.286934Z","iopub.execute_input":"2023-05-03T23:37:13.288167Z","iopub.status.idle":"2023-05-03T23:37:13.393756Z","shell.execute_reply.started":"2023-05-03T23:37:13.288114Z","shell.execute_reply":"2023-05-03T23:37:13.391243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_audio, sample_rate = librosa.load(np.random.choice(FILE_PATHS_TRAIN))\n\ndisplay(Audio(sample_audio, rate=sample_rate))","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:37:13.396264Z","iopub.execute_input":"2023-05-03T23:37:13.396866Z","iopub.status.idle":"2023-05-03T23:37:13.429088Z","shell.execute_reply.started":"2023-05-03T23:37:13.396832Z","shell.execute_reply":"2023-05-03T23:37:13.428195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_audio","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:37:13.430395Z","iopub.execute_input":"2023-05-03T23:37:13.430947Z","iopub.status.idle":"2023-05-03T23:37:13.438225Z","shell.execute_reply.started":"2023-05-03T23:37:13.430914Z","shell.execute_reply":"2023-05-03T23:37:13.437158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_audio.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:39:17.430294Z","iopub.execute_input":"2023-05-03T23:39:17.430690Z","iopub.status.idle":"2023-05-03T23:39:17.438329Z","shell.execute_reply.started":"2023-05-03T23:39:17.430645Z","shell.execute_reply":"2023-05-03T23:39:17.437102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_rate","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:37:13.439687Z","iopub.execute_input":"2023-05-03T23:37:13.440038Z","iopub.status.idle":"2023-05-03T23:37:13.449886Z","shell.execute_reply.started":"2023-05-03T23:37:13.440010Z","shell.execute_reply":"2023-05-03T23:37:13.448936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h2 style=\"text-align:left; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:15px;\n            line-height:1.4; font-size:28px; font-weight:bold; text-transform: title; padding: 9px;\">\n            <strong>2.4. Audio Exploration - Most Popular Birds</strong></h2>   \n</div>","metadata":{}},{"cell_type":"markdown","source":"**Audio Files**: An audio file is a type of digital file format that stores recorded sound or music. It can be played back through speakers or headphones and is commonly used in a variety of applications, such as music, film, television, radio, and other forms of media. Audio files come in many different formats, including **MP3, WAV, OGG, AAC, and FLAC**.\n\n**How to Visualize Audio Files?**\n\n- There are many ways in which we can view audio in 2D like :\n\n1. **Waveforms**: In audio processing, a waveform is a graphical representation of a sound signal that shows how the signal varies over time. It is a plot of the amplitude of the sound wave on the y-axis versus time on the x-axis. Waveforms can be used to visualize and analyze the properties of audio signals, such as frequency, amplitude, phase, and duration.\n\n2. **Spectograms**: A spectrogram is a visual representation of the spectrum of frequencies of a sound or other signal as it varies with time. It is a way of analyzing the sound signal to understand how the signal changes over time and what frequencies are present in the signal at any given point in time. The x-axis of a spectrogram represents time, while the y-axis represents frequency. The intensity of the color or brightness of each point in the spectrogram represents the amplitude or energy of the frequency at that point in time.\n\n3. **Mel Spectograms**: Mel spectrograms are a type of spectrogram that use the Mel scale to determine the spacing of frequency bins. The Mel scale is a perceptual scale of pitches judged by listeners to be equal in distance from one another. By using the Mel scale, the frequency bins are spaced closer together at lower frequencies and farther apart at higher frequencies, which better matches the way humans perceive sound.\n\n4. **Chromagram**: A chromagram is a visual representation of the pitch content of an audio signal, where each column corresponds to a particular frequency band or pitch class. It is a 2D representation of the distribution of energy in a sound, plotted as a function of time and pitch class.The chromagram is calculated by taking the short-time Fourier transform (STFT) of the audio signal and projecting the resulting spectrogram onto a pitch-class basis, typically using a bank of 12 log-spaced filters to represent the 12 pitch classes of the Western musical scale\n\n5. **MFCC**: MFCC stands for Mel Frequency Cepstral Coefficients, which is a feature extraction technique commonly used in speech processing, music information retrieval, and other related fields.MFCCs are derived from a short-time Fourier transform (STFT) of the audio signal, which involves breaking the signal down into small, overlapping segments and performing a Fourier transform on each segment to obtain its frequency spectrum. The Mel scale is then applied to the frequency spectrum to transform it into a logarithmic scale that approximates the human auditory system's perception of sound.","metadata":{}},{"cell_type":"markdown","source":"<h4>Waveforms</h4>\n\n- **What are audio waveforms?**\n\nAudio waveforms are graphical representations of sound waves. Sound waves are vibrations of air molecules that propagate through the air and are detected by our ears. When we record sound using a microphone, the sound wave is converted into an electrical signal that can be stored and analyzed. The waveform represents how loud or quiet the sound is at any given moment.\n\n- **Waveform plot**\n\nIn a waveform, the x-axis typically represents time, while the y-axis represents the amplitude or strength of the signal at each point in time.\n\n- **What is Passive Acoustic Monitoring?**\n\nPassive acoustic monitoring (PAM) is a technique used to detect and analyze sounds in the environment. It involves the use of specialized equipment to record and analyze acoustic signals, without actively emitting any sound waves.\n\n\n<h4>Spectograms</h4>\n\n- **What are spectograms?**\n\nA spectrogram is a visual representation of the frequency content of a signal over time. It is commonly used in audio signal processing to analyze and display the spectral characteristics of a sound wave.\n\n- **What are spectograms used for?**\n\nSpectrograms are often used to analyze and visualize the frequency content of sounds, such as speech, music, and environmental noise. They can be useful in a wide range of applications, from music production and sound engineering to speech recognition and acoustic analysis.\n\n- **Uses of spectograms?**\n\nSpectrograms can reveal information about the frequency composition of a signal, including the presence of harmonics and overtones, changes in pitch or frequency over time, and the spectral characteristics of noise or other background sounds. By analyzing spectrograms, researchers can gain insights into the underlying structure and patterns of complex signals, and use this information to develop more effective signal processing algorithms and machine learning models.\n\n- **Spectogram plot**\n\nIn a spectrogram, the x-axis represents time, the y-axis represents frequency, and the color or intensity of each point in the plot represents the magnitude or strength of the frequency component at that particular time and frequency. The resulting image is a two-dimensional representation of the three-dimensional frequency-time domain of the signal.","metadata":{}},{"cell_type":"code","source":"def audio_eda(audio_path):   \n    \n    # create subplots and plot data on each subplot\n    fig, axs = plt.subplots(nrows=5, ncols=1, sharex=False, figsize=(12, 4*5))\n    # add labels and title to the figure\n    fig.suptitle(f\"{audio_path.split('/')[-2].upper()}\", y=1, x=.502)\n   \n    # Load an audio file\n    # samples, sample_rate = torchaudio.load(audio_path)\n    samples, sample_rate   = librosa.load(audio_path)    \n    \n    # Visualize the waveform\n    librosa.display.waveshow(samples, sr=sample_rate, ax=axs[0])\n    axs[0].set_title('Waveform')\n    axs[0].set_ylabel('Amplitude')\n    \n    # Compute the spectrogram, then Visualize the spectrogram\n    spectrogram    = librosa.stft(samples)\n    spectrogram_db = librosa.amplitude_to_db(abs(spectrogram))\n    im1 = librosa.display.specshow(spectrogram_db, sr=sample_rate, x_axis='time', y_axis='log', ax=axs[1])\n    axs[1].set_title('Spectrogram (dB)')\n\n    # Compute the mel spectrogram, then Visualize mel spectrogram\n    S   = librosa.feature.melspectrogram(y=samples, sr=sample_rate, fmax=8192)\n    im2 = librosa.display.specshow(librosa.power_to_db(S, ref=np.max), y_axis='mel', fmax=8192, x_axis='time', ax=axs[2])\n    axs[2].set_title('Mel spectrogram')\n\n    # Compute the chromagram, then Visualize the chromagram\n    chromagram = librosa.feature.chroma_stft( y = samples , sr = sample_rate)\n    \n    im3 = librosa.display.specshow(chromagram, sr=sample_rate, x_axis='time', y_axis='chroma', ax=axs[3])\n    axs[3].set_title('Chromagram')\n\n    # Compute the MFCCs, then Visualize the MFCCs\n    mfccs = librosa.feature.mfcc(y=samples, sr=sample_rate, n_mfcc=13)\n    im4   = librosa.display.specshow(mfccs, sr=sample_rate, x_axis='time', ax=axs[4])\n    axs[4].set_title('MFCCs')\n\n    # add color bar with formatted tick labels\n    cb1 = fig.colorbar(im1, ax=axs[1], format='%+2.0f dB', pad=0.01, fraction=0.015)\n    cb2 = fig.colorbar(im2, ax=axs[2], format='%+2.0f dB', pad=0.01, fraction=0.015)\n    cb3 = fig.colorbar(im3, ax=axs[3], pad=0.01, fraction=0.015)\n    cb4 = fig.colorbar(im4, ax=axs[4], pad=0.01, fraction=0.015)   \n    \n    # Show the plots\n    display(Audio(samples, rate=sample_rate))\n    fig.tight_layout( pad=1, h_pad=3)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:39:46.754808Z","iopub.execute_input":"2023-05-03T23:39:46.755269Z","iopub.status.idle":"2023-05-03T23:39:46.769086Z","shell.execute_reply.started":"2023-05-03T23:39:46.755235Z","shell.execute_reply":"2023-05-03T23:39:46.767571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <h3 align='center'>Black-and-white Mannikin</h3>\n    <br>\n    <img src=\"https://cdn.download.ams.birds.cornell.edu/api/v1/asset/212028731/1800\" style=\"display:block; margin:auto; width:55%; height:450px;\">\n</div>","metadata":{}},{"cell_type":"code","source":"audio_eda(\"/kaggle/input/birdclef-2023/train_audio/bawman1/XC115075.ogg\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:41:26.551905Z","iopub.execute_input":"2023-05-03T23:41:26.552296Z","iopub.status.idle":"2023-05-03T23:41:31.301288Z","shell.execute_reply.started":"2023-05-03T23:41:26.552270Z","shell.execute_reply":"2023-05-03T23:41:31.300170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <h3 align='center'>African Black-headed Oriole</h3>\n    <br>\n    <img src=\"https://cdn.download.ams.birds.cornell.edu/api/v1/asset/183432491/1800\" style=\"display:block; margin:auto; width:55%; height:450px;\">\n</div>","metadata":{}},{"cell_type":"code","source":"audio_eda(\"/kaggle/input/birdclef-2023/train_audio/abhori1/XC120251.ogg\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:41:59.673853Z","iopub.execute_input":"2023-05-03T23:41:59.674930Z","iopub.status.idle":"2023-05-03T23:42:05.205861Z","shell.execute_reply.started":"2023-05-03T23:41:59.674893Z","shell.execute_reply":"2023-05-03T23:42:05.204751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <h3 align='center'>African Bare-eyed Thrush</h3>\n    <br>\n    <img src=\"https://cdn.download.ams.birds.cornell.edu/api/v1/asset/247097241/1800\" style=\"display:block; margin:auto; width:55%; height:450px;\">\n</div>","metadata":{}},{"cell_type":"code","source":"audio_eda(\"/kaggle/input/birdclef-2023/train_audio/abethr1/XC128013.ogg\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:42:43.567080Z","iopub.execute_input":"2023-05-03T23:42:43.567485Z","iopub.status.idle":"2023-05-03T23:42:48.013897Z","shell.execute_reply.started":"2023-05-03T23:42:43.567453Z","shell.execute_reply":"2023-05-03T23:42:48.012627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <h3 align='center'>African Gray Flycatcher</h3>\n    <br>\n    <img src=\"https://cdn.download.ams.birds.cornell.edu/api/v1/asset/247103901/1800\" style=\"display:block; margin:auto; width:55%; height:450px;\">\n</div>","metadata":{}},{"cell_type":"code","source":"audio_eda(\"/kaggle/input/birdclef-2023/train_audio/afgfly1/XC134487.ogg\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:43:30.236274Z","iopub.execute_input":"2023-05-03T23:43:30.236685Z","iopub.status.idle":"2023-05-03T23:43:33.281060Z","shell.execute_reply.started":"2023-05-03T23:43:30.236639Z","shell.execute_reply":"2023-05-03T23:43:33.279807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <h3 align='center'>African Goshawk</h3>\n    <br>\n    <img src=\"https://upload.wikimedia.org/wikipedia/commons/c/c7/African_Goshawk_RWD2.jpg\" style=\"display:block; margin:auto; width:55%; height:450px;\">\n</div>","metadata":{}},{"cell_type":"code","source":"audio_eda(\"/kaggle/input/birdclef-2023/train_audio/afrgos1/XC115984.ogg\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T23:44:08.992000Z","iopub.execute_input":"2023-05-03T23:44:08.992404Z","iopub.status.idle":"2023-05-03T23:44:12.923407Z","shell.execute_reply.started":"2023-05-03T23:44:08.992375Z","shell.execute_reply":"2023-05-03T23:44:12.922259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# End of the Project","metadata":{}}]}