{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <center style=\"font-family: consolas; font-size: 32px; font-weight: bold;\">BirdCLEF Data and problem investigation</center>\n<p><center style=\"color:#949494; font-family: consolas; font-size: 20px;\">BirdCLEF 2023 - Identify bird calls in soundscapes</center></p>\n\n***","metadata":{}},{"cell_type":"markdown","source":"\n\n<p style=\"font-family: consolas; font-size: 16px;\">⚪ The goal of this competition is to use machine learning to <b>identify Eastern African bird species by sound</b>.</p>\n\n<p style=\"font-family: consolas; font-size: 16px;\">⚪ The purpose of this is to <b>provide a more cost-effective and logistically feasible method</b> of conducting bird biodiversity surveys, which can be challenging and expensive when done through traditional observer-based methods.</p>\n\n<p style=\"font-family: consolas; font-size: 16px;\">⚪ By using passive acoustic monitoring (PAM) combined with new analytical tools based on machine learning, conservationists can sample much larger spatial scales with higher temporal resolution, allowing for a more comprehensive exploration of the relationship between restoration interventions and biodiversity.</p>\n\n<p style=\"font-family: consolas; font-size: 16px;\">⚪ The best entries in the competition will be able to develop reliable classifiers with limited training data, which will help advance ongoing efforts to protect avian biodiversity in Africa, including those led by the Kenyan conservation organization NATURAL STATE.</p>","metadata":{}},{"cell_type":"markdown","source":"\n\n<div style=\"background-color: rgba(60, 121, 245, 0.03); padding:30px; font-size:15px; font-family: consolas;\">\n\n* [0. Import all dependencies](#0)\n* [1. Overview directories](#1)\n    * [1.1 Overview train_audio/ directory](#1.1)\n    * [1.2 Overview test_soundscapes/ directory](#1.2)\n* [2. Overview train_metadata.csv file](#2)\n    * [2.1 Check for missing data](#2.1)\n    * [2.2 Consider how many classes are present in the training set](#2.2)\n    * [2.3 Consider the column secondary labels](#2.3)\n    * [2.4 Consider the column type](#2.4)\n    * [2.5 Consider the column scientific name](#2.5)\n    * [2.6 Consider the columns latitude & longitude](#2.6)\n* [3. Overview eBird_Taxonomy_v2021.csv file](#3)\n    * [3.1 Check for missing data](#3.1)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"0\"></a>\n# <div style=\"box-shadow: rgba(0, 0, 0, 0.16) 0px 1px 4px inset, rgb(51, 51, 51) 0px 0px 0px 3px inset; padding:20px; font-size:32px; font-family: consolas; text-align:center; display:fill; border-radius:15px;  color:rgb(34, 34, 34);\"> <b> 0. Import all dependencies </b></div>","metadata":{}},{"cell_type":"code","source":"import os\nimport random\nimport cv2\nimport folium\nimport pandas as pd\nimport numpy as np\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nfrom folium.plugins import HeatMap\nfrom folium.features import DivIcon\nfrom IPython.display import Audio, display","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:17.298176Z","iopub.execute_input":"2023-03-10T17:30:17.298616Z","iopub.status.idle":"2023-03-10T17:30:22.221484Z","shell.execute_reply.started":"2023-03-10T17:30:17.298577Z","shell.execute_reply":"2023-03-10T17:30:22.219423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class color:\n   PURPLE = '\\033[95m'\n   CYAN = '\\033[96m'\n   DARKCYAN = '\\033[36m'\n   BLUE = '\\033[94m'\n   GREEN = '\\033[92m'\n   YELLOW = '\\033[93m'\n   RED = '\\033[91m'\n   BOLD = '\\033[1m'\n   UNDERLINE = '\\033[4m'\n   END = '\\033[0m'","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:22.224409Z","iopub.execute_input":"2023-03-10T17:30:22.226246Z","iopub.status.idle":"2023-03-10T17:30:22.234049Z","shell.execute_reply.started":"2023-03-10T17:30:22.226174Z","shell.execute_reply":"2023-03-10T17:30:22.232476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_audio(\n    dir_path: str, label: str, example: str\n) -> None:\n    \n    if label == \"\":\n        filename = f\"{dir_path}/{example}.ogg\"\n        label = \"None\"\n    else:\n        filename = f\"{dir_path}/{label}/{example}.ogg\"\n    \n    print(f\"\\nLabel - {color.BOLD}{color.PURPLE}{label}{color.END}, example - {example}:\")\n    display(Audio(filename=filename))","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:22.236043Z","iopub.execute_input":"2023-03-10T17:30:22.236602Z","iopub.status.idle":"2023-03-10T17:30:22.274170Z","shell.execute_reply.started":"2023-03-10T17:30:22.236547Z","shell.execute_reply":"2023-03-10T17:30:22.271082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# <div style=\"box-shadow: rgba(0, 0, 0, 0.16) 0px 1px 4px inset, rgb(51, 51, 51) 0px 0px 0px 3px inset; padding:20px; font-size:32px; font-family: consolas; text-align:center; display:fill; border-radius:15px;  color:rgb(34, 34, 34);\"> <b> 1. Overview directories </b></div>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.1\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 1.1 Overview <i>train_audio</i> directory</b></div>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ The provided training data for this competition includes brief recordings of separate bird calls that have been contributed by users of <a href=\"https://xeno-canto.org/\"><strong>xenocanto.org</strong></a>. To ensure compatibility with the test set audio, these files have been converted to the ogg format and downsampled to 32 kHz where appropriate. It is expected that the training data comprises almost all of the pertinent files, and it is not necessary to search for additional ones on <a href=\"https://xeno-canto.org/\"><strong>xenocanto.org</strong></a>:</p>\n\n<p style=\"text-align:center;\"><img src=\"https://user-images.githubusercontent.com/45982614/223520408-82b31ee8-3733-4ed6-b62d-46a88b9def3b.png\" width=\"90%\" height=\"90%\"></p>\n\n","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Number of entries in the directory:</p>","metadata":{}},{"cell_type":"code","source":"len(os.listdir(\"/kaggle/input/birdclef-2023/train_audio\"))","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:22.277855Z","iopub.execute_input":"2023-03-10T17:30:22.279160Z","iopub.status.idle":"2023-03-10T17:30:22.330702Z","shell.execute_reply.started":"2023-03-10T17:30:22.279088Z","shell.execute_reply":"2023-03-10T17:30:22.328053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Let's listen to a few samples:</p>","metadata":{}},{"cell_type":"code","source":"display_audio(\"/kaggle/input/birdclef-2023/train_audio\", \"abethr1\", \"XC128013\")\ndisplay_audio(\"/kaggle/input/birdclef-2023/train_audio\", \"abhori1\", \"XC120250\")\ndisplay_audio(\"/kaggle/input/birdclef-2023/train_audio\", \"abythr1\", \"XC115981\")\ndisplay_audio(\"/kaggle/input/birdclef-2023/train_audio\", \"afbfly1\", \"XC200995\")\ndisplay_audio(\"/kaggle/input/birdclef-2023/train_audio\", \"afdfly1\", \"XC115969\")","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:22.332806Z","iopub.execute_input":"2023-03-10T17:30:22.333352Z","iopub.status.idle":"2023-03-10T17:30:22.546629Z","shell.execute_reply.started":"2023-03-10T17:30:22.333312Z","shell.execute_reply":"2023-03-10T17:30:22.545467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.2\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 1.2 Overview <i>test_soundscapes/</i> directory</b></div>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ When you submit a notebook, the test_soundscapes directory will be populated with approximately 200 recordings to be used for scoring. These recordings are 10 minutes in duration and are saved in the ogg audio format, with their file names randomized. Your submission notebook should take approximately five minutes to load all of the test soundscapes.</p>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ This directory has only one audiofile as an example.</p>","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/birdclef-2023/test_soundscapes","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:22.548286Z","iopub.execute_input":"2023-03-10T17:30:22.549086Z","iopub.status.idle":"2023-03-10T17:30:23.723145Z","shell.execute_reply.started":"2023-03-10T17:30:22.549039Z","shell.execute_reply":"2023-03-10T17:30:23.721492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -lh /kaggle/input/birdclef-2023/test_soundscapes/soundscape_29201.ogg","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:23.725892Z","iopub.execute_input":"2023-03-10T17:30:23.726448Z","iopub.status.idle":"2023-03-10T17:30:24.902238Z","shell.execute_reply.started":"2023-03-10T17:30:23.726393Z","shell.execute_reply":"2023-03-10T17:30:24.900523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Let's listen to it:</p>","metadata":{}},{"cell_type":"code","source":"display_audio(\"/kaggle/input/birdclef-2023/test_soundscapes\", \"\", \"soundscape_29201\")","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:24.904822Z","iopub.execute_input":"2023-03-10T17:30:24.905293Z","iopub.status.idle":"2023-03-10T17:30:25.158612Z","shell.execute_reply.started":"2023-03-10T17:30:24.905239Z","shell.execute_reply":"2023-03-10T17:30:25.157494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# <div style=\"box-shadow: rgba(0, 0, 0, 0.16) 0px 1px 4px inset, rgb(51, 51, 51) 0px 0px 0px 3px inset; padding:20px; font-size:32px; font-family: consolas; text-align:center; display:fill; border-radius:15px;  color:rgb(34, 34, 34);\"> <b> 2. Overview <i>train_metadata.csv</i> file</b></div>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">A wide range of metadata is provided for the training data. The most directly relevant fields are:</p>\n\n* <p style=\"font-family: consolas; font-size: 16px;\"> <b><i><code>primary_label</code></i></b> - a code for the bird species. You can review detailed information about the bird codes by appending the <a href=\"https://ebird.org/species/\"><strong>code</strong></a>, such as <a href=\"https://ebird.org/species/amecro\"><strong>American Crow</strong></a>.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <b><i><code>latitude </code></i></b> & <b><i><code>longitude</code></i></b>: coordinates for where the recording was taken. Some bird species may have local call 'dialects,' so you may want to seek geographic diversity in your training data.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <b><i><code>author</code></i></b> - The user who provided the recording.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <b><i><code>filename</code></i></b>: the name of the associated audio file.</p>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Read .csv file.</p>","metadata":{}},{"cell_type":"code","source":"train_metadata_df = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.160890Z","iopub.execute_input":"2023-03-10T17:30:25.161554Z","iopub.status.idle":"2023-03-10T17:30:25.348155Z","shell.execute_reply.started":"2023-03-10T17:30:25.161510Z","shell.execute_reply":"2023-03-10T17:30:25.346525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.350268Z","iopub.execute_input":"2023-03-10T17:30:25.350716Z","iopub.status.idle":"2023-03-10T17:30:25.401600Z","shell.execute_reply.started":"2023-03-10T17:30:25.350674Z","shell.execute_reply":"2023-03-10T17:30:25.400171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Examples count:\", len(train_metadata_df))","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.407669Z","iopub.execute_input":"2023-03-10T17:30:25.408382Z","iopub.status.idle":"2023-03-10T17:30:25.416155Z","shell.execute_reply.started":"2023-03-10T17:30:25.408309Z","shell.execute_reply":"2023-03-10T17:30:25.414488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata_df.describe()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.418197Z","iopub.execute_input":"2023-03-10T17:30:25.418715Z","iopub.status.idle":"2023-03-10T17:30:25.467998Z","shell.execute_reply.started":"2023-03-10T17:30:25.418663Z","shell.execute_reply":"2023-03-10T17:30:25.466848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.1\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 2.1 Check for missing data</b></div>","metadata":{}},{"cell_type":"code","source":"train_metadata_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.470811Z","iopub.execute_input":"2023-03-10T17:30:25.471358Z","iopub.status.idle":"2023-03-10T17:30:25.493619Z","shell.execute_reply.started":"2023-03-10T17:30:25.471302Z","shell.execute_reply":"2023-03-10T17:30:25.491630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.2\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 2.2 Consider how many classes are present in the training set</b></div>","metadata":{}},{"cell_type":"code","source":"primary_label_counts = train_metadata_df.primary_label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.496364Z","iopub.execute_input":"2023-03-10T17:30:25.496996Z","iopub.status.idle":"2023-03-10T17:30:25.508247Z","shell.execute_reply.started":"2023-03-10T17:30:25.496940Z","shell.execute_reply":"2023-03-10T17:30:25.506456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Primary labels count:\", len(primary_label_counts.index))","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.511964Z","iopub.execute_input":"2023-03-10T17:30:25.513018Z","iopub.status.idle":"2023-03-10T17:30:25.527555Z","shell.execute_reply.started":"2023-03-10T17:30:25.512954Z","shell.execute_reply":"2023-03-10T17:30:25.525627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Let's build a bar plot to see the ratio of the number of instances for each of the classes. Since the number of labels exceeds the plot limit, an interactive graph was built, with which you can fully examine the distribution.</p>","metadata":{}},{"cell_type":"code","source":"primary_label_counts = train_metadata_df.primary_label.value_counts()\n\nfig = px.bar(x=primary_label_counts.index, y=primary_label_counts.values)\nfig.update_layout(xaxis_title=\"Label\", yaxis_title=\"Count\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:25.529557Z","iopub.execute_input":"2023-03-10T17:30:25.530489Z","iopub.status.idle":"2023-03-10T17:30:28.141552Z","shell.execute_reply.started":"2023-03-10T17:30:25.530442Z","shell.execute_reply":"2023-03-10T17:30:28.139927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.3\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 2.3 Consider the column <code><i>secondary labels</i></code></b></div>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\"> ⚪ Consider how common the secondary column is.</p>","metadata":{"execution":{"iopub.status.busy":"2023-03-07T19:37:28.447183Z","iopub.execute_input":"2023-03-07T19:37:28.448055Z","iopub.status.idle":"2023-03-07T19:37:28.457364Z","shell.execute_reply.started":"2023-03-07T19:37:28.447992Z","shell.execute_reply":"2023-03-07T19:37:28.455326Z"}}},{"cell_type":"code","source":"print(\"All secondary column occurrences:\", sum(train_metadata_df.secondary_labels != \"[]\"))","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:28.143586Z","iopub.execute_input":"2023-03-10T17:30:28.144108Z","iopub.status.idle":"2023-03-10T17:30:28.158703Z","shell.execute_reply.started":"2023-03-10T17:30:28.144053Z","shell.execute_reply":"2023-03-10T17:30:28.154927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Let's combine all secondary labels in one array and plot their distribution. Since the secondary labels are a list that is represented as a string, we can convert the string back to a list using the <b>eval</b> method.</p>","metadata":{}},{"cell_type":"code","source":"all_secondary_labels = sum([eval(x) for x in train_metadata_df.secondary_labels], [])\nall_secondary_labels_counts = pd.value_counts(all_secondary_labels)","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:28.160590Z","iopub.execute_input":"2023-03-10T17:30:28.161130Z","iopub.status.idle":"2023-03-10T17:30:28.441883Z","shell.execute_reply.started":"2023-03-10T17:30:28.161074Z","shell.execute_reply":"2023-03-10T17:30:28.440585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(x=all_secondary_labels_counts.index, y=all_secondary_labels_counts.values)\nfig.update_layout(xaxis_title=\"Secondary Label\", yaxis_title=\"Count\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:28.443851Z","iopub.execute_input":"2023-03-10T17:30:28.444431Z","iopub.status.idle":"2023-03-10T17:30:28.515770Z","shell.execute_reply.started":"2023-03-10T17:30:28.444338Z","shell.execute_reply":"2023-03-10T17:30:28.514173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.4\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 2.4 Consider the column <code><i>type</i></code></b></div>","metadata":{}},{"cell_type":"code","source":"print(\"Type column occurrences:\", sum(train_metadata_df.type != \"[]\"))","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:28.517764Z","iopub.execute_input":"2023-03-10T17:30:28.518434Z","iopub.status.idle":"2023-03-10T17:30:28.529791Z","shell.execute_reply.started":"2023-03-10T17:30:28.518393Z","shell.execute_reply":"2023-03-10T17:30:28.528239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_labels = sum([eval(x) for x in train_metadata_df.type], [])\ntype_counts = pd.value_counts(type_labels)","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:28.531449Z","iopub.execute_input":"2023-03-10T17:30:28.531840Z","iopub.status.idle":"2023-03-10T17:30:30.071570Z","shell.execute_reply.started":"2023-03-10T17:30:28.531803Z","shell.execute_reply":"2023-03-10T17:30:30.070018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(x=type_counts.index, y=type_counts.values)\nfig.update_layout(xaxis_title=\"Audio type\", yaxis_title=\"Count\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:30.073261Z","iopub.execute_input":"2023-03-10T17:30:30.073643Z","iopub.status.idle":"2023-03-10T17:30:30.141613Z","shell.execute_reply.started":"2023-03-10T17:30:30.073608Z","shell.execute_reply":"2023-03-10T17:30:30.140195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.5\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 2.5 Consider the column <code><i>scientific name</i></code></b></div>","metadata":{}},{"cell_type":"code","source":"scientific_name_counts = train_metadata_df.scientific_name.value_counts()\n\nfig = px.bar(x=scientific_name_counts.index, y=scientific_name_counts.values)\nfig.update_layout(xaxis_title=\"Scientific name\", yaxis_title=\"Count\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:30.143080Z","iopub.execute_input":"2023-03-10T17:30:30.143518Z","iopub.status.idle":"2023-03-10T17:30:30.214804Z","shell.execute_reply.started":"2023-03-10T17:30:30.143481Z","shell.execute_reply":"2023-03-10T17:30:30.213473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.6\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 2.6 Consider the columns <code><i>latitude</i></code> & <code><i>longitude</i></code></b></div>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">🔴 Each record has data about the place of its creation (its latitude and longitude). Let's visualize all this data on a map using <b>folio</b>.</p>","metadata":{}},{"cell_type":"code","source":"# As we considired before we have some NaN values in the data, let's drop it\nfiltered_train_metadata_df = train_metadata_df.dropna()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:30.216535Z","iopub.execute_input":"2023-03-10T17:30:30.217750Z","iopub.status.idle":"2023-03-10T17:30:30.238021Z","shell.execute_reply.started":"2023-03-10T17:30:30.217683Z","shell.execute_reply":"2023-03-10T17:30:30.236671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.density_mapbox(\n    filtered_train_metadata_df, \n    lat='latitude', lon='longitude', \n    radius=7, zoom=2,\n    mapbox_style=\"stamen-terrain\", \n    center=dict(\n        lat=filtered_train_metadata_df['latitude'].mean(), \n        lon=filtered_train_metadata_df['longitude'].mean()\n    )\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:30.239687Z","iopub.execute_input":"2023-03-10T17:30:30.240169Z","iopub.status.idle":"2023-03-10T17:30:30.377499Z","shell.execute_reply.started":"2023-03-10T17:30:30.240116Z","shell.execute_reply":"2023-03-10T17:30:30.376478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">🔴 Let's combine latitute and longitude with a class label. When plotting the entire dataframe, the visualization lags a lot, so I sample 10% of the dataframe.</p>\n\n<p style=\"font-family: consolas; font-size: 16px;\">⚪ Each of the labels has its own unique color, and if you want to know the label on the map, you can simply click on the icon you are interested in and annotation will be shown.</p>","metadata":{}},{"cell_type":"code","source":"sample_10 = len(filtered_train_metadata_df) // 10\n\nfig = px.scatter_mapbox(\n    filtered_train_metadata_df.sample(sample_10), \n    lat='latitude', lon='longitude',\n    color='primary_label', zoom=2,\n    mapbox_style=\"stamen-terrain\", \n    center=dict(\n        lat=filtered_train_metadata_df['latitude'].mean(), \n        lon=filtered_train_metadata_df['longitude'].mean()\n    )\n)\n\nfig.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-10T17:30:30.378970Z","iopub.execute_input":"2023-03-10T17:30:30.379839Z","iopub.status.idle":"2023-03-10T17:30:31.063641Z","shell.execute_reply.started":"2023-03-10T17:30:30.379799Z","shell.execute_reply":"2023-03-10T17:30:31.062247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# <div style=\"box-shadow: rgba(0, 0, 0, 0.16) 0px 1px 4px inset, rgb(51, 51, 51) 0px 0px 0px 3px inset; padding:20px; font-size:32px; font-family: consolas; text-align:center; display:fill; border-radius:15px;  color:rgb(34, 34, 34);\"> <b> 3. Overview <i>eBird_Taxonomy_v2021.csv</i> file</b></div>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">🔴 In this .csv file represented the data on the relationships between different species. This data may be used to identify relationships between different species of birds based on their taxonomic classification.</p>\n\n<p style=\"font-family: consolas; font-size: 16px;\"> Description of the columns:</p>\n\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>TAXON_ORDER</code>: The taxonomic order of the species.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>CATEGORY</code>: The taxonomic category of the species (e.g., species, subspecies, genus, family, etc.)</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>SPECIES_CODE</code>: A unique code assigned to each species.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>PRIMARY_COM_NAME</code>: The common name of the species.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>SCI_NAME</code>: The scientific name of the species.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>ORDER1</code>: The taxonomic order of the species.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>FAMILY</code>: The taxonomic family of the species.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>SPECIES_GROUP</code>: The taxonomic group that the species belongs to.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\"> <code>REPORT_AS</code>: A code indicating how the species should be reported.</p>","metadata":{}},{"cell_type":"code","source":"ebt_df = pd.read_csv(\"/kaggle/input/birdclef-2023/eBird_Taxonomy_v2021.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:31.065427Z","iopub.execute_input":"2023-03-10T17:30:31.065892Z","iopub.status.idle":"2023-03-10T17:30:31.170770Z","shell.execute_reply.started":"2023-03-10T17:30:31.065855Z","shell.execute_reply":"2023-03-10T17:30:31.169254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ebt_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:31.172197Z","iopub.execute_input":"2023-03-10T17:30:31.172579Z","iopub.status.idle":"2023-03-10T17:30:31.190680Z","shell.execute_reply.started":"2023-03-10T17:30:31.172543Z","shell.execute_reply":"2023-03-10T17:30:31.189249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\"> ⚪ Let's get the len of this dataframe.</p>","metadata":{}},{"cell_type":"code","source":"len(ebt_df)","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:31.191871Z","iopub.execute_input":"2023-03-10T17:30:31.192220Z","iopub.status.idle":"2023-03-10T17:30:31.200984Z","shell.execute_reply.started":"2023-03-10T17:30:31.192187Z","shell.execute_reply":"2023-03-10T17:30:31.199731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.1\"></a>\n## <div style=\"box-shadow: rgba(0, 0, 0, 0.18) 0px 2px 4px inset; padding:20px; font-size:24px; font-family: consolas; text-align:center; display:fill; border-radius:15px; color:rgb(67, 66, 66)\"> <b> 3.1 Check for missing data</b></div>","metadata":{}},{"cell_type":"code","source":"ebt_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-03-10T17:30:31.202553Z","iopub.execute_input":"2023-03-10T17:30:31.203001Z","iopub.status.idle":"2023-03-10T17:30:31.226703Z","shell.execute_reply.started":"2023-03-10T17:30:31.202961Z","shell.execute_reply":"2023-03-10T17:30:31.224479Z"},"trusted":true},"execution_count":null,"outputs":[]}]}