{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Project Overview\n* Goal: Develop a model that can accuractely predict the occurance of bird species given an audio sample\n\n* Purpose: Providing a cheaper and more efficient alternative to rigorous, manual obvervation of species prevelance and population change.\n\n* Evaluation: Model performance is graded with a system dervied from the 'macro-averaged average precision score'. This grading scale accounts for imbalanced class frequency. (i.e. some species of birds will appear more than others)","metadata":{}},{"cell_type":"markdown","source":"# Data Organization and File Analysis\nFirst, let's look at the different files provided and understand their relationships to one another.","metadata":{}},{"cell_type":"code","source":"#importing dependancies\nimport numpy as np\nimport pandas as pd\nimport os","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:33:21.110882Z","iopub.execute_input":"2023-03-08T02:33:21.111317Z","iopub.status.idle":"2023-03-08T02:33:21.115806Z","shell.execute_reply.started":"2023-03-08T02:33:21.111278Z","shell.execute_reply":"2023-03-08T02:33:21.114938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 'train_metadata.csv'","metadata":{}},{"cell_type":"code","source":"#train_metatdata.csv\ntrain_metadata = pd.read_csv('/kaggle/input/birdclef-2023/train_metadata.csv')\ntrain_metadata.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:35:30.994336Z","iopub.execute_input":"2023-03-08T02:35:30.994835Z","iopub.status.idle":"2023-03-08T02:35:31.087112Z","shell.execute_reply.started":"2023-03-08T02:35:30.994798Z","shell.execute_reply":"2023-03-08T02:35:31.086116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#visualizing amount of missing data\nnull_values = train_metadata.isna().sum()\nprint(null_values)","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:37:14.690623Z","iopub.execute_input":"2023-03-08T02:37:14.691211Z","iopub.status.idle":"2023-03-08T02:37:14.709417Z","shell.execute_reply.started":"2023-03-08T02:37:14.691165Z","shell.execute_reply":"2023-03-08T02:37:14.707974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#visualizing frequency of species in dataset\nimport matplotlib.pyplot as plt\nspecies_frequency = train_metadata['primary_label'].value_counts()\nplt.bar(species_frequency.index, species_frequency.values)","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:39:36.714176Z","iopub.execute_input":"2023-03-08T02:39:36.714651Z","iopub.status.idle":"2023-03-08T02:39:39.341837Z","shell.execute_reply.started":"2023-03-08T02:39:36.714611Z","shell.execute_reply":"2023-03-08T02:39:39.340699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 'train_metadata.csv' Analysis\nEach row contains data to accompany a specific audio recording in the 'train_audio' folder. The linking point between them is the 'filename' column which points to the audio recording in the training data. The 'primary_label' column denotes the species of bird in the recording. As more supplimentary metadata, columns pertaining to the type of sound and the recording location are also provided.\n\nAs far as missing data goes, the only columns with holes are 'longitude' and 'latitude'. Further analysis will provide insights into whether those columns significantly benifit model training.\n\nSome species of birds have significantly more data than others; the most frequently recorded species have 500 audio recordings, while the least frequently recorded species have below 10 recordings.\n\nIn conclusion, the only overtly important columns in this dataset are the 'primary_label' column and the 'filename' column.","metadata":{}},{"cell_type":"markdown","source":"### 'eBird_Taxonomy_v2021.csv'","metadata":{}},{"cell_type":"code","source":"eBird_Taxonomy_v2021 = pd.read_csv('/kaggle/input/birdclef-2023/eBird_Taxonomy_v2021.csv')\neBird_Taxonomy_v2021.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:51:49.761103Z","iopub.execute_input":"2023-03-08T02:51:49.761635Z","iopub.status.idle":"2023-03-08T02:51:49.848896Z","shell.execute_reply.started":"2023-03-08T02:51:49.761594Z","shell.execute_reply":"2023-03-08T02:51:49.847734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null_values = eBird_Taxonomy_v2021.isna().sum()\nprint(null_values)","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:52:30.247526Z","iopub.execute_input":"2023-03-08T02:52:30.248230Z","iopub.status.idle":"2023-03-08T02:52:30.264027Z","shell.execute_reply.started":"2023-03-08T02:52:30.248166Z","shell.execute_reply":"2023-03-08T02:52:30.262730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 'eBird_Taxonomy_v2021.csv' Analysis\nThis file contains additional information on each species of bird in the dataset. At a glance, this data will not be useful for model training. Possible further analysis may be conducted during the feature selection phase to determine if there is a relationship between certain sounds/calls in audio files and species' family, order, etc.","metadata":{}},{"cell_type":"markdown","source":"### 'sample_submission.csv'","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('/kaggle/input/birdclef-2023/sample_submission.csv')\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-08T02:55:50.529398Z","iopub.execute_input":"2023-03-08T02:55:50.529879Z","iopub.status.idle":"2023-03-08T02:55:50.570028Z","shell.execute_reply.started":"2023-03-08T02:55:50.529845Z","shell.execute_reply":"2023-03-08T02:55:50.568122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 'sample_submission.csv' Analysis\nThis file contains an example of the formatting expected for competition evaluation and grading. Apart from the 'row_id' column, all of the subsequent columns contain the model's confidence that a given species of bird appeared in the associated audio recording.\n\nGeneral note: During project submission, the test dataset will automatically populate with 200 randomized/shuffled audio recordings for grading.","metadata":{}},{"cell_type":"markdown","source":"### 'train_audio' and 'test_soundscapes' Analysis\n'train_audio' is a folder that contains all of the audio recordings provided for training a model. Within this folder there are subfolders categorized by bird species name. Note that there is a direct link between these recordings and their associated metadata via the 'filename' column in the 'train_metadata.csv' file discussed previously.\n\n'test_soundscapes' contains one audio file that mimics the format that will be used in the testing dataset. This audio file seems to be significantly longer than the audio recordings used to populate the training audio data.","metadata":{}},{"cell_type":"markdown","source":"# Conclusions\nI think that a CNN would be a good option for this project. My next step(s) will be to organize the data such that a given row in a new dataframe will contain an audio file, and columns for each species of bird where the species of bird actually heard in the recording has a value to represent 100% confidence in its appearance.\n\nI'm looking forward to continuing this project and I hope that you'll follow along with my progress! If you are also contributing a submission to this project best of luck to you!\n\nCheers","metadata":{}}]}