{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BirdCLEF 2021 - BirdCall Identification","metadata":{"ExecuteTime":{"end_time":"2021-04-04T07:11:38.236859Z","start_time":"2021-04-04T07:11:38.230844Z"}}},{"cell_type":"markdown","source":"## Problem statement","metadata":{"heading_collapsed":true}},{"cell_type":"markdown","source":"- In this competition, you’ll automate the acoustic identification of birds in soundscape recordings. You'll examine an acoustic dataset to build detectors and classifiers to extract the signals of interest (bird calls).\n\n- With proper sound detection and classification—aided by machine learning—researchers can improve their ability to track the status and trends of biodiversity in important ecosystems, enabling them to better support global conservation efforts.","metadata":{"hidden":true}},{"cell_type":"markdown","source":"## Provided Data","metadata":{"heading_collapsed":true}},{"cell_type":"markdown","source":"1. **train_short_audio**\n     1. These are short recording of individual bird call as recorded by users of [xenocanto](www.xenocanto.org). These \n     contitutes bulk of your training data.\n\n2. **train_soundscapes** \n     1. These files are similar to your test data. Audio files are of ~10 mins long duration.\n    \n3. **train_metadata.csv** \n     1. These file contains metadata of recordings of train_short_audio like site, date, filename, recordist etc\n4. **train_soundscape_labels**\n     1. contains the labels for auido file present in train_sounscapes folder.Labels are given for each 5-second window of auido file. \n     2. For example - row_id - 7019_COR_5 means and auido file with id 7019, \n     3. COR is the site at which audio was recorded and 5 indicates the 0 to 5 second window of the complete audio file of id 7019. \n     4.\"birds\" column is the labels i.e. birds(separated by space) which were heard during that time window. ","metadata":{"hidden":true}},{"cell_type":"markdown","source":"## Submission File","metadata":{"heading_collapsed":true}},{"cell_type":"markdown","source":"For each row_id i.e.(audio id)_(site name)_(5 second time_window), you need to list the birds which are being heard in that window. which will be evaluated based on F1 score.","metadata":{"hidden":true}},{"cell_type":"markdown","source":"# Data Exploration","metadata":{}},{"cell_type":"markdown","source":"## Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport re\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport geopandas as gpd\nimport descartes\nfrom shapely.geometry import Point,Polygon\n\nfrom collections import Counter","metadata":{"ExecuteTime":{"end_time":"2021-04-05T07:08:19.343361Z","start_time":"2021-04-05T07:08:19.327943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Lets analyse the train_metadata.csv file","metadata":{}},{"cell_type":"code","source":"base_dir = '../input/birdclef-2021/'\ntrain_metadata = pd.read_csv(base_dir + '/train_metadata.csv')","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:26.158274Z","start_time":"2021-04-05T05:44:25.79908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Shape and columns of the train_metadata","metadata":{}},{"cell_type":"code","source":"train_metadata.shape","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:26.167165Z","start_time":"2021-04-05T05:44:26.160183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata.head()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:26.205067Z","start_time":"2021-04-05T05:44:26.169157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. We have Labels for **62874** individual xenocanto audio files\n2. Each file_name has 2 labels - \n    1. Primary_label - Primary bird in the audio file\n    2. Secondary_labels - bird audio(if present) in the background\n3. Type - Type of call made by the bird in the audio file\n4. Latitude & Longitude - Co-ordinates of location where audio was captured\n5. Scientific_name - Scientific name of primary bird in the audio file\n6. Common_name - Common_name of primary bird in the audio file\n7. Author - Recordist name\n8. Filename - Filename as in **train_short_audio** folder\n9. date - Date on which recording was captured\n10. Rating - Rating based on quality of audio captured\n11. Time - Time when recording was captured\n12. url - url of <https://xenocanto.org>","metadata":{"ExecuteTime":{"end_time":"2021-04-04T09:37:27.095173Z","start_time":"2021-04-04T09:37:27.089402Z"}}},{"cell_type":"markdown","source":"### Is there any null values in the this file?","metadata":{"heading_collapsed":true}},{"cell_type":"code","source":"train_metadata.isnull().sum()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:26.239119Z","start_time":"2021-04-05T05:44:26.20609Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **No Null value; Good to Go.**","metadata":{"hidden":true}},{"cell_type":"markdown","source":"### Individual Column wise analysis","metadata":{}},{"cell_type":"markdown","source":"#### Primary labels","metadata":{"heading_collapsed":true}},{"cell_type":"markdown","source":"- **Number of unique Primary labels**","metadata":{"hidden":true}},{"cell_type":"code","source":"train_metadata.primary_label.nunique()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:26.249967Z","start_time":"2021-04-05T05:44:26.241963Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Distribution of differnt primary labels**","metadata":{"hidden":true}},{"cell_type":"code","source":"primary_label_dist = train_metadata.primary_label.value_counts().reset_index().rename(columns={'primary_label':'count',\n                                                                                               'index':'primary_label'})\n\nprimary_label_dist","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:26.274953Z","start_time":"2021-04-05T05:44:26.251936Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(24,12))\nax = sns.barplot(x = 'primary_label', y='count', data= primary_label_dist[primary_label_dist['count']>=200])\nax.set_xticklabels(ax.get_xticklabels(), rotation=90)\nplt.plot()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T08:37:38.275482Z","start_time":"2021-04-05T08:37:36.787917Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **So, we have count within groups ranging from maximum of 500 to minimum of 8**. Ok, lets see the distribution of these counts","metadata":{"hidden":true}},{"cell_type":"code","source":"plt.figure(figsize=(24,12))\nsns.distplot(primary_label_dist['count'], kde=False)\nplt.title(\"Primary Label Distribution\")","metadata":{"ExecuteTime":{"end_time":"2021-04-05T08:37:46.447704Z","start_time":"2021-04-05T08:37:46.218518Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Looks like most of the bird species has  ~75 to 200 samples corresponding to it.**","metadata":{"hidden":true}},{"cell_type":"markdown","source":"#### Secondary label","metadata":{"heading_collapsed":true}},{"cell_type":"markdown","source":"- Number of unique secondary labels","metadata":{"hidden":true}},{"cell_type":"code","source":"train_metadata.secondary_labels.nunique()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:27.7946Z","start_time":"2021-04-05T05:44:27.782597Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"seocondary_label_dist = train_metadata.secondary_labels.value_counts().reset_index().rename(columns={'secondary_labels':'count',\n                                                                             'index':'secondary_labels'})\n\nseocondary_label_dist","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:27.825482Z","start_time":"2021-04-05T05:44:27.796559Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- secondarylabels contains more than one species.\n- Majority of the these labels is empty i.e. there is no other bird heard in that audio clip\n\n- **ok, so lets see how these counts is actually distributed.**","metadata":{"hidden":true}},{"cell_type":"code","source":"plt.figure(figsize=(24,12))\nsns.distplot(seocondary_label_dist['count'],kde= False)","metadata":{"ExecuteTime":{"end_time":"2021-04-05T08:37:57.857619Z","start_time":"2021-04-05T08:37:57.572343Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Majority of these unique values in secondary_labels is basically repeated very few number of time.** But what about the repeatation of unique individual bird?","metadata":{"hidden":true}},{"cell_type":"code","source":"all_secondary_labels = train_metadata.secondary_labels.apply(lambda x: eval(x)).sum()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:32.565404Z","start_time":"2021-04-05T05:44:28.063845Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"individual_secondary_label = dict(Counter(all_secondary_labels))\nindividual_secondary_label_df = pd.DataFrame({'label':list(individual_secondary_label.keys()),\n                                              'count':list(individual_secondary_label.values())}).sort_values('count',ascending=False)","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:32.574685Z","start_time":"2021-04-05T05:44:32.566354Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"individual_secondary_label_df","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:44:32.596279Z","start_time":"2021-04-05T05:44:32.575297Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- lets see the count of these bird labels where count is more than 200","metadata":{"hidden":true}},{"cell_type":"code","source":"plt.figure(figsize=(24,12))\nax = sns.barplot(x = 'label', y='count', data= individual_secondary_label_df[individual_secondary_label_df['count']>=200])\nax.set_xticklabels(ax.get_xticklabels(), rotation=90)\nplt.plot()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T08:38:19.10269Z","start_time":"2021-04-05T08:38:18.193398Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Geolocation column","metadata":{"heading_collapsed":true}},{"cell_type":"markdown","source":"- As there is no null in the co-ordinates lets plot these in on a world map","metadata":{"ExecuteTime":{"end_time":"2021-04-05T05:45:18.24909Z","start_time":"2021-04-05T05:45:18.242087Z"},"hidden":true}},{"cell_type":"code","source":"gdf = gpd.GeoDataFrame(train_metadata, geometry=gpd.points_from_xy(train_metadata.longitude, \n                                                                   train_metadata.latitude))","metadata":{"ExecuteTime":{"end_time":"2021-04-05T07:16:19.967551Z","start_time":"2021-04-05T07:16:19.189185Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"world = gpd.read_file(gpd.datasets.get_path('naturalearth_lowres'))\nfig,ax = plt.subplots(figsize=(24,12))\nworld.plot(ax= ax, color='black', edgecolor='black')\ngdf.plot(ax=ax, color='red', markersize=2)\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T07:37:30.33959Z","start_time":"2021-04-05T07:37:26.92785Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Looks like, though the data has been collected across the globe, but majority of it has been collected in North America, South America and European countries**","metadata":{"ExecuteTime":{"end_time":"2021-04-05T07:14:59.860075Z","start_time":"2021-04-05T07:14:59.76155Z"},"hidden":true}},{"cell_type":"markdown","source":"#### Date","metadata":{"ExecuteTime":{"end_time":"2021-04-05T06:53:01.937162Z","start_time":"2021-04-05T06:53:01.758347Z"},"heading_collapsed":true}},{"cell_type":"code","source":"train_metadata['year'] = train_metadata['date'].astype('str').str[0:4]\nyear_dist = train_metadata.year.value_counts().reset_index().rename(columns={'index':'year',\n                                                                 'year':'count'}).sort_values('year',ascending=True).reset_index(drop=True)\n\nplt.figure(figsize=(24,12))\nax = sns.barplot(x='year',y='count',data = year_dist)\nax.set_xticklabels(labels= ax.get_xticklabels(), rotation= 90)\nplt.title('Distribution across year of auido recordings')\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T08:31:10.474735Z","start_time":"2021-04-05T08:31:09.733042Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Some garbage date value is present where year is in (0000, 0199, 0201,0202,2104)\n- Most of the audio has been captured in last decade","metadata":{"hidden":true}},{"cell_type":"code","source":"train_metadata['month'] = train_metadata['date'].astype('str').str[5:7]\nmonth_dist = train_metadata.month.value_counts().reset_index().rename(columns={'index':'month',\n                                                                              'month':'count'}).sort_values('month',ascending=True).reset_index(drop=True)\n\nplt.figure(figsize=(24,12))\nax = sns.barplot(x='month',y='count',data = month_dist[month_dist['month']!='00'])\nax.set_xticklabels(labels= ax.get_xticklabels(), rotation= 90)\nplt.title('Distribution across month of auido recordings')\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2021-04-05T08:36:31.829362Z","start_time":"2021-04-05T08:36:31.533583Z"},"hidden":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Most of the audio has been captured between March to July timeline","metadata":{"hidden":true}},{"cell_type":"markdown","source":"#### ratings","metadata":{}},{"cell_type":"code","source":"rating_dist = train_metadata.rating.value_counts().reset_index().rename(columns={'index':'rating',\n                                                                                 'rating':'count'}).sort_values('rating',ascending=True).reset_index(drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(24,12))\nax = sns.barplot(x='rating',y='count',data = rating_dist)\nplt.title('Distribution across rating of auido recordings')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Most of the audio clip has rating more than 3.5","metadata":{}},{"cell_type":"markdown","source":"#### time","metadata":{}},{"cell_type":"code","source":"def hour_extractor(time_object):\n    if re.match(r'^\\d\\d:\\d\\d$',time_object.strip('')):\n        return time_object.strip('')[0:2]\n    else:\n        return 'NA'\ntrain_metadata['hour'] = train_metadata['time'].apply(lambda x: hour_extractor(x))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hour_dist = train_metadata['hour'].value_counts().reset_index().rename(columns={'index':'hour',\n                                                                                'hour':'count'}).sort_values('hour',ascending=True).reset_index(drop=True)\n\nplt.figure(figsize=(24,12))\nax = sns.barplot(x='hour',y='count',data = hour_dist)\nplt.title('Distribution across hour of auido recordings')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Singnificant number of audio recoding does not have time captured. However, looks like majority of audio recording has been captured in first half.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}