{"cells":[{"metadata":{},"cell_type":"markdown","source":"<div align=\"center\">\n<font size=\"6\"> Cornell Birdcall Identification </font>  \n</div>\n\n<div align=\"center\">\n<font size=\"4\"> Build tools for bird population monitoring </font>  \n</div>\n\n<img align=\"right\" src=\"https://www.birds.cornell.edu/ccb/wp-content/uploads/2020/02/Raven-1.-Joseph-Westgate-20.png\" data-canonical-src=\"https://www.birds.cornell.edu/ccb/wp-content/uploads/2020/02/Raven-1.-Joseph-Westgate-20.png\" width=\"400\" height=\"400\" />\n\n<!-- <font size=\"2\"> -->\n    \nOver 10,000 bird species occur in the world, and they can be found in nearly every environment, from untouched rainforests to suburbs and even cities. However, it is often easier to hear birds than see them. With proper sound detection and classification, researchers could automatically intuit factors about an area’s quality of life based on a changing bird population.  \n\n\nThere are already many projects underway to extensively monitor birds by continuously recording natural soundscapes over long periods. To unlock the full potential of these extensive and information-rich sound archives, researchers need good machine listeners to reliably extract as much information as possible to aid data-driven conservation.  \n\n\nThe [Cornell Lab of Ornithology’s Center for Conservation Bioacoustics](https://www.birds.cornell.edu/ccb/) (CCB)’s mission is to collect/interpret sounds in nature.  \n\n\n* In this competition, you will identify a wide variety of bird vocalizations in soundscape recordings. \n* Due to the complexity of the recordings, they contain weak labels.\n\n<!-- </font>  -->\n\n<!-- ![](https://www.birds.cornell.edu/ccb/wp-content/uploads/2020/02/Raven-1.-Joseph-Westgate-20.png) -->\n","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        #print(os.path.join(dirname, filename))\n        continue\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data files","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!ls ../input/birdsong-recognition","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/birdsong-recognition/train.csv')\ntest = pd.read_csv('/kaggle/input/birdsong-recognition/test.csv')\nsub = pd.read_csv('/kaggle/input/birdsong-recognition/sample_submission.csv')\n\nprint(\"Shape train: \", train.shape)\nprint(\"Shape  test: \", test.shape)\nprint(\"Shape   sub: \", sub.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train.copy()\nprint(\"Long min:\", df.longitude.min())\nprint(\"Long min:\", df.longitude.max())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axs = plt.subplots(2,4, figsize=(23,12))\n\ndf['rating'].value_counts().plot(kind='bar', legend=True, ax=axs[0,0])\ndf['playback_used'].value_counts().plot(kind='bar', legend=True, ax=axs[0,1])\ndf['speed'].value_counts().plot(kind='bar', legend=True, ax=axs[0,2])\ndf['channels'].value_counts().plot(kind='bar', legend=True, ax=axs[0,3])\n\ndf['pitch'].value_counts().plot(kind='bar', legend=True, ax=axs[1,0])\ndf['length'].value_counts().plot(kind='bar', legend=True, ax=axs[1,1])\ndf['bird_seen'].value_counts().plot(kind='bar', legend=True, ax=axs[1,2])\ndf['license'].value_counts().plot(kind='bar', legend=True, ax=axs[1,3])\n\nplt.savefig('data_exploration.png',dpi=300)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Bird samples","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"df['ebird_code'].unique(), len(df['ebird_code'].unique())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Audio\n\nWe have 264 directories for 264 birds","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"list_dirs = []\nfor dirs in next(os.walk('../input/birdsong-recognition/train_audio/'))[1]:\n    list_dirs.append(dirs)\n#list_dirs\n\narray_dirs = np.array(list_dirs)\n\nprint('List:\\n', list_dirs,'\\n')\nprint('Total: {} directores'.format(len(array_dirs)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#!ls ../input/birdsong-recognition/train_audio/\n#!ls ../input/birdsong-recognition/train_audio/astfly/\n\n# For sound within notebook\nimport IPython.display as ipd  \nipd.Audio('../input/birdsong-recognition/train_audio/astfly/XC109920.mp3')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Duration\n\nMost samples are <5 min","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(12,6))\n\ndf['duration'].hist(bins=360)\n\nplt.xlim(-10,600)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Print the longest ones separately. We have just few","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dflong =  df[df['duration'] > 600]\ndflong['duration'].hist()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Country\n* North America and Europe have the main contribution, expecially few countries  \n* Later see on the map","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"countries = df['country'].value_counts()\ncountries","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(23,6))\n\ncountries.plot.bar()\n\nplt.savefig('countries.png',dpi=200)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Location","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df.replace(['Not specified'], [0], inplace=True)\ndf_longitude= df['longitude'].astype(float)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_longitude.hist(bins=360)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_latitude= df['latitude'].astype(float)\ndf_latitude.hist(bins=180)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Map","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#!pip install basemap","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Map\n### Distribution of longitudes \nMost data is from North America and Europe. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"%matplotlib inline\nfrom mpl_toolkits.basemap import Basemap\n\nimport matplotlib.gridspec as gridspec\nfrom itertools import chain\n\n\ndef draw_map(m, scale):\n    # draw a shaded-relief image\n    m.shadedrelief(scale=scale)\n    \n    # lats and longs are returned as a dictionary\n    lats = m.drawparallels(np.linspace(-90, 90, 13))\n    lons = m.drawmeridians(np.linspace(-180, 180, 13))\n\n    # keys contain the plt.Line2D instances\n    lat_lines = chain(*(tup[1][0] for tup in lats.items()))\n    lon_lines = chain(*(tup[1][0] for tup in lons.items()))\n    all_lines = chain(lat_lines, lon_lines)\n    \n    # cycle through these lines and set the desired style\n    for line in all_lines:\n        line.set(linestyle='--', alpha=0.5, color='w')\n\n    \nfig = plt.figure(figsize=(15, 10), edgecolor='w')\ngs = gridspec.GridSpec(4,100) # 5 rows, 3 column\n\n## ---------------------------------------------------------------------------------------------------------------\n## World Map\nax1 = plt.subplot(gs[:2, 1:99]) \nm = Basemap(ax=ax1,projection='cyl', resolution='l', llcrnrlat=-90, urcrnrlat=90,llcrnrlon=-180, urcrnrlon=180, ) # 'mill'; None, 'c'\ndraw_map(m, scale=0.5)\n## ---------------------------------------------------------------------------------------------------------------\n\n## ---------------------------------------------------------------------------------------------------------------\n## Our data\nax2 = plt.subplot(gs[2, 19:81]) #Third row, span all columns by :\nax2.hist(df_longitude, 360, histtype='bar', orientation='vertical', color='blue',alpha=0.5)\nplt.xticks([-180, -150, -120, -90, -60, -30, 0, 30, 60 , 90, 120, 150, 180], [-180, -150, -120, -90, -60, -30, 0, 30, 60 , 90, 120, 150, 180],rotation=0)\nplt.yscale('symlog')\nplt.ylim(1,2000)\nplt.grid(True)\n## ---------------------------------------------------------------------------------------------------------------\n\n#plt.tight_layout()\nplt.savefig('world_plus_distribution.png',dpi=200)\n\nplt.show()\n\n## For world map, can be cutted via\n#lons, lats = np.meshgrid(m.drawmeridians(np.linspace(-180, 180, 13)),m.drawparallels(np.linspace(-90, 90, 13)))\n#x, y = m(lons, lats)\n#plt.xlim(-90, 90)\n#plt.ylim(-45, 45)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Geo","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"lon = df_longitude\nlat = df_latitude","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"duration = df['duration'].values\nrating = df['rating'].values\n\n# Draw the map background\nfig = plt.figure(figsize=(23, 23))\nm = Basemap(projection='cyl', resolution='l', lat_0=90, lon_0=0) # high resolution basemap-data-hires is needed, no internet is used in this competiton\nm.shadedrelief()\nm.drawcoastlines(color='gray')\nm.drawcountries(color='gray')\nm.drawstates(color='gray')\n\n# Scatter city data, with color reflecting rating and size reflecting duration\nm.scatter(lon, lat, latlon=True, c=rating, s=np.log10(duration+2)**3, cmap='Reds', alpha=0.5) # +2 to avoid log10 error, xx2 to increase visibility\n\n# Create colorbar and legend\n#plt.colorbar(label=r'rating')\n##plt.clim(3, 7)\n\n# create an axes on the right side of ax. The width of cax will be 5%\n# of ax and the padding between cax and ax will be fixed at 0.05 inch.\nfrom mpl_toolkits.axes_grid1 import make_axes_locatable\nax = plt.gca()\ndivider = make_axes_locatable(ax)\ncax = divider.append_axes(\"right\", size=\"5%\", pad=0.05)\nplt.colorbar(cax=cax)\n\n# Legend with dummy points\nfor a in [100, 500, 1000, 2000]:\n    plt.scatter([], [], c='k', alpha=0.5, s=np.log10(a+2)**3,label=str(a) + ' sec')\n    \nplt.legend(scatterpoints=1, frameon=False, labelspacing=1, loc='lower left');\n\nplt.savefig('geo_duration_rating.png',dpi=300)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(15, 9))\nplt.scatter(x=df['longitude'].astype(float), y=df['latitude'].astype(float))\n\nplt.xticks([-180, -150, -120, -90, -60, -30, 0, 30, 60 , 90, 120, 150, 180], [-180, -150, -120, -90, -60, -30, 0, 30, 60 , 90, 120, 150, 180],rotation=0)\nplt.ylim(-90,90)\nplt.grid(True)\n\nplt.savefig('birds_location.png',dpi=300)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from shapely.geometry import Point\nimport geopandas as gpd\nfrom geopandas import GeoDataFrame\n\n#fig = plt.figure(figsize=(15, 9))\n\ngeometry = [Point(xy) for xy in zip(df['longitude'].astype(float), df['latitude'].astype(float))]\ngdf = GeoDataFrame(df, geometry=geometry)  \n\n#this is a simple map that goes with geopandas\nworld = gpd.read_file(gpd.datasets.get_path('naturalearth_lowres'))\ngdf.plot(ax=world.plot(figsize=(23, 16)), marker='o', color='red', markersize=10);\n\nplt.savefig('birds_location_world.png',dpi=300)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Elevation","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['elevation']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['elevation'].value_counts().unique","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"elev_list = list(df['elevation'])\nelev_list,len(elev_list)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Take num values from text","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df['elevation'].head(5)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"import re\n\nelev_list_num = []\n\nfor i in range(len(elev_list)):\n    taken_num = re.findall(r\"[-+]?\\d*\\.\\d+|\\d+\", elev_list[i])\n    print(taken_num)\n    elev_list_num.append(taken_num)\n    \nelev_list_num","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(elev_list_num)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### some elevation are with  missing data, we will fill it by 0","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"elev_list_num[449]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"elev_list_num2 = []\nfor i in range(len(elev_list_num)):\n    if elev_list_num[i] == []:\n        elev_list_num2.append(0.000)\n    else:\n        elev_list_num2.append(float(elev_list_num[i][0]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(elev_list_num2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_elevation = pd.DataFrame(elev_list_num2)\ndf_elevation.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_elevation.rename(columns={0: 'elevation'}, inplace=True)\ndf_elevation.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(23, 8))\n\nplt.ylabel('Counts')\nplt.xlabel('Elevation value');\n\ndegrees = 90\nplt.xticks(rotation=degrees)\n\nplt.hist(elev_list_num, density=False, bins=30)  # density=True\n\nplt.savefig('birds_elevation.png',dpi=300)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['longitude'].shape,df['longitude'].shape,df_elevation['elevation'].shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_elevation['elevation'].max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(20,20))\nax = fig.add_subplot(111, projection='3d')\n\nax.scatter((1)*df['latitude'].astype(float),(1)*df['longitude'].astype(float),df_elevation['elevation']/1000, s = 0.5, color = 'r')\n\nax.set_xlim(-90,90)\nax.set_ylim(-180,180)\nax.set_zlim(0,5)\n\nax.set_xlabel('lat')\nax.set_ylabel('lon')\nax.set_zlabel('elev [km]')\nax.set_title('Birds on Sky :) ',fontsize=18)\n\nax.view_init(90, 90)\n\nplt.savefig('birds_3D_elevation_topview.png',dpi=300)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(20,20))\nax = fig.add_subplot(111, projection='3d')\n\nax.scatter((1)*df['latitude'].astype(float),(-1)*df['longitude'].astype(float),df_elevation['elevation']/1000, s = 0.5, color = 'r')\n\nax.set_xlim(-90,90)\nax.set_ylim(-180,180)\nax.set_zlim(0,5)\n\nax.set_xlabel('lat')\nax.set_ylabel('lon')\nax.set_zlabel('elev [km]')\nax.set_title('Birds on Sky :) ',fontsize=18)\n\nax.view_init(40, 135)\n\nplt.savefig('birds_3D_elevation.png',dpi=300)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(rc={'figure.figsize':(10,10)})\n\nsns_plot = sns.jointplot(x='latitude', y='longitude', data=df, kind='kde')\n\nsns_plot.savefig('birds_2D_world.png', dpi=300)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from IPython.display import Image\nImage(filename='../working/birds_2D_world.png') ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Date","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(23, 6), edgecolor='w')\ndf['date'].value_counts().sort_index().plot(c='green', linewidth=1)\nplt.savefig('date.png',dpi=200)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(12, 10), edgecolor='w')\n\ntime_series = df['date'].value_counts().reset_index()\ntime_series.columns = ['date', 'count']\n\ntime_series.plot(kind='kde')\ntime_series.plot(kind='hist')\n\n#plt.savefig('date2.png',dpi=200)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Submission","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}