{"cells":[{"metadata":{},"cell_type":"markdown","source":"# 🐦 Birds over the years ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![viz](https://github.com/syborg91/kaggle/blob/master/cornell-birdcall-identification/map.png?raw=true)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"In this notebook, we visually explore the rich geographic distribution of various species of birds over time. This would allow us to potentially trace:\n1. Migration patterns, and\n2. Prevalence of certain species in specific regions\n\nTo that end, we will create an animated map with a time slider.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Libraries","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import random\nimport plotly\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\nimport plotly.graph_objs as go\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For the mapbox visualization, please paste the `mapbox_access_token` in kaggle environment and retrieve as follows,","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nmapbox_access_token = user_secrets.get_secret(\"mapbox_access_token\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* Lets start by loading `train.csv` and checking some basic information as follows,","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('/kaggle/input/birdsong-recognition')\ntrain = pd.read_csv(path/'train.csv')","execution_count":null,"outputs":[]},{"metadata":{"executionInfo":{"elapsed":848,"status":"ok","timestamp":1593585037192,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"fbfURsYdRrU_","outputId":"d59e2ca3-5bf2-4e1c-e8fe-3e6271db0c85","trusted":true},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For our exploration we are primarily interested in the following geo-spatial and temporal features :\n- `latitude`\n- `longitude`\n- `elevation`\n- `time`\n- `date`","execution_count":null},{"metadata":{"id":"ZewuoBkQRrVs"},"cell_type":"markdown","source":"## Temporal features","execution_count":null},{"metadata":{"id":"zPFon-IqRrVt"},"cell_type":"markdown","source":"We will handle the instances with 12 hour format by converting them to 24 hours. We will also convert all datetime to `%Y-%m-%d %H:%M:%S` format.\n\n> Note : All new columns created would be prefixed by a underscore.","execution_count":null},{"metadata":{"executionInfo":{"elapsed":671,"status":"ok","timestamp":1593585058348,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"FacXyzIyRrVt","trusted":true},"cell_type":"code","source":"train['_time'] = pd.to_datetime(train.time, errors='coerce').dt.strftime('%H:%M:%S')\ntrain['_date'] = pd.to_datetime(train.date, format='%Y-%m-%d %H:%M:%S', errors='coerce').dt.strftime('%Y-%m-%d')\n# creating a new column: _datetime\ntrain['_datetime'] = pd.to_datetime(train['_date'] + ' ' + train['_time'], errors='coerce').dt.strftime('%Y-%m-%d %H:%M:%S')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let's check the `NaN` columns for `_datetime` i.e. where either the date or time is not in proper format,","execution_count":null},{"metadata":{"executionInfo":{"elapsed":597,"status":"ok","timestamp":1593585059534,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"9OK0R4lGRrVx","outputId":"b90e5c64-3dee-429b-8c12-4f732baae144","trusted":true},"cell_type":"code","source":"train[train._datetime.isna()][['date', 'time', '_datetime']].head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see either the date or time (or both) are invalid in these cases. ","execution_count":null},{"metadata":{"id":"oJ2kg1S3RrV5"},"cell_type":"markdown","source":"### Time distribution\nLets check the distribution of `time`, rounded to every quarter of an hour i.e. every 15 mins, as follows ","execution_count":null},{"metadata":{"executionInfo":{"elapsed":1720,"status":"ok","timestamp":1593585066034,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"Lsa23gBuRrV7","outputId":"f6fa0028-a869-4f1a-dca2-fcfe34d1c5fe","trusted":true},"cell_type":"code","source":"fig = go.Figure(data=[go.Histogram(x=pd.to_datetime(train._time, format='%H:%M:%S').dt.round('15min'))]) # rounding to nearest quarter of an hour\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"id":"HRsJvP1jRrV_"},"cell_type":"markdown","source":"Also checking unique invalid `time` as follows,","execution_count":null},{"metadata":{"executionInfo":{"elapsed":808,"status":"ok","timestamp":1593585068982,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"w6T0VT6hRrWA","outputId":"cee18537-c99e-4e83-83eb-aa607171c90a","trusted":true},"cell_type":"code","source":"print(train._time.isna().sum())\ntrain[train._time.isna()]['time'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see the above instances do not conform to any known standard of time.","execution_count":null},{"metadata":{"id":"dpuP7wPIRrWE"},"cell_type":"markdown","source":"Now checking the invalid dates as follows,","execution_count":null},{"metadata":{"executionInfo":{"elapsed":453,"status":"ok","timestamp":1593585070681,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"gmzgW2E2RrWF","outputId":"7bc6cf9e-219d-43d1-94d3-aab4181264cb","trusted":true},"cell_type":"code","source":"print(train._date.isna().sum())\ntrain[train._date.isna()]['date'].unique()","execution_count":null,"outputs":[]},{"metadata":{"id":"mzF2wohTRrWI"},"cell_type":"markdown","source":"> Most of these invalid dates and times have:\n- `0000-00-00`, or\n- Either `00` as the date or month","execution_count":null},{"metadata":{"id":"PegIC9SeRrWJ"},"cell_type":"markdown","source":"### Coarse-grained dates\n\nNow lets consider the dates which have a valid `YYYY-MM` format and ignore `dd` for now","execution_count":null},{"metadata":{"executionInfo":{"elapsed":627,"status":"ok","timestamp":1593585085700,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"7MuMw5ItRrWJ","outputId":"61c3b813-4eff-4c6e-8dec-af44f74f8185","trusted":true},"cell_type":"code","source":"train['_year_month'] = train.date.apply(lambda x : '-'.join(x.split('-')[:2])) # 'keeping only year-month and excluding date'\ntrain['_year_month'] = pd.to_datetime(train._year_month, format='%Y-%m', errors='coerce')\ntrain._year_month.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We managed to reduce the invalid dates from 152 to 37. Now lets plot a few histograms as follows ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"1. Year-month histogram","execution_count":null},{"metadata":{"executionInfo":{"elapsed":1623,"status":"ok","timestamp":1593585088320,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"47_WiAfHRrWM","outputId":"53a94cd8-9705-4135-d518-5f30e45ef953","trusted":true},"cell_type":"code","source":"fig = go.Figure(data=[go.Histogram(x=train._year_month)])\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"2. Month histogram","execution_count":null},{"metadata":{"executionInfo":{"elapsed":1069,"status":"ok","timestamp":1593585090486,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"b-Fb8S_rRrWP","outputId":"076dbe8f-f7d3-4b3e-b42b-c1d6a4bb8539","trusted":true},"cell_type":"code","source":"fig = go.Figure(data=[go.Histogram(x=pd.DatetimeIndex(train._year_month).month)])\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"3. Year histogram","execution_count":null},{"metadata":{"executionInfo":{"elapsed":1585,"status":"ok","timestamp":1593585091215,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"gH3VSTR2RrWS","outputId":"2269c4a8-58b8-4fe6-d662-a74101649e17","trusted":true},"cell_type":"code","source":"fig = go.Figure(data=[go.Histogram(x=pd.DatetimeIndex(train._year_month).year)])\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"id":"ni6NQ6tHRrWV"},"cell_type":"markdown","source":"Observations :\n- Most of the birdcalls were recorded between `Apr - May` for most of the years\n- Highest recorded audio peaked between `Apr 2014 - Jun 2014`","execution_count":null},{"metadata":{"id":"iGHM7M-qRrX9"},"cell_type":"markdown","source":"## Geo-spatial features","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"We perform the following transformations :\n1. Replace `m`, `~`, `,` and `?` with empty string\n2. Replace `1650-1900`, `930-990`, `Unknown` and `-` with empty string\n3. Only consider rows which have a valid longitude and latitude i.e. dropping `Not specified`\n4. Replacing elevation with empty string as `0.0` and scaling the values for the size of marker on the map later ","execution_count":null},{"metadata":{"executionInfo":{"elapsed":692,"status":"ok","timestamp":1593585173352,"user":{"displayName":"Satya Borgohain","photoUrl":"","userId":"04897513417763409229"},"user_tz":-600},"id":"HvHnJkwNRrYO","outputId":"e49e51eb-c085-40f0-9cec-c8d6e48f9a7a","trusted":true},"cell_type":"code","source":"train['_year_month'] = train._year_month.dt.strftime('%Y-%m') # converting to string \ntrain['_elevation'] = train.elevation.apply(lambda x : x.replace('m', '').replace('~', '').replace(',', '').replace('?', '').strip()) # replace\ntrain.loc[train._elevation.isin(['1650-1900', '930-990', 'Unknown', '-']), '_elevation'] = '' # assign empty string \ndf = train.loc[(train.longitude != 'Not specified') & (train.latitude != 'Not specified'), ['country', 'latitude', 'longitude', '_elevation', '_year_month', 'ebird_code', 'elevation']]\ndf.loc[df._elevation == '', '_elevation'] = None # empty string with None\ndf['_elevation'] = df._elevation.astype(float) # convert to float\ndf['_elevation'].fillna(0.0, inplace=True) # replace NaN with 0.0\ndf['_elevation'] = (df._elevation + 100.0)/80.0 # scale values ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Our new dataframe is as follows,","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now dropping all rows with invalid dates","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df.loc[~df._year_month.isna(), :] # dropping all NaN dates","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And setting the `_date` as the index of the dataframe (convenient for creating frames later)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df.set_index('_year_month') # setting date as the dataframe index","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Map","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"First, lets assign a unique id and color to the different species of birds as follows","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# total no of birds\nnumber_of_colors = 264\n\n# list of random hex-valued colors \ncolor = [\"#\"+''.join([random.choice('0123456789ABCDEF') for j in range(6)])\n             for i in range(number_of_colors)]\n\nebird_code = df.ebird_code.unique().tolist()\n# get ID and color for each bird\nEBIRD_CODE = {k : color[i] for i, k in enumerate(ebird_code)}\n# assign them to the dataframe\ndf['_color'] = df.ebird_code.apply(lambda x : EBIRD_CODE[x])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now lets get all the unique dates (in ascending order)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"months = sorted(df.index.unique().tolist())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"First we will create a list of dicts which will contain all the individual frames for our map. The tooltip will display:\n- `ebird_code`\n- `elevation`, and\n- `country`\n\nAlong with this each bird would is assigned the color as per the `_color` column.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"frames = [{   \n    'name':'frame_{}'.format(x),\n    'data':[{\n        'type':'scattermapbox',\n        'lat':np.array(df.xs(x)['latitude']),\n        'lon':np.array(df.xs(x)['longitude']),\n        'marker':go.scattermapbox.Marker(\n            size= 9 + df.xs(x)['_elevation'],\n            color=df.xs(x)['_color']\n        ),\n        'customdata': np.stack((df.xs(x)['ebird_code'], df.xs(x)['elevation'], df.xs(x)['country']), axis=-1),\n        'hovertemplate': \"<extra></extra> 🐦 <em>%{customdata[0]}</em><br> 📏 %{customdata[1]}<br> 🗺️ %{customdata[2]}<br>\",\n    }],           \n} for x in months]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Next lets create our slider and assign all the neccesary configuration as follows,","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sliders = [{\n    'transition':{'duration': 0},\n    'x':0.08, \n    'len':0.88,\n    'currentvalue':{'font':{'size':15}, 'prefix':'📅 ', 'visible':True, 'xanchor':'center'},  \n    'steps':[\n        {\n            'label':x,\n            'method':'animate',\n            'args':[\n                ['frame_{}'.format(x)],\n                {'mode':'immediate', 'frame':{'duration':100, 'redraw': True}, 'transition':{'duration':50}}\n              ],\n        } for x in months]\n}]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Next we define the play and pause button which would allow us to play all the frames over time as follows,","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"play_button = [\n    {\n        \"buttons\": [\n            {\n                \"args\": [None, {\"frame\": {\"duration\": 100, \"redraw\": True},\n                                \"fromcurrent\": True, \"transition\": {\"duration\": 50}}],\n                \"label\": \"Play\",\n                \"method\": \"animate\"\n            },\n            {\n                \"args\": [[None], {\"frame\": {\"duration\": 0, \"redraw\": False},\n                                  \"mode\": \"immediate\",\n                                  \"transition\": {\"duration\": 0}}],\n                \"label\": \"Pause\",\n                \"method\": \"animate\"\n            }\n        ],\n        \"direction\": \"left\",\n        \"pad\": {\"r\": 10, \"t\": 87},\n        \"showactive\": True,\n        \"type\": \"buttons\",\n        \"x\": 0.1,\n        \"xanchor\": \"right\",\n        \"y\": 0,\n        \"yanchor\": \"top\"\n    }\n]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And finally, lets display our map as follows","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# defining the initial state\ndata = frames[0]['data']\n\n# adding all sliders and play button to the layout\nlayout = go.Layout(\n    sliders=sliders,\n    updatemenus=play_button,\n    title=\"Birds over the years\",\n    mapbox={\n        'accesstoken':mapbox_access_token,\n        'center':{\"lat\": 37.86, \"lon\": 2.15},\n        'zoom':1.7,\n        'style':'dark', # choose from: dark or light\n    },\n    height=1000\n)\n\n# creating the figure\nfig = go.Figure(data=data, layout=layout, frames=frames)\n\n# displaying the figure\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And there you have it! Feel free to tinker the settings as required and explore away the different birds in their habitats through the years. \n\nThis notebook hopefully enables people to understand how some of the species are more prevalent than others in specific geographic locations (and in particular seasons). Encoding this information while training our models could be an interesting avenue to explore.\n\n🐦 Happy birding!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"markdown","source":"## References\n- [Intro to Animations in Python](https://plotly.com/python/animations/)\n- [How to create outstanding animated scatter maps with Plotly and Dash](https://towardsdatascience.com/how-to-create-animated-scatter-maps-with-plotly-and-dash-f10bb82d357a)","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}