{"cells":[{"metadata":{},"cell_type":"markdown","source":"## General information\n\nThere are already many projects underway to extensively monitor birds by continuously recording natural soundscapes over long periods. However, as many living and nonliving things make noise, the analysis of these datasets is often done manually by domain experts. These analyses are painstakingly slow, and results are often incomplete.\n\nIn this competition we predict which bird are in the audio\n\nWork, obviously, is in progress :)\n![](https://i.imgur.com/30Eqq6Y.png)","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport os\nimport IPython.display as ipd\npd.set_option('max_columns', 50)\npd.set_option('max_rows', 150)\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data overview\n\nLet's have a look at the data!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('/kaggle/input/birdsong-recognition')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have a lot of data - audiofiles and metadata","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"len(os.listdir('/kaggle/input/birdsong-recognition/train_audio'))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 264 folders with audio files - for each class of birds","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here is a sample of the recording!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"ipd.Audio('/kaggle/input/birdsong-recognition/train_audio/nutwoo/XC462016.mp3')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/birdsong-recognition/train.csv')\ntrain.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Target","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['ebird_code'].nunique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12, 8))\ntrain['ebird_code'].value_counts().plot(kind='hist')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 264 different birds in train data!\n\nInteresting to notice that max count of rcordings per a bird is 100.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Location\n\nWhere was the recording made","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['location'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['location'].apply(lambda x: x.split(',')[-1]).value_counts().head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that most recordings were done in Canada and America","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['location'].value_counts().plot(kind='hist')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['location'].nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are a lot of different locations! More than 6 thousands, with some locations having more than 100 recordings","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Country","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12, 8))\ntrain['country'].value_counts().head(20).plot(kind='barh');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Date\n\nDate of recording","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(20, 8))\ntrain['date'].value_counts().sort_index().plot();","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It would be interesting to understand why these peaks happen...","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['date'].sort_values()[15:30].values","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Be aware that there are some missing values (aka 0000-00-00) and several strange old dates. Also there are some wrong values like `1992-12-00`. If we want to use the data, we will have to fix such values.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Rating","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['rating'].value_counts().plot(kind='barh')\nplt.title('Counts of different ratings');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(24, 6))\nplt.subplot(1, 2, 1)\ntrain.groupby(['ebird_code']).agg({'rating': ['mean', 'std']}).reset_index().sort_values(('rating', 'mean'), ascending=False).set_index('ebird_code')['rating']['mean'].plot(kind='bar')\nplt.subplot(1, 2, 2)\ntrain.groupby(['ebird_code']).agg({'rating': ['mean', 'std']}).reset_index().sort_values(('rating', 'mean'), ascending=False).set_index('ebird_code')['rating']['mean'][:20].plot(kind='barh')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It is quite interesting to see that some birds have full 5 rating and some are less loved.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### duration","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['duration'].plot(kind='hist')\nplt.title('Distribution of durations');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(50, 100, 5):\n    perc = np.percentile(train['duration'], i)\n    print(f\"{i} percentile of duration is {perc}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that most recording have quite a duration of less than 2 minutes.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Test data\n\nLet's have a look at test data!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/birdsong-recognition/test.csv')\ntest","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Wait, what? We have only 3 rows in open test data. The rest is hidden and available only when submitting.\n\nOpen test is 27%.\n\nit seems we will have a  shakeup!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"At least we have some more data!","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_metadata = pd.read_csv('/kaggle/input/birdsong-recognition/example_test_audio_metadata.csv')\ntest_metadata.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_metadata.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_summary = pd.read_csv('/kaggle/input/birdsong-recognition/example_test_audio_summary.csv')\ntest_summary.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_summary.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/birdsong-recognition/sample_submission.csv')\nsub","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's submit it.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}