{"cells":[{"metadata":{},"cell_type":"markdown","source":"![](https://resources.tidal.com/images/76e0cb4c/bde6/4fca/983b/d5e4891b63dc/640x640.jpg)\n- [1. Basic Setup](#1)\n  -[1.1 Loading Libraries and Dataframes](#1-1)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<a name='1'></a>\n# 1. Basic Setup","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<a name='1-1'></a>\n## 1.1  Loading Libraries and Dataframes","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom os import listdir\nimport matplotlib.pyplot as plt\nimport IPython.display as ipydisplay ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true},"cell_type":"code","source":"PATH_AUDIO = '../input/birdsong-recognition/train_audio/'\ntrain_df = pd.read_csv('../input/birdsong-recognition/train.csv')\ntest_df = pd.read_csv('../input/birdsong-recognition/test.csv')\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that most of the data is recorded from North America (USA, CA, MEX)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10, 5))\ntrain_df.country.value_counts()[:10].plot.bar()\nplt.title('Top 10 Countries where data is recorded')\nplt.grid('True')\nplt.xlabel('Courntries')\nplt.ylabel('No of audio samples taken')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ebird_name_map = train_df[['ebird_code','species', 'primary_label']].drop_duplicates()\nno_of_categories = ebird_name_map.shape[0]\nprint('No of bird categories: ', no_of_categories)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.columns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Uncomment the following code cell to hear the sounds. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# for ebird_code in listdir(PATH_AUDIO)[:10]:\n#     sample_audio_path = f'{PATH_AUDIO}{ebird_code}/{listdir(PATH_AUDIO + ebird_code)[0]}'\n#     species = ebird_name_map[ebird_name_map.ebird_code == ebird_code].species.values[0]\n#     ipydisplay.display(ipydisplay.HTML(f\"<h3>{ebird_code} ({species})</h3>\"))    \n#     ipydisplay.display(ipydisplay.Audio(sample_audio_path))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# train_df.groupby(\"species\")[\"filename\"].count().reset_index().rename(columns = {\"filename\": \"recordings\"}).sort_values(\"recordings\", ascending=False).plot.barh(figsize = (5,100))\noutput = train_df.groupby(\"species\")[\"filename\"].count().reset_index().rename(columns = {\"filename\": \"recordings\"}).sort_values(\"recordings\", ascending=False).reset_index(drop=True, inplace=False).copy()\nplt.figure(figsize = (7, 60))\nplt.xticks(rotation = 90)\nplt.barh(output.species, output.recordings)\nplt.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Categories with 100 recordings: ', output[output.recordings ==100].species.count())\nprint('Categories with less than 100 recordings:', output[output.recordings <100].species.count())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that almost **(130/264) = 49.2%** of the categories of the birds have less than the maximum number of samples present(100) in the training folder. \n* However in the extended datasets are available in [dataset part 1](https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m) and [dataset part 2](https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-n-z) (thanks to [Vopani](https://www.kaggle.com/rohanrao)) which contains almost 20k samples each categories. \n* **Therefore it will be a good idea to take samples from that original external dataset to fill out the categories which has less than 100 samples** ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# I will keep updating this notebook as I walk through more things.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}