{"cells":[{"metadata":{},"cell_type":"markdown","source":"### 1. Introduction\n\nThis Notebook shows how you can download additional metadata for the birds songs recording so that you enhance your training data.  \nAs an application, we will just show how to download data from specific countries. As well, you can modify the scripts to download data per family or species.  \n\nWe will run requests for **xeno-canto** public API and use them to download the data, as specified in our request.  \nThe scripts download first page and then, based on the content of the first page response, will also run repeatedly, to download all pages of response.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"\n\n### 2. Load Packages\n\nWe will need Python packages for **json** (to read the requested data), **requests** (to build the request to the **xeno-canto** API). We also add **tqdm** package, to show progress of using the API.","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport json\nimport requests\nfrom tqdm import tqdm","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 2. Functions\n\nThe functions used to download content from **xeno-canto-org** are the following:  \n\n* **get_first_page_per_country** - get, for a specific country, the first page of response, as well as metadata for the next pages;  \n* **get_page_per_country** - get a specific page per country - this is called after a first page was downloaded and is called for each subsequent page;  \n* **inspect_json** - print metadata for the first response;  \n* **get_recordings** - retrieve payload from a downloaded page;   \n* **download_suite_from_country** - end-to-end suite for downloading content for a specific country - call the above described functions.\n","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"def get_first_page_per_country(country):\n    \"\"\"\n    @country: the country for which we download metadata content \n    @returns: the content downloaded\n    \"\"\"\n    api_search = f\"https://www.xeno-canto.org/api/2/recordings?query=cnt:{country}\"\n    response = requests.get(api_search)\n    if response.status_code == 200:\n        response_payload = json.loads(response.content)\n        return response_payload\n    else:\n        return None\n\ndef get_page_per_country(country, page):\n    \"\"\"\n    @country: the country for which we download metadata content \n    @page: the current page to be downloaded\n    @returns: the content downloaded\n    \"\"\"\n    api_search = f\"https://www.xeno-canto.org/api/2/recordings?query=cnt:{country}&page={page}\"\n    response = requests.get(api_search)\n    if response.status_code == 200:\n        response_payload = json.loads(response.content)\n        return response_payload\n    else:\n        return None\n\ndef inspect_json(json_data):\n    \"\"\"\n    @json_data: json data to be inspected\n    \"\"\"\n    print(f\"recordings: {json_data['numRecordings']}\")\n    print(f\"species: {json_data['numSpecies']}\")\n    print(f\"page: {json_data['page']}\")\n    print(f\"number pages: {json_data['numPages']}\")\n\ndef get_recordings(payload):\n    \"\"\"\n    @payload: json data from which we extract the bird recordings metadata collection\n    @returns: birds recordings metadata collection\n    \"\"\"\n    return payload[\"recordings\"]\n\ndef download_suite_from_country(country, country_initial_payload):\n    \"\"\"\n    @country: the country for which we download metadata content \n    @country_initial_payload: the initial downloaded payload for the country (1st page). We download all the other pages.\n    @returns: the content recordings (all pages, including the original one)\n    \"\"\"\n    pages = country_initial_payload[\"numPages\"]\n    \n    all_recordings = []\n    all_recordings = all_recordings + get_recordings(country_initial_payload)\n    for page in tqdm(range(2,pages+1)):\n        payload = get_page_per_country(country, page)\n        recordings = get_recordings(payload)\n        all_recordings = all_recordings + recordings\n    \n    return all_recordings","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 3. Application: download all metadata of recordings from a country\n\nWe are using the utility funtions to download and save the meta information for birdsongs recording for a specific country.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def download_save_all_meta_for_country(country):\n    # download first batch. From here we extract the number of pages\n    birds = get_first_page_per_country(country)\n    # let's inspect the first batch\n    inspect_json(birds)\n    print(f\"recordings in first batch: {len(get_recordings(birds))}\")\n    # download entire suite (all pages)\n    suite = download_suite_from_country(country, birds)\n    # convert the collection in a dataFrame\n    data_df = pd.DataFrame.from_records(suite)\n    # export the dataframe as a csv\n    data_df.to_csv(f\"birds_{country}.csv\", index=False)\n    print(f\"suite length: {data_df.shape[0]}\")\n    return data_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.1. Download France data\n\nDownload and save all data for France.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df = download_save_all_meta_for_country('france')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.set_option('max_columns', 30)\npd.set_option('max_colwidth', 100)\ndata_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.2. Download Romania data\n\nDownload and save all data for Romania.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df = download_save_all_meta_for_country('romania')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.3. Download Bulgaria data\n\nDownload and save all data for Bulgaria.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df = download_save_all_meta_for_country('bulgaria')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.4. Download Italy data\n\nDownload and save all data for Italy.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df = download_save_all_meta_for_country('italy')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.5. Download India data\n\nDownload and save all data for India.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df = download_save_all_meta_for_country('india')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.6. Download Brazil data\n\nDownload and save all data for Brazil. Brazil is one of the countries with largest collection of recordings.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df = download_save_all_meta_for_country('brazil')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_df.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}