{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n\n<h1><center>🦉BirdCEF2022🦉</center></h1>\n\n# 1. Introduction\n\n大家好,这是我们对于 BirdCLEF-2022 的数据分析, 思路和版式借鉴了Andrada的\nhttps://www.kaggle.com/code/andradaolteanu/birdcall-recognition-eda-and-audio-fe/notebook\n\n\n\n### Libraries 📚⬇","metadata":{}},{"cell_type":"code","source":"import os\n\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\n\n\nimport folium\n\n# Map 1 library\nimport plotly.express as px\n\n# Map 2 libraries\nimport descartes\nimport geopandas as gpd\nfrom shapely.geometry import Point, Polygon\n\n# Librosa Libraries\nimport librosa\nimport librosa.display\nimport IPython.display as ipd\n\nimport sklearn\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-05-01T03:31:00.960609Z","iopub.execute_input":"2022-05-01T03:31:00.961369Z","iopub.status.idle":"2022-05-01T03:31:00.982214Z","shell.execute_reply.started":"2022-05-01T03:31:00.961316Z","shell.execute_reply":"2022-05-01T03:31:00.981217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. csv都说了些什么呢 📁\n\n## 2.1 train.csv\n\n> 📌**Note**:\n* `train.csv` 表示了`train_audio`的统计信息. 总计14852个样本, 13个特征列, 152类小鸟.\n\n* <div class=\"特征列分析\"> \n    <b>特征列分析</b><br>\n    训练集中, 每条样本的特征为</p>\n    <code>primary_label</code> - (主要的鸟鸣), <code>secondary_label</code> - (次要鸟鸣);</p>\n    <code>type</code> - (鸣叫类型/内容);</p>\n    <code>latitude</code>, <code>longitude</code> - (录音的经纬度);</p>\n    <code>scientific_name</code>, <code>common_name</code> - (鸟类的学名和常用名); <code>author</code> - (录音作者);</p>\n    <code>license</code>, <code>rating</code>, <code>time</code>, <code>url</code> - (许可证, 录音等级, 录音时间, 网址);</p>\n    <code>filename</code> - (文件名);</p>\n</div>","metadata":{}},{"cell_type":"code","source":"train_csv=pd.read_csv(\"../input/birdclef-2022/train_metadata.csv\")\ntrain_csv.head(3)\n#metadata = train_csv['primary_label'].reset_index().explode(\"primary_label\")\n","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:00.983839Z","iopub.execute_input":"2022-05-01T03:31:00.984331Z","iopub.status.idle":"2022-05-01T03:31:01.102044Z","shell.execute_reply.started":"2022-05-01T03:31:00.984260Z","shell.execute_reply":"2022-05-01T03:31:01.101021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 test.csv\n* 测试集表示test.csv的统计信息, 共3条, 其余在隐藏测试集中.","metadata":{}},{"cell_type":"code","source":"test_csv = pd.read_csv(\"../input/birdclef-2022/test.csv\")\ntest_csv","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:01.103656Z","iopub.execute_input":"2022-05-01T03:31:01.104001Z","iopub.status.idle":"2022-05-01T03:31:01.120999Z","shell.execute_reply.started":"2022-05-01T03:31:01.103967Z","shell.execute_reply":"2022-05-01T03:31:01.119798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 eBird_Taxonomy_v2021.csv\n* 是样本鸟类的分类文件</p>\n* 以African SilverBill为例\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"../input/birdclef-2022/eBird_Taxonomy_v2021.csv\")\ndf[df['PRIMARY_COM_NAME'].isin(['African Silverbill'])]","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:01.124631Z","iopub.execute_input":"2022-05-01T03:31:01.125008Z","iopub.status.idle":"2022-05-01T03:31:01.215077Z","shell.execute_reply.started":"2022-05-01T03:31:01.124965Z","shell.execute_reply":"2022-05-01T03:31:01.214156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* TAXON_ORDER - 30031 分类编号\n* CATEGORY - species 物种\n* SCI_NAME - Euodice cantans https://es.wikipedia.org/wiki/Euodice_cantans\n* ORDER1 - Passeriformes 雀形目\n* FAMILY - Estrildidae 梅花雀 ","metadata":{}},{"cell_type":"markdown","source":"## 2.4 scored_bird.json\n#### 测试集会出现的小鸟种类","metadata":{}},{"cell_type":"code","source":"scored_bird = [\"akiapo\", \"aniani\", \"apapan\", \"barpet\", \"crehon\", \"elepai\", \"ercfra\",\n               \"hawama\", \"hawcre\", \"hawgoo\", \"hawhaw\", \"hawpet1\", \"houfin\", \"iiwi\",\n               \"jabwar\", \"maupar\", \"omao\", \"puaioh\", \"skylar\", \"warwhe1\", \"yefcan\"]\n#a=['aniani']\n# 如果要分析全部小鸟 请注释掉下面这行代码\n#train_csv = train_csv[train_csv[\"primary_label\"].isin(scored_bird)]\n\n#print(train_csv)","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:01.217604Z","iopub.execute_input":"2022-05-01T03:31:01.217957Z","iopub.status.idle":"2022-05-01T03:31:01.224243Z","shell.execute_reply.started":"2022-05-01T03:31:01.217922Z","shell.execute_reply":"2022-05-01T03:31:01.222923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.1 来看看歌曲内容吧 🎼\n\n","metadata":{}},{"cell_type":"markdown","source":"**Type Column**:\n\n> 📌**鸣叫类型/内容**: 这本身就是一个比较难分类的特征:\n* **alarm call** is: alarm call | alarm call, call \n* **flight call** is: flight call | call, flight call etc.","metadata":{}},{"cell_type":"code","source":"# Create a new variable type by exploding all the values\nadjusted_type = train_csv['type'].apply(lambda x: x[1:-1].split(',')).reset_index().explode(\"type\")\n\n# Strip of white spaces and convert to lower chars\nadjusted_type = adjusted_type['type'].apply(lambda x: x.strip().lower()).reset_index()\n#adjusted_type['type'] = adjusted_type['type'].replace('calls', 'call')\n\n# Create Top 15 list with song types\ntop_15 = list(adjusted_type['type'].value_counts().head(15).reset_index()['index'])\ndata = adjusted_type[adjusted_type['type'].isin(top_15)]\n\n# === PLOT ===\n\nplt.figure(figsize=(16, 6))\nax = sns.countplot(data['type'], palette=\"hls\", order = data['type'].value_counts().index)\n\nplt.title(\"Top 15 Song Types\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:01.225526Z","iopub.execute_input":"2022-05-01T03:31:01.225841Z","iopub.status.idle":"2022-05-01T03:31:01.856010Z","shell.execute_reply.started":"2022-05-01T03:31:01.225809Z","shell.execute_reply":"2022-05-01T03:31:01.854782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 按物种的音频样本统计 \n","metadata":{}},{"cell_type":"code","source":"# Create a new variable type by exploding all the values\nmetadata = train_csv['primary_label'].reset_index().explode(\"primary_label\")\n#metadata\n# Create Top 15 list with species types\ntop_15 = list(metadata['primary_label'].value_counts().head(8).reset_index()['index'])\n#print(top_15)\ndata = metadata[metadata['primary_label'].isin(top_15)]#这是找出了top15所在的行\n\nprint(data['primary_label'].value_counts())\n#print(data)\n# === PLOT ===\n#这个画图有必要吗 看看大致的数据趋势吧\n#print(len(metadata[\"primary_label\"].value_counts()))\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = metadata['primary_label'], palette=\"hls\", order = metadata['primary_label'].value_counts().index)\n\nplt.title(\"Label Counts = \"+ str(len(metadata[\"primary_label\"].value_counts())), fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks()\n#plt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:01.858042Z","iopub.execute_input":"2022-05-01T03:31:01.858682Z","iopub.status.idle":"2022-05-01T03:31:03.653901Z","shell.execute_reply.started":"2022-05-01T03:31:01.858628Z","shell.execute_reply":"2022-05-01T03:31:03.652959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 当然 我们往往关心那些较少的数据\n#### 我们需要对这几个缺少的数据做增强吗?","metadata":{}},{"cell_type":"code","source":"# Create Top 15 list with species types\nleast_15 = list(metadata['primary_label'].value_counts().tail(15).reset_index()['index'])\n#print(least_15)\ndata = metadata[metadata['primary_label'].isin(least_15)]#这是找出了least15所在的行\nprint(data['primary_label'].value_counts())\n# === PLOT ===\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = data['primary_label'], palette=\"hls\", order = data['primary_label'].value_counts().index)\n\nplt.title(\"Data Counts\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:03.655195Z","iopub.execute_input":"2022-05-01T03:31:03.655727Z","iopub.status.idle":"2022-05-01T03:31:03.920072Z","shell.execute_reply.started":"2022-05-01T03:31:03.655679Z","shell.execute_reply":"2022-05-01T03:31:03.918976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 音频质量","metadata":{}},{"cell_type":"code","source":"# Create Top 15 list with species types\ndata_rating = train_csv.loc[:,[\"primary_label\",\"rating\"]].reset_index().explode(\"rating\")\n#print(data_rating)\nleast_15 = list(data_rating['rating'].value_counts().tail(15).reset_index()['index'])\n#print(top_15)\ndata = data_rating[data_rating['rating'].isin(least_15)]#这是找出了least15所在的行\n#print(data_rating.loc[data_rating['rating']==0, :])\n# === PLOT ===\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = data_rating['rating'], palette=\"hls\")\n\nplt.title(\"Data Rating Counts\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:03.921881Z","iopub.execute_input":"2022-05-01T03:31:03.923511Z","iopub.status.idle":"2022-05-01T03:31:04.161707Z","shell.execute_reply.started":"2022-05-01T03:31:03.923454Z","shell.execute_reply":"2022-05-01T03:31:04.160827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 录音时间","metadata":{}},{"cell_type":"code","source":"# Create Top 15 list with species types\ndata_rating = train_csv['time'].reset_index().explode('time')\nfreq = '60min'\n#print(len(data_rating))\ndata_rating['time']= pd.to_datetime(data_rating['time'],errors = 'coerce')\ndata_rating['time'] = data_rating['time'].dt.floor(freq)\ndata_rating['time'] = data_rating['time'].dt.hour\n#print(len(data_rating[\"time\"]))\n#print(data_rating)\n#least_15 = list(data_rating['time'].value_counts().tail(15).reset_index()['index'])\n#print(top_15)\n#data = data_rating[data_rating['time'].isin(least_15)]#这是找出了least15所在的行\n#print(data_rating.loc[data_rating['time']==0, :])\n# === PLOT ===\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = data_rating['time'], palette=\"hls\")\n\nplt.title(\"Data Rating Counts\", fontsize=16)\nplt.ylabel(\"Counts\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks()\nplt.xlabel(\"hours\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:04.163084Z","iopub.execute_input":"2022-05-01T03:31:04.163749Z","iopub.status.idle":"2022-05-01T03:31:04.596002Z","shell.execute_reply.started":"2022-05-01T03:31:04.163700Z","shell.execute_reply":"2022-05-01T03:31:04.593875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 研究一下你的小鸟? 📸🔭\n\n","metadata":{}},{"cell_type":"markdown","source":"## World View of the Species 🧭🌏\n\n","metadata":{}},{"cell_type":"markdown","source":"### 2.3.1 Where are our birds? 🦜\n#### 2.3.1.1 完整数据集 按目分类 总计17个目","metadata":{}},{"cell_type":"code","source":"# SHP file\nworld_map = gpd.read_file(\"../input/world-shapefile/world_shapefile.shp\")\n\n# Coordinate reference system\ncrs = {\"init\" : \"epsg:4326\"}\ntrain_species_csv=pd.read_csv(\"../input//2022-birdcleftrain-data-with-species//train_metadata_with_species.csv\")\n#请删除接下来这一行\n# 就是这一行 train_species_csv= train_species_csv[train_species_csv[\"primary_label\"].isin(scored_bird)]\n\n# Lat and Long need to be of type float, not object\ndata = train_species_csv[train_species_csv[\"latitude\"] != \"Not specified\"]\ndata[\"latitude\"] = data[\"latitude\"].astype(float)\ndata[\"longitude\"] = data[\"longitude\"].astype(float)\n\n# Create geometry\ngeometry = [Point(xy) for xy in zip(data[\"longitude\"], data[\"latitude\"])]\n\n# Geo Dataframe\ngeo_df = gpd.GeoDataFrame(data, crs=crs, geometry=geometry)\n\n# Create ID for species\nspecies_id = geo_df[\"order1\"].value_counts().reset_index()\nspecies_id.insert(0, 'ID', range(0, 0 + len(species_id)))\n\nspecies_id.columns = [\"ID\", \"order1\", \"count\"]\nprint(species_id)\n# Add ID to geo_df\ngeo_df = pd.merge(geo_df, species_id, how=\"left\", on=\"order1\")\n\n# === PLOT ===\nfig, ax = plt.subplots(figsize = (16, 10))\nworld_map.plot(ax=ax, alpha=0.4, color=\"grey\")\n\npalette = iter(sns.hls_palette(len(species_id)))\n\nfor i in range(17):\n    geo_df[geo_df[\"ID\"] == i].plot(ax=ax, markersize=20, color=next(palette), marker=\"o\", label = \"test\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:04.597784Z","iopub.execute_input":"2022-05-01T03:31:04.598257Z","iopub.status.idle":"2022-05-01T03:31:08.471632Z","shell.execute_reply.started":"2022-05-01T03:31:04.598210Z","shell.execute_reply":"2022-05-01T03:31:08.470527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.3.1.2 完整数据集 按科分类 总计41个科","metadata":{}},{"cell_type":"code","source":"# SHP file\nworld_map = gpd.read_file(\"../input/world-shapefile/world_shapefile.shp\")\n\n# Coordinate reference system\ncrs = {\"init\" : \"epsg:4326\"}\n#train_species_csv=pd.read_csv(\"../input//2022-birdcleftrain-data-with-species//train_metadata_with_species.csv\")\n# Lat and Long need to be of type float, not object\ndata = train_species_csv[train_species_csv[\"latitude\"] != \"Not specified\"]\ndata[\"latitude\"] = data[\"latitude\"].astype(float)\ndata[\"longitude\"] = data[\"longitude\"].astype(float)\n\n# Create geometry\ngeometry = [Point(xy) for xy in zip(data[\"longitude\"], data[\"latitude\"])]\n\n# Geo Dataframe\ngeo_df = gpd.GeoDataFrame(data, crs=crs, geometry=geometry)\n\n# Create ID for species\nspecies_id = geo_df[\"family\"].value_counts().reset_index()\nspecies_id.insert(0, 'ID', range(0, 0 + len(species_id)))\n\nspecies_id.columns = [\"ID\", \"family\", \"count\"]\nprint(species_id)\n# Add ID to geo_df\ngeo_df = pd.merge(geo_df, species_id, how=\"left\", on=\"family\")\n\n# === PLOT ===\nfig, ax = plt.subplots(figsize = (16, 10))\nworld_map.plot(ax=ax, alpha=0.4, color=\"grey\")\n\npalette = iter(sns.hls_palette(len(species_id)))\n\nfor i in range(41):\n    geo_df[geo_df[\"ID\"] == i].plot(ax=ax, markersize=20, color=next(palette), marker=\"o\", label = \"test\");","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:08.473327Z","iopub.execute_input":"2022-05-01T03:31:08.473672Z","iopub.status.idle":"2022-05-01T03:31:13.840543Z","shell.execute_reply.started":"2022-05-01T03:31:08.473640Z","shell.execute_reply":"2022-05-01T03:31:13.839182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.3 Listening to some Recordings\n\n","metadata":{}},{"cell_type":"markdown","source":"### Ok, let's hear some songs! 🕊🎶","metadata":{}},{"cell_type":"code","source":"# 混音\nipd.Audio(\"../input/join-voice-of-birdclef/joinVoice.ogg\")","metadata":{"execution":{"iopub.status.busy":"2022-05-01T03:31:13.842223Z","iopub.execute_input":"2022-05-01T03:31:13.842738Z","iopub.status.idle":"2022-05-01T03:31:14.076807Z","shell.execute_reply.started":"2022-05-01T03:31:13.842676Z","shell.execute_reply":"2022-05-01T03:31:14.075428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Work in Progress ... ⏳","metadata":{}}]}