{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n\n<h1><center>🦉BirdCEF2022🦉</center></h1>\n\n# 1. Introduction\n\n大家好,这是我们对于 BirdCLEF-2022 的数据分析, 思路和版式借鉴了Andrada的\nhttps://www.kaggle.com/code/andradaolteanu/birdcall-recognition-eda-and-audio-fe/notebook\n\n\n\n### Libraries 📚⬇","metadata":{}},{"cell_type":"code","source":"import os\n\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\n\n\nimport folium\n\n# Map 1 library\nimport plotly.express as px\n\n# Map 2 libraries\nimport descartes\nimport geopandas as gpd\nfrom shapely.geometry import Point, Polygon\n\n# Librosa Libraries\nimport librosa\nimport librosa.display\nimport IPython.display as ipd\n\nimport sklearn\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-04-24T14:15:32.156469Z","iopub.execute_input":"2022-04-24T14:15:32.156755Z","iopub.status.idle":"2022-04-24T14:15:32.165543Z","shell.execute_reply.started":"2022-04-24T14:15:32.156728Z","shell.execute_reply":"2022-04-24T14:15:32.164682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. csv都说了些什么呢 📁\n\n## 2.1 train.csv\n\n> 📌**Note**:\n* `train.csv` 表示了`train_audio`的统计信息. 总计14852个样本, 13个特征列.\n\n* <div class=\"特征列分析\"> \n    <b>特征列分析</b><br>\n    训练集中, 每条样本的特征为</p>\n    <code>primary_label</code> - (主要的鸟鸣), <code>secondary_label</code> - (次要鸟鸣);</p>\n    <code>type</code> - (鸣叫类型/内容);</p>\n    <code>latitude</code>, <code>longitude</code> - (录音的经纬度);</p>\n    <code>scientific_name</code>, <code>common_name</code> - (鸟类的学名和常用名); <code>author</code> - (录音作者);</p>\n    <code>license</code>, <code>rating</code>, <code>time</code>, <code>url</code> - (许可证, 录音等级, 录音时间, 网址);</p>\n    <code>filename</code> - (文件名);</p>\n</div>","metadata":{}},{"cell_type":"code","source":"train_csv=pd.read_csv(\"../input/birdclef-2022/train_metadata.csv\")\ntrain_csv.head(3)\n#metadata = train_csv['primary_label'].reset_index().explode(\"primary_label\")\n","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:32.293139Z","iopub.execute_input":"2022-04-24T14:15:32.293485Z","iopub.status.idle":"2022-04-24T14:15:32.370200Z","shell.execute_reply.started":"2022-04-24T14:15:32.293455Z","shell.execute_reply":"2022-04-24T14:15:32.369258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 test.csv\n* 测试集表示test.csv的统计信息, 共3条, 其余在隐藏测试集中.","metadata":{}},{"cell_type":"code","source":"test_csv = pd.read_csv(\"../input/birdclef-2022/test.csv\")\ntest_csv","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:32.396324Z","iopub.execute_input":"2022-04-24T14:15:32.396701Z","iopub.status.idle":"2022-04-24T14:15:32.410064Z","shell.execute_reply.started":"2022-04-24T14:15:32.396677Z","shell.execute_reply":"2022-04-24T14:15:32.408805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 eBird_Taxonomy_v2021.csv\n* 是样本鸟类的分类文件</p>\n* 以African SilverBill为例\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"../input/birdclef-2022/eBird_Taxonomy_v2021.csv\")\ndf[df['PRIMARY_COM_NAME'].isin(['African Silverbill'])]","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:32.496845Z","iopub.execute_input":"2022-04-24T14:15:32.497099Z","iopub.status.idle":"2022-04-24T14:15:32.557593Z","shell.execute_reply.started":"2022-04-24T14:15:32.497077Z","shell.execute_reply":"2022-04-24T14:15:32.557011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* TAXON_ORDER - 30031 分类编号\n* CATEGORY - species 物种\n* SCI_NAME - Euodice cantans https://es.wikipedia.org/wiki/Euodice_cantans\n* ORDER1 - Passeriformes 雀形目\n* FAMILY - Estrildidae 梅花雀 ","metadata":{}},{"cell_type":"markdown","source":"## 2.4 scored_bird.json\n#### 测试集会出现的小鸟种类","metadata":{}},{"cell_type":"code","source":"scored_bird = [\"akiapo\", \"aniani\", \"apapan\", \"barpet\", \"crehon\", \"elepai\", \"ercfra\",\n               \"hawama\", \"hawcre\", \"hawgoo\", \"hawhaw\", \"hawpet1\", \"houfin\", \"iiwi\",\n               \"jabwar\", \"maupar\", \"omao\", \"puaioh\", \"skylar\", \"warwhe1\", \"yefcan\"]\n#a=['aniani']\n# 如果要分析全部小鸟 请注释掉下面这行代码\ntrain_csv = train_csv[train_csv[\"primary_label\"].isin(scored_bird)]\n\n#print(train_csv)","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:32.574861Z","iopub.execute_input":"2022-04-24T14:15:32.575290Z","iopub.status.idle":"2022-04-24T14:15:32.583753Z","shell.execute_reply.started":"2022-04-24T14:15:32.575251Z","shell.execute_reply":"2022-04-24T14:15:32.582496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.1 来看看歌曲内容吧 🎼\n\n","metadata":{}},{"cell_type":"markdown","source":"**Type Column**:\n\n> 📌**鸣叫类型/内容**: 这本身就是一个比较难分类的特征:\n* **alarm call** is: alarm call | alarm call, call \n* **flight call** is: flight call | call, flight call etc.","metadata":{}},{"cell_type":"code","source":"# Create a new variable type by exploding all the values\nadjusted_type = train_csv['type'].apply(lambda x: x[1:-1].split(',')).reset_index().explode(\"type\")\n\n# Strip of white spaces and convert to lower chars\nadjusted_type = adjusted_type['type'].apply(lambda x: x.strip().lower()).reset_index()\n#adjusted_type['type'] = adjusted_type['type'].replace('calls', 'call')\n\n# Create Top 15 list with song types\ntop_15 = list(adjusted_type['type'].value_counts().head(15).reset_index()['index'])\ndata = adjusted_type[adjusted_type['type'].isin(top_15)]\n\n# === PLOT ===\n\nplt.figure(figsize=(16, 6))\nax = sns.countplot(data['type'], palette=\"hls\", order = data['type'].value_counts().index)\n\nplt.title(\"Top 15 Song Types\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:32.656973Z","iopub.execute_input":"2022-04-24T14:15:32.657374Z","iopub.status.idle":"2022-04-24T14:15:32.893267Z","shell.execute_reply.started":"2022-04-24T14:15:32.657350Z","shell.execute_reply":"2022-04-24T14:15:32.892199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 按物种的音频样本统计 \n","metadata":{}},{"cell_type":"code","source":"# Create a new variable type by exploding all the values\nmetadata = train_csv['primary_label'].reset_index().explode(\"primary_label\")\n#metadata\n# Create Top 15 list with species types\ntop_15 = list(metadata['primary_label'].value_counts().head(8).reset_index()['index'])\n#print(top_15)\ndata = metadata[metadata['primary_label'].isin(top_15)]#这是找出了top15所在的行\n\nprint(data['primary_label'].value_counts())\n#print(data)\n# === PLOT ===\n#这个画图有必要吗 看看大致的数据趋势吧\n#print(len(metadata[\"primary_label\"].value_counts()))\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = metadata['primary_label'], palette=\"hls\", order = metadata['primary_label'].value_counts().index)\n\nplt.title(\"Label Counts = \"+ str(len(metadata[\"primary_label\"].value_counts())), fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks()\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:32.894965Z","iopub.execute_input":"2022-04-24T14:15:32.895312Z","iopub.status.idle":"2022-04-24T14:15:33.183624Z","shell.execute_reply.started":"2022-04-24T14:15:32.895280Z","shell.execute_reply":"2022-04-24T14:15:33.182341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 当然 我们往往关心那些较少的数据\n#### 我们需要对这几个缺少的数据做增强吗?","metadata":{}},{"cell_type":"code","source":"# Create Top 15 list with species types\nleast_15 = list(metadata['primary_label'].value_counts().tail(15).reset_index()['index'])\n#print(least_15)\ndata = metadata[metadata['primary_label'].isin(least_15)]#这是找出了least15所在的行\nprint(data['primary_label'].value_counts())\n# === PLOT ===\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = data['primary_label'], palette=\"hls\", order = data['primary_label'].value_counts().index)\n\nplt.title(\"Data Counts\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:33.185764Z","iopub.execute_input":"2022-04-24T14:15:33.186099Z","iopub.status.idle":"2022-04-24T14:15:33.431606Z","shell.execute_reply.started":"2022-04-24T14:15:33.186070Z","shell.execute_reply":"2022-04-24T14:15:33.430478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 音频质量","metadata":{}},{"cell_type":"code","source":"# Create Top 15 list with species types\ndata_rating = train_csv.loc[:,[\"primary_label\",\"rating\"]].reset_index().explode(\"rating\")\n#print(data_rating)\nleast_15 = list(data_rating['rating'].value_counts().tail(15).reset_index()['index'])\n#print(top_15)\ndata = data_rating[data_rating['rating'].isin(least_15)]#这是找出了least15所在的行\n#print(data_rating.loc[data_rating['rating']==0, :])\n# === PLOT ===\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = data_rating['rating'], palette=\"hls\")\n\nplt.title(\"Data Rating Counts\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:33.433347Z","iopub.execute_input":"2022-04-24T14:15:33.433620Z","iopub.status.idle":"2022-04-24T14:15:33.635791Z","shell.execute_reply.started":"2022-04-24T14:15:33.433592Z","shell.execute_reply":"2022-04-24T14:15:33.634892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 录音时间","metadata":{}},{"cell_type":"code","source":"# Create Top 15 list with species types\ndata_rating = train_csv['time'].reset_index().explode('time')\nfreq = '60min'\n#print(len(data_rating))\ndata_rating['time']= pd.to_datetime(data_rating['time'],errors = 'coerce')\ndata_rating['time'] = data_rating['time'].dt.floor(freq)\ndata_rating['time'] = data_rating['time'].dt.hour\n#print(len(data_rating[\"time\"]))\n#print(data_rating)\n#least_15 = list(data_rating['time'].value_counts().tail(15).reset_index()['index'])\n#print(top_15)\n#data = data_rating[data_rating['time'].isin(least_15)]#这是找出了least15所在的行\n#print(data_rating.loc[data_rating['time']==0, :])\n# === PLOT ===\nplt.figure(figsize=(16, 6))\nax = sns.countplot(x = data_rating['time'], palette=\"hls\")\n\nplt.title(\"Data Rating Counts\", fontsize=16)\nplt.ylabel(\"Counts\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks()\nplt.xlabel(\"hours\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:33.637216Z","iopub.execute_input":"2022-04-24T14:15:33.637480Z","iopub.status.idle":"2022-04-24T14:15:33.954568Z","shell.execute_reply.started":"2022-04-24T14:15:33.637453Z","shell.execute_reply":"2022-04-24T14:15:33.953607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 研究一下你的小鸟? 📸🔭\n\n","metadata":{}},{"cell_type":"markdown","source":"## World View of the Species 🧭🌏\n\n","metadata":{}},{"cell_type":"markdown","source":"### 2.3.1 Where are our birds? 🦜\n#### 2.3.1.1 完整数据集 按目分类 总计17个目","metadata":{}},{"cell_type":"code","source":"# SHP file\nworld_map = gpd.read_file(\"../input/world-shapefile/world_shapefile.shp\")\n\n# Coordinate reference system\ncrs = {\"init\" : \"epsg:4326\"}\ntrain_species_csv=pd.read_csv(\"../input//2022-birdcleftrain-data-with-species//train_metadata_with_species.csv\")\n#请删除接下来这一行\n# 就是这一行 train_species_csv= train_species_csv[train_species_csv[\"primary_label\"].isin(scored_bird)]\n\n# Lat and Long need to be of type float, not object\ndata = train_species_csv[train_species_csv[\"latitude\"] != \"Not specified\"]\ndata[\"latitude\"] = data[\"latitude\"].astype(float)\ndata[\"longitude\"] = data[\"longitude\"].astype(float)\n\n# Create geometry\ngeometry = [Point(xy) for xy in zip(data[\"longitude\"], data[\"latitude\"])]\n\n# Geo Dataframe\ngeo_df = gpd.GeoDataFrame(data, crs=crs, geometry=geometry)\n\n# Create ID for species\nspecies_id = geo_df[\"order1\"].value_counts().reset_index()\nspecies_id.insert(0, 'ID', range(0, 0 + len(species_id)))\n\nspecies_id.columns = [\"ID\", \"order1\", \"count\"]\nprint(species_id)\n# Add ID to geo_df\ngeo_df = pd.merge(geo_df, species_id, how=\"left\", on=\"order1\")\n\n# === PLOT ===\nfig, ax = plt.subplots(figsize = (16, 10))\nworld_map.plot(ax=ax, alpha=0.4, color=\"grey\")\n\npalette = iter(sns.hls_palette(len(species_id)))\n\nfor i in range(17):\n    geo_df[geo_df[\"ID\"] == i].plot(ax=ax, markersize=20, color=next(palette), marker=\"o\", label = \"test\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:33.955659Z","iopub.execute_input":"2022-04-24T14:15:33.955998Z","iopub.status.idle":"2022-04-24T14:15:37.828351Z","shell.execute_reply.started":"2022-04-24T14:15:33.955971Z","shell.execute_reply":"2022-04-24T14:15:37.827045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2.3.1.2 完整数据集 按科分类 总计41个科","metadata":{}},{"cell_type":"code","source":"# SHP file\nworld_map = gpd.read_file(\"../input/world-shapefile/world_shapefile.shp\")\n\n# Coordinate reference system\ncrs = {\"init\" : \"epsg:4326\"}\n#train_species_csv=pd.read_csv(\"../input//2022-birdcleftrain-data-with-species//train_metadata_with_species.csv\")\n# Lat and Long need to be of type float, not object\ndata = train_species_csv[train_species_csv[\"latitude\"] != \"Not specified\"]\ndata[\"latitude\"] = data[\"latitude\"].astype(float)\ndata[\"longitude\"] = data[\"longitude\"].astype(float)\n\n# Create geometry\ngeometry = [Point(xy) for xy in zip(data[\"longitude\"], data[\"latitude\"])]\n\n# Geo Dataframe\ngeo_df = gpd.GeoDataFrame(data, crs=crs, geometry=geometry)\n\n# Create ID for species\nspecies_id = geo_df[\"family\"].value_counts().reset_index()\nspecies_id.insert(0, 'ID', range(0, 0 + len(species_id)))\n\nspecies_id.columns = [\"ID\", \"family\", \"count\"]\nprint(species_id)\n# Add ID to geo_df\ngeo_df = pd.merge(geo_df, species_id, how=\"left\", on=\"family\")\n\n# === PLOT ===\nfig, ax = plt.subplots(figsize = (16, 10))\nworld_map.plot(ax=ax, alpha=0.4, color=\"grey\")\n\npalette = iter(sns.hls_palette(len(species_id)))\n\nfor i in range(41):\n    geo_df[geo_df[\"ID\"] == i].plot(ax=ax, markersize=20, color=next(palette), marker=\"o\", label = \"test\");","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:37.830351Z","iopub.execute_input":"2022-04-24T14:15:37.830649Z","iopub.status.idle":"2022-04-24T14:15:42.494859Z","shell.execute_reply.started":"2022-04-24T14:15:37.830621Z","shell.execute_reply":"2022-04-24T14:15:42.493263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.3 Listening to some Recordings\n\n","metadata":{}},{"cell_type":"markdown","source":"### Ok, let's hear some songs! 🕊🎶","metadata":{}},{"cell_type":"code","source":"# 混音\nipd.Audio(\"../input/join-voice-of-birdclef/joinVoice.ogg\")","metadata":{"execution":{"iopub.status.busy":"2022-04-24T14:15:42.496378Z","iopub.execute_input":"2022-04-24T14:15:42.496668Z","iopub.status.idle":"2022-04-24T14:15:42.718026Z","shell.execute_reply.started":"2022-04-24T14:15:42.496639Z","shell.execute_reply":"2022-04-24T14:15:42.715777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Work in Progress ... ⏳","metadata":{}}]}