{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":19596,"databundleVersionId":1292430,"sourceType":"competition"},{"sourceId":236698,"sourceType":"datasetVersion","datasetId":100139},{"sourceId":1308479,"sourceType":"datasetVersion","datasetId":741657},{"sourceId":1487019,"sourceType":"datasetVersion","datasetId":726237},{"sourceId":1487116,"sourceType":"datasetVersion","datasetId":726312}],"dockerImageVersionId":29962,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🐦 Bird Call Recognition Project\n\nThis project focuses on analyzing bird call audio data to recognize bird species from sound recordings. The dataset includes training and test audio samples accompanied by metadata. Using libraries like `librosa` for audio processing, `matplotlib` for visualization, and `geopandas` for geospatial data, we aim to understand, visualize, and eventually classify bird calls.\n\n---\n\n### 🎯 Objective\n\nThe goal is to prepare audio data, explore its properties, and lay the foundation for building a machine learning model that can recognize multiple bird species in audio clips.\n\n---\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import os\n\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\n\n# Map 1 library\nimport plotly.express as px\n\n# Map 2 libraries\nimport descartes\nimport geopandas as gpd\nfrom shapely.geometry import Point, Polygon\n\n# Librosa Libraries\nimport librosa\nimport librosa.display\nimport IPython.display as ipd\n\nimport sklearn\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:36:34.793458Z","iopub.execute_input":"2025-04-29T21:36:34.793765Z","iopub.status.idle":"2025-04-29T21:36:34.805055Z","shell.execute_reply.started":"2025-04-29T21:36:34.793733Z","shell.execute_reply":"2025-04-29T21:36:34.804206Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import plotly.express as px","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:51:16.519997Z","iopub.execute_input":"2025-04-29T21:51:16.520321Z","iopub.status.idle":"2025-04-29T21:51:16.527647Z","shell.execute_reply.started":"2025-04-29T21:51:16.520293Z","shell.execute_reply":"2025-04-29T21:51:16.526629Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. Loading CSV Data & Extracting Date Features 📁\n\n> 📌**Note**:\n* `train.csv` contains information about the audio files available in `train_audio`. It contains 21,375 datapoints in 35 unique columns.\n* `test.csv` contains only 3 observations (the rest are available in the *hidden test set*).\n* To enable temporal analysis, we extract the `year`, `month`, and `day_of_month` from the `date` column in `train.csv`. Finally, we print the number of unique bird species in the dataset.\n\n<div class=\"alert alert-block alert-info\">\n<b>Note:</b> The TRAIN data has 1 labeled bird species per recording. However, in nature usually you can hear tens (even hundreds) of birds in one go, so in TEST set we need to predict 0, 1 or multiple species for one recording. Because of this, in TRAIN we have <code>species</code> column - or <code>primary_label</code> - (main bird), <code>secondary label</code> (other birds heard) and <code>background</code> (background noises, other birds etc.)\n</div>\n \n### Discussions💬\n* [Few questions about test data](https://www.kaggle.com/c/birdsong-recognition/discussion/159123)\n* [Confusion about test set](https://www.kaggle.com/c/birdsong-recognition/discussion/158987)\n\n","metadata":{}},{"cell_type":"code","source":"# Import data\ntrain_csv = pd.read_csv(\"../input/birdsong-recognition/train.csv\")\ntest_csv = pd.read_csv(\"../input/birdsong-recognition/test.csv\")\n\n# Create some time features\ntrain_csv['year'] = train_csv['date'].apply(lambda x: x.split('-')[0])\ntrain_csv['month'] = train_csv['date'].apply(lambda x: x.split('-')[1])\ntrain_csv['day_of_month'] = train_csv['date'].apply(lambda x: x.split('-')[2])\n\nprint(\"There are {:,} unique bird species in the dataset.\".format(len(train_csv['species'].unique())))","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:37:10.63873Z","iopub.execute_input":"2025-04-29T21:37:10.639037Z","iopub.status.idle":"2025-04-29T21:37:11.064089Z","shell.execute_reply.started":"2025-04-29T21:37:10.63901Z","shell.execute_reply":"2025-04-29T21:37:11.063312Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔍 Previewing Test CSV\n\nWe display the contents of `test_csv` to understand the structure and available metadata of the test set, which contains only a few labeled rows.\n\n### TEST.csv - let's take a look here as well before going further\n\n> 📌**Note**:\n* only 3 rows available (rest are in the hidden set)\n* `site`: there are 3 sites in total, with first 2 having labeles every 5 seconds, while site_3 has labels at file level.\n* `row_id`: this is the unique ID that will be used for the submission\n* `seconds`: how long the clip is\n* `audio_id`: `row_id` without site\n\n*PS: \"nocall\" can be also one of the labels (hearing no bird).*","metadata":{}},{"cell_type":"code","source":"# Inspect text_csv before checking train data\ntest_csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:39:01.112186Z","iopub.execute_input":"2025-04-29T21:39:01.112485Z","iopub.status.idle":"2025-04-29T21:39:01.131178Z","shell.execute_reply.started":"2025-04-29T21:39:01.11246Z","shell.execute_reply":"2025-04-29T21:39:01.130413Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Time of the Recording ⏰\n\n> 📌**Note**: Majority of the data was registered between 2013 and 2019, during Spring and Summer months (`00` is for the dates 0000-00-00, which are unknown).\n","metadata":{}},{"cell_type":"markdown","source":"## 📈 Audio File Registrations per Year\n\nThis visualization shows how audio recordings are distributed across different years. An image of a bird is annotated on the plot to enhance presentation.\n","metadata":{}},{"cell_type":"code","source":"\nab = AnnotationBbox(imagebox, xy, frameon=False, pad=1, xybox=(6.5, 2000))\n\nplt.figure(figsize=(20, 6))\nax = sns.countplot(train_csv['year'], palette=\"Blues_d\")\n\nplt.title(\"Audio Files Registration per Year Made\", fontsize=16)\nplt.xticks(rotation=90, fontsize=13)\nplt.yticks(fontsize=13)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:44:46.059069Z","iopub.execute_input":"2025-04-29T21:44:46.059482Z","iopub.status.idle":"2025-04-29T21:44:46.411587Z","shell.execute_reply.started":"2025-04-29T21:44:46.059435Z","shell.execute_reply":"2025-04-29T21:44:46.41087Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nplt.figure(figsize=(20, 6))\nax = sns.countplot(train_csv['month'], palette=\"Blues_d\")\n\n\nplt.title(\"Audio Files Registration per Month Made\", fontsize=16)\nplt.xticks(fontsize=13)\nplt.yticks(fontsize=13)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:46:31.159237Z","iopub.execute_input":"2025-04-29T21:46:31.159528Z","iopub.status.idle":"2025-04-29T21:46:31.350863Z","shell.execute_reply.started":"2025-04-29T21:46:31.159502Z","shell.execute_reply":"2025-04-29T21:46:31.349884Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.2 The Songs 🎼\n\n> 📌**Note**: Pitch is usually unspecified. This is one of the more *miscellaneous* columns, that we need to be careful how we interpret. Most Song Types are *call, song or flight*.","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize=(20, 6))\nax = sns.countplot(train_csv['pitch'], palette=\"hls\", order = train_csv['pitch'].value_counts().index)\n\nplt.title(\"Pitch (quality of sound - how high/low the tone is)\", fontsize=16)\nplt.xticks(fontsize=13)\nplt.yticks(fontsize=13)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:47:04.491499Z","iopub.execute_input":"2025-04-29T21:47:04.49179Z","iopub.status.idle":"2025-04-29T21:47:04.628589Z","shell.execute_reply.started":"2025-04-29T21:47:04.491763Z","shell.execute_reply":"2025-04-29T21:47:04.627764Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Type Column**:\n\n> 📌**Note**: This column is a bit messy, as the same description can be found under multiple names. Also, there can be multiple descriptions for multiple sounds (one bird song can mean a different thing from another one in the same recording). Some examples are:\n* **alarm call** is: alarm call | alarm call, call \n* **flight call** is: flight call | call, flight call etc.","metadata":{}},{"cell_type":"code","source":"# Create a new variable type by exploding all the values\nadjusted_type = train_csv['type'].apply(lambda x: x.split(',')).reset_index().explode(\"type\")\n\n# Strip of white spaces and convert to lower chars\nadjusted_type = adjusted_type['type'].apply(lambda x: x.strip().lower()).reset_index()\nadjusted_type['type'] = adjusted_type['type'].replace('calls', 'call')\n\n# Create Top 15 list with song types\ntop_15 = list(adjusted_type['type'].value_counts().head(15).reset_index()['index'])\ndata = adjusted_type[adjusted_type['type'].isin(top_15)]\n\n# === PLOT ===\n\nplt.figure(figsize=(20, 6))\nax = sns.countplot(data['type'], palette=\"hls\", order = data['type'].value_counts().index)\n\nplt.title(\"Top 15 Song Types\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:47:42.844763Z","iopub.execute_input":"2025-04-29T21:47:42.845125Z","iopub.status.idle":"2025-04-29T21:47:43.280239Z","shell.execute_reply.started":"2025-04-29T21:47:42.845091Z","shell.execute_reply":"2025-04-29T21:47:43.279325Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.3 Where is the bird? 📸🔭\n\n> 📌**Note**: In most recordings the birds were seen, usually at an altitude between 0m and 10m.","metadata":{}},{"cell_type":"code","source":"# Top 15 most common elevations\ntop_15 = list(train_csv['elevation'].value_counts().head(15).reset_index()['index'])\ndata = train_csv[train_csv['elevation'].isin(top_15)]\n\n# === PLOT ===\n\nplt.figure(figsize=(20, 6))\nax = sns.countplot(data['elevation'], palette=\"Blues_d\", order = data['elevation'].value_counts().index)\n\nplt.title(\"Top 15 Elevation Types\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:48:20.770783Z","iopub.execute_input":"2025-04-29T21:48:20.771221Z","iopub.status.idle":"2025-04-29T21:48:20.97723Z","shell.execute_reply.started":"2025-04-29T21:48:20.771181Z","shell.execute_reply":"2025-04-29T21:48:20.976505Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create data\ndata = train_csv['bird_seen'].value_counts().reset_index()\n\n# === PLOT ===\n\nplt.figure(figsize=(20, 6))\nax = sns.barplot(x = 'bird_seen', y = 'index', data = data, palette=\"hls\")\n\nplt.title(\"Song was Heard, but was Bird Seen?\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:48:55.79414Z","iopub.execute_input":"2025-04-29T21:48:55.79443Z","iopub.status.idle":"2025-04-29T21:48:55.922689Z","shell.execute_reply.started":"2025-04-29T21:48:55.794405Z","shell.execute_reply":"2025-04-29T21:48:55.92192Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.4 World View of the Species 🧭🌏\n\n### #1. Countries 🚩\n\n> 📌**Note**: Let's look at **top 15** countries with most recordings. The majority of recordings are located in the US, followed by Canada and Mexico.","metadata":{}},{"cell_type":"code","source":"# Top 15 most common elevations\ntop_15 = list(train_csv['country'].value_counts().head(15).reset_index()['index'])\ndata = train_csv[train_csv['country'].isin(top_15)]\n\n# === PLOT ===\n\nplt.figure(figsize=(20, 6))\nax = sns.countplot(data['country'], palette='hls', order = data['country'].value_counts().index)\n\nplt.title(\"Top 15 Countries with most Recordings\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T21:49:29.689288Z","iopub.execute_input":"2025-04-29T21:49:29.689587Z","iopub.status.idle":"2025-04-29T21:49:29.906437Z","shell.execute_reply.started":"2025-04-29T21:49:29.689562Z","shell.execute_reply":"2025-04-29T21:49:29.90553Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #2. Map View 🧭","metadata":{}},{"cell_type":"code","source":"# Import necessary libraries\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.colors as colors\nfrom matplotlib.colors import LinearSegmentedColormap\nimport matplotlib.patches as mpatches\n\n# Load the plot utils (this is built into Kaggle)\nfrom urllib.request import urlopen\nimport json\nwith urlopen('https://raw.githubusercontent.com/plotly/datasets/master/geojson-counties-fips.json') as response:\n    counties = json.load(response)\n\n# Load world geojson data\nworld_geojson_url = 'https://raw.githubusercontent.com/python-visualization/folium/master/examples/data/world-countries.json'\nwith urlopen(world_geojson_url) as response:\n    world_geojson = json.load(response)\n\n# Create a dictionary mapping country names to ISO codes\ncountry_codes = {}\nfor feature in world_geojson['features']:\n    country_codes[feature['properties']['name']] = feature['id']\n\n# Process your data\ndata = pd.merge(left=train_csv, right=df, how=\"inner\", on=\"country\")\ndata = data.groupby(by=[\"country\", \"iso_alpha\"]).count()[\"species\"].reset_index()\n\n# Prepare the visualization data\nworld_data = {}\nfor _, row in data.iterrows():\n    country = row['country']\n    species_count = row['species']\n    iso_code = row['iso_alpha']\n    world_data[iso_code] = species_count\n\n# Create a figure\nplt.figure(figsize=(16, 10))\n\n# Create a custom colormap (teal-like)\ncmap = plt.cm.Blues\n\n# Set up the map\nfrom mpl_toolkits.basemap import Basemap\nm = Basemap(projection='robin', resolution='l', area_thresh=1000.0, lat_0=0, lon_0=0)\nm.drawcoastlines()\nm.drawcountries()\nm.fillcontinents(color='lightgray', lake_color='white')\nm.drawmapboundary(fill_color='white')\n\n# Color countries based on data\nmax_value = max(world_data.values()) if world_data else 1\nfor feature in world_geojson['features']:\n    iso_code = feature['id']\n    if iso_code in world_data:\n        # Get the geometry and the country name\n        geometry = feature['geometry']\n        \n        # Skip countries with invalid geometry\n        if geometry['type'] != 'Polygon' and geometry['type'] != 'MultiPolygon':\n            continue\n            \n        # Set color based on species count\n        count = world_data[iso_code]\n        color_intensity = count / max_value\n        country_color = cmap(color_intensity)\n        \n        # Draw polygons for the country\n        if geometry['type'] == 'Polygon':\n            coords = geometry['coordinates'][0]\n            xs = [coord[0] for coord in coords]\n            ys = [coord[1] for coord in coords]\n            m.plot(xs, ys, 'k-', linewidth=0.5)\n            m.fillcontinents(color=country_color)\n        elif geometry['type'] == 'MultiPolygon':\n            for polygon in geometry['coordinates']:\n                coords = polygon[0]\n                xs = [coord[0] for coord in coords]\n                ys = [coord[1] for coord in coords]\n                m.plot(xs, ys, 'k-', linewidth=0.5)\n                m.fillcontinents(color=country_color)\n\n# Add a colorbar\nsm = plt.cm.ScalarMappable(cmap=cmap, norm=plt.Normalize(0, max_value))\nsm._A = []  # Fake up the array of the scalar mappable\ncbar = plt.colorbar(sm, orientation='horizontal', fraction=0.046, pad=0.04)\ncbar.set_label('Number of Species')\n\n# Set title\nplt.title('World Map: Recordings per Country', fontsize=16)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:02:00.752104Z","iopub.execute_input":"2025-04-29T22:02:00.752431Z","iopub.status.idle":"2025-04-29T22:04:48.760154Z","shell.execute_reply.started":"2025-04-29T22:02:00.752398Z","shell.execute_reply":"2025-04-29T22:04:48.759464Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Columns in train_csv:\", train_csv.columns.tolist())\nprint(\"Columns in gapminder:\", px.data.gapminder().columns.tolist())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:05:16.437991Z","iopub.execute_input":"2025-04-29T22:05:16.438279Z","iopub.status.idle":"2025-04-29T22:05:16.451241Z","shell.execute_reply.started":"2025-04-29T22:05:16.438255Z","shell.execute_reply":"2025-04-29T22:05:16.450426Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #3. Another Map! 😄 Where are our birds? 🦜","metadata":{}},{"cell_type":"code","source":"# SHP file\nworld_map = gpd.read_file(\"../input/world-shapefile/world_shapefile.shp\")\n\n# Coordinate reference system\ncrs = {\"init\" : \"epsg:4326\"}\n\n# Lat and Long need to be of type float, not object\ndata = train_csv[train_csv[\"latitude\"] != \"Not specified\"]\ndata[\"latitude\"] = data[\"latitude\"].astype(float)\ndata[\"longitude\"] = data[\"longitude\"].astype(float)\n\n# Create geometry\ngeometry = [Point(xy) for xy in zip(data[\"longitude\"], data[\"latitude\"])]\n\n# Geo Dataframe\ngeo_df = gpd.GeoDataFrame(data, crs=crs, geometry=geometry)\n\n# Create ID for species\nspecies_id = geo_df[\"species\"].value_counts().reset_index()\nspecies_id.insert(0, 'ID', range(0, 0 + len(species_id)))\n\nspecies_id.columns = [\"ID\", \"species\", \"count\"]\n\n# Add ID to geo_df\ngeo_df = pd.merge(geo_df, species_id, how=\"left\", on=\"species\")\n\n# === PLOT ===\nfig, ax = plt.subplots(figsize = (16, 10))\nworld_map.plot(ax=ax, alpha=0.4, color=\"grey\")\n\npalette = iter(sns.hls_palette(len(species_id)))\n\nfor i in range(264):\n    geo_df[geo_df[\"ID\"] == i].plot(ax=ax, markersize=20, color=next(palette), marker=\"o\", label = \"test\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:05:19.992662Z","iopub.execute_input":"2025-04-29T22:05:19.992979Z","iopub.status.idle":"2025-04-29T22:05:41.143227Z","shell.execute_reply.started":"2025-04-29T22:05:19.992946Z","shell.execute_reply":"2025-04-29T22:05:41.142514Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. The Audio Files🔈🔉🔊\n\n## 3.1 Description\n\n> 📌**Note**: \n* **train_audio**: short recording (majority in mp3 format) of INDIVIDUAL birds.\n* **test_audio**: recordings took in 3 locations:\n    * Site 1 and Site 2: recordings 10 mins long (mp3) that have labeled a bird every 5 seconds. This is meant to mimic the *real life scenario*, when you would usually have more than 1 bird (or no bird) singing.\n    * Site 3: recordings labeled at file level (because it is especially hard to have coders trained to label these kind of files)\n    \n## 3.2 Duration and File Types 📁","metadata":{}},{"cell_type":"code","source":"# Creating Interval for *duration* variable\ntrain_csv['duration_interval'] = \">500\"\ntrain_csv.loc[train_csv['duration'] <= 100, 'duration_interval'] = \"<=100\"\ntrain_csv.loc[(train_csv['duration'] > 100) & (train_csv['duration'] <= 200), 'duration_interval'] = \"100-200\"\ntrain_csv.loc[(train_csv['duration'] > 200) & (train_csv['duration'] <= 300), 'duration_interval'] = \"200-300\"\ntrain_csv.loc[(train_csv['duration'] > 300) & (train_csv['duration'] <= 400), 'duration_interval'] = \"300-400\"\ntrain_csv.loc[(train_csv['duration'] > 400) & (train_csv['duration'] <= 500), 'duration_interval'] = \"400-500\"\n\n\nplt.figure(figsize=(20, 6))\nax = sns.countplot(train_csv['duration_interval'], palette=\"hls\")\n\nplt.title(\"Distribution of Recordings Duration\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:08:41.178872Z","iopub.execute_input":"2025-04-29T22:08:41.179178Z","iopub.status.idle":"2025-04-29T22:08:41.341042Z","shell.execute_reply.started":"2025-04-29T22:08:41.179153Z","shell.execute_reply":"2025-04-29T22:08:41.340197Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def show_values_on_bars(axs, h_v=\"v\", space=0.4):\n    def _show_on_single_plot(ax):\n        if h_v == \"v\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() / 2\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_height())\n                ax.text(_x, _y, value, ha=\"center\") \n        elif h_v == \"h\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() + float(space)\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_width())\n                ax.text(_x, _y, value, ha=\"left\")\n\n    if isinstance(axs, np.ndarray):\n        for idx, ax in np.ndenumerate(axs):\n            _show_on_single_plot(ax)\n    else:\n        _show_on_single_plot(axs)","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:08:48.066763Z","iopub.execute_input":"2025-04-29T22:08:48.067064Z","iopub.status.idle":"2025-04-29T22:08:48.074505Z","shell.execute_reply.started":"2025-04-29T22:08:48.067039Z","shell.execute_reply":"2025-04-29T22:08:48.073598Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# bird = mpimg.imread('../input/birdcall-recognition-data/yellow birds.jpg')\nimagebox = OffsetImage(bird, zoom=0.6)\nxy = (0.5, 0.7)\nab = AnnotationBbox(imagebox, xy, frameon=False, pad=1, xybox=(2.7, 12000))\n\nplt.figure(figsize=(16, 6))\nax = sns.countplot(train_csv['file_type'], palette = \"hls\", order = train_csv['file_type'].value_counts().index)\nax.add_artist(ab)\n\nshow_values_on_bars(ax, \"v\", 0)\n\nplt.title(\"Recording File Types\", fontsize=16)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(rotation=45, fontsize=13)\nplt.xlabel(\"\");","metadata":{"execution":{"iopub.status.busy":"2025-04-02T22:41:48.87202Z","iopub.execute_input":"2025-04-02T22:41:48.872238Z","iopub.status.idle":"2025-04-02T22:41:49.052924Z","shell.execute_reply.started":"2025-04-02T22:41:48.872217Z","shell.execute_reply":"2025-04-02T22:41:49.052013Z"}}},{"cell_type":"markdown","source":"## 3.3 Listening to some Recordings\n\n> 📌**Note**: What is sound? [In physics, sound is a vibration that propagates as an acoustic wave, through a transmission medium such as a gas, liquid or solid.](https://en.wikipedia.org/wiki/Sound)\n<img src=\"https://i.imgur.com/jic0QY3.png\" width=400>","metadata":{}},{"cell_type":"code","source":"# Create Full Path so we can access data more easily\nbase_dir = '../input/birdsong-recognition/train_audio/'\ntrain_csv['full_path'] = base_dir + train_csv['ebird_code'] + '/' + train_csv['filename']\n\n# Now let's sample a fiew audio files\namered = train_csv[train_csv['ebird_code'] == \"amered\"].sample(1, random_state = 33)['full_path'].values[0]\ncangoo = train_csv[train_csv['ebird_code'] == \"cangoo\"].sample(1, random_state = 33)['full_path'].values[0]\nhaiwoo = train_csv[train_csv['ebird_code'] == \"haiwoo\"].sample(1, random_state = 33)['full_path'].values[0]\npingro = train_csv[train_csv['ebird_code'] == \"pingro\"].sample(1, random_state = 33)['full_path'].values[0]\nvesspa = train_csv[train_csv['ebird_code'] == \"vesspa\"].sample(1, random_state = 33)['full_path'].values[0]\n\nbird_sample_list = [\"amered\", \"cangoo\", \"haiwoo\", \"pingro\", \"vesspa\"]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:08:59.611986Z","iopub.execute_input":"2025-04-29T22:08:59.612299Z","iopub.status.idle":"2025-04-29T22:08:59.668396Z","shell.execute_reply.started":"2025-04-29T22:08:59.612274Z","shell.execute_reply":"2025-04-29T22:08:59.667637Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Ok, let's hear some songs! 🕊🎶","metadata":{}},{"cell_type":"code","source":"def viz_from_path(filepath):\n    # Load the audio file\n    data, sr = librosa.load(filepath, sr=SAMPLE_RATE)\n\n    signals = get_signal(data)\n    melsp_features = get_melsp_img(data, sr=sr)\n\n    fig, ax = plt.subplots(nrows=CHUNKS, ncols=2, figsize=(30, 5 * CHUNKS))\n\n    if CHUNKS == 1:\n        ax = np.array([ax])  # Convert to 2D array for consistent indexing\n\n    for i, (signal, melsp_feature) in enumerate(zip(signals, melsp_features)):\n        ax[i][0].plot(signal, color='crimson')\n        ax[i][0].set_title(\"Signal\", fontsize=16)\n\n        ax[i][1].imshow(cv2.resize(melsp_feature, (4096, 1024)), aspect='auto', origin='lower')\n        ax[i][1].set_title(\"Melspectrogram\", fontsize=16)\n\n    display(ipd.Audio(filepath))\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:20:51.831581Z","iopub.execute_input":"2025-04-29T22:20:51.831906Z","iopub.status.idle":"2025-04-29T22:20:51.840018Z","shell.execute_reply.started":"2025-04-29T22:20:51.831872Z","shell.execute_reply":"2025-04-29T22:20:51.839244Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"viz_from_path(amered)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:20:55.210288Z","iopub.execute_input":"2025-04-29T22:20:55.210563Z","iopub.status.idle":"2025-04-29T22:20:56.457113Z","shell.execute_reply.started":"2025-04-29T22:20:55.210539Z","shell.execute_reply":"2025-04-29T22:20:56.456229Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Cangoo\nviz_from_path(cangoo)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:21:18.720725Z","iopub.execute_input":"2025-04-29T22:21:18.721091Z","iopub.status.idle":"2025-04-29T22:21:21.661748Z","shell.execute_reply.started":"2025-04-29T22:21:18.721055Z","shell.execute_reply":"2025-04-29T22:21:21.660943Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Haiwoo\nviz_from_path(haiwoo)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:21:40.35761Z","iopub.execute_input":"2025-04-29T22:21:40.357958Z","iopub.status.idle":"2025-04-29T22:21:44.397949Z","shell.execute_reply.started":"2025-04-29T22:21:40.357921Z","shell.execute_reply":"2025-04-29T22:21:44.397024Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Pingro\nviz_from_path(pingro)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:21:56.317235Z","iopub.execute_input":"2025-04-29T22:21:56.317552Z","iopub.status.idle":"2025-04-29T22:22:01.911968Z","shell.execute_reply.started":"2025-04-29T22:21:56.31752Z","shell.execute_reply":"2025-04-29T22:22:01.910926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Vesspa\nviz_from_path(vesspa)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:22:15.700305Z","iopub.execute_input":"2025-04-29T22:22:15.700622Z","iopub.status.idle":"2025-04-29T22:22:19.891634Z","shell.execute_reply.started":"2025-04-29T22:22:15.700591Z","shell.execute_reply":"2025-04-29T22:22:19.890931Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.4 Extracting Features from Sounds 🔓\n\n> 📌The audio data is composed by:\n1. **Sound**: sequence of vibrations in varying pressure strengths (`y`)\n2. **Sample Rate**: (`sr`) is the number of samples of audio carried per second, measured in Hz or kHz","metadata":{}},{"cell_type":"code","source":"# Importing 1 file\ny, sr = librosa.load(vesspa)\n\nprint('y:', y, '\\n')\nprint('y shape:', np.shape(y), '\\n')\nprint('Sample Rate (KHz):', sr, '\\n')\n\n# Verify length of the audio\nprint('Check Len of Audio:', np.shape(y)[0]/sr)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:22:50.181896Z","iopub.execute_input":"2025-04-29T22:22:50.18221Z","iopub.status.idle":"2025-04-29T22:22:53.161356Z","shell.execute_reply.started":"2025-04-29T22:22:50.18218Z","shell.execute_reply":"2025-04-29T22:22:53.160501Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Trim leading and trailing silence from an audio signal (silence before and after the actual audio)\naudio_file, _ = librosa.effects.trim(y)\n\n# the result is an numpy ndarray\nprint('Audio File:', audio_file, '\\n')\nprint('Audio File shape:', np.shape(audio_file))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:23:45.288998Z","iopub.execute_input":"2025-04-29T22:23:45.289324Z","iopub.status.idle":"2025-04-29T22:23:45.310587Z","shell.execute_reply.started":"2025-04-29T22:23:45.28929Z","shell.execute_reply":"2025-04-29T22:23:45.309733Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Importing the 5 files\ny_amered, sr_amered = librosa.load(amered)\naudio_amered, _ = librosa.effects.trim(y_amered)\n\ny_cangoo, sr_cangoo = librosa.load(cangoo)\naudio_cangoo, _ = librosa.effects.trim(y_cangoo)\n\ny_haiwoo, sr_haiwoo = librosa.load(haiwoo)\naudio_haiwoo, _ = librosa.effects.trim(y_haiwoo)\n\ny_pingro, sr_pingro = librosa.load(pingro)\naudio_pingro, _ = librosa.effects.trim(y_pingro)\n\ny_vesspa, sr_vesspa = librosa.load(vesspa)\naudio_vesspa, _ = librosa.effects.trim(y_vesspa)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:23:49.747576Z","iopub.execute_input":"2025-04-29T22:23:49.747885Z","iopub.status.idle":"2025-04-29T22:24:02.061084Z","shell.execute_reply.started":"2025-04-29T22:23:49.747854Z","shell.execute_reply":"2025-04-29T22:24:02.060012Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #1. Sound Waves🌊 (2D Representation)","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(5, figsize = (16, 9))\nfig.suptitle('Sound Waves', fontsize=16)\n\nlibrosa.display.waveplot(y = audio_amered, sr = sr_amered, color = \"#A300F9\", ax=ax[0])\nlibrosa.display.waveplot(y = audio_cangoo, sr = sr_cangoo, color = \"#4300FF\", ax=ax[1])\nlibrosa.display.waveplot(y = audio_haiwoo, sr = sr_haiwoo, color = \"#009DFF\", ax=ax[2])\nlibrosa.display.waveplot(y = audio_pingro, sr = sr_pingro, color = \"#00FFB0\", ax=ax[3])\nlibrosa.display.waveplot(y = audio_vesspa, sr = sr_vesspa, color = \"#D9FF00\", ax=ax[4]);\n\nfor i, name in zip(range(5), bird_sample_list):\n    ax[i].set_ylabel(name, fontsize=13)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:02.063191Z","iopub.execute_input":"2025-04-29T22:24:02.063446Z","iopub.status.idle":"2025-04-29T22:24:02.86377Z","shell.execute_reply.started":"2025-04-29T22:24:02.063421Z","shell.execute_reply":"2025-04-29T22:24:02.862891Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #2. Fourier Transform 🥁\n\n> 📌**Note**: Function that gets a signal in the time domain as input, and outputs its decomposition into frequencies. Transform both the y-axis (frequency) to log scale, and the “color” axis (amplitude) to Decibels, which is approx. the log scale of amplitudes.","metadata":{}},{"cell_type":"code","source":"# Default FFT window size\nn_fft = 2048 # FFT window size\nhop_length = 512 # number audio of frames between STFT columns (looks like a good default)\n\n# Short-time Fourier transform (STFT)\nD_amered = np.abs(librosa.stft(audio_amered, n_fft = n_fft, hop_length = hop_length))\nD_cangoo = np.abs(librosa.stft(audio_cangoo, n_fft = n_fft, hop_length = hop_length))\nD_haiwoo = np.abs(librosa.stft(audio_haiwoo, n_fft = n_fft, hop_length = hop_length))\nD_pingro = np.abs(librosa.stft(audio_pingro, n_fft = n_fft, hop_length = hop_length))\nD_vesspa = np.abs(librosa.stft(audio_vesspa, n_fft = n_fft, hop_length = hop_length))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:02.864948Z","iopub.execute_input":"2025-04-29T22:24:02.865182Z","iopub.status.idle":"2025-04-29T22:24:03.15706Z","shell.execute_reply.started":"2025-04-29T22:24:02.865159Z","shell.execute_reply":"2025-04-29T22:24:03.15637Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Shape of D object:', np.shape(D_amered))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:11.837692Z","iopub.execute_input":"2025-04-29T22:24:11.838115Z","iopub.status.idle":"2025-04-29T22:24:11.842838Z","shell.execute_reply.started":"2025-04-29T22:24:11.838073Z","shell.execute_reply":"2025-04-29T22:24:11.842116Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #3. Spectrogram 🎷\n\n> 📌**Note**: \n* What is a spectrogram? A spectrogram is a visual representation of the spectrum of frequencies of a signal as it varies with time. When applied to an audio signal, spectrograms are sometimes called sonographs, voiceprints, or voicegrams (wiki).\n* Here we convert the frequency axis to a logarithmic one.","metadata":{}},{"cell_type":"code","source":"# Convert an amplitude spectrogram to Decibels-scaled spectrogram.\nDB_amered = librosa.amplitude_to_db(D_amered, ref = np.max)\nDB_cangoo = librosa.amplitude_to_db(D_cangoo, ref = np.max)\nDB_haiwoo = librosa.amplitude_to_db(D_haiwoo, ref = np.max)\nDB_pingro = librosa.amplitude_to_db(D_pingro, ref = np.max)\nDB_vesspa = librosa.amplitude_to_db(D_vesspa, ref = np.max)\n\n# === PLOT ===\nfig, ax = plt.subplots(2, 3, figsize=(16, 9))\nfig.suptitle('Spectrogram', fontsize=16)\nfig.delaxes(ax[1, 2])\n\nlibrosa.display.specshow(DB_amered, sr = sr_amered, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'cool', ax=ax[0, 0])\nlibrosa.display.specshow(DB_cangoo, sr = sr_cangoo, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'cool', ax=ax[0, 1])\nlibrosa.display.specshow(DB_haiwoo, sr = sr_haiwoo, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'cool', ax=ax[0, 2])\nlibrosa.display.specshow(DB_pingro, sr = sr_pingro, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'cool', ax=ax[1, 0])\nlibrosa.display.specshow(DB_vesspa, sr = sr_vesspa, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'cool', ax=ax[1, 1]);\n\nfor i, name in zip(range(0, 2*3), bird_sample_list):\n    x = i // 3\n    y = i % 3\n    ax[x, y].set_title(name, fontsize=13) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:16.369707Z","iopub.execute_input":"2025-04-29T22:24:16.370047Z","iopub.status.idle":"2025-04-29T22:24:23.118505Z","shell.execute_reply.started":"2025-04-29T22:24:16.370016Z","shell.execute_reply":"2025-04-29T22:24:23.117566Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #4. Mel Spectrogram 🎷\n> 📌**Note**: The Mel Scale, mathematically speaking, is the result of some non-linear transformation of the frequency scale. The Mel Spectrogram is a normal Spectrogram, but with a Mel Scale on the y axis.","metadata":{}},{"cell_type":"code","source":"# Create the Mel Spectrograms\nS_amered = librosa.feature.melspectrogram(y_amered, sr=sr_amered)\nS_DB_amered = librosa.amplitude_to_db(S_amered, ref=np.max)\n\nS_cangoo = librosa.feature.melspectrogram(y_cangoo, sr=sr_cangoo)\nS_DB_cangoo = librosa.amplitude_to_db(S_cangoo, ref=np.max)\n\nS_haiwoo = librosa.feature.melspectrogram(y_haiwoo, sr=sr_haiwoo)\nS_DB_haiwoo = librosa.amplitude_to_db(S_haiwoo, ref=np.max)\n\nS_pingro = librosa.feature.melspectrogram(y_pingro, sr=sr_pingro)\nS_DB_pingro = librosa.amplitude_to_db(S_pingro, ref=np.max)\n\nS_vesspa = librosa.feature.melspectrogram(y_vesspa, sr=sr_vesspa)\nS_DB_vesspa = librosa.amplitude_to_db(S_vesspa, ref=np.max)\n\n# === PLOT ====\nfig, ax = plt.subplots(2, 3, figsize=(16, 9))\nfig.suptitle('Mel Spectrogram', fontsize=16)\nfig.delaxes(ax[1, 2])\n\nlibrosa.display.specshow(S_DB_amered, sr = sr_amered, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'rainbow', ax=ax[0, 0])\nlibrosa.display.specshow(S_DB_cangoo, sr = sr_cangoo, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'rainbow', ax=ax[0, 1])\nlibrosa.display.specshow(S_DB_haiwoo, sr = sr_haiwoo, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'rainbow', ax=ax[0, 2])\nlibrosa.display.specshow(S_DB_pingro, sr = sr_pingro, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'rainbow', ax=ax[1, 0])\nlibrosa.display.specshow(S_DB_vesspa, sr = sr_vesspa, hop_length = hop_length, x_axis = 'time', \n                         y_axis = 'log', cmap = 'rainbow', ax=ax[1, 1]);\n\nfor i, name in zip(range(0, 2*3), bird_sample_list):\n    x = i // 3\n    y = i % 3\n    ax[x, y].set_title(name, fontsize=13)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:29.270484Z","iopub.execute_input":"2025-04-29T22:24:29.27078Z","iopub.status.idle":"2025-04-29T22:24:30.973477Z","shell.execute_reply.started":"2025-04-29T22:24:29.270753Z","shell.execute_reply":"2025-04-29T22:24:30.972667Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #5. Zero Crossing Rate 🚷\n\n> 📌**Note**: the rate at which the signal changes from positive to negative or back.","metadata":{}},{"cell_type":"code","source":"# Total zero_crossings in our 1 song\nzero_amered = librosa.zero_crossings(audio_amered, pad=False)\nzero_cangoo = librosa.zero_crossings(audio_cangoo, pad=False)\nzero_haiwoo = librosa.zero_crossings(audio_haiwoo, pad=False)\nzero_pingro = librosa.zero_crossings(audio_pingro, pad=False)\nzero_vesspa = librosa.zero_crossings(audio_vesspa, pad=False)\n\nzero_birds_list = [zero_amered, zero_cangoo, zero_haiwoo, zero_pingro, zero_vesspa]\n\nfor bird, name in zip(zero_birds_list, bird_sample_list):\n    print(\"{} change rate is {:,}\".format(name, sum(bird)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:34.524693Z","iopub.execute_input":"2025-04-29T22:24:34.525045Z","iopub.status.idle":"2025-04-29T22:24:47.877592Z","shell.execute_reply.started":"2025-04-29T22:24:34.525012Z","shell.execute_reply":"2025-04-29T22:24:47.876638Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #6. Harmonics and Perceptrual 🎹\n\n> 📌**Note**: \n* Harmonics are characteristichs that represent the sound *color*\n* Perceptrual shock wave represents the sound *rhythm and emotion*","metadata":{}},{"cell_type":"code","source":"y_harm_haiwoo, y_perc_haiwoo = librosa.effects.hpss(audio_haiwoo)\n\nplt.figure(figsize = (16, 6))\nplt.plot(y_perc_haiwoo, color = '#FFB100')\nplt.plot(y_harm_haiwoo, color = '#A300F9')\nplt.legend((\"Perceptrual\", \"Harmonics\"))\nplt.title(\"Harmonics and Perceptrual : Haiwoo Bird\", fontsize=16);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:47.879205Z","iopub.execute_input":"2025-04-29T22:24:47.879428Z","iopub.status.idle":"2025-04-29T22:24:54.506684Z","shell.execute_reply.started":"2025-04-29T22:24:47.879406Z","shell.execute_reply":"2025-04-29T22:24:54.505926Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #7. Spectral Centroid 🎯\n\n> 📌**Note**: \nIndicates where the ”centre of mass” for a sound is located and is calculated as the weighted mean of the frequencies present in the sound.","metadata":{}},{"cell_type":"code","source":"# Calculate the Spectral Centroids\nspectral_centroids = librosa.feature.spectral_centroid(audio_cangoo, sr=sr)[0]\n\n# Shape is a vector\nprint('Centroids:', spectral_centroids, '\\n')\nprint('Shape of Spectral Centroids:', spectral_centroids.shape, '\\n')\n\n# Computing the time variable for visualization\nframes = range(len(spectral_centroids))\n\n# Converts frame counts to time (seconds)\nt = librosa.frames_to_time(frames)\n\nprint('frames:', frames, '\\n')\nprint('t:', t)\n\n# Function that normalizes the Sound Data\ndef normalize(x, axis=0):\n    return sklearn.preprocessing.minmax_scale(x, axis=axis)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:54.508251Z","iopub.execute_input":"2025-04-29T22:24:54.508599Z","iopub.status.idle":"2025-04-29T22:24:54.573899Z","shell.execute_reply.started":"2025-04-29T22:24:54.508563Z","shell.execute_reply":"2025-04-29T22:24:54.572861Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Plotting the Spectral Centroid along the waveform\nplt.figure(figsize = (16, 6))\nlibrosa.display.waveplot(audio_cangoo, sr=sr, alpha=0.4, color = '#A300F9', lw=3)\nplt.plot(t, normalize(spectral_centroids), color='#FFB100', lw=2)\nplt.legend([\"Spectral Centroid\", \"Wave\"])\nplt.title(\"Spectral Centroid: Cangoo Bird\", fontsize=16);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:24:59.318475Z","iopub.execute_input":"2025-04-29T22:24:59.318994Z","iopub.status.idle":"2025-04-29T22:24:59.576935Z","shell.execute_reply.started":"2025-04-29T22:24:59.318953Z","shell.execute_reply":"2025-04-29T22:24:59.575929Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #8. Chroma Frequencies\n\n> 📌**Note**: Chroma features are an interesting and powerful representation for music audio in which the entire spectrum is projected onto 12 bins representing the 12 distinct semitones (or chromas) of the musical octave.","metadata":{}},{"cell_type":"code","source":"# Increase or decrease hop_length to change how granular you want your data to be\nhop_length = 5000\n\n# Chromogram Vesspa\nchromagram = librosa.feature.chroma_stft(audio_vesspa, sr=sr_vesspa, hop_length=hop_length)\nprint('Chromogram Vesspa shape:', chromagram.shape)\n\nplt.figure(figsize=(16, 6))\nlibrosa.display.specshow(chromagram, x_axis='time', y_axis='chroma', hop_length=hop_length, cmap='twilight')\n\nplt.title(\"Chromogram: Vesspa\", fontsize=16);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:25:01.924703Z","iopub.execute_input":"2025-04-29T22:25:01.925046Z","iopub.status.idle":"2025-04-29T22:25:02.137462Z","shell.execute_reply.started":"2025-04-29T22:25:01.925017Z","shell.execute_reply":"2025-04-29T22:25:02.136757Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #9. Tempo BPM (beats per minute)🎤\n> 📌**Note**: Dynamic programming beat tracker.","metadata":{}},{"cell_type":"code","source":"# Create Tempo BPM variable\ntempo_amered, _ = librosa.beat.beat_track(y_amered, sr = sr_amered)\ntempo_cangoo, _ = librosa.beat.beat_track(y_cangoo, sr = sr_cangoo)\ntempo_haiwoo, _ = librosa.beat.beat_track(y_haiwoo, sr = sr_haiwoo)\ntempo_pingro, _ = librosa.beat.beat_track(y_pingro, sr = sr_pingro)\ntempo_vesspa, _ = librosa.beat.beat_track(y_vesspa, sr = sr_vesspa)\n\ndata = pd.DataFrame({\"Type\": bird_sample_list , \n                     \"BPM\": [tempo_amered, tempo_cangoo, tempo_haiwoo, tempo_pingro, tempo_vesspa] })\n\n\n# Plot\nplt.figure(figsize = (20, 6))\nax = sns.barplot(y = data[\"BPM\"], x = data[\"Type\"], palette=\"hls\")\n\nplt.ylabel(\"BPM\", fontsize=14)\nplt.yticks(fontsize=13)\nplt.xticks(fontsize=13)\nplt.xlabel(\"\")\nplt.title(\"BPM for 5 Different Bird Species\", fontsize=16);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T22:25:28.171056Z","iopub.execute_input":"2025-04-29T22:25:28.171357Z","iopub.status.idle":"2025-04-29T22:25:29.431885Z","shell.execute_reply.started":"2025-04-29T22:25:28.171331Z","shell.execute_reply":"2025-04-29T22:25:29.431136Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### #10. Spectral Rolloff 🥏\n> 📌**Note**: Is a measure of the *shape of the signal*. It represents the frequency below which a specified percentage of the total spectral energy (e.g. 85%) lies.","metadata":{}},{"cell_type":"code","source":"# Spectral RollOff Vector\nspectral_rolloff = librosa.feature.spectral_rolloff(audio_amered, sr=sr_amered)[0]\n\n# Computing the time variable for visualization\nframes = range(len(spectral_rolloff))\n# Converts frame counts to time (seconds)\nt = librosa.frames_to_time(frames)\n\n# The plot\nplt.figure(figsize = (16, 6))\nlibrosa.display.waveplot(audio_amered, sr=sr_amered, alpha=0.4, color = '#A300F9', lw=3)\nplt.plot(t, normalize(spectral_rolloff), color='#FFB100', lw=3)\nplt.legend([\"Spectral Rolloff\", \"Wave\"])\nplt.title(\"Spectral Rolloff: Amered Bird\", fontsize=16);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T08:25:50.293251Z","iopub.execute_input":"2025-04-29T08:25:50.293566Z","iopub.status.idle":"2025-04-29T08:25:50.456359Z","shell.execute_reply.started":"2025-04-29T08:25:50.293542Z","shell.execute_reply":"2025-04-29T08:25:50.455546Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"[There are many more features that librosa can extract from sound. Check them all here.](https://librosa.org/librosa/0.6.0/feature.html)\n\n# 4. Additional Data ➕➕➕\n\n**[Why stop at 100? thread here](https://www.kaggle.com/c/birdsong-recognition/discussion/159970)**\n* [@Vopani](https://www.kaggle.com/rohanrao) kindly scraped the remaining of the available bird audios from Xeno-Canto site.\n* The data:\n    * doesn't contain already available train audios from competition data\n    * only MP3 format\n    * same license as the original data\n    * no corrupt audio present\n    \n**However, the Competition Hosts replied with:**\n* limiting to 100 audio files had the reason to not overload the memory\n* only 100 top rated audios were used\n* excluded videos that did not alow derivatives\n\n**They also recommend the following:**\n* the upper limit is usually 500 recordings per species\n* how many recordings are needed to train? (maybe even less than 100?)\n\n> 📌**Note**: So, should you use the extended data? Should you stick only with original data? *I have no clue, try multiple ideas and see what performs better*","metadata":{}},{"cell_type":"code","source":"# Bird Call Recognition - Model Implementation\n\nimport os\nimport numpy as np\nimport pandas as pd\nimport librosa\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score, classification_report, confusion_matrix\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Dropout, Conv2D, MaxPooling2D, Flatten\nfrom tensorflow.keras.utils import to_categorical\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-29T08:20:20.098005Z","iopub.execute_input":"2025-04-29T08:20:20.098371Z","iopub.status.idle":"2025-04-29T08:20:21.27489Z","shell.execute_reply.started":"2025-04-29T08:20:20.098326Z","shell.execute_reply":"2025-04-29T08:20:21.273845Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ------------------- Feature Extraction Function -------------------\n\ndef extract_features(file_path, max_pad_len=300):\n    \"\"\"\n    Extract features from audio file:\n    - MFCC (Mel-frequency cepstral coefficients)\n    - Zero Crossing Rate\n    - Spectral Centroid\n    - Spectral Rolloff\n    \"\"\"\n    try:\n        audio, sample_rate = librosa.load(file_path, res_type='kaiser_fast') \n        \n        # Extract features\n        mfccs = librosa.feature.mfcc(y=audio, sr=sample_rate, n_mfcc=40)\n        zcr = librosa.feature.zero_crossing_rate(audio)\n        spectral_centroid = librosa.feature.spectral_centroid(y=audio, sr=sample_rate)\n        spectral_rolloff = librosa.feature.spectral_rolloff(y=audio, sr=sample_rate)\n        \n        # Pad or trim to ensure consistent dimensions\n        pad_width = max_pad_len - mfccs.shape[1]\n        if pad_width > 0:\n            mfccs = np.pad(mfccs, pad_width=((0, 0), (0, pad_width)), mode='constant')\n            zcr = np.pad(zcr, pad_width=((0, 0), (0, pad_width)), mode='constant')\n            spectral_centroid = np.pad(spectral_centroid, pad_width=((0, 0), (0, pad_width)), mode='constant')\n            spectral_rolloff = np.pad(spectral_rolloff, pad_width=((0, 0), (0, pad_width)), mode='constant')\n        else:\n            mfccs = mfccs[:, :max_pad_len]\n            zcr = zcr[:, :max_pad_len]\n            spectral_centroid = spectral_centroid[:, :max_pad_len]\n            spectral_rolloff = spectral_rolloff[:, :max_pad_len]\n            \n        # Stack all features\n        features = np.vstack((mfccs, zcr, spectral_centroid, spectral_rolloff))\n        return features\n    except Exception as e:\n        print(f\"Error extracting features from {file_path}: {str(e)}\")\n        return None\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:18:46.215403Z","iopub.execute_input":"2025-04-25T10:18:46.215724Z","iopub.status.idle":"2025-04-25T10:18:46.225769Z","shell.execute_reply.started":"2025-04-25T10:18:46.215699Z","shell.execute_reply":"2025-04-25T10:18:46.224905Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------- Data Loading & Preprocessing -------------------\n\n# Define base directory (update this to your actual path)\nbase_dir = '../input/birdsong-recognition/train_audio/'\n\n# Load the training CSV\ntrain_csv = pd.read_csv(\"../input/birdsong-recognition/train.csv\")\n\n# Add full path for audio files\ntrain_csv['full_path'] = base_dir + train_csv['ebird_code'] + '/' + train_csv['filename']\n\n# Limit to a subset of species to make the initial model building faster\n# You can increase this number or remove this filter for a more comprehensive model\ntop_species = train_csv['species'].value_counts().head(10).index.tolist()\nfiltered_df = train_csv[train_csv['species'].isin(top_species)].copy()\n\nprint(f\"Working with {len(filtered_df)} audio files from {len(top_species)} species\")\n\n# Encode labels\nlabel_encoder = LabelEncoder()\nfiltered_df['encoded_species'] = label_encoder.fit_transform(filtered_df['species'])\n\n# Sample a smaller dataset for quick model building\n# Increase this number or remove this sampling for better results\nsample_size = min(1000, len(filtered_df))\nsampled_df = filtered_df.sample(sample_size, random_state=42)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:19:07.558223Z","iopub.execute_input":"2025-04-25T10:19:07.558562Z","iopub.status.idle":"2025-04-25T10:19:07.847501Z","shell.execute_reply.started":"2025-04-25T10:19:07.55853Z","shell.execute_reply":"2025-04-25T10:19:07.846629Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------- Feature Extraction -------------------\n\nprint(\"Extracting features from audio files...\")\nfeatures = []\nlabels = []\n\nfor i, row in tqdm(sampled_df.iterrows(), total=len(sampled_df)):\n    file_path = row['full_path']\n    feature = extract_features(file_path)\n    \n    if feature is not None:\n        features.append(feature)\n        labels.append(row['encoded_species'])\n\n# Convert to numpy arrays\nX = np.array(features)\ny = np.array(labels)\n\n# Reshape for CNN: (samples, height, width, channels)\nX = X.reshape(X.shape[0], X.shape[1], X.shape[2], 1)\n\n# One-hot encode the labels\ny = to_categorical(y, num_classes=len(top_species))\n\n# Split into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)\n\nprint(f\"Training on {X_train.shape[0]} samples, validating on {X_val.shape[0]} samples\")\nprint(f\"Input shape: {X_train.shape}\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:19:30.615566Z","iopub.execute_input":"2025-04-25T10:19:30.615866Z","iopub.status.idle":"2025-04-25T10:37:26.22495Z","shell.execute_reply.started":"2025-04-25T10:19:30.61584Z","shell.execute_reply":"2025-04-25T10:37:26.224036Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------- Model Building -------------------\n\n# Create CNN model\nmodel = Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(X_train.shape[1], X_train.shape[2], 1)),\n    MaxPooling2D((2, 2)),\n    Conv2D(64, (3, 3), activation='relu'),\n    MaxPooling2D((2, 2)),\n    Conv2D(128, (3, 3), activation='relu'),\n    MaxPooling2D((2, 2)),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dropout(0.5),\n    Dense(len(top_species), activation='softmax')\n])\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:47:25.167694Z","iopub.execute_input":"2025-04-25T10:47:25.168107Z","iopub.status.idle":"2025-04-25T10:47:28.120717Z","shell.execute_reply.started":"2025-04-25T10:47:25.168066Z","shell.execute_reply":"2025-04-25T10:47:28.119772Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compile model\nmodel.compile(\n    optimizer='adam',\n    loss='categorical_crossentropy',\n    metrics=['accuracy']\n)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:47:28.122627Z","iopub.execute_input":"2025-04-25T10:47:28.122962Z","iopub.status.idle":"2025-04-25T10:47:28.136278Z","shell.execute_reply.started":"2025-04-25T10:47:28.122927Z","shell.execute_reply":"2025-04-25T10:47:28.135624Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Summary of model\nmodel.summary()\n\n# Train model\nhistory = model.fit(\n    X_train, y_train,\n    epochs=15,  # Increase for better results\n    batch_size=32,\n    validation_data=(X_val, y_val),\n    verbose=1\n)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:47:32.352725Z","iopub.execute_input":"2025-04-25T10:47:32.353065Z","iopub.status.idle":"2025-04-25T10:47:39.885996Z","shell.execute_reply.started":"2025-04-25T10:47:32.353034Z","shell.execute_reply":"2025-04-25T10:47:39.885181Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------- Evaluation -------------------\n\n# Evaluate the model\ntest_loss, test_acc = model.evaluate(X_val, y_val, verbose=2)\nprint(f'\\nTest accuracy: {test_acc:.4f}')\n\n# Get predictions\ny_pred = model.predict(X_val)\ny_pred_classes = np.argmax(y_pred, axis=1)\ny_true = np.argmax(y_val, axis=1)\n\n# Classification report\nprint(\"\\nClassification Report:\")\nprint(classification_report(\n    y_true, \n    y_pred_classes, \n    target_names=[label_encoder.inverse_transform([i])[0] for i in range(len(top_species))]\n))\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:47:41.44688Z","iopub.execute_input":"2025-04-25T10:47:41.447218Z","iopub.status.idle":"2025-04-25T10:47:41.663064Z","shell.execute_reply.started":"2025-04-25T10:47:41.447183Z","shell.execute_reply":"2025-04-25T10:47:41.66225Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Confusion matrix\nplt.figure(figsize=(12, 10))\ncm = confusion_matrix(y_true, y_pred_classes)\nsns.heatmap(\n    cm, \n    annot=True, \n    fmt='d', \n    cmap='Blues',\n    xticklabels=[label_encoder.inverse_transform([i])[0] for i in range(len(top_species))],\n    yticklabels=[label_encoder.inverse_transform([i])[0] for i in range(len(top_species))]\n)\nplt.title('Confusion Matrix')\nplt.ylabel('True Label')\nplt.xlabel('Predicted Label')\nplt.xticks(rotation=45, ha='right')\nplt.tight_layout()\nplt.savefig('confusion_matrix.png')\nplt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:47:45.274704Z","iopub.execute_input":"2025-04-25T10:47:45.274993Z","iopub.status.idle":"2025-04-25T10:47:46.026415Z","shell.execute_reply.started":"2025-04-25T10:47:45.274968Z","shell.execute_reply":"2025-04-25T10:47:46.025522Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot training history\nplt.figure(figsize=(12, 5))\n\n# Plot accuracy\nplt.subplot(1, 2, 1)\nplt.plot(history.history['accuracy'], label='Training Accuracy')\nplt.plot(history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:48:03.034556Z","iopub.execute_input":"2025-04-25T10:48:03.034859Z","iopub.status.idle":"2025-04-25T10:48:03.183683Z","shell.execute_reply.started":"2025-04-25T10:48:03.034833Z","shell.execute_reply":"2025-04-25T10:48:03.182862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Plot loss\nplt.subplot(1, 2, 2)\nplt.plot(history.history['loss'], label='Training Loss')\nplt.plot(history.history['val_loss'], label='Validation Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.tight_layout()\nplt.savefig('training_history.png')\nplt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:48:07.684609Z","iopub.execute_input":"2025-04-25T10:48:07.684924Z","iopub.status.idle":"2025-04-25T10:48:07.864383Z","shell.execute_reply.started":"2025-04-25T10:48:07.684898Z","shell.execute_reply":"2025-04-25T10:48:07.8635Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------- Predict on Test Data -------------------\n\n# Function to predict on test data\ndef predict_test(test_csv, model, label_encoder):\n    # Create prediction dataframe\n    predictions_df = pd.DataFrame()\n    predictions_df['row_id'] = test_csv['row_id']\n    \n    # Extract features and predict for each test file\n    for i, row in tqdm(test_csv.iterrows(), total=len(test_csv)):\n        file_path = f\"../input/birdsong-recognition/test_audio/{row['audio_id']}.mp3\"\n        \n        if os.path.exists(file_path):\n            feature = extract_features(file_path)\n            \n            if feature is not None:\n                # Reshape for prediction\n                feature = feature.reshape(1, feature.shape[0], feature.shape[1], 1)\n                \n                # Predict\n                pred = model.predict(feature)\n                pred_class = np.argmax(pred, axis=1)[0]\n                pred_species = label_encoder.inverse_transform([pred_class])[0]\n                \n                # Add to dataframe\n                predictions_df.loc[i, 'birds'] = pred_species\n            else:\n                predictions_df.loc[i, 'birds'] = 'nocall'  # Default if feature extraction fails\n        else:\n            predictions_df.loc[i, 'birds'] = 'nocall'  # Default if file doesn't exist\n    \n    return predictions_df\n\n# Load test CSV (this would be used for test predictions if we had test data)\n# test_csv = pd.read_csv(\"../input/birdsong-recognition/test.csv\")\n# predictions_df = predict_test(test_csv, model, label_encoder)\n# predictions_df.to_csv('submission.csv', index=False)\n\nprint(\"Model training and evaluation complete!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T10:48:13.808575Z","iopub.execute_input":"2025-04-25T10:48:13.808861Z","iopub.status.idle":"2025-04-25T10:48:13.8173Z","shell.execute_reply.started":"2025-04-25T10:48:13.808838Z","shell.execute_reply":"2025-04-25T10:48:13.816469Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ✅ Bird Call Recognition - Full Pipeline with Complete Results\n\nimport os\nimport numpy as np\nimport pandas as pd\nimport librosa\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score, classification_report, confusion_matrix\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Dropout, Conv2D, MaxPooling2D, Flatten\nfrom tensorflow.keras.utils import to_categorical\n\n# ------------------- Feature Extraction -------------------\ndef extract_features(file_path, max_pad_len=300):\n    try:\n        audio, sample_rate = librosa.load(file_path, res_type='kaiser_fast') \n        mfccs = librosa.feature.mfcc(y=audio, sr=sample_rate, n_mfcc=40)\n        zcr = librosa.feature.zero_crossing_rate(audio)\n        spectral_centroid = librosa.feature.spectral_centroid(y=audio, sr=sample_rate)\n        spectral_rolloff = librosa.feature.spectral_rolloff(y=audio, sr=sample_rate)\n\n        pad_width = max_pad_len - mfccs.shape[1]\n        if pad_width > 0:\n            mfccs = np.pad(mfccs, ((0, 0), (0, pad_width)), mode='constant')\n            zcr = np.pad(zcr, ((0, 0), (0, pad_width)), mode='constant')\n            spectral_centroid = np.pad(spectral_centroid, ((0, 0), (0, pad_width)), mode='constant')\n            spectral_rolloff = np.pad(spectral_rolloff, ((0, 0), (0, pad_width)), mode='constant')\n        else:\n            mfccs = mfccs[:, :max_pad_len]\n            zcr = zcr[:, :max_pad_len]\n            spectral_centroid = spectral_centroid[:, :max_pad_len]\n            spectral_rolloff = spectral_rolloff[:, :max_pad_len]\n\n        features = np.vstack((mfccs, zcr, spectral_centroid, spectral_rolloff))\n        return features\n    except Exception as e:\n        print(f\"Error extracting features from {file_path}: {e}\")\n        return None\n\n# ------------------- Data Loading -------------------\nbase_dir = '../input/birdsong-recognition/train_audio/'\ntrain_csv = pd.read_csv(\"../input/birdsong-recognition/train.csv\")\ntrain_csv['full_path'] = base_dir + train_csv['ebird_code'] + '/' + train_csv['filename']\ntop_species = train_csv['species'].value_counts().head(10).index.tolist()\nfiltered_df = train_csv[train_csv['species'].isin(top_species)].copy()\nfiltered_df['encoded_species'] = LabelEncoder().fit_transform(filtered_df['species'])\nlabel_encoder = LabelEncoder()\nfiltered_df['encoded_species'] = label_encoder.fit_transform(filtered_df['species'])\n\nsampled_df = filtered_df.sample(min(1000, len(filtered_df)), random_state=42)\n\n# ------------------- Feature Extraction -------------------\nfeatures, labels = [], []\nfor _, row in tqdm(sampled_df.iterrows(), total=len(sampled_df)):\n    feature = extract_features(row['full_path'])\n    if feature is not None:\n        features.append(feature)\n        labels.append(row['encoded_species'])\n\nX = np.array(features).reshape(len(features), features[0].shape[0], features[0].shape[1], 1)\ny = to_categorical(labels, num_classes=len(top_species))\n\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# ------------------- Model Definition -------------------\nmodel = Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(X.shape[1], X.shape[2], 1)),\n    MaxPooling2D(2, 2),\n    Conv2D(64, (3, 3), activation='relu'),\n    MaxPooling2D(2, 2),\n    Conv2D(128, (3, 3), activation='relu'),\n    MaxPooling2D(2, 2),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dropout(0.5),\n    Dense(len(top_species), activation='softmax')\n])\n\nmodel.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])\nmodel.summary()\n\n# ------------------- Model Training -------------------\nhistory = model.fit(X_train, y_train, epochs=15, batch_size=32, validation_data=(X_val, y_val), verbose=1)\n\n# ------------------- Model Evaluation -------------------\ntest_loss, test_acc = model.evaluate(X_val, y_val)\nprint(f\"\\nTest Accuracy: {test_acc:.4f}\")\n\ny_pred = model.predict(X_val)\ny_pred_classes = np.argmax(y_pred, axis=1)\ny_true = np.argmax(y_val, axis=1)\n\nprint(\"\\nClassification Report:\")\nprint(classification_report(y_true, y_pred_classes, target_names=label_encoder.inverse_transform(range(len(top_species)))))\n\n# ------------------- Confusion Matrix -------------------\ncm = confusion_matrix(y_true, y_pred_classes)\nplt.figure(figsize=(12, 10))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues', \n            xticklabels=label_encoder.inverse_transform(range(len(top_species))),\n            yticklabels=label_encoder.inverse_transform(range(len(top_species))))\nplt.title('Confusion Matrix')\nplt.ylabel('True Label')\nplt.xlabel('Predicted Label')\nplt.xticks(rotation=45, ha='right')\nplt.tight_layout()\nplt.savefig('confusion_matrix.png')\nplt.show()\n\n# ------------------- Plot Accuracy and Loss -------------------\nplt.figure(figsize=(12, 5))\nplt.subplot(1, 2, 1)\nplt.plot(history.history['accuracy'], label='Train Accuracy')\nplt.plot(history.history['val_accuracy'], label='Val Accuracy')\nplt.title('Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\nplt.subplot(1, 2, 2)\nplt.plot(history.history['loss'], label='Train Loss')\nplt.plot(history.history['val_loss'], label='Val Loss')\nplt.title('Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.tight_layout()\nplt.savefig('training_history.png')\nplt.show()\n\n# ------------------- Prediction Function -------------------\ndef predict_test(test_csv, model, label_encoder):\n    predictions_df = pd.DataFrame()\n    predictions_df['row_id'] = test_csv['row_id']\n    for i, row in tqdm(test_csv.iterrows(), total=len(test_csv)):\n        file_path = f\"../input/birdsong-recognition/test_audio/{row['audio_id']}.mp3\"\n        if os.path.exists(file_path):\n            feature = extract_features(file_path)\n            if feature is not None:\n                feature = feature.reshape(1, feature.shape[0], feature.shape[1], 1)\n                pred = model.predict(feature)\n                pred_class = np.argmax(pred)\n                pred_species = label_encoder.inverse_transform([pred_class])[0]\n                predictions_df.loc[i, 'birds'] = pred_species\n            else:\n                predictions_df.loc[i, 'birds'] = 'nocall'\n        else:\n            predictions_df.loc[i, 'birds'] = 'nocall'\n    return predictions_df\n\n# test_csv = pd.read_csv(\"../input/birdsong-recognition/test.csv\")\n# predictions = predict_test(test_csv, model, label_encoder)\n# predictions.to_csv(\"submission.csv\", index=False)\n\nprint(\"\\n✅ All steps completed: Model defined, trained, evaluated, and ready for inference.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T11:36:25.36564Z","iopub.execute_input":"2025-04-25T11:36:25.365959Z","iopub.status.idle":"2025-04-25T11:55:46.493194Z","shell.execute_reply.started":"2025-04-25T11:36:25.365924Z","shell.execute_reply":"2025-04-25T11:55:46.491603Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}