{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Introduction**  \nThis notebook provides a comprehensive exploratory data analysis (EDA) of the BirdCLEF+ 2025 dataset. We examine missing values, duplicates, species distribution, audio quality ratings, geographical distribution, correlations, taxonomy consistency, file counts, outlier detection (for both ratings and geography), and data imbalances (species and collection bias).\n\n**Key Findings:**  \n- The dataset comprises 28,564 entries with minimal missing data (only in geographical coordinates) and no duplicates.  \n- There is a pronounced imbalance in species counts and collection sources.  \n- Audio quality ratings are consistent without extreme outliers.  \n- Geographical analysis reveals several outlier locations.  \n- Merging taxonomy data confirms complete species classification.","metadata":{}},{"cell_type":"markdown","source":"### Import Libraries\nIn this cell, we import all necessary libraries.","metadata":{}},{"cell_type":"code","source":"import ast\nimport folium\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport os\nimport pandas as pd\nimport seaborn as sns\n\nfrom collections import Counter\nfrom folium.plugins import MarkerCluster\n\n# Configure plots for inline display (if using Jupyter)\n%matplotlib inline","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:33.012760Z","iopub.execute_input":"2025-03-16T04:55:33.013140Z","iopub.status.idle":"2025-03-16T04:55:36.720410Z","shell.execute_reply.started":"2025-03-16T04:55:33.013074Z","shell.execute_reply":"2025-03-16T04:55:36.719210Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Load the Data\nHere we load the CSV files and the location text file from the Kaggle dataset.","metadata":{}},{"cell_type":"code","source":"# Load CSV files\ntrain_df = pd.read_csv(\"/kaggle/input/birdclef-2025/train.csv\")\ntaxonomy_df = pd.read_csv(\"/kaggle/input/birdclef-2025/taxonomy.csv\")\nsample_submission = pd.read_csv(\"/kaggle/input/birdclef-2025/sample_submission.csv\")\n\n# Load location metadata from the text file\nwith open(\"/kaggle/input/birdclef-2025/recording_location.txt\", \"r\") as f:\n    recording_location = f.read()\n\nprint(\"Train data shape:\", train_df.shape)\nprint(\"Taxonomy data shape:\", taxonomy_df.shape)\nprint(\"Sample submission shape:\", sample_submission.shape)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:36.721657Z","iopub.execute_input":"2025-03-16T04:55:36.722372Z","iopub.status.idle":"2025-03-16T04:55:36.976284Z","shell.execute_reply.started":"2025-03-16T04:55:36.722331Z","shell.execute_reply":"2025-03-16T04:55:36.974933Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Key Findings:**  \n- The train dataset has 28,564 rows and 13 columns.  \n- Taxonomy data includes 206 species across 5 columns.  \n- The sample submission file contains 3 rows and 207 columns.  \n- Location metadata is loaded successfully from the text file.","metadata":{}},{"cell_type":"markdown","source":"### Missing Values & Duplicates\nWe check the dataset for missing values and duplicates.","metadata":{}},{"cell_type":"code","source":"# Check for missing values\nprint(train_df.info())\nprint('\\n',train_df.isnull().sum())\n\n# Check for duplicates\nprint('\\n',\"Duplicates in train.csv:\", train_df.duplicated().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:36.977633Z","iopub.execute_input":"2025-03-16T04:55:36.978114Z","iopub.status.idle":"2025-03-16T04:55:37.095009Z","shell.execute_reply.started":"2025-03-16T04:55:36.978050Z","shell.execute_reply":"2025-03-16T04:55:37.093690Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Missing Values & Duplicates**  \n*Purpose:* Check for missing values and duplicate entries in the train dataset.\n\n**Key Findings:**  \n- All columns are complete except for `latitude` and `longitude`, which have 809 missing values each.  \n- No duplicate records were found, ensuring data integrity.","metadata":{}},{"cell_type":"markdown","source":"### Primary Species Distribution\nThis cell visualizes the frequency distribution of the primary species labels.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,6))\ntrain_df['primary_label'].value_counts().plot(kind='bar')\nplt.title(\"Distribution of Primary Species Labels\")\nplt.xlabel(\"Species\")\nplt.ylabel(\"Count\")\nplt.xticks(rotation=90)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:37.097982Z","iopub.execute_input":"2025-03-16T04:55:37.098355Z","iopub.status.idle":"2025-03-16T04:55:38.942320Z","shell.execute_reply.started":"2025-03-16T04:55:37.098323Z","shell.execute_reply":"2025-03-16T04:55:38.940844Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Primary Species Distribution**  \n*Purpose:* Visualize the frequency distribution of the primary species labels using a bar chart.\n\n**Key Findings:**  \n- The distribution is highly imbalanced with a few species dominating the counts.  - This imbalance may require special handling during modeling to account for rare classes.","metadata":{}},{"cell_type":"markdown","source":"### Audio Quality Ratings Distribution\nWe plot a histogram to understand the distribution of audio quality ratings.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,6))\nsns.histplot(train_df['rating'], bins=30, kde=True)\nplt.title(\"Distribution of Audio Quality Ratings\")\nplt.xlabel(\"Rating\")\nplt.ylabel(\"Frequency\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:38.944337Z","iopub.execute_input":"2025-03-16T04:55:38.944840Z","iopub.status.idle":"2025-03-16T04:55:39.366798Z","shell.execute_reply.started":"2025-03-16T04:55:38.944788Z","shell.execute_reply":"2025-03-16T04:55:39.365616Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Audio Quality Ratings Distribution**  \n*Purpose:* Plot a histogram of the audio quality ratings to understand their distribution.\n\n**Key Findings:**  \n- Ratings are mostly concentrated within a specific range.  \n- The kernel density estimate (KDE) confirms a consistent distribution, indicating similar audio quality across recordings.","metadata":{}},{"cell_type":"markdown","source":"### Correlation Matrix\nWe compute and visualize the correlation matrix for numerical features (currently only 'rating' is available).","metadata":{}},{"cell_type":"code","source":"# Compute correlation matrix (extend list if additional numerical features are added)\ncorr_matrix = train_df[['rating']].copy()\nprint(corr_matrix.corr())\n\nsns.heatmap(corr_matrix.corr(), annot=True, cmap='coolwarm')\nplt.title(\"Correlation Matrix\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:39.694429Z","iopub.execute_input":"2025-03-16T04:55:39.694823Z","iopub.status.idle":"2025-03-16T04:55:39.919753Z","shell.execute_reply.started":"2025-03-16T04:55:39.694794Z","shell.execute_reply":"2025-03-16T04:55:39.918474Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Cell 7: Correlation Matrix**  \n*Purpose:* Compute and visualize the correlation matrix for numerical features (currently only the audio rating).\n\n**Key Findings:**  \n- The correlation matrix shows a perfect self-correlation (1.0) for the 'rating' variable, as it is the only numerical feature at this stage.\n","metadata":{}},{"cell_type":"markdown","source":"### Average Rating by Species\nHere, we group by species to see the average audio quality rating for each species.","metadata":{}},{"cell_type":"code","source":"avg_rating_by_species = train_df.groupby('primary_label')['rating'].mean().sort_values(ascending=False)\nplt.figure(figsize=(12,6))\navg_rating_by_species.plot(kind='bar')\nplt.title(\"Average Rating by Species\")\nplt.xlabel(\"Species\")\nplt.ylabel(\"Average Rating\")\nplt.xticks(rotation=90)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:39.920886Z","iopub.execute_input":"2025-03-16T04:55:39.921326Z","iopub.status.idle":"2025-03-16T04:55:41.660697Z","shell.execute_reply.started":"2025-03-16T04:55:39.921287Z","shell.execute_reply":"2025-03-16T04:55:41.659451Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Average Rating by Species**  \n*Purpose:* Group the data by species and calculate the average audio quality rating for each species.\n\n**Key Findings:**  \n- Average ratings differ among species, indicating which species tend to have higher or lower quality recordings.  \n- This insight is valuable for prioritizing species during further analysis or model development.","metadata":{}},{"cell_type":"markdown","source":"### Recording Types Distribution\nWe plot the frequency of different recording types.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,6))\ntrain_df['type'].value_counts().plot(kind='bar')\nplt.title(\"Distribution of Recording Types\")\nplt.xlabel(\"Type\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:41.868961Z","iopub.execute_input":"2025-03-16T04:55:41.869303Z","iopub.status.idle":"2025-03-16T04:55:48.408859Z","shell.execute_reply.started":"2025-03-16T04:55:41.869274Z","shell.execute_reply":"2025-03-16T04:55:48.407342Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Recording Types Distribution**  \n*Purpose:* Display the frequency of different recording types (e.g., \"song\").\n\n**Key Findings:**  \n- The dataset predominantly consists of one recording type, which simplifies analysis but may limit the diversity of sound events.","metadata":{}},{"cell_type":"markdown","source":"### Secondary Labels Analysis\nWe convert the secondary_labels column from string representation to lists and then analyze the frequency of additional species.","metadata":{}},{"cell_type":"code","source":"# Convert secondary_labels from string to list and flatten the results\nsecondary_labels_list = []\nfor label in train_df['secondary_labels']:\n    try:\n        parsed = ast.literal_eval(label) if label and label != \"[]\" else []\n    except:\n        parsed = []\n    secondary_labels_list.extend(parsed)\n\nsecondary_counter = Counter(secondary_labels_list)\nprint(\"Most common secondary labels:\", secondary_counter.most_common(10))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:48.410616Z","iopub.execute_input":"2025-03-16T04:55:48.411168Z","iopub.status.idle":"2025-03-16T04:55:48.660902Z","shell.execute_reply.started":"2025-03-16T04:55:48.411112Z","shell.execute_reply":"2025-03-16T04:55:48.659586Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Secondary Labels Analysis**  \n*Purpose:* Convert the `secondary_labels` (stored as a string representation of a list) into actual lists and analyze the frequency of additional species labels.\n\n**Key Findings:**  \n- The most common secondary labels are identified.  \n- A high frequency of empty entries suggests secondary labels might be under-reported or inconsistently provided.","metadata":{}},{"cell_type":"markdown","source":"### Taxonomic Classes Distribution\nThis cell shows the distribution of species across different taxonomic classes.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,6))\ntaxonomy_df['class_name'].value_counts().plot(kind='bar')\nplt.title(\"Distribution of Taxonomic Classes\")\nplt.xlabel(\"Class\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:48.662226Z","iopub.execute_input":"2025-03-16T04:55:48.662667Z","iopub.status.idle":"2025-03-16T04:55:48.874041Z","shell.execute_reply.started":"2025-03-16T04:55:48.662623Z","shell.execute_reply":"2025-03-16T04:55:48.872637Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Taxonomic Classes Distribution**  \n*Purpose:* Visualize the distribution of species across different taxonomic classes using taxonomy data.\n\n**Key Findings:**  \n- The bar chart illustrates the diversity of taxonomic classes in the dataset.  \n- This helps understand the overall taxonomic spread and focus areas.","metadata":{}},{"cell_type":"markdown","source":"### Merging Data\nWe merge the train dataset with taxonomy data to verify taxonomy consistency and check for missing taxonomy information.","metadata":{}},{"cell_type":"code","source":"merged_df = pd.merge(train_df, taxonomy_df, left_on=\"primary_label\", right_on=\"primary_label\", how=\"left\")\nprint(\"Merged data shape:\", merged_df.shape)\nprint(\"Missing taxonomy info:\", merged_df['class_name'].isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:48.877972Z","iopub.execute_input":"2025-03-16T04:55:48.878350Z","iopub.status.idle":"2025-03-16T04:55:48.913523Z","shell.execute_reply.started":"2025-03-16T04:55:48.878314Z","shell.execute_reply":"2025-03-16T04:55:48.912383Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Merging Data with Taxonomy**  \n*Purpose:* Merge the train dataset with the taxonomy dataset to verify taxonomy consistency and completeness.\n\n**Key Findings:**  \n- The merged dataset has 28,564 rows and 17 columns.  \n- No missing taxonomy information was found, confirming complete species classification.","metadata":{}},{"cell_type":"markdown","source":"### Count Audio Files in Directories\nHere, we define a function to count files in a given directory and then count the files for train_audio, train_soundscapes, and test_soundscapes.\n","metadata":{}},{"cell_type":"code","source":"def count_files(directory):\n    return sum([len(files) for _, _, files in os.walk(directory)])\n\nprint(\"Train Audio files:\", count_files(\"/kaggle/input/birdclef-2025/train_audio\"))\nprint(\"Train Soundscapes files:\", count_files(\"/kaggle/input/birdclef-2025/train_soundscapes\"))\nprint(\"Test Soundscapes files:\", count_files(\"/kaggle/input/birdclef-2025/test_soundscapes\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:55:48.914982Z","iopub.execute_input":"2025-03-16T04:55:48.915355Z","iopub.status.idle":"2025-03-16T04:56:21.736358Z","shell.execute_reply.started":"2025-03-16T04:55:48.915322Z","shell.execute_reply":"2025-03-16T04:56:21.735142Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Count Audio Files in Directories**  \n*Purpose:* Define a function to count the number of files in a directory and check file counts for train_audio, train_soundscapes, and test_soundscapes.\n\n**Key Findings:**  \n- There are 28,564 train audio files, 9,726 train soundscape files, and 1 test soundscape file.  \n- These counts match the dataset description.","metadata":{}},{"cell_type":"markdown","source":"### Outlier Detection – Rating Outliers\nWe examine the distribution of ratings using a boxplot, compute the IQR, and identify potential outliers.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.boxplot(x=train_df['rating'])\nplt.title(\"Boxplot of Ratings\")\nplt.xlabel(\"Rating\")\nplt.show()\n\nprint(\"Rating Summary Statistics:\")\nprint(train_df['rating'].describe())\n\n# Calculate the interquartile range (IQR)\nQ1 = train_df['rating'].quantile(0.25)\nQ3 = train_df['rating'].quantile(0.75)\nIQR = Q3 - Q1\nlower_bound = Q1 - 1.5 * IQR\nupper_bound = Q3 + 1.5 * IQR\nprint(f\"Rating outlier bounds: Lower = {lower_bound:.2f}, Upper = {upper_bound:.2f}\")\n\n# Identify potential rating outliers\nrating_outliers = train_df[(train_df['rating'] < lower_bound) | (train_df['rating'] > upper_bound)]\nprint(\"Number of rating outliers:\", rating_outliers.shape[0])\nprint(\"Sample rating outliers:\")\nprint(rating_outliers[['rating']].head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:56:21.737620Z","iopub.execute_input":"2025-03-16T04:56:21.737999Z","iopub.status.idle":"2025-03-16T04:56:21.889800Z","shell.execute_reply.started":"2025-03-16T04:56:21.737968Z","shell.execute_reply":"2025-03-16T04:56:21.888475Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Outlier Detection – Rating Outliers**  \n*Purpose:* Use a boxplot and the IQR method to detect potential outliers in the audio quality ratings.\n\n**Key Findings:**  \n- The boxplot and statistics show ratings range between 0 and 5.  \n- No rating outliers were detected using the IQR method, indicating ratings are within expected bounds.","metadata":{}},{"cell_type":"markdown","source":"### Outlier Detection – Geographical Outliers\nWe remove rows with missing coordinates, plot the locations, compute the IQR for both latitude and longitude, and identify geographical outliers.","metadata":{}},{"cell_type":"code","source":"# Remove rows with missing coordinate values\nlocation_data = train_df.dropna(subset=['latitude', 'longitude'])\n\n# Scatterplot of recording locations\nplt.figure(figsize=(8, 6))\nsns.scatterplot(x='longitude', y='latitude', data=location_data, alpha=0.5)\nplt.title(\"Scatter Plot of Recording Locations\")\nplt.xlabel(\"Longitude\")\nplt.ylabel(\"Latitude\")\nplt.show()\n\n# Calculate IQR for latitude\nlat_Q1 = location_data['latitude'].quantile(0.25)\nlat_Q3 = location_data['latitude'].quantile(0.75)\nlat_IQR = lat_Q3 - lat_Q1\nlat_lower_bound = lat_Q1 - 1.5 * lat_IQR\nlat_upper_bound = lat_Q3 + 1.5 * lat_IQR\n\n# Calculate IQR for longitude\nlon_Q1 = location_data['longitude'].quantile(0.25)\nlon_Q3 = location_data['longitude'].quantile(0.75)\nlon_IQR = lon_Q3 - lon_Q1\nlon_lower_bound = lon_Q1 - 1.5 * lon_IQR\nlon_upper_bound = lon_Q3 + 1.5 * lon_IQR\n\nprint(f\"Latitude bounds: {lat_lower_bound:.2f} - {lat_upper_bound:.2f}\")\nprint(f\"Longitude bounds: {lon_lower_bound:.2f} - {lon_upper_bound:.2f}\")\n\n# Identify geographical outliers\ngeo_outliers = location_data[\n    (location_data['latitude'] < lat_lower_bound) | (location_data['latitude'] > lat_upper_bound) |\n    (location_data['longitude'] < lon_lower_bound) | (location_data['longitude'] > lon_upper_bound)\n]\nprint(\"Number of geographical outliers:\", geo_outliers.shape[0])\nprint(\"Sample geographical outliers:\")\nprint(geo_outliers[['latitude', 'longitude']].head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:56:21.890914Z","iopub.execute_input":"2025-03-16T04:56:21.891351Z","iopub.status.idle":"2025-03-16T04:56:22.225594Z","shell.execute_reply.started":"2025-03-16T04:56:21.891308Z","shell.execute_reply":"2025-03-16T04:56:22.224336Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Geographical Distribution of Recordings**  \n*Purpose:* \n- Use a scatter plot to visualize the geographic distribution (longitude vs. latitude) of recordings.\n- Remove records with missing coordinates, plot the geographic distribution, calculate the IQR for latitude and longitude, and identify geographical outliers.\n\n**Key Findings:**  \n- Recordings are spread over a wide geographic area.  \n- Some outlier locations are visible, suggesting further investigation may be needed.\n- The scatterplot shows a broad geographic spread.  \n- The IQR method identifies 514 geographical outlier points, suggesting some recordings are far outside the main cluster.","metadata":{}},{"cell_type":"markdown","source":"### Interactive Map\nWe create an interactive map with Folium to visualize recording locations, using a marker cluster.","metadata":{}},{"cell_type":"code","source":"# Calculate the center of the map using the mean coordinates\nmean_lat = location_data['latitude'].mean()\nmean_lon = location_data['longitude'].mean()\n\n# Initialize the folium map\nm = folium.Map(location=[mean_lat, mean_lon], zoom_start=8)\n\n# Add a marker cluster to group nearby points\nmarker_cluster = MarkerCluster().add_to(m)\n\n# Add markers for each recording location with a popup showing the primary species label\nfor idx, row in location_data.iterrows():\n    folium.Marker(\n        location=[row['latitude'], row['longitude']],\n        popup=f\"Species: {row['primary_label']}\"\n    ).add_to(marker_cluster)\n\n# Optionally, save the map to an HTML file\nm.save(\"recording_locations_map.html\")\nprint(\"Interactive map saved as 'recording_locations_map.html'.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:56:22.226886Z","iopub.execute_input":"2025-03-16T04:56:22.227362Z","iopub.status.idle":"2025-03-16T04:56:58.076956Z","shell.execute_reply.started":"2025-03-16T04:56:22.227317Z","shell.execute_reply":"2025-03-16T04:56:58.075758Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Interactive Map of Recording Locations**  \n*Purpose:* Create an interactive map with Folium to visualize recording locations with clustered markers.\n\n**Key Findings:**  \n- An interactive map displays all recording locations with markers showing the primary species on click.  \n- The map is saved as an HTML file for further exploration.","metadata":{}},{"cell_type":"markdown","source":"### Data Imbalance – Species Imbalance\nWe visualize the distribution of species counts and print summary statistics.","metadata":{}},{"cell_type":"code","source":"species_counts = train_df['primary_label'].value_counts()\nplt.figure(figsize=(12, 6))\nspecies_counts.plot(kind='bar')\nplt.title(\"Distribution of Primary Species Labels\")\nplt.xlabel(\"Species\")\nplt.ylabel(\"Count\")\nplt.xticks(rotation=90)\nplt.show()\n\nprint(\"Species count statistics:\")\nprint(species_counts.describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:56:58.078024Z","iopub.execute_input":"2025-03-16T04:56:58.078476Z","iopub.status.idle":"2025-03-16T04:57:00.175489Z","shell.execute_reply.started":"2025-03-16T04:56:58.078443Z","shell.execute_reply":"2025-03-16T04:57:00.174338Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Data Imbalance – Species Imbalance**  \n*Purpose:* Visualize the distribution of species counts using a bar chart and display descriptive statistics.\n\n**Key Findings:**  \n- The bar chart confirms significant species imbalance.  \n- Descriptive statistics reveal a wide range of counts, indicating a need for class imbalance handling in modeling.","metadata":{}},{"cell_type":"markdown","source":"### Data Imbalance – Collection Bias\nThis cell visualizes and prints the distribution of recordings by their collection source.","metadata":{}},{"cell_type":"code","source":"collection_counts = train_df['collection'].value_counts()\nplt.figure(figsize=(8, 6))\ncollection_counts.plot(kind='bar')\nplt.title(\"Distribution of Recordings by Collection Source\")\nplt.xlabel(\"Collection\")\nplt.ylabel(\"Count\")\nplt.show()\n\nprint(\"Collection counts:\")\nprint(collection_counts)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:57:00.176793Z","iopub.execute_input":"2025-03-16T04:57:00.177249Z","iopub.status.idle":"2025-03-16T04:57:00.387983Z","shell.execute_reply.started":"2025-03-16T04:57:00.177213Z","shell.execute_reply":"2025-03-16T04:57:00.386744Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Data Imbalance – Collection Bias**  \n*Purpose:* Visualize and print the distribution of recordings by collection source.\n\n**Key Findings:**  \n- Average ratings differ among species, indicating which species tend to have higher or lower quality recordings.  \n- This insight is valuable for prioritizing species during further analysis or model development.\n- The \"XC\" collection dominates the dataset with over 21,000 recordings.  \n- This imbalance may affect model performance and should be accounted for during preprocessing.","metadata":{}},{"cell_type":"markdown","source":"### Sample Submission","metadata":{}},{"cell_type":"code","source":"sample_submission.to_csv('submission.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-16T04:57:00.389027Z","iopub.execute_input":"2025-03-16T04:57:00.389408Z","iopub.status.idle":"2025-03-16T04:57:00.402320Z","shell.execute_reply.started":"2025-03-16T04:57:00.389380Z","shell.execute_reply":"2025-03-16T04:57:00.401137Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}