{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"BirdCLEF 2024 - Exploratory Data Analysis\nDataset: https://www.kaggle.com/competitions/birdclef-2024/data","metadata":{}},{"cell_type":"markdown","source":"Exploratory Data Analysis (EDA) is a fundamental process in understanding and preparing your data for analysis or machine learning tasks. Below is a simplified guide on how to perform EDA for your dataset:\n\n1. **Overview**:\n   Understand the purpose of the dataset, including the goal of prediction and any specific metrics involved.\n\n2. **Data Collection**:\n   Import the dataset into your analysis environment (e.g., Python) using libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n3. **Initial Exploration**:\n   - Check the first few rows of the dataset to understand its structure.\n   - Inspect data types and identify missing values using `df.info()`.\n   - Calculate basic statistics to grasp the data's distribution using `df.describe()`.\n\n4. **Data Cleaning**:\n   - Handle missing values through imputation or removal.\n   - Detect and handle outliers appropriately.\n   - Remove duplicate rows if necessary.\n\n5. **Data Visualization**:\n   - Use visualizations like histograms, box plots, and bar plots to understand data distributions and relationships.\n   - Explore correlations between variables using correlation matrices and scatter plots.\n\n6. **Feature Analysis**:\n   - Investigate relationships between features and the target variable(s) using visualizations.\n   - Compare feature distributions across different classes or subtypes.\n\n7. **Outlier Detection**:\n   - If relevant, identify and visualize outliers using appropriate plots.\n   - Apply statistical methods or machine learning techniques to detect outliers.\n\n8. **Dimensionality Reduction (optional)**:\n   - Consider techniques like Principal Component Analysis (PCA) if dealing with high-dimensional data.\n\n9. **Summary and Insights**:\n   - Summarize your findings, including any patterns, trends, or anomalies observed during EDA.\n   - Document data preprocessing steps undertaken.\n\n10. **Next Steps**:\n    - Plan further actions based on EDA findings, such as feature engineering or selecting appropriate models.\n    - Prepare for iterative exploration and refinement of your analysis.\n\nRemember, EDA is a dynamic process, and you may need to revisit these steps as you gain more insights or refine your analysis goals.","metadata":{}},{"cell_type":"markdown","source":"The task in this competition is to identify bird species present in recordings made in a biodiversity hotspot in the Western Ghats. This is important for conservation efforts and requires accurate identification of bird calls.\n\nHere's a breakdown of the files provided:\n\n1. **train_audio/**: Contains short recordings of individual bird calls sourced from xenocanto.org. These recordings are downsampled to 32 kHz and converted to the ogg format.\n\n2. **test_soundscapes/**: This directory will be populated with approximately 1,100 recordings when you submit your notebook. These recordings are 4 minutes long and in ogg audio format. They are used for scoring your predictions.\n\n3. **unlabeled_soundscapes/**: Contains unlabeled audio data from the same recording locations as the test soundscapes.\n\n4. **train_metadata.csv**: Provides metadata for the training data, including:\n   - `primary_label`: Code for the bird species.\n   - `latitude` & `longitude`: Coordinates of the recording location.\n   - `author`: User who provided the recording.\n   - `filename`: Name of the associated audio file.\n\n5. **sample_submission.csv**: A sample submission file containing row IDs and columns for predicting the presence of each bird species.\n\n6. **eBird_Taxonomy_v2021.csv**: Data on the relationships between different bird species.\n\nTo tackle this challenge effectively, you'll likely need to:\n- Analyze the provided audio recordings.\n- Utilize the metadata provided to understand recording locations and potentially account for geographic variations in bird calls.\n- Make predictions about the presence of different bird species in the test soundscapes based on your analysis.\n\nYour approach might involve techniques such as audio signal processing, machine learning, and possibly leveraging information from the provided metadata. It's essential to thoroughly understand the dataset and explore relationships between variables before developing your model.","metadata":{}},{"cell_type":"code","source":"pip install --upgrade pip","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:06.749028Z","iopub.execute_input":"2024-04-04T17:06:06.749450Z","iopub.status.idle":"2024-04-04T17:06:21.641727Z","shell.execute_reply.started":"2024-04-04T17:06:06.749415Z","shell.execute_reply":"2024-04-04T17:06:21.640065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\n\nimport pandas as pd\npd.options.mode.chained_assignment = None # avoids assignment warning\nimport numpy as np\nimport random\nfrom glob import glob\nfrom tqdm import tqdm\ntqdm.pandas()  # enable progress bars in pandas operations\nimport gc\n\nimport librosa\nimport sklearn\nimport json\n\n# Import for visualization\nimport matplotlib as mpl\n#cmap = mpl.cm.get_cmap('coolwarm')\nimport matplotlib.pyplot as plt\nimport librosa.display as lid\nimport IPython.display as ipd\nimport cv2\n\n# Import KaggleDatasets for accessing Kaggle datasets\nfrom kaggle_datasets import KaggleDatasets\n\n# WandB for experiment tracking\nimport wandb\n\nimport torchaudio\nimport plotly.express as px\nfrom IPython.display import Audio\nfrom shapely.geometry import Point\n\nimport plotly.express as px","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-04T17:06:21.644377Z","iopub.execute_input":"2024-04-04T17:06:21.644757Z","iopub.status.idle":"2024-04-04T17:06:21.658196Z","shell.execute_reply.started":"2024-04-04T17:06:21.644719Z","shell.execute_reply":"2024-04-04T17:06:21.657030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the train_metadata.csv file\nmetadata_df = pd.read_csv('/kaggle/input/birdclef-2024/train_metadata.csv')\n\n# Display the first few rows of the metadata dataframe\nmetadata_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:21.659664Z","iopub.execute_input":"2024-04-04T17:06:21.660039Z","iopub.status.idle":"2024-04-04T17:06:21.828332Z","shell.execute_reply.started":"2024-04-04T17:06:21.660008Z","shell.execute_reply":"2024-04-04T17:06:21.827432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = '/kaggle/input/birdclef-2024/train_audio/'\ndata, rate = torchaudio.load(train_path + metadata_df.filename[0])\ndisplay(Audio(data[0, :rate*5], rate=rate))\npx.line(y=data[0, :rate*5], title=metadata_df.common_name[0])","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:21.830776Z","iopub.execute_input":"2024-04-04T17:06:21.831150Z","iopub.status.idle":"2024-04-04T17:06:23.065626Z","shell.execute_reply.started":"2024-04-04T17:06:21.831119Z","shell.execute_reply":"2024-04-04T17:06:23.064071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assuming metadata_df contains the data\n# Create scatter plot on a map\nfig = px.scatter_mapbox(metadata_df, lat='latitude', lon='longitude', color='primary_label', \n                        hover_name='primary_label', hover_data=['latitude', 'longitude'], \n                        title='Geographical Distribution of Bird Species',\n                        zoom=1, height=600)\nfig.update_layout(mapbox_style=\"open-street-map\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:23.067000Z","iopub.execute_input":"2024-04-04T17:06:23.067362Z","iopub.status.idle":"2024-04-04T17:06:24.025606Z","shell.execute_reply.started":"2024-04-04T17:06:23.067333Z","shell.execute_reply":"2024-04-04T17:06:24.024643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the train_metadata.csv file\neBird_Taxonomy_df = pd.read_csv('/kaggle/input/birdclef-2024/eBird_Taxonomy_v2021.csv')\n\n# Display the first few rows of the metadata dataframe\neBird_Taxonomy_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:24.027303Z","iopub.execute_input":"2024-04-04T17:06:24.027654Z","iopub.status.idle":"2024-04-04T17:06:24.098250Z","shell.execute_reply.started":"2024-04-04T17:06:24.027624Z","shell.execute_reply":"2024-04-04T17:06:24.096965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Summary statistics\nmetadata_df.describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:24.100025Z","iopub.execute_input":"2024-04-04T17:06:24.100515Z","iopub.status.idle":"2024-04-04T17:06:24.125828Z","shell.execute_reply.started":"2024-04-04T17:06:24.100472Z","shell.execute_reply":"2024-04-04T17:06:24.124917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for missing values\nmetadata_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:24.127141Z","iopub.execute_input":"2024-04-04T17:06:24.127639Z","iopub.status.idle":"2024-04-04T17:06:24.161416Z","shell.execute_reply.started":"2024-04-04T17:06:24.127611Z","shell.execute_reply":"2024-04-04T17:06:24.160330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of bird species in the training data\nplt.figure(figsize=(12, 6))\nsns.countplot(x='primary_label', data=metadata_df, order=metadata_df['primary_label'].value_counts().index)\nplt.xticks(rotation=45)\nplt.title('Distribution of Bird Species')\nplt.xlabel('Bird Species')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:24.162739Z","iopub.execute_input":"2024-04-04T17:06:24.163116Z","iopub.status.idle":"2024-04-04T17:06:26.166834Z","shell.execute_reply.started":"2024-04-04T17:06:24.163088Z","shell.execute_reply":"2024-04-04T17:06:26.165509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Relationship between latitude and longitude\nplt.figure(figsize=(10, 8))\nsns.scatterplot(x='longitude', y='latitude', data=metadata_df, hue='primary_label', palette='viridis', alpha=0.5)\nplt.title('Geographical Distribution of Bird Species')\nplt.xlabel('Longitude')\nplt.ylabel('Latitude')\nplt.legend(title='Bird Species', loc='upper right', bbox_to_anchor=(1.25, 1))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:26.170039Z","iopub.execute_input":"2024-04-04T17:06:26.170512Z","iopub.status.idle":"2024-04-04T17:06:32.253240Z","shell.execute_reply.started":"2024-04-04T17:06:26.170480Z","shell.execute_reply":"2024-04-04T17:06:32.251811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of recordings by authors\nplt.figure(figsize=(12, 6))\nsns.countplot(x='author', data=metadata_df, order=metadata_df['author'].value_counts().index[:10])\nplt.xticks(rotation=45)\nplt.title('Top 10 Authors by Number of Recordings')\nplt.xlabel('Author')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:32.254832Z","iopub.execute_input":"2024-04-04T17:06:32.255987Z","iopub.status.idle":"2024-04-04T17:06:32.629200Z","shell.execute_reply.started":"2024-04-04T17:06:32.255921Z","shell.execute_reply":"2024-04-04T17:06:32.628021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Submission:**\nWe train the model using all available training data and generate a submission file for evaluation.","metadata":{}},{"cell_type":"code","source":"df_sub = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\")\ndisplay(df_sub)","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:32.630670Z","iopub.execute_input":"2024-04-04T17:06:32.631129Z","iopub.status.idle":"2024-04-04T17:06:32.666513Z","shell.execute_reply.started":"2024-04-04T17:06:32.631080Z","shell.execute_reply":"2024-04-04T17:06:32.665526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-04T17:06:32.667848Z","iopub.execute_input":"2024-04-04T17:06:32.668186Z","iopub.status.idle":"2024-04-04T17:06:32.677134Z","shell.execute_reply.started":"2024-04-04T17:06:32.668159Z","shell.execute_reply":"2024-04-04T17:06:32.675663Z"},"trusted":true},"execution_count":null,"outputs":[]}]}