{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"},{"sourceId":5668986,"sourceType":"datasetVersion","datasetId":3012613}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <mark>BirdCLEF 2024 - EDA</mark>\n<span style=\"font-size:22px;color:purple\"> Thank you for having a look at my notebook - advice and feedback always welcomed!</span>\n\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌 Dataset Link: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/data\">https://www.kaggle.com/competitions/birdclef-2024/data</a>\n</div>\n\n\n## **Exploratory Data Analysis (EDA)** \nEDA is a crucial step in understanding and preparing your data for any data analysis or machine learning project. Here's a step-by-step guide on how to perform an EDA for this dataset:\n\n**Overview**\n\nIn this competition, you will be predicting default of clients based on internal and external information that are available for each client. Scoring is performed using custom metric that not only evaluates the AUC of predictions but also considers the stability of predictions model across the data range of the test set. To better understand this metric.\n\n\n**Data Collection:**\n\nThis dataset contains a large number of tables as a result of utilizing diverse data sources and the varying levels of data aggregation used while preparing the dataset. Note: All files listed below are found in both .csv and .parquet formats.\n\n\n**Data Loading:**\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n\n\n### **Initial Exploration:**\n\n**1 - Start by examining the basic characteristics of the data:**\n\n    Check the first few rows using df.head().\n    Check the data types and missing values using df.info().\n    Calculate basic statistics using df.describe().\n    Data Cleaning:\n\n**2 - Handle missing values, outliers, and duplicates:**\n\n    Use techniques like imputation for missing values.\n    Identify and deal with outliers appropriately.\n    Remove duplicate rows if necessary.\n    \n**3 - Data Visualization:**\n\n    Create visualizations to gain insights into the data:\n    Histograms and box plots for numerical features.\n    Bar plots for categorical features.\n    Correlation matrix and scatter plots to understand relationships between variables.\n    \n**4 - Feature Analysis:**\n\n    Explore relationships between features and the target variable(s) for classification and outlier detection.\n    Visualize how different features vary across different subtypes or classes.\n    Use box plots, violin plots, or swarm plots to compare feature distributions.\n\n**5 - Outlier Detection:**\n\n    If your dataset contains information related to outlier detection, perform a dedicated EDA for this aspect:\n    Visualize outliers using scatter plots or box plots.\n    Apply statistical methods or machine learning techniques to identify outliers.\n\n**6 - Dimensionality Reduction (optional):**\n\n    If the dataset has many features, consider dimensionality reduction techniques like Principal Component Analysis (PCA) to reduce the number of variables while preserving important information.\n\n**8 - Summary and Insights:**\n\n    Summarize your findings from the EDA, including any patterns, trends, or anomalies observed.\n    Document any data preprocessing steps applied.\n\n**7 - Next Steps:**\n\n    Based on your EDA findings, plan your next steps, which may include feature engineering, model selection, and further data preprocessing.\n\n\nRemember that EDA is an iterative process, and you may need to revisit these steps as you delve deeper into the dataset and develop your machine learning or data analysis models.\n\n\n## **Dataset Description**\nYour challenge in this competition is to identify which birds are calling in recordings made in a Global Biodiversity Hotspot in the Western Ghats. This is an important task for scientists who monitor bird populations for conservation purposes. More accurate solutions could enable more comprehensive monitoring.\n\nThis competition uses a hidden test set. When your submitted notebook is scored, the actual test data will be made available to your notebook.\n\n## **Files**\n**train_audio/** The training data consists of short recordings of individual bird calls generously uploaded by users of xenocanto.org. These files have been downsampled to 32 kHz where applicable to match the test set audio and converted to the ogg format. The training data should have nearly all relevant files; we expect there is no benefit to looking for more on xenocanto.org and appreciate your cooperation in limiting the burden on their servers.\n\n**test_soundscapes/** When you submit a notebook, the test_soundscapes directory will be populated with approximately 1,100 recordings to be used for scoring. They are 4 minutes long and in ogg audio format. The file names are randomized. It should take your submission notebook approximately five minutes to load all of the test soundscapes.\n\n**unlabeled_soundscapes/** Unlabeled audio data from the same recording locations as the test soundscapes.\n\n**train_metadata.csv** A wide range of metadata is provided for the training data. The most directly relevant fields are:\n\n**primary_label** - a code for the bird species. You can review detailed information about the bird codes by appending the code to https://ebird.org/species/, such as https://ebird.org/species/amecro for the American Crow.\n\n**latitude & longitude:** coordinates for where the recording was taken. Some bird species may have local call 'dialects,' so you may want to seek geographic diversity in your training data.\n\n**author** - The user who provided the recording.\n\n**filename:** the name of the associated audio file.\n\n**sample_submission.csv** A valid sample submission.\n\nrow_id: A slug of [soundscape_id]_[end_time] for the prediction.\n[bird_id]: There are 182 bird ID columns. You will need to predict the probability of the presence of each bird for each row.\neBird_Taxonomy_v2021.csv - Data on the relationships between different species.","metadata":{}},{"cell_type":"code","source":"pip install --upgrade pip","metadata":{"execution":{"iopub.status.busy":"2024-04-04T11:18:29.093716Z","iopub.execute_input":"2024-04-04T11:18:29.094117Z","iopub.status.idle":"2024-04-04T11:19:01.084243Z","shell.execute_reply.started":"2024-04-04T11:18:29.094087Z","shell.execute_reply":"2024-04-04T11:19:01.082449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\n\nimport pandas as pd\npd.options.mode.chained_assignment = None # avoids assignment warning\nimport numpy as np\nimport random\nfrom glob import glob\nfrom tqdm import tqdm\ntqdm.pandas()  # enable progress bars in pandas operations\nimport gc\n\nimport librosa\nimport sklearn\nimport json\n\n# Import for visualization\nimport matplotlib as mpl\n#cmap = mpl.cm.get_cmap('coolwarm')\nimport matplotlib.pyplot as plt\nimport librosa.display as lid\nimport IPython.display as ipd\nimport cv2\n\n# Import KaggleDatasets for accessing Kaggle datasets\nfrom kaggle_datasets import KaggleDatasets\n\n# WandB for experiment tracking\nimport wandb\n\nimport torchaudio\nimport plotly.express as px\nfrom IPython.display import Audio\nfrom shapely.geometry import Point\n\nimport plotly.express as px","metadata":{"execution":{"iopub.status.busy":"2024-04-04T11:19:41.596788Z","iopub.execute_input":"2024-04-04T11:19:41.597219Z","iopub.status.idle":"2024-04-04T11:19:42.423868Z","shell.execute_reply.started":"2024-04-04T11:19:41.597190Z","shell.execute_reply":"2024-04-04T11:19:42.422769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the train_metadata.csv file\nmetadata_df = pd.read_csv('/kaggle/input/birdclef-2024/train_metadata.csv')\n\n# Display the first few rows of the metadata dataframe\nmetadata_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T11:20:30.033372Z","iopub.execute_input":"2024-04-04T11:20:30.034312Z","iopub.status.idle":"2024-04-04T11:20:30.245720Z","shell.execute_reply.started":"2024-04-04T11:20:30.034261Z","shell.execute_reply":"2024-04-04T11:20:30.244398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = '/kaggle/input/birdclef-2024/train_audio/'\ndata, rate = torchaudio.load(train_path + metadata_df.filename[0])\ndisplay(Audio(data[0, :rate*5], rate=rate))\npx.line(y=data[0, :rate*5], title=metadata_df.common_name[0])","metadata":{"execution":{"iopub.status.busy":"2024-04-04T11:20:42.676231Z","iopub.execute_input":"2024-04-04T11:20:42.676678Z","iopub.status.idle":"2024-04-04T11:20:45.954432Z","shell.execute_reply.started":"2024-04-04T11:20:42.676645Z","shell.execute_reply":"2024-04-04T11:20:45.952425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assuming metadata_df contains the data\n# Create scatter plot on a map\nfig = px.scatter_mapbox(metadata_df, lat='latitude', lon='longitude', color='primary_label', \n                        hover_name='primary_label', hover_data=['latitude', 'longitude'], \n                        title='Geographical Distribution of Bird Species',\n                        zoom=1, height=600)\nfig.update_layout(mapbox_style=\"open-street-map\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-04T11:24:59.509431Z","iopub.execute_input":"2024-04-04T11:24:59.510385Z","iopub.status.idle":"2024-04-04T11:25:00.377415Z","shell.execute_reply.started":"2024-04-04T11:24:59.510346Z","shell.execute_reply":"2024-04-04T11:25:00.376291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the train_metadata.csv file\neBird_Taxonomy_df = pd.read_csv('/kaggle/input/birdclef-2024/eBird_Taxonomy_v2021.csv')\n\n# Display the first few rows of the metadata dataframe\neBird_Taxonomy_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:49:42.574505Z","iopub.execute_input":"2024-04-03T20:49:42.575288Z","iopub.status.idle":"2024-04-03T20:49:42.637953Z","shell.execute_reply.started":"2024-04-03T20:49:42.575248Z","shell.execute_reply":"2024-04-03T20:49:42.636798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Summary statistics\nmetadata_df.describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:02:52.703642Z","iopub.execute_input":"2024-04-03T20:02:52.704038Z","iopub.status.idle":"2024-04-03T20:02:52.729515Z","shell.execute_reply.started":"2024-04-03T20:02:52.704009Z","shell.execute_reply":"2024-04-03T20:02:52.728482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for missing values\nmetadata_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:02:53.013571Z","iopub.execute_input":"2024-04-03T20:02:53.015919Z","iopub.status.idle":"2024-04-03T20:02:53.036628Z","shell.execute_reply.started":"2024-04-03T20:02:53.015881Z","shell.execute_reply":"2024-04-03T20:02:53.035047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of bird species in the training data\nplt.figure(figsize=(12, 6))\nsns.countplot(x='primary_label', data=metadata_df, order=metadata_df['primary_label'].value_counts().index)\nplt.xticks(rotation=45)\nplt.title('Distribution of Bird Species')\nplt.xlabel('Bird Species')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:02:53.258556Z","iopub.execute_input":"2024-04-03T20:02:53.258968Z","iopub.status.idle":"2024-04-03T20:02:54.747202Z","shell.execute_reply.started":"2024-04-03T20:02:53.258939Z","shell.execute_reply":"2024-04-03T20:02:54.746331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Relationship between latitude and longitude\nplt.figure(figsize=(10, 8))\nsns.scatterplot(x='longitude', y='latitude', data=metadata_df, hue='primary_label', palette='viridis', alpha=0.5)\nplt.title('Geographical Distribution of Bird Species')\nplt.xlabel('Longitude')\nplt.ylabel('Latitude')\nplt.legend(title='Bird Species', loc='upper right', bbox_to_anchor=(1.25, 1))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:03:04.073852Z","iopub.execute_input":"2024-04-03T20:03:04.074299Z","iopub.status.idle":"2024-04-03T20:03:09.412474Z","shell.execute_reply.started":"2024-04-03T20:03:04.074264Z","shell.execute_reply":"2024-04-03T20:03:09.411371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of recordings by authors\nplt.figure(figsize=(12, 6))\nsns.countplot(x='author', data=metadata_df, order=metadata_df['author'].value_counts().index[:10])\nplt.xticks(rotation=45)\nplt.title('Top 10 Authors by Number of Recordings')\nplt.xlabel('Author')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:03:09.414478Z","iopub.execute_input":"2024-04-03T20:03:09.414808Z","iopub.status.idle":"2024-04-03T20:03:09.721938Z","shell.execute_reply.started":"2024-04-03T20:03:09.414779Z","shell.execute_reply":"2024-04-03T20:03:09.721026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Submission**\nWe retrain the model on the full training data and create a submission file.\n\n","metadata":{}},{"cell_type":"code","source":"df_sub = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\")\ndisplay(df_sub)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:47:38.227183Z","iopub.execute_input":"2024-04-03T20:47:38.227614Z","iopub.status.idle":"2024-04-03T20:47:38.264875Z","shell.execute_reply.started":"2024-04-03T20:47:38.227561Z","shell.execute_reply":"2024-04-03T20:47:38.263567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T20:48:09.843480Z","iopub.execute_input":"2024-04-03T20:48:09.843878Z","iopub.status.idle":"2024-04-03T20:48:09.853206Z","shell.execute_reply.started":"2024-04-03T20:48:09.843848Z","shell.execute_reply":"2024-04-03T20:48:09.852126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **🚀 The best way to BirdCLEF 2024 competition 🚀**\n\nDears,\n\nI hope this message finds you well.\n\n**1. Cast Your Vote:**\n\nVisit the competition platform and find my EDA submission. Click on the \"Vote\" or \"Support\" button to cast your vote.\n\n**2. Share with Your Network:**\n\nSpread the word among your friends, family, and colleagues who may be interested in supporting my work.\n\n1**3. Provide Feedback:**\n\nIf you have any feedback or suggestions on my EDA, please feel free to share them with me. Your input is valuable and can help me improve. I am committed to making a positive impact in the field of cancer research, and your support will bring me one step closer to achieving that goal.\n\nThank you for taking the time to read this message, and I genuinely appreciate your support in this competition. Together, we can contribute to the fight against ovarian cancer and advance the field of data-driven healthcare.\n\n1If you have any questions or need more information about my EDA, please don't hesitate to reach out to me. Your support means the world to me!\n\n","metadata":{}}]}