{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"papermill":{"duration":14.880861,"end_time":"2024-04-24T14:46:03.585575","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-04-24T14:45:48.704714","version":"2.1.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #8ABEB9; padding: 10px;\">\n    <h2 style=\"font-weight: bold;\">BIRDCLEF 2024 DATA ANALYSIS</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #8ABEB9; padding: 10px;\">\n    <h2 style=\"font-weight: bold;\">IMPORTING VARIOUS MODULES</h2>\n</div>","metadata":{"execution":{"iopub.status.busy":"2024-05-08T22:59:31.989078Z","iopub.execute_input":"2024-05-08T22:59:31.989500Z","iopub.status.idle":"2024-05-08T22:59:31.996157Z","shell.execute_reply.started":"2024-05-08T22:59:31.989447Z","shell.execute_reply":"2024-05-08T22:59:31.994890Z"}}},{"cell_type":"code","source":"# packages\n\nimport numpy as np\nimport pandas as pd\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport folium\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom scipy import stats\nplt.rcParams['figure.figsize']=[15,8]","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":1.828736,"end_time":"2024-04-24T14:45:56.477764","exception":false,"start_time":"2024-04-24T14:45:54.649028","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:17.560035Z","iopub.execute_input":"2024-05-08T23:27:17.560466Z","iopub.status.idle":"2024-05-08T23:27:19.317379Z","shell.execute_reply.started":"2024-05-08T23:27:17.560417Z","shell.execute_reply":"2024-05-08T23:27:19.316088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# file overview\n!ls -l '../input/birdclef-2024/'","metadata":{"papermill":{"duration":1.192107,"end_time":"2024-04-24T14:45:57.695400","exception":false,"start_time":"2024-04-24T14:45:56.503293","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:19.319501Z","iopub.execute_input":"2024-05-08T23:27:19.320092Z","iopub.status.idle":"2024-05-08T23:27:20.452831Z","shell.execute_reply.started":"2024-05-08T23:27:19.320057Z","shell.execute_reply":"2024-05-08T23:27:20.451371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read metedata file\ndf = pd.read_csv('../input/birdclef-2024/train_metadata.csv')","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","papermill":{"duration":0.227329,"end_time":"2024-04-24T14:45:57.948960","exception":false,"start_time":"2024-04-24T14:45:57.721631","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:20.454772Z","iopub.execute_input":"2024-05-08T23:27:20.455156Z","iopub.status.idle":"2024-05-08T23:27:20.671801Z","shell.execute_reply.started":"2024-05-08T23:27:20.455122Z","shell.execute_reply":"2024-05-08T23:27:20.670493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# First glance","metadata":{"papermill":{"duration":0.025905,"end_time":"2024-04-24T14:45:57.999700","exception":false,"start_time":"2024-04-24T14:45:57.973795","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# preview of data\ndf.head()","metadata":{"papermill":{"duration":0.064117,"end_time":"2024-04-24T14:45:58.088475","exception":false,"start_time":"2024-04-24T14:45:58.024358","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:20.674890Z","iopub.execute_input":"2024-05-08T23:27:20.675995Z","iopub.status.idle":"2024-05-08T23:27:20.717464Z","shell.execute_reply.started":"2024-05-08T23:27:20.675950Z","shell.execute_reply":"2024-05-08T23:27:20.715648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.tail()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.719330Z","iopub.execute_input":"2024-05-08T23:27:20.720061Z","iopub.status.idle":"2024-05-08T23:27:20.742072Z","shell.execute_reply.started":"2024-05-08T23:27:20.720015Z","shell.execute_reply":"2024-05-08T23:27:20.740834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.743831Z","iopub.execute_input":"2024-05-08T23:27:20.744291Z","iopub.status.idle":"2024-05-08T23:27:20.751875Z","shell.execute_reply.started":"2024-05-08T23:27:20.744248Z","shell.execute_reply":"2024-05-08T23:27:20.750641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# structure details\ndf.info()","metadata":{"papermill":{"duration":0.073572,"end_time":"2024-04-24T14:45:58.184828","exception":false,"start_time":"2024-04-24T14:45:58.111256","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:20.753633Z","iopub.execute_input":"2024-05-08T23:27:20.754065Z","iopub.status.idle":"2024-05-08T23:27:20.813055Z","shell.execute_reply.started":"2024-05-08T23:27:20.754026Z","shell.execute_reply":"2024-05-08T23:27:20.811830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Summary Statistics of numeric variables:","metadata":{}},{"cell_type":"code","source":"df.describe().round(2).T","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.814435Z","iopub.execute_input":"2024-05-08T23:27:20.814822Z","iopub.status.idle":"2024-05-08T23:27:20.855754Z","shell.execute_reply.started":"2024-05-08T23:27:20.814791Z","shell.execute_reply":"2024-05-08T23:27:20.854204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_vars = df.select_dtypes(include=['int64', 'float64']).columns.tolist()\ncategorical_vars = df.select_dtypes(include=['object']).columns.tolist()                           \nprint('Numerical variables:', numerical_vars)\nprint('Categorical variables:', categorical_vars)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.857421Z","iopub.execute_input":"2024-05-08T23:27:20.857941Z","iopub.status.idle":"2024-05-08T23:27:20.876668Z","shell.execute_reply.started":"2024-05-08T23:27:20.857893Z","shell.execute_reply":"2024-05-08T23:27:20.874481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of categorical and numerical variables\ncategorical_count = df.select_dtypes(include='object').shape[1]\nnumerical_count = df.select_dtypes(exclude='object').shape[1]\n\nprint(f\"Number of categorical variables: {categorical_count}\")\nprint(f\"Number of numerical variables: {numerical_count}\")","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.883937Z","iopub.execute_input":"2024-05-08T23:27:20.884320Z","iopub.status.idle":"2024-05-08T23:27:20.898079Z","shell.execute_reply.started":"2024-05-08T23:27:20.884290Z","shell.execute_reply":"2024-05-08T23:27:20.896623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Unique values for categorical features\nprint(df.select_dtypes(include=['object']).nunique())","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.899763Z","iopub.execute_input":"2024-05-08T23:27:20.900235Z","iopub.status.idle":"2024-05-08T23:27:20.958476Z","shell.execute_reply.started":"2024-05-08T23:27:20.900196Z","shell.execute_reply":"2024-05-08T23:27:20.956928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Missing Value","metadata":{}},{"cell_type":"code","source":"missing_df =  df.isnull().sum().to_frame().rename(columns={0:\"Total No. of Missing Values\"})\nmissing_df[\"% of Missing Values\"] = round((missing_df[\"Total No. of Missing Values\"]/len( df))*100,2)\nmissing_df","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:20.960075Z","iopub.execute_input":"2024-05-08T23:27:20.960510Z","iopub.status.idle":"2024-05-08T23:27:21.005467Z","shell.execute_reply.started":"2024-05-08T23:27:20.960468Z","shell.execute_reply":"2024-05-08T23:27:21.004081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the percentage of missing values for each column\nmissing_values_percentage = df.isnull().mean() * 100\n\n# Now you can sort and visualize the missing values\nmissing_values_percentage_sorted = missing_values_percentage.sort_values()\n\n# Visualization code\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nplt.figure(figsize=(10, 6))\nsns.barplot(x=missing_values_percentage_sorted, y=missing_values_percentage_sorted.index)\nplt.title('Percentage of Missing Values in Each Column (Ascending Order)')\nplt.xlabel('Percentage of Missing Values')\nplt.ylabel('Columns')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:21.007186Z","iopub.execute_input":"2024-05-08T23:27:21.007637Z","iopub.status.idle":"2024-05-08T23:27:21.502050Z","shell.execute_reply.started":"2024-05-08T23:27:21.007598Z","shell.execute_reply":"2024-05-08T23:27:21.500694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.heatmap(df.isnull(),cbar=False)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:21.504123Z","iopub.execute_input":"2024-05-08T23:27:21.504577Z","iopub.status.idle":"2024-05-08T23:27:22.715933Z","shell.execute_reply.started":"2024-05-08T23:27:21.504519Z","shell.execute_reply":"2024-05-08T23:27:22.713016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Handling missing values\n# Imputing missing values with the mean for continuous variables and mode for categorical variables\nfor col in df.columns:\n    if df[col].dtype == 'object':\n        df[col].fillna(df[col].mode()[0], inplace=True)\n    else:\n        df[col].fillna(df[col].mean(), inplace=True)\n\n# Checking for missing values before imputation\nmissing_values = df.isnull().sum()\n\n# Rechecking for missing values after imputation\nmissing_values_after = df.isnull().sum()\n\n(missing_values, missing_values_after)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:22.718071Z","iopub.execute_input":"2024-05-08T23:27:22.718416Z","iopub.status.idle":"2024-05-08T23:27:22.898524Z","shell.execute_reply.started":"2024-05-08T23:27:22.718388Z","shell.execute_reply":"2024-05-08T23:27:22.897282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_df =  df.isnull().sum().to_frame().rename(columns={0:\"Total No. of Missing Values\"})\nmissing_df[\"% of Missing Values\"] = round((missing_df[\"Total No. of Missing Values\"]/len( df))*100,2)\nmissing_df","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:22.900294Z","iopub.execute_input":"2024-05-08T23:27:22.900697Z","iopub.status.idle":"2024-05-08T23:27:22.944221Z","shell.execute_reply.started":"2024-05-08T23:27:22.900658Z","shell.execute_reply":"2024-05-08T23:27:22.942809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Duplicate Value","metadata":{}},{"cell_type":"code","source":"df[df.duplicated(keep=False)]","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:22.945672Z","iopub.execute_input":"2024-05-08T23:27:22.946049Z","iopub.status.idle":"2024-05-08T23:27:22.998901Z","shell.execute_reply.started":"2024-05-08T23:27:22.946019Z","shell.execute_reply":"2024-05-08T23:27:22.997642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.000202Z","iopub.execute_input":"2024-05-08T23:27:23.000622Z","iopub.status.idle":"2024-05-08T23:27:23.046260Z","shell.execute_reply.started":"2024-05-08T23:27:23.000590Z","shell.execute_reply":"2024-05-08T23:27:23.044991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop duplicate rows from the DataFrame\ndf.drop_duplicates(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.047467Z","iopub.execute_input":"2024-05-08T23:27:23.047880Z","iopub.status.idle":"2024-05-08T23:27:23.095782Z","shell.execute_reply.started":"2024-05-08T23:27:23.047851Z","shell.execute_reply":"2024-05-08T23:27:23.094233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.097456Z","iopub.execute_input":"2024-05-08T23:27:23.097976Z","iopub.status.idle":"2024-05-08T23:27:23.106310Z","shell.execute_reply.started":"2024-05-08T23:27:23.097932Z","shell.execute_reply":"2024-05-08T23:27:23.105027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the list of categorical columns\ncat_cols = df.select_dtypes(include='object').columns.tolist()\n\n# Create a DataFrame containing counts of unique values for each categorical column\ncat_df = pd.DataFrame(df[cat_cols].melt(var_name='column', value_name='value')\n                      .value_counts()).rename(columns={0: 'count'}).sort_values(by=['column', 'count'])\n\n# Display summary statistics of categorical variables\ndisplay(df[cat_cols].describe())\n\n# Display counts of unique values for each categorical column\ndisplay(cat_df)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.107934Z","iopub.execute_input":"2024-05-08T23:27:23.108370Z","iopub.status.idle":"2024-05-08T23:27:23.429724Z","shell.execute_reply.started":"2024-05-08T23:27:23.108330Z","shell.execute_reply":"2024-05-08T23:27:23.428357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.describe(include='O').T","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.431425Z","iopub.execute_input":"2024-05-08T23:27:23.432048Z","iopub.status.idle":"2024-05-08T23:27:23.553087Z","shell.execute_reply.started":"2024-05-08T23:27:23.432011Z","shell.execute_reply":"2024-05-08T23:27:23.551735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Inspect useless features\ndf.nunique().sort_values()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.554416Z","iopub.execute_input":"2024-05-08T23:27:23.554887Z","iopub.status.idle":"2024-05-08T23:27:23.604205Z","shell.execute_reply.started":"2024-05-08T23:27:23.554846Z","shell.execute_reply":"2024-05-08T23:27:23.602942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #8ABEB9; padding: 10px;\">\n    <h2 style=\"font-weight: bold;\">EXPLORATOTY DATA ANALYSIS</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Univariate Analysis","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\nfrom scipy.stats import skew\n\n# Calculate skewness for numerical columns\nskewness = df.select_dtypes(include=['int64', 'float64']).skew()\n\n# Count the number of numerical columns\nnum_cols_count = len(df.select_dtypes(include=['int64', 'float64']).columns)\n\n# Determine the layout for subplots\nnum_rows = (num_cols_count + 3) // 4  # Adjust the number of columns in each row\nnum_cols = min(4, num_cols_count)  # Maximum of 4 columns in each row\n\n# Plot histograms for numerical columns to visualize distributions and identify anomalies\nfig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 10))\n\n# Convert axes to a 2D array if it's 1D\nif axes.ndim == 1:\n    axes = np.array([axes])\n\nfor i in range(num_rows):\n    for j in range(num_cols):\n        col_idx = i * num_cols + j\n        if col_idx < num_cols_count:\n            col = df.select_dtypes(include=['int64', 'float64']).columns[col_idx]\n            axes[i, j].hist(df[col], bins=15, color='green', alpha=0.7)\n            axes[i, j].set_title(f'{col}')\n            axes[i, j].set_xlabel(col)\n            axes[i, j].set_ylabel('Frequency')\n            \n            # Compute skewness\n            skew_val = skewness[col]\n            \n            # Plot skewness value in the center of plot\n            axes[i, j].text(0.5, 0.5, f'Skewness: {skew_val:.2f}', horizontalalignment='center',\n                            verticalalignment='center', transform=axes[i, j].transAxes, fontsize=10, color='red')\n\n# Hide empty subplots\nfor ax in axes.flat[num_cols_count:]:\n    ax.set_visible(False)\n\nplt.tight_layout()\nplt.show()\n\n# Print skewness values\nprint(\"Skewness:\")\nprint(skewness)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:23.605671Z","iopub.execute_input":"2024-05-08T23:27:23.605992Z","iopub.status.idle":"2024-05-08T23:27:24.518184Z","shell.execute_reply.started":"2024-05-08T23:27:23.605966Z","shell.execute_reply":"2024-05-08T23:27:24.516904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the boxplot with rotated text labels\ndf.plot(kind='box', rot=45,color='green')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:24.519638Z","iopub.execute_input":"2024-05-08T23:27:24.520028Z","iopub.status.idle":"2024-05-08T23:27:24.832696Z","shell.execute_reply.started":"2024-05-08T23:27:24.519998Z","shell.execute_reply":"2024-05-08T23:27:24.831392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter numeric columns\nnumeric_cols = df.select_dtypes(include=['int64', 'float64']).columns\n\n# Plotting boxplots for each numerical feature to identify outliers\nfor column in numeric_cols:\n    plt.figure(figsize=(10, 6))\n    sns.boxplot(x=df[column],palette='rainbow')\n    plt.title(f'Boxplot of {column}')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:24.834094Z","iopub.execute_input":"2024-05-08T23:27:24.834434Z","iopub.status.idle":"2024-05-08T23:27:25.459512Z","shell.execute_reply.started":"2024-05-08T23:27:24.834400Z","shell.execute_reply":"2024-05-08T23:27:25.458228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Multivariate Analysis","metadata":{}},{"cell_type":"code","source":"# Correlation matrix\n\n# Select only the numeric columns from the DataFrame\nnumeric_columns = df.select_dtypes(include=['number'])\n\n# Calculate the correlation matrix\ncorrelation_matrix = numeric_columns.corr()\n\n# Create a heatmap to visualize the correlations\nplt.figure(figsize=(12, 8))\nsns.heatmap(correlation_matrix, annot=True, cmap='rainbow', fmt=\".2f\", linewidths=0.5)\nplt.title('Correlation Matrix')","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:25.460977Z","iopub.execute_input":"2024-05-08T23:27:25.461332Z","iopub.status.idle":"2024-05-08T23:27:25.894842Z","shell.execute_reply.started":"2024-05-08T23:27:25.461301Z","shell.execute_reply":"2024-05-08T23:27:25.893529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Heatmap Plotting\n# Select only numeric columns\nnumeric_columns = df.select_dtypes(include=['int64', 'float64'])\n\n# Calculate the correlation matrix\ncorr_matrix = numeric_columns.corr()\n\n# Filter correlation matrix to include values greater than 0.5 or less than -0.5\ncorr_matrix_filtered = corr_matrix[(corr_matrix > 0.5) | (corr_matrix < -0.5)]\n\n# Plot the heatmap with filtered correlation values\nplt.figure(figsize=(12, 10))\nsns.heatmap(corr_matrix_filtered, annot=True, cmap='rainbow', fmt=\".2f\", linewidths=0.5)\nplt.title('Correlation Heatmap of Numeric Features (|Correlation| > 0.5)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:25.901679Z","iopub.execute_input":"2024-05-08T23:27:25.902079Z","iopub.status.idle":"2024-05-08T23:27:26.269353Z","shell.execute_reply.started":"2024-05-08T23:27:25.902048Z","shell.execute_reply":"2024-05-08T23:27:26.267918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Outlier Treatment","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Sample DataFrame (replace this with your actual DataFrame)\n# df = ...\n\n# Function to remove outliers using the IQR method\ndef remove_outliers_iqr(df):\n    # Select only numeric columns\n    numeric_df = df.select_dtypes(include=['int64', 'float64'])\n    \n    # Calculate the first quartile (Q1) and third quartile (Q3)\n    Q1 = numeric_df.quantile(0.25)\n    Q3 = numeric_df.quantile(0.75)\n    \n    # Interquartile range (IQR)\n    IQR = Q3 - Q1\n    \n    # Define the lower and upper bounds for outlier detection\n    lower_bound = Q1 - 1.5 * IQR\n    upper_bound = Q3 + 1.5 * IQR\n    \n    # Identify outliers\n    outliers = ((numeric_df < lower_bound) | (numeric_df > upper_bound)).any(axis=1)\n    \n    # Count the number of outliers removed\n    num_outliers_removed = outliers.sum()\n    \n    # Filter DataFrame based on rows without outliers\n    df_no_outliers = df[~outliers]\n    \n    return df_no_outliers, num_outliers_removed\n\n# Remove outliers using IQR method and get the number of outliers removed\ndf_no_outliers, num_outliers_removed = remove_outliers_iqr(df)\n\nprint(\"Number of outliers removed:\", num_outliers_removed)\n\n# Function to plot boxplots before and after removing outliers\ndef plot_boxplots_before_after(df_before, df_after):\n    # Set up the figure\n    fig, axes = plt.subplots(nrows=1, ncols=2, figsize=(12, 6))\n    \n    # Boxplot before removing outliers (blue color)\n    sns.boxplot(data=df_before, ax=axes[0], color='blue')\n    axes[0].set_title('Before Removing Outliers')\n    \n    # Boxplot after removing outliers (green color)\n    sns.boxplot(data=df_after, ax=axes[1], color='green')\n    axes[1].set_title('After Removing Outliers')\n    \n    # Adjust layout\n    plt.tight_layout()\n    plt.show()\n\n# Plot boxplots before and after outlier removal\nplot_boxplots_before_after(df, df_no_outliers)\nprint(\"Number of outliers removed:\", num_outliers_removed)","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:26.270837Z","iopub.execute_input":"2024-05-08T23:27:26.271192Z","iopub.status.idle":"2024-05-08T23:27:26.937268Z","shell.execute_reply.started":"2024-05-08T23:27:26.271161Z","shell.execute_reply":"2024-05-08T23:27:26.935820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# first simple plot of locations\nplt.figure(figsize=(12,6))\nsns.scatterplot(x='longitude', y='latitude', \n                data=df,\n                color='green')\nplt.grid()\nplt.show()","metadata":{"papermill":{"duration":0.398663,"end_time":"2024-04-24T14:45:58.661762","exception":false,"start_time":"2024-04-24T14:45:58.263099","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:26.938769Z","iopub.execute_input":"2024-05-08T23:27:26.939131Z","iopub.status.idle":"2024-05-08T23:27:27.270101Z","shell.execute_reply.started":"2024-05-08T23:27:26.939100Z","shell.execute_reply":"2024-05-08T23:27:27.268839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# eval frequencies\nbird_freq = df.primary_label.value_counts()\nbird_freq","metadata":{"papermill":{"duration":0.049158,"end_time":"2024-04-24T14:45:58.736387","exception":false,"start_time":"2024-04-24T14:45:58.687229","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:27.271918Z","iopub.execute_input":"2024-05-08T23:27:27.272355Z","iopub.status.idle":"2024-05-08T23:27:27.287917Z","shell.execute_reply.started":"2024-05-08T23:27:27.272317Z","shell.execute_reply":"2024-05-08T23:27:27.286681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# secondary labels\ndf.secondary_labels.value_counts()","metadata":{"papermill":{"duration":0.046909,"end_time":"2024-04-24T14:45:58.810242","exception":false,"start_time":"2024-04-24T14:45:58.763333","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:27.289355Z","iopub.execute_input":"2024-05-08T23:27:27.290373Z","iopub.status.idle":"2024-05-08T23:27:27.308630Z","shell.execute_reply.started":"2024-05-08T23:27:27.290339Z","shell.execute_reply":"2024-05-08T23:27:27.307386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.cm as cm\n\n# ratings\nplt.figure(figsize=(10,4))\ndf.rating.value_counts().sort_index().plot(kind='bar', color=cm.viridis(np.linspace(0, 1, len(df.rating.unique()))))\nplt.title('Ratings')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:27:27.310223Z","iopub.execute_input":"2024-05-08T23:27:27.310693Z","iopub.status.idle":"2024-05-08T23:27:27.662228Z","shell.execute_reply.started":"2024-05-08T23:27:27.310651Z","shell.execute_reply":"2024-05-08T23:27:27.660872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# select first 10 categories and plot in color\ndf_select = df[df.primary_label.isin(bird_freq[0:9+1].index)]\nplt.figure(figsize=(12,6))\nsns.scatterplot(x='longitude', y='latitude', hue='primary_label', data=df_select, palette='rainbow')\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.) # move legend out of the plot area\nplt.grid()\nplt.show()","metadata":{"papermill":{"duration":0.556016,"end_time":"2024-04-24T14:45:59.681987","exception":false,"start_time":"2024-04-24T14:45:59.125971","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:27.663846Z","iopub.execute_input":"2024-05-08T23:27:27.664299Z","iopub.status.idle":"2024-05-08T23:27:28.366306Z","shell.execute_reply.started":"2024-05-08T23:27:27.664253Z","shell.execute_reply":"2024-05-08T23:27:28.365057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# select next 10 categories and plot in color\ndf_select = df[df.primary_label.isin(bird_freq[10:19+1].index)]\nplt.figure(figsize=(12,6))\nsns.scatterplot(x='longitude', y='latitude', hue='primary_label', data=df_select, palette='rainbow')\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.)\nplt.grid()\nplt.show()","metadata":{"papermill":{"duration":0.455096,"end_time":"2024-04-24T14:46:00.167412","exception":false,"start_time":"2024-04-24T14:45:59.712316","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:28.368264Z","iopub.execute_input":"2024-05-08T23:27:28.368689Z","iopub.status.idle":"2024-05-08T23:27:29.093812Z","shell.execute_reply.started":"2024-05-08T23:27:28.368655Z","shell.execute_reply":"2024-05-08T23:27:29.092633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# select next 10 categories and plot in color\ndf_select = df[df.primary_label.isin(bird_freq[20:29+1].index)]\nplt.figure(figsize=(12,6))\nsns.scatterplot(x='longitude', y='latitude', hue='primary_label', data=df_select, palette='rainbow')\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.)\nplt.grid()\nplt.show()","metadata":{"papermill":{"duration":0.454872,"end_time":"2024-04-24T14:46:00.660346","exception":false,"start_time":"2024-04-24T14:46:00.205474","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:29.095438Z","iopub.execute_input":"2024-05-08T23:27:29.095888Z","iopub.status.idle":"2024-05-08T23:27:29.794839Z","shell.execute_reply.started":"2024-05-08T23:27:29.095851Z","shell.execute_reply":"2024-05-08T23:27:29.792836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# select next 10 categories and plot in color\ndf_select = df[df.primary_label.isin(bird_freq[30:39+1].index)]\nplt.figure(figsize=(12,6))\nsns.scatterplot(x='longitude', y='latitude', hue='primary_label', data=df_select, palette='rainbow')\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.)\nplt.grid()\nplt.show()","metadata":{"papermill":{"duration":0.427775,"end_time":"2024-04-24T14:46:01.127192","exception":false,"start_time":"2024-04-24T14:46:00.699417","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-08T23:27:29.796680Z","iopub.execute_input":"2024-05-08T23:27:29.797090Z","iopub.status.idle":"2024-05-08T23:27:30.489629Z","shell.execute_reply.started":"2024-05-08T23:27:29.797055Z","shell.execute_reply":"2024-05-08T23:27:30.488275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #8ABEB9; padding: 10px;\">\n    <h2 style=\"font-weight: bold;\">INTERACTIVE MAP</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"import folium\n\nbirdclef2024 = 'greegr'\ndf_example = df[df.primary_label.isin([birdclef2024])]\ndf_example = df_example.dropna(axis=0, subset=['latitude','longitude'])\n\n# interactive map\nzoom_factor = 1.9\nmy_map = folium.Map(location=[0,0], zoom_start=zoom_factor)\n\nfor i in range(0,df_example.shape[0]):\n    folium.Circle(\n        location=[df_example.iloc[i]['latitude'], df_example.iloc[i]['longitude']],\n        radius=np.sqrt(df_example.iloc[i]['rating'])*25000,\n        color='red',\n        weight=1,\n        popup='label: ' + df_example.iloc[i]['primary_label'] + '<br>' +\n              'sec_labels: ' + df_example.iloc[i]['secondary_labels'] + '<br>' +\n              'type: ' + df_example.iloc[i]['type'] + '<br>' +\n              'URL: ' + df_example.iloc[i]['url'],\n        fill=True,\n        fill_color='red').add_to(my_map)\n\nmy_map # display","metadata":{"execution":{"iopub.status.busy":"2024-05-08T23:32:48.753329Z","iopub.execute_input":"2024-05-08T23:32:48.753846Z","iopub.status.idle":"2024-05-08T23:32:49.457776Z","shell.execute_reply.started":"2024-05-08T23:32:48.753815Z","shell.execute_reply":"2024-05-08T23:32:49.456414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}