{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":46105,"databundleVersionId":5087314,"isSourceIdPinned":false,"sourceType":"competition"},{"sourceId":11789986,"sourceType":"datasetVersion","datasetId":7402874}],"dockerImageVersionId":31012,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Introduction\n\nThis project is based on the Kaggle competition \"Isolated Sign Language Recognition.\"  \nThe objective is to explore and understand the provided dataset for isolated sign recognition.  \n\n- **Competition Link**: [ASL Recognition - Kaggle](https://www.kaggle.com/competitions/asl-signs)\n- **Dataset Link**: [ASL Dataset](https://www.kaggle.com/competitions/asl-signs/data)\n\nIn this notebook, we focus on:\n- Documentation of the dataset structure\n- Understanding and key observations about the data","metadata":{}},{"cell_type":"markdown","source":"# 2. Dataset description\n\n## 2.1 Files and Directories\n\n- **train_landmark_files/**:  \n  A directory containing subfolders organized by `participant_id`. Each subfolder contains `.parquet` files storing the landmark data for sequences.\n\n  The structure is:\n  Each `.parquet` file contains multiple rows with the following columns:\n\n- `frame`: The frame number in the raw video.\n- `row_id`: A unique identifier for the row.\n- `type`: The type of landmark — one of ['face', 'left_hand', 'pose', 'right_hand'].\n- `landmark_index`: The landmark index number.\n- `x`, `y`, `z`: Normalized spatial coordinates of the landmark.\n\n> Notes:\n> - The landmarks were extracted using the **MediaPipe Holistic** model.\n> - Not all frames necessarily contain detected hands or face landmarks.\n> - The `z` coordinate (depth) is less reliable and can be ignored if needed.\n> - Landmark data should not be used to identify or re-identify individuals.\n\n- **train.csv**:  \n  Contains metadata for each training sample. It has the following columns:\n  - `path`: The path to the landmark file (relative path).\n  - `participant_id`: A unique identifier for the data contributor.\n  - `sequence_id`: A unique identifier for the landmark sequence.\n  - `sign`: The label for the landmark sequence.\n\n\n- **sign_to_prediction_index_map.json**:  \n  A JSON file that provides a mapping between each sign label and a prediction index (integer value).  ","metadata":{}},{"cell_type":"markdown","source":"# 3. Load and Understand train.csv\n\nIn this section, we load the `train.csv` file and explore its structure.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:27.059754Z","iopub.execute_input":"2025-05-13T16:15:27.060098Z","iopub.status.idle":"2025-05-13T16:15:31.741125Z","shell.execute_reply.started":"2025-05-13T16:15:27.060072Z","shell.execute_reply":"2025-05-13T16:15:31.739874Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_csv_path = \"/kaggle/input/asl-signs/train.csv\"\ntrain_df = pd.read_csv(train_csv_path)\n\nprint('Shape of train.csv ',train_df.shape)\n\nprint('Data types and not null count')\ntrain_df.info()\n\nprint(\"\\nMissing Values in Each Column:\")\nprint(train_df.isnull().sum())\n\nprint(\"\\nNumber of Unique Signs (Classes):\", train_df['sign'].nunique())\n\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:31.743795Z","iopub.execute_input":"2025-05-13T16:15:31.744316Z","iopub.status.idle":"2025-05-13T16:15:32.095991Z","shell.execute_reply.started":"2025-05-13T16:15:31.744287Z","shell.execute_reply":"2025-05-13T16:15:32.094801Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Summary of train.csv\n\n- **Columns**: `path`, `participant_id`, `sequence_id`, `sign`\n- **Number of Samples**: [94477 rows] and [4 columns]\n- **Missing Values**: None\n- **Number of Unique Signs**: [250 classes]\n\nEach row in the `train.csv` corresponds to one landmark sequence and its associated sign label.  \nThe `path` points to a `.parquet` file containing detailed landmark data for each frame in the sequence.","metadata":{}},{"cell_type":"markdown","source":"## 3.1 Exploratory Data Analysis using 'ydata-profiling'","metadata":{}},{"cell_type":"code","source":"from ydata_profiling import ProfileReport\n\nprofile = ProfileReport(train_df, title=\"ASL Sign Language Profile Report\", explorative=True)\nprofile.to_notebook_iframe()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:32.097253Z","iopub.execute_input":"2025-05-13T16:15:32.097658Z","iopub.status.idle":"2025-05-13T16:15:45.212360Z","shell.execute_reply.started":"2025-05-13T16:15:32.097624Z","shell.execute_reply":"2025-05-13T16:15:45.210890Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.2 Visualizing Class Imbalance: Top 20 and Bottom 20 Most Frequent Signs","metadata":{}},{"cell_type":"code","source":"sign_classes = train_df['sign'].value_counts()\n\n#Barplot for TOP 20 sign counts\nplt.figure(figsize=(15,5))\ntop_20 = sign_classes.head(20)\nsns.barplot(x=top_20.index, y=top_20.values)\nplt.title('Top 20 most frequent signs')\nplt.xlabel('Sign class')\nplt.ylabel('Frequency')\nplt.show()\n\n#Barplot for Bottom 20 sign counts\nplt.figure(figsize=(15,5))\nbottom_20 = sign_classes.tail(20)\nsns.barplot(x=bottom_20.index, y=bottom_20.values)\nplt.title('Bottom 20 least frequent signs')\nplt.xlabel('Sign class')\nplt.ylabel('Frequency')\nplt.show()\n\nprint('Minimum samples per sign: ',sign_classes.min())\nprint('Maximum samples per sign: ',sign_classes.max())\nprint('Median samples per sign: ',sign_classes.median())\nprint('Mean samples per sign: ',sign_classes.mean())\nprint('\\nTop 20 signs: \\n',top_20)\nprint('Bottom 20 signs: \\n',bottom_20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:45.214070Z","iopub.execute_input":"2025-05-13T16:15:45.215392Z","iopub.status.idle":"2025-05-13T16:15:46.175967Z","shell.execute_reply.started":"2025-05-13T16:15:45.215360Z","shell.execute_reply":"2025-05-13T16:15:46.174928Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.3 Participant Distribution","metadata":{}},{"cell_type":"code","source":"participant_counts = train_df['participant_id'].value_counts()\n\n#Barplot for number of sequences per participant\nplt.figure(figsize=(15,5))\nsns.barplot(x=participant_counts.index, y=participant_counts.values)\nplt.title('Number of sequences per participant')\nplt.xlabel('Participant ID')\nplt.ylabel('Number of sequences')\nplt.xticks(rotation=90)\nplt.show()\n\n# Number of unique signs per participant\nparticipant_signs = train_df.groupby('participant_id')['sign'].nunique()\nplt.figure(figsize=(15,5))\nsns.barplot(x=participant_signs.index, y=participant_signs.values)\nplt.title('Number of unique signs per participant')\nplt.xlabel('Participant_ID')\nplt.ylabel('Number of unique signs')\nplt.show()\n\nprint('Number of unique participant_ids:', train_df['participant_id'].nunique())\nprint('Min sequences per participant:', participant_counts.min())\nprint('Max sequences per participant:', participant_counts.max())\nprint('Mean sequences per participant:', participant_counts.mean())\nprint('Median sequences per participant:', participant_counts.median())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:46.177763Z","iopub.execute_input":"2025-05-13T16:15:46.178157Z","iopub.status.idle":"2025-05-13T16:15:46.797592Z","shell.execute_reply.started":"2025-05-13T16:15:46.178121Z","shell.execute_reply":"2025-05-13T16:15:46.796637Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4. Load and Understand train_landmark_files\n\n### In this section we combine all the `.parquet` files to create a single dataframe for further analysis","metadata":{}},{"cell_type":"markdown","source":"## 4.1 Combining all `.parquet` files.","metadata":{}},{"cell_type":"code","source":"landmark_directory = '/kaggle/input/asl-signs/train_landmark_files'\ndirectory_list = os.listdir(landmark_directory)\nprint('List of directories inside train_landmark_files: \\n',directory_list)\n\nparticipant_dirs = []\nfor directory in directory_list:\n    directory_path = os.path.join(landmark_directory, directory)\n\n    if os.path.isdir(directory_path):\n        participant_dirs.append(directory)\n\nprint(f'Number of participant folders: {len(participant_dirs)}')\nprint('Participant IDs: ',participant_dirs)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:46.798939Z","iopub.execute_input":"2025-05-13T16:15:46.799376Z","iopub.status.idle":"2025-05-13T16:15:46.822706Z","shell.execute_reply.started":"2025-05-13T16:15:46.799350Z","shell.execute_reply":"2025-05-13T16:15:46.821469Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4.1.1 Below is the code to create a combined dataframe from all the participants. We have taken 100 random `.parquet` from each participants to perform EDA.","metadata":{}},{"cell_type":"code","source":"# import os\n# import pandas as pd\n# import random\n\n# # Set the target sample size:  \n# # We want to load a maximum of 4000 files in total\n# FILE_LIMIT = 210\n\n# # Set the per participant limit:  \n# # From each participant, we want at most 200 files\n# PER_PARTICIPANT_LIMIT = 10\n\n# # Initialize a counter to track how many files have been loaded in total so far\n# files_loaded = 0\n\n# # Set the output file path where we will save the final combined dataset\n# output_file_path = '/kaggle/working/sample_landmarks_data.parquet'\n\n# # Initialize an empty DataFrame to hold all combined participant data\n# landmarks_df = pd.DataFrame()\n\n# # Start looping through each participant folder\n# for participant in participant_dirs:\n    \n#     # Construct the full path to the current participant's folder\n#     participant_path = os.path.join(landmark_directory, participant)\n    \n#     # List all files in the current participant's folder\n#     file_list = os.listdir(participant_path)\n\n#     # Initialize an empty list to store only `.parquet` file names\n#     parquet_files = []\n    \n#     # Loop through each file in the folder and filter `.parquet` files\n#     for file in file_list:\n#         if file.endswith('.parquet'):\n#             parquet_files.append(file)\n\n#     # Determine how many files we can sample from this participant\n#     if len(parquet_files) >= PER_PARTICIPANT_LIMIT:\n#         number_of_files_to_sample = PER_PARTICIPANT_LIMIT\n#     else:\n#         number_of_files_to_sample = len(parquet_files)\n\n#     # If there are files to sample, randomly select them using random.sample\n#     if number_of_files_to_sample > 0:\n#         sample_files = random.sample(parquet_files, number_of_files_to_sample)\n#     else:\n#         sample_files = []  # In case participant folder is empty\n\n#     # Initialize a DataFrame to hold this participant’s sampled files\n#     participant_df = pd.DataFrame()\n\n#     # Loop over each sampled `.parquet` file for this participant\n#     for parquet_file in sample_files:\n        \n#         # Construct the full path to the current `.parquet` file\n#         file_path = os.path.join(participant_path, parquet_file)\n        \n#         # Read the `.parquet` file into a DataFrame\n#         df = pd.read_parquet(file_path)\n\n#         # Add a column indicating the participant ID\n#         df['participant_id'] = participant\n\n#         # Add a column indicating the file name\n#         df['file_name'] = parquet_file\n\n#         # Concatenate this file’s data to the participant's DataFrame\n#         participant_df = pd.concat([participant_df, df], ignore_index=True)\n\n#         # Increment the **global file counter** since we loaded one more file\n#         files_loaded = files_loaded + 1\n\n#         # Check if we've reached the global file limit (FILE_LIMIT)\n#         if files_loaded >= FILE_LIMIT:\n#             break  # Stop loading more files\n\n#     # After processing this participant, append their data to the main DataFrame\n#     landmarks_df = pd.concat([landmarks_df, participant_df], ignore_index=True)\n\n#     # Delete the participant DataFrame to free memory\n#     del participant_df\n\n#     # Check again if we've reached the global file limit\n#     if files_loaded >= FILE_LIMIT:\n#         print(f'Reached file limit of {FILE_LIMIT}. Stopping...')\n#         break\n\n# # After processing, save the combined DataFrame to a `.parquet` file on disk\n# landmarks_df.to_parquet(output_file_path, index=False)\n\n# print(f'Shape of combined sampled landmark data: {landmarks_df.shape}')\n# landmarks_df.head()","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-05-13T16:15:46.826590Z","iopub.execute_input":"2025-05-13T16:15:46.826898Z","iopub.status.idle":"2025-05-13T16:15:46.833719Z","shell.execute_reply.started":"2025-05-13T16:15:46.826876Z","shell.execute_reply":"2025-05-13T16:15:46.832569Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 4.1.2 Read the `sample_landmarks_data.parquet` file","metadata":{}},{"cell_type":"code","source":"landmarks_df = pd.read_parquet('/kaggle/input/sample-landmarks-dataset/sample_landmarks_data.parquet')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:46.834505Z","iopub.execute_input":"2025-05-13T16:15:46.834843Z","iopub.status.idle":"2025-05-13T16:15:49.659739Z","shell.execute_reply.started":"2025-05-13T16:15:46.834821Z","shell.execute_reply":"2025-05-13T16:15:49.658729Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2 EDA of sample_landmarks_data","metadata":{}},{"cell_type":"code","source":"print('Shape of landmarks_df: ',landmarks_df.shape)\nprint('---------------------------------------------')\nprint(landmarks_df.info())\nprint('---------------------------------------------')\nprint('Missing values in landmarks_df: ',landmarks_df.isna().sum())\nprint('---------------------------------------------')\nlandmarks_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:49.660965Z","iopub.execute_input":"2025-05-13T16:15:49.661279Z","iopub.status.idle":"2025-05-13T16:15:50.593816Z","shell.execute_reply.started":"2025-05-13T16:15:49.661258Z","shell.execute_reply":"2025-05-13T16:15:50.592800Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"landmarks_df[['frame','landmark_index','x','y','z']].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:50.594924Z","iopub.execute_input":"2025-05-13T16:15:50.595286Z","iopub.status.idle":"2025-05-13T16:15:51.621783Z","shell.execute_reply.started":"2025-05-13T16:15:50.595262Z","shell.execute_reply":"2025-05-13T16:15:51.620643Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2.1 Univariate Analysis of `frame`, `x`, `y`, `z` columns\n\n### Distribution of `frame`, `x`, `y`, `z` columns","metadata":{}},{"cell_type":"code","source":"# Histogram to see distribution of 'x' coordinate\nplt.figure(figsize=(20,6))\n\nplt.subplot(1, 3, 1)\nsns.histplot(landmarks_df['x'], bins=50, kde=True, color='skyblue')\nplt.title(\"Distribution of 'x'\")\nplt.xlabel('x')\n\nplt.subplot(1, 3, 2)\nsns.histplot(landmarks_df['y'], bins=50, kde=True, color='lightgreen')\nplt.title(\"Distribution of 'y'\")\nplt.xlabel('y')\n\nplt.subplot(1, 3, 3)\nsns.histplot(landmarks_df['z'], bins=50, kde=True, color='salmon')\nplt.title(\"Distribution of 'z'\")\nplt.xlabel('z')\n\nplt.tight_layout()\nplt.show()\n\n# Boxplot to see outliers\nplt.figure(figsize=(25, 8))\n\nplt.subplot(1, 3, 1)\nsns.boxplot(x=landmarks_df['x'], color='lightblue')\nplt.title(\"Boxplot of x\")\n\nplt.subplot(1, 3, 2)\nsns.boxplot(x=landmarks_df['y'], color='lightgreen')\nplt.title(\"Boxplot of y\")\n\nplt.subplot(1, 3, 3)\nsns.boxplot(x=landmarks_df['z'], color='salmon')\nplt.title(\"Boxplot of z\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:15:51.622823Z","iopub.execute_input":"2025-05-13T16:15:51.623164Z","iopub.status.idle":"2025-05-13T16:16:48.828854Z","shell.execute_reply.started":"2025-05-13T16:15:51.623141Z","shell.execute_reply":"2025-05-13T16:16:48.827610Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Skewness of `x`, `y`, `z` columns\n\n`x` is fairly symmetric\n`y` is positive skewed\n`z` is negative skewed","metadata":{}},{"cell_type":"code","source":"print(\"Skewness of x:\", landmarks_df['x'].skew())\nprint(\"Skewness of y:\", landmarks_df['y'].skew())\nprint(\"Skewness of z:\", landmarks_df['z'].skew())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:48.830020Z","iopub.execute_input":"2025-05-13T16:16:48.830311Z","iopub.status.idle":"2025-05-13T16:16:49.002034Z","shell.execute_reply.started":"2025-05-13T16:16:48.830288Z","shell.execute_reply":"2025-05-13T16:16:49.000743Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2.2 Finding Outliers using IQR method","metadata":{}},{"cell_type":"code","source":"Q1 = landmarks_df['x'].quantile(0.25)  # 25th percentile\nQ3 = landmarks_df['x'].quantile(0.75)  # 75th percentile\nIQR = Q3 - Q1\nprint(f\"Q1: {Q1}, Q3: {Q3}, IQR: {IQR}\")\n\nlower_bound = Q1 - 1.5 * IQR\nupper_bound = Q3 + 1.5 * IQR\nprint(f\"Lower Bound: {lower_bound}\")\nprint(f\"Upper Bound: {upper_bound}\")\n\noutliers_lower = landmarks_df[landmarks_df['x'] < lower_bound]\noutliers_upper = landmarks_df[landmarks_df['x'] > upper_bound]\n\noutliers = pd.concat([outliers_lower, outliers_upper])\nprint(f\"Outliers detected: {outliers.shape[0]} rows\")\noutliers.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:49.003697Z","iopub.execute_input":"2025-05-13T16:16:49.004044Z","iopub.status.idle":"2025-05-13T16:16:49.210724Z","shell.execute_reply.started":"2025-05-13T16:16:49.004015Z","shell.execute_reply":"2025-05-13T16:16:49.209827Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\nsns.histplot(landmarks_df['frame'], bins=100, kde=False, color='teal')\nplt.title('Distribution of Frame Numbers')\nplt.xlabel('Frame Number')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:49.211755Z","iopub.execute_input":"2025-05-13T16:16:49.212099Z","iopub.status.idle":"2025-05-13T16:16:51.625241Z","shell.execute_reply.started":"2025-05-13T16:16:49.212070Z","shell.execute_reply":"2025-05-13T16:16:51.624173Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2.3 Univariate Analysis for Categorical columns\n### Categorical columns : `type` and `landmark_index`","metadata":{}},{"cell_type":"code","source":"print(\"Value counts for 'type' column: \",landmarks_df['type'].value_counts())\nplt.figure(figsize=(10, 5))\nsns.countplot(data=landmarks_df, x='type', palette='Set1')\nplt.title('Distribution of Landmark Types')\nplt.xlabel('Type of Landmark')\nplt.ylabel('Count')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:51.626477Z","iopub.execute_input":"2025-05-13T16:16:51.626816Z","iopub.status.idle":"2025-05-13T16:16:54.062498Z","shell.execute_reply.started":"2025-05-13T16:16:51.626781Z","shell.execute_reply":"2025-05-13T16:16:54.061507Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Count NaNs per column for each type","metadata":{}},{"cell_type":"code","source":"missing_by_type = landmarks_df.groupby('type')[['x', 'y', 'z']].apply(lambda group: group.isna().sum())\nprint(missing_by_type,end='\\n')\n\ntotal_values_by_type = landmarks_df.groupby('type')[['x', 'y', 'z']].count()\nprint('Total values by type:\\n',total_values_by_type)\n\n\nmissing_percent = (missing_by_type / (missing_by_type + total_values_by_type)) * 100\nprint(\"Missing percentage:\\n\", missing_percent)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:54.063629Z","iopub.execute_input":"2025-05-13T16:16:54.063979Z","iopub.status.idle":"2025-05-13T16:16:54.802821Z","shell.execute_reply.started":"2025-05-13T16:16:54.063956Z","shell.execute_reply":"2025-05-13T16:16:54.801874Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"missing_by_type.plot(kind='bar', figsize=(10,6), colormap='Set2')\nplt.title(\"Missing Values per Coordinate by Type\")\nplt.xlabel(\"Landmark Type\")\nplt.ylabel(\"Number of Missing Values\")\nplt.xticks(rotation=45)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:54.803868Z","iopub.execute_input":"2025-05-13T16:16:54.804213Z","iopub.status.idle":"2025-05-13T16:16:55.065746Z","shell.execute_reply.started":"2025-05-13T16:16:54.804183Z","shell.execute_reply":"2025-05-13T16:16:55.064652Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Top 20 and Bottom 20 Landmark index barplot","metadata":{}},{"cell_type":"code","source":"top_landmarks = landmarks_df['landmark_index'].value_counts().head(20)\nplt.figure(figsize=(10,5))\nsns.barplot(x=top_landmarks.index, y=top_landmarks.values, palette='viridis')\nplt.title('Top 20 Most Frequent Landmark Indexes')\nplt.xlabel('Landmark Index')\nplt.ylabel('Count')\nplt.show()\n\nbottom_landmarks = landmarks_df['landmark_index'].value_counts().tail(20)\nplt.figure(figsize=(10,5))\nsns.barplot(x=bottom_landmarks.index, y=bottom_landmarks.values, palette='viridis')\nplt.title('Bottom 20 Least Frequent Landmark Indexes')\nplt.xlabel('Landmark Index')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:55.066793Z","iopub.execute_input":"2025-05-13T16:16:55.067176Z","iopub.status.idle":"2025-05-13T16:16:56.045907Z","shell.execute_reply.started":"2025-05-13T16:16:55.067147Z","shell.execute_reply":"2025-05-13T16:16:56.044876Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.3 Bivariate Analysis","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.3.1 Scatter plots for `x`, `y`, `z` columns\n","metadata":{}},{"cell_type":"code","source":"sample_df = landmarks_df[['x', 'y', 'z']].dropna().sample(10000, random_state=42)\nsns.pairplot(sample_df, plot_kws={'alpha':0.4, 's':10})\nplt.suptitle(\"Scatter plots among x, y, z\", y=1.02)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:16:56.046825Z","iopub.execute_input":"2025-05-13T16:16:56.047104Z","iopub.status.idle":"2025-05-13T16:17:00.927445Z","shell.execute_reply.started":"2025-05-13T16:16:56.047069Z","shell.execute_reply":"2025-05-13T16:17:00.926295Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##4.3.2 Correlation Matrix (for `x`, `y`, `z`)","metadata":{}},{"cell_type":"code","source":"numerical_cols = ['x','y','z']\nplt.figure(figsize=(8,4))\nsns.heatmap(landmarks_df[numerical_cols].corr(), annot=True, cmap='coolwarm', fmt='.2f')\nplt.title('Correlation Matrix (for x, y, z)')\nplt.show","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:17:00.928647Z","iopub.execute_input":"2025-05-13T16:17:00.928960Z","iopub.status.idle":"2025-05-13T16:17:01.385237Z","shell.execute_reply.started":"2025-05-13T16:17:00.928935Z","shell.execute_reply":"2025-05-13T16:17:01.384192Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.3.3 Distribution of (`x`, `y`, `z`) with respect to `type`.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.boxplot(x='type', y='x', data=landmarks_df, palette='Set2')\nplt.title(\"Distribution of 'x' coordinate by Landmark Type\")\nplt.xlabel('Landmark Type')\nplt.ylabel('X Coordinate')\nplt.show()\n\nplt.figure(figsize=(12, 6))\nsns.boxplot(x='type', y='y', data=landmarks_df, palette='Set2')\nplt.title(\"Distribution of 'y' coordinate by Landmark Type\")\nplt.xlabel('Landmark Type')\nplt.ylabel('Y Coordinate')\nplt.show()\n\nplt.figure(figsize=(12, 6))\nsns.boxplot(x='type', y='z', data=landmarks_df, palette='Set2')\nplt.title(\"Distribution of 'z' coordinate by Landmark Type\")\nplt.xlabel('Landmark Type')\nplt.ylabel('Z Coordinate')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:22:58.919911Z","iopub.execute_input":"2025-05-13T16:22:58.920347Z","iopub.status.idle":"2025-05-13T16:23:05.355709Z","shell.execute_reply.started":"2025-05-13T16:22:58.920320Z","shell.execute_reply":"2025-05-13T16:23:05.354526Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### IQR of `x`, `y`, `z` with respect to `type`.","metadata":{}},{"cell_type":"code","source":"iqr_stats = landmarks_df.groupby('type')[['x', 'y', 'z']].quantile([0.25, 0.75]).unstack()\niqr_stats.columns = ['Q1_x', 'Q3_x', 'Q1_y', 'Q3_y', 'Q1_z', 'Q3_z']\n\niqr_stats['IQR_x'] = iqr_stats['Q3_x'] - iqr_stats['Q1_x']\niqr_stats['IQR_y'] = iqr_stats['Q3_y'] - iqr_stats['Q1_y']\niqr_stats['IQR_z'] = iqr_stats['Q3_z'] - iqr_stats['Q1_z']\n\nprint(iqr_stats[['IQR_x', 'IQR_y', 'IQR_z']])\n\nprint('Upper and Lower bounds for each coordinates: ')\niqr_stats['lower_x'] = iqr_stats['Q1_x'] - 1.5 * iqr_stats['IQR_x']\niqr_stats['upper_x'] = iqr_stats['Q3_x'] + 1.5 * iqr_stats['IQR_x']\n\niqr_stats['lower_y'] = iqr_stats['Q1_y'] - 1.5 * iqr_stats['IQR_y']\niqr_stats['upper_y'] = iqr_stats['Q3_y'] + 1.5 * iqr_stats['IQR_y']\n\niqr_stats['lower_z'] = iqr_stats['Q1_z'] - 1.5 * iqr_stats['IQR_z']\niqr_stats['upper_z'] = iqr_stats['Q3_z'] + 1.5 * iqr_stats['IQR_z']\n\nprint('Upper bound for \"X\" coordinate')\nprint(iqr_stats['upper_x'],'\\n')\nprint('Lower bound for \"X\" coordinate')\nprint(iqr_stats['lower_x'])\n\nprint('---------------------------------')\nprint('Upper bound for \"Y\" coordinate')\nprint(iqr_stats['upper_y'],'\\n')\nprint('Lower bound for \"Y\" coordinate')\nprint(iqr_stats['lower_y'])\n\nprint('---------------------------------')\nprint('Upper bound for \"Z\" coordinate')\nprint(iqr_stats['upper_z'],'\\n')\nprint('Lower bound for \"Z\" coordinate')\nprint(iqr_stats['lower_y'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:49:11.726027Z","iopub.execute_input":"2025-05-13T16:49:11.726791Z","iopub.status.idle":"2025-05-13T16:49:13.937434Z","shell.execute_reply.started":"2025-05-13T16:49:11.726761Z","shell.execute_reply":"2025-05-13T16:49:13.936405Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"landmarks_with_bounds = landmarks_df.merge(iqr_stats, on='type')\nlandmarks_with_bounds.sample(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:50:00.301447Z","iopub.execute_input":"2025-05-13T16:50:00.301829Z","iopub.status.idle":"2025-05-13T16:50:02.866059Z","shell.execute_reply.started":"2025-05-13T16:50:00.301807Z","shell.execute_reply":"2025-05-13T16:50:02.865138Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Total number of outliers in each `x`, `y`, `z` coordinate with respect to `type`.","metadata":{}},{"cell_type":"code","source":"landmarks_with_bounds['x_outlier'] = (\n    (landmarks_with_bounds['x'] < landmarks_with_bounds['lower_x']) |\n    (landmarks_with_bounds['x'] > landmarks_with_bounds['upper_x'])\n)\n\nlandmarks_with_bounds['y_outlier'] = (\n    (landmarks_with_bounds['y'] < landmarks_with_bounds['lower_y']) |\n    (landmarks_with_bounds['y'] > landmarks_with_bounds['upper_y'])\n)\n\nlandmarks_with_bounds['z_outlier'] = (\n    (landmarks_with_bounds['z'] < landmarks_with_bounds['lower_z']) |\n    (landmarks_with_bounds['z'] > landmarks_with_bounds['upper_z'])\n)\noutlier_summary = landmarks_with_bounds.groupby('type')[['x_outlier', 'y_outlier', 'z_outlier']].sum()\nprint(outlier_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T16:50:59.006186Z","iopub.execute_input":"2025-05-13T16:50:59.006478Z","iopub.status.idle":"2025-05-13T16:50:59.455918Z","shell.execute_reply.started":"2025-05-13T16:50:59.006459Z","shell.execute_reply":"2025-05-13T16:50:59.454868Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"outlier_summary.plot(kind='bar', figsize=(10, 6))\nplt.title('Number of Outliers per Type and Coordinate Axis')\nplt.ylabel('Outlier Count')\nplt.xlabel('Landmark Type')\nplt.legend(title='Coordinate')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-13T17:05:44.391966Z","iopub.execute_input":"2025-05-13T17:05:44.392339Z","iopub.status.idle":"2025-05-13T17:05:44.698691Z","shell.execute_reply.started":"2025-05-13T17:05:44.392315Z","shell.execute_reply":"2025-05-13T17:05:44.697618Z"}},"outputs":[],"execution_count":null}]}