{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":46105,"databundleVersionId":5087314}],"dockerImageVersionId":31259,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://1.bp.blogspot.com/-QdumdOczeco/X9Fx3wTEVdI/AAAAAAAAG4Q/efRt8CO65iUyLlBIyp5MaaTNUk4Qy0xFQCLcBGAsYHQ/s762/image4.gif\">\n","metadata":{}},{"cell_type":"markdown","source":"#  The Research problem.\n- Most of deaf children born to hearing parents with no knowldge of ASL are highly at risk of language deprivation syndrome.learning ASL is time consuming and most parents work long hours and some of them do not have resources to attend classes or access learning material.To assist the language and communication abilities of deaf children and their families it is crucial to create efficient and available ASL learning resources. About 94477 vidoes were collected and the language signs was defined according to it's accurate video. This collected data will be used on development of effective and accessible ASL learning tools to support language development and communication skills for deaf children.\n\n## Understanding the problem.\n- About 94477 videos where collected and language was defined according to its video. This collected data will be used onthe development of effective and accessible ASL learning tools to support the language development and communication skills of deaf children.\n\n### The aim\n- To developa high perfoming model for PopSign game that accuratly classifies isolated ASL signs using the videos given inthe data\n\n## Hypothesis\n- LEVEL OF PROFICIENCY .Develop the level of proficiency in ASL.This will cater to learners of all levels and allow them to progress in their own pace _Socio_econimic status Currently we are not sure\n\n- Vocabulary .we need to enrich the app with more words and improve vocabulary\n\n- Graphic .improve graphic perfomance so that videoscan be clear so it will cater for also people who color_blind\n\n- Social interaction .social interaction between learners ,this can include features such as chat rooms or forums where learners can communicate with each other and practice their signing skills","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Understanding the data\n- This data is related to American sign language and the specific sign video of a particular word\n\n- Path-------------------- The path to the landmark file.\n\n- Participant_id-----------A unique identifier for the data contributor.\n\n- Sequence_id -------------A unique identifier for the landmark sequence.\n\n- Sign---------------------The label for the landmark sequenc","metadata":{}},{"cell_type":"markdown","source":"# Library ","metadata":{}},{"cell_type":"code","source":"#  Core Libraries\nimport os                      # For file and directory operations\nimport random                  # For random sampling and reproducibility control\nfrom pathlib import Path       # For handling file system paths in a clean way\n\n# Data Handling & Processing\nimport numpy as np             # Numerical operations and array manipulation\nimport pandas as pd            # Data manipulation and analysis (DataFrames)\nimport polars as pl            # High-performance DataFrame library (optional alternative to pandas)\n\n# Computer Vision\nimport cv2                     # OpenCV for image and video processing\n\n# Visualization\nimport matplotlib.pyplot as plt        # Static plotting\nimport matplotlib as mpl               # Matplotlib configurations\nfrom matplotlib.animation import FuncAnimation, ArtistAnimation  # For animations\nimport plotly.graph_objects as go      # Interactive plots (low-level API)\nimport plotly.express as px            # Quick interactive plots\n\n# Interactive & Notebook Display\nfrom ipywidgets import interact, IntSlider  # Interactive sliders in Jupyter\nfrom IPython.display import HTML, display   # Render animations and HTML in notebooks","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:49.873991Z","iopub.execute_input":"2026-02-20T13:14:49.874966Z","iopub.status.idle":"2026-02-20T13:14:49.882241Z","shell.execute_reply.started":"2026-02-20T13:14:49.874923Z","shell.execute_reply":"2026-02-20T13:14:49.880679Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Inspaction ","metadata":{}},{"cell_type":"code","source":"# Loop through all folders inside the Kaggle input directory\nfor folder in os.listdir(\"/kaggle/input\"):\n    # Print the folder name (usually the dataset name)\n    print(folder)\n    # Print the files and subfolders inside each dataset folder\n    print(os.listdir(f\"/kaggle/input/{folder}\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:49.918552Z","iopub.execute_input":"2026-02-20T13:14:49.918944Z","iopub.status.idle":"2026-02-20T13:14:49.928646Z","shell.execute_reply.started":"2026-02-20T13:14:49.918905Z","shell.execute_reply":"2026-02-20T13:14:49.927354Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import polars as pl\n\ntrain = pl.read_csv(\"/kaggle/input/asl-signs/train.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:18:58.889288Z","iopub.execute_input":"2026-02-20T13:18:58.889724Z","iopub.status.idle":"2026-02-20T13:18:59.100823Z","shell.execute_reply.started":"2026-02-20T13:18:58.889690Z","shell.execute_reply":"2026-02-20T13:18:59.099892Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Count the number of missing (null) values in each column of the Polars DataFrame\ntrain.select(\n    pl.all().null_count()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:19:00.946553Z","iopub.execute_input":"2026-02-20T13:19:00.946918Z","iopub.status.idle":"2026-02-20T13:19:00.955269Z","shell.execute_reply.started":"2026-02-20T13:19:00.946889Z","shell.execute_reply":"2026-02-20T13:19:00.954327Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#checking the datatype of column\ntrain.dtypes","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:49.952184Z","iopub.execute_input":"2026-02-20T13:14:49.952623Z","iopub.status.idle":"2026-02-20T13:14:49.976796Z","shell.execute_reply.started":"2026-02-20T13:14:49.952556Z","shell.execute_reply":"2026-02-20T13:14:49.975392Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n**OBSERVATION**\n\n    The data show that there are two data types which are objects and integers.\n    participant_id and sequence_id they are of integer type and sign and path they are of object type\n\n","metadata":{}},{"cell_type":"code","source":"# Print the shape of the dataset (rows, columns)\nprint(\"Shape:\", train.shape)\n\n# Print all column names\nprint(\"\\nColumns:\", train.columns)\n\n# Print the schema (column names with their data types)\nprint(\"\\nSchema:\", train.schema)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:49.978089Z","iopub.execute_input":"2026-02-20T13:14:49.978395Z","iopub.status.idle":"2026-02-20T13:14:49.998451Z","shell.execute_reply.started":"2026-02-20T13:14:49.978364Z","shell.execute_reply":"2026-02-20T13:14:49.997451Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"INPUT = Path(\"/kaggle/input/asl-signs\")\n\ntrain = pl.read_csv(INPUT / \"train.csv\")\n\n# Convert to pandas for easier aggre/gation with plotly\ntrain_pd = train.to_pandas()\n\n# Count number of sequences per sign\nseq_count = train_pd.groupby(\"sign\")[\"sequence_id\"].count()\n\nmeta_data_df = pd.DataFrame({\n    \"Sequences\": seq_count\n})\n\nmeta_data_df = meta_data_df.sort_values(\"Sequences\", ascending=False)\nmeta_data_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:50.001316Z","iopub.execute_input":"2026-02-20T13:14:50.001710Z","iopub.status.idle":"2026-02-20T13:14:50.117339Z","shell.execute_reply.started":"2026-02-20T13:14:50.001677Z","shell.execute_reply":"2026-02-20T13:14:50.116037Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Checking the name of the data frame different sign.\ntrain['sign'].unique()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:50.118636Z","iopub.execute_input":"2026-02-20T13:14:50.119002Z","iopub.status.idle":"2026-02-20T13:14:50.131124Z","shell.execute_reply.started":"2026-02-20T13:14:50.118960Z","shell.execute_reply":"2026-02-20T13:14:50.129629Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**OBSERVATION**\n\n    Base onthe data there are 250 unique sign that where recorded\n","metadata":{}},{"cell_type":"code","source":"# Create an empty Plotly figure\nfig = go.Figure()\n\n# Get column names from metadata DataFrame\ncols_name = meta_data_df.columns.values.tolist()\n\n# Define custom colors for each trace\ncolors = [\"#0F9D58\", \"#4285F4\", \"#F4B400\"]\n\n# Loop through each column and corresponding color\nfor color, col in zip(colors, cols_name):\n    \n    # Sort values by the current column\n    tmp = meta_data_df.sort_values(col)\n    \n    # Add a bar trace for each column\n    fig.add_trace(\n        go.Bar(\n            x=tmp.index,          # Sign labels (index)\n            y=tmp[col],           # Corresponding values\n            width=0.5,\n            name=col,\n            marker_color=color\n        )\n    )\n\n# Update layout settings\nfig.update_layout(\n    \n    # Main title configuration\n    title={\n        'text': \"ISL Distribution: All Traces\",\n        'font': dict(size=20, family=\"Georgia\", color=colors[1]),\n        'y': 0.87,\n        'x': 0.035,\n        'xanchor': 'left',\n        'yanchor': 'top'\n    },\n    \n    template=\"plotly_white\",     # Clean white theme\n    xaxis_tickangle=-45,         # Rotate x-axis labels\n    width=4000,                  # Large figure width\n    \n    # Axis labels and locking zoom\n    xaxis=dict(title='Sign', fixedrange=True),\n    yaxis=dict(title='Count', fixedrange=True),\n    \n    showlegend=True,\n    \n    # Dropdown menu for interactive trace control\n    updatemenus=[\n        dict(\n            active=0,\n            direction=\"down\",\n            pad={\"r\": 50, \"t\": 25},\n            showactive=True,\n            x=0.005,\n            xanchor=\"right\",\n            y=1.2,\n            yanchor=\"top\",\n            \n            buttons=list([\n                \n                # Show all traces\n                dict(\n                    label=\"All\",\n                    method=\"update\",\n                    args=[\n                        {\"visible\": [True, True, True]},\n                        {\"title\": \"ISL Distribution: All Traces\"}\n                    ]\n                ),\n                \n                # Show only Sequences\n                dict(\n                    label=\"Sequences\",\n                    method=\"update\",\n                    args=[\n                        {\"visible\": [True, False, False]},\n                        {\"title\": \"Distribution of Number of Sequences per Sign\"}\n                    ]\n                ),\n                \n                # Show only Frames\n                dict(\n                    label=\"Frames\",\n                    method=\"update\",\n                    args=[\n                        {\"visible\": [False, True, False]},\n                        {\"title\": \"Distribution of Number of Frames per Sign\"}\n                    ]\n                ),\n                \n                # Show only Average Frames\n                dict(\n                    label=\"Avg frames\",\n                    method=\"update\",\n                    args=[\n                        {\"visible\": [False, False, True]},\n                        {\"title\": \"Distribution of Average Frames per Sequence per Sign\"}\n                    ]\n                ),\n            ]),\n        ),\n    ]\n)\n\n# Display the interactive figure without the mode bar\nfig.show(config=dict(displayModeBar=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:52.059167Z","iopub.execute_input":"2026-02-20T13:14:52.059611Z","iopub.status.idle":"2026-02-20T13:14:52.112337Z","shell.execute_reply.started":"2026-02-20T13:14:52.059579Z","shell.execute_reply":"2026-02-20T13:14:52.110981Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Observations – ISL Distribution Plot\n\n- The dataset appears **well balanced across signs**, with each sign having a very similar number of samples.\n- Most signs range approximately between **300 and 370 sequences**, indicating low class imbalance.\n- There is no extreme outlier class with significantly higher or lower frequency.\n- The distribution is slightly increasing from left to right due to sorting, not due to imbalance.\n- This balanced structure is beneficial for training classification models, as it reduces bias toward dominant classes.\n- Minimal need for aggressive resampling techniques (e.g., oversampling or class weighting).\n- The dataset is suitable for fair evaluation using standard metrics like accuracy or macro F1-score.","metadata":{}},{"cell_type":"markdown","source":"# Landmark Detection","metadata":{}},{"cell_type":"code","source":"# Use ggplot style for better visualization aesthetics\nplt.style.use(\"ggplot\")\n\n# Define dataset path\nINPUT = Path(\"/kaggle/input/asl-signs\")\n\n# Read the training metadata file (train.csv)\ntrain = pl.read_csv(INPUT / \"train.csv\")\n\n# Extract first sample row as dictionary (contains parquet path)\nlm_data = {k: v for k, v in zip(train.columns, train.row(0))}\n\n# Read corresponding landmark parquet file\ndf_landmark = pl.read_parquet(INPUT / lm_data[\"path\"])\n\n# Get the first actual frame ID (important to avoid missing frames)\nfirst_frame_id = df_landmark[\"frame\"].unique().sort()[0]\n\n# Filter landmarks for that first frame only\nlm_first_frame = df_landmark.filter(pl.col(\"frame\") == first_frame_id)\n\n# Split landmarks into groups (face, pose, left_hand, right_hand)\nlms = lm_first_frame.partition_by(\"type\")\n\n\n# Define landmark connections (skeleton structure)\nedges = {\n    \"left_hand\": [\n        (0,1),(1,2),(2,3),(3,4),\n        (0,5),(5,6),(6,7),(7,8),\n        (0,9),(9,10),(10,11),(11,12),\n        (0,13),(13,14),(14,15),(15,16),\n        (0,17),(17,18),(18,19),(19,20)\n    ],\n    \"right_hand\": [\n        (0,1),(1,2),(2,3),(3,4),\n        (0,5),(5,6),(6,7),(7,8),\n        (0,9),(9,10),(10,11),(11,12),\n        (0,13),(13,14),(14,15),(15,16),\n        (0,17),(17,18),(18,19),(19,20)\n    ],\n    \"pose\": [\n        (11,13),(13,15),(12,14),(14,16),\n        (11,12),(11,23),(12,24),(23,24)\n    ]\n}\n\n\n# Plot landmarks for first frame\n\n# Create 2x2 subplot grid\nfig, axes = plt.subplots(2, 2, figsize=(8, 8))\naxes = axes.ravel()\n\n# Loop through each landmark type (face, hands, pose)\nfor lm, ax in zip(lms, axes):\n\n    # Get landmark type name\n    lm_type = lm[\"type\"][0]\n\n    # Scatter landmark points\n    ax.scatter(lm[\"x\"], lm[\"y\"], s=10)\n\n    # Annotate landmark index numbers (except face for clarity)\n    if lm_type != \"face\":\n        for row in lm.iter_rows(named=True):\n            x, y, idx = row[\"x\"], row[\"y\"], row[\"landmark_index\"]\n            if x is not None and y is not None:\n                ax.text(x, y, str(idx), fontsize=6)\n\n    # Draw skeletal connections between landmarks\n    if lm_type in edges:\n        for i, j in edges[lm_type]:\n            if i < len(lm) and j < len(lm):\n                x1, x2 = lm[\"x\"][i], lm[\"x\"][j]\n                y1, y2 = lm[\"y\"][i], lm[\"y\"][j]\n                if None not in (x1, x2, y1, y2):\n                    ax.plot((x1, x2), (y1, y2))\n\n    # Set subplot title\n    ax.set_title(lm_type)\n\n    # Invert y-axis to match image coordinate system\n    ax.invert_yaxis()\n\n# Adjust spacing between plots\nplt.tight_layout()\n\n# Display final figure\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:52.114762Z","iopub.execute_input":"2026-02-20T13:14:52.115223Z","iopub.status.idle":"2026-02-20T13:14:53.231548Z","shell.execute_reply.started":"2026-02-20T13:14:52.115183Z","shell.execute_reply":"2026-02-20T13:14:53.230340Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Observations\n\n## Face\n- Facial landmarks are densely distributed and form a clear facial structure.\n- Key regions (eyes, nose, mouth) are properly aligned.\n- The face appears centered with no major distortion.\n- Detection is stable and consistent across the facial area.\n\n## Left Hand\n- No landmarks detected for the left hand.\n- This may indicate that the left hand is either:\n  - Outside the frame, or  \n  - Not confidently detected by the model.\n- Missing left-hand data could impact tasks requiring both hands (e.g., sign language recognition).\n\n## Pose\n- Full-body pose landmarks are successfully detected.\n- Upper body joints (shoulders, elbows) are clearly connected.\n- Lower body joints (knees, ankles) are visible and structurally aligned.\n- The skeletal connections reflect a natural human posture.\n- Minor positional variations may indicate motion or detection noise.\n\n## Right Hand\n- All 21 right-hand landmarks are detected.\n- Finger structures are clearly defined and anatomically consistent.\n- Landmark connectivity is correct and suitable for gesture modeling.\n- Detection quality appears reliable for downstream classification tasks.\n\n## Overall Assessment\n- Strong detection performance for face, pose, and right hand.\n- Left-hand landmarks are missing.\n- Output is suitable for feature extraction and gesture-based modeling tasks.","metadata":{}},{"cell_type":"markdown","source":"## Overview of train.csv (Metadata Structure)","metadata":{}},{"cell_type":"code","source":"train_csv = pd.read_csv(\"/kaggle/input/asl-signs/train.csv\")\ntrain_csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:53.233006Z","iopub.execute_input":"2026-02-20T13:14:53.233922Z","iopub.status.idle":"2026-02-20T13:14:53.378846Z","shell.execute_reply.started":"2026-02-20T13:14:53.233882Z","shell.execute_reply":"2026-02-20T13:14:53.377803Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"critical_emergency_signs = [\n    \"bad\", \"boy\", \"child\", \"dad\", \"fireman\", \"girl\", \"go\", \"home\", \n    \"hot\", \"listen\", \"look\", \"man\", \"mom\", \"no\", \"outside\", \"owie\", \n    \"person\", \"police\", \"sick\", \"there\", \"wait\", \"water\", \"where\", \"yes\",\n    \"airplane\", \"all\", \"arm\", \"black\", \"blow\", \"blue\", \"boat\", \"car\", \n    \"close\", \"cry\", \"cut\", \"down\", \"dry\", \"ear\", \"eye\", \"face\", \"fall\", \n    \"fast\", \"feet\", \"finger\", \"green\", \"happy\", \"head\", \"hear\", \n    \"helicopter\", \"later\", \"loud\", \"mad\", \"many\", \"mouth\", \"nose\", \n    \"now\", \"open\", \"quiet\", \"red\", \"sad\", \"say\", \"see\", \"talk\", \n    \"touch\", \"up\", \"wet\", \"white\", \"who\", \"why\", \"yellow\", \"alligator\", \"animal\", \"backyard\", \"bed\", \"because\", \"after\", \"another\", \"any\", \"bedroom\", \"bee\", \"before\", \"beside\", \"bug\", \"can\", \"cat\", \"cheek\", \"chin\", \"dog\", \"drop\", \"find\", \n    \"for\", \"give\", \"glasswindow\", \"grandma\", \"grandpa\", \"hair\", \"have\", \"haveto\", \"hesheit\", \"jump\",\"if\", \n    \"into\", \"hide\", \"high\",\"like\",  \"night\", \"noisy\", \n    \"not\", \"please\", \"pool\", \"room\", \"stairs\", \"stuck\", \"think\", \"that\", \"time\", \"tree\", \"will\", \n]\n\n# filter out critical emergency signs\ndf_filtered = train_csv[train_csv['sign'].isin(critical_emergency_signs)]\ndf_filtered","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:53.380111Z","iopub.execute_input":"2026-02-20T13:14:53.380434Z","iopub.status.idle":"2026-02-20T13:14:53.406284Z","shell.execute_reply.started":"2026-02-20T13:14:53.380403Z","shell.execute_reply":"2026-02-20T13:14:53.405129Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"parquet_files = df_filtered.loc[df_filtered[\"participant_id\"] == 32319][\"path\"].to_list()\nfolder_path = \"/kaggle/input/asl-signs/\"\n\n# Read, label with path, and concatenate\ndf_parquet = pd.concat(\n    [\n        pd.read_parquet(os.path.join(folder_path, path)).assign(path=path)\n        for path in parquet_files\n    ],\n    ignore_index=True\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:14:53.408292Z","iopub.execute_input":"2026-02-20T13:14:53.408712Z","iopub.status.idle":"2026-02-20T13:15:41.059241Z","shell.execute_reply.started":"2026-02-20T13:14:53.408681Z","shell.execute_reply":"2026-02-20T13:15:41.057317Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"merged_df = df_parquet.merge(df_filtered, on=\"path\", how=\"left\")\nmerged_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:15:41.060597Z","iopub.execute_input":"2026-02-20T13:15:41.061001Z","iopub.status.idle":"2026-02-20T13:15:57.033816Z","shell.execute_reply.started":"2026-02-20T13:15:41.060968Z","shell.execute_reply":"2026-02-20T13:15:57.032511Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Interactive Frame Slider – Landmark Visualization","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\npf = pd.read_parquet(\n    \"/kaggle/input/asl-signs/train_landmark_files/26734/1000035562.parquet\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:24:10.409399Z","iopub.execute_input":"2026-02-20T13:24:10.409940Z","iopub.status.idle":"2026-02-20T13:24:10.439503Z","shell.execute_reply.started":"2026-02-20T13:24:10.409903Z","shell.execute_reply":"2026-02-20T13:24:10.438443Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Hand skeleton connections based on MediaPipe hand topology\nHAND_CONNECTIONS = [\n    (0,1),(1,2),(2,3),(3,4),\n    (0,5),(5,6),(6,7),(7,8),\n    (5,9),(9,10),(10,11),(11,12),\n    (9,13),(13,14),(14,15),(15,16),\n    (13,17),(17,18),(18,19),(19,20),\n    (0,17)\n]\n\n# Pose skeleton connections (upper & lower body joints)\nPOSE_CONNECTIONS = [\n    (11,13),(13,15),\n    (12,14),(14,16),\n    (11,12),\n    (23,24),\n    (11,23),(12,24),\n    (23,25),(24,26),\n    (25,27),(26,28)\n]\n\n\n# Function: Draw Landmarks\n\ndef draw_part(image, df_part, connections=None, color=(0,255,0)):\n    \"\"\"\n    Draw landmark points and optional skeletal connections\n    for a specific body part onto an image canvas.\n    \"\"\"\n\n    h, w, _ = image.shape\n    coords = {}\n\n    # Remove missing coordinates (NaN values)\n    df_part = df_part.dropna(subset=['x','y'])\n\n    # Draw each landmark point\n    for _, row in df_part.iterrows():\n        idx = int(row['landmark_index'])\n\n        # Convert normalized coordinates (0–1) to pixel space\n        x = int(row['x'] * w)\n        y = int(row['y'] * h)\n\n        # Ensure coordinates are inside image boundaries\n        if 0 <= x < w and 0 <= y < h:\n            coords[idx] = (x, y)\n            cv2.circle(image, (x, y), 2, color, -1)\n\n    # Draw skeletal connections if provided\n    if connections:\n        for i, j in connections:\n            if i in coords and j in coords:\n                cv2.line(image, coords[i], coords[j], color, 1)\n\n    return image\n\n\n# Function: Render Frame\n\ndef render_frame(frame_index):\n    \"\"\"\n    Render a selected frame using an interactive slider.\n    \"\"\"\n\n    # Get sorted list of unique frame IDs\n    frames = sorted(pf.frame.unique())\n\n    # Select frame number based on slider index\n    frame_number = frames[frame_index]\n\n    # Filter data for selected frame\n    f = pf[pf.frame == frame_number]\n\n    # Create blank black image canvas\n    image = np.zeros((800, 800, 3), dtype=np.uint8)\n\n    # Draw different body parts\n    image = draw_part(image, f[f.type=='face'], None, (0,255,0))\n    image = draw_part(image, f[f.type=='pose'], POSE_CONNECTIONS, (255,0,0))\n    image = draw_part(image, f[f.type=='left_hand'], HAND_CONNECTIONS, (0,0,255))\n    image = draw_part(image, f[f.type=='right_hand'], HAND_CONNECTIONS, (255,255,0))\n\n    # Display the frame\n    plt.figure(figsize=(6,6))\n    plt.imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))\n    plt.axis(\"off\")\n    plt.title(f\"Frame {frame_number}\")\n    plt.show()\n\n\n# Interactive Frame Slider\n\ninteract(\n    render_frame,\n    frame_index=IntSlider(\n        min=0,\n        max=len(pf.frame.unique()) - 1,\n        step=1,\n        value=0\n    )\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:24:11.719899Z","iopub.execute_input":"2026-02-20T13:24:11.720287Z","iopub.status.idle":"2026-02-20T13:24:11.921916Z","shell.execute_reply.started":"2026-02-20T13:24:11.720253Z","shell.execute_reply":"2026-02-20T13:24:11.921128Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Animated Sign Visualization","metadata":{}},{"cell_type":"code","source":"# Enable JavaScript-based animation rendering inside Jupyter\nplt.rcParams[\"animation.html\"] = \"jshtml\"\n\n\n# Detect Dataset Directory\n\n# Automatically detect the first folder inside Kaggle input\nfolders = os.listdir(\"/kaggle/input\")\nDATA_DIR = f\"/kaggle/input/{folders[0]}\"\nprint(\"Using DATA_DIR:\", DATA_DIR)\n\n# Load training metadata\ntrain_df = pd.read_csv(f\"{DATA_DIR}/train.csv\")\n\n\n# Select a Sign Sample\n\n# Try selecting a specific sign\nsign_name = \"thankyou\"\n\nsample_rows = train_df[train_df[\"sign\"] == sign_name]\n\n# Fallback: if sign does not exist, use first available sign\nif len(sample_rows) == 0:\n    sign_name = train_df[\"sign\"].iloc[0]\n    sample_rows = train_df[train_df[\"sign\"] == sign_name]\n\n# Get path to corresponding parquet file\npath_to_sign = sample_rows.iloc[0].path\nprint(\"Showing sign:\", sign_name)\n\n# Load landmark sequence\nsign = pd.read_parquet(f\"{DATA_DIR}/{path_to_sign}\")\n\n# Flip Y-axis to match typical Cartesian orientation\nsign[\"y\"] = sign[\"y\"] * -1\n\n\n# Define Graph Limits\n\n# Add small margin around min/max values\nxmin, xmax = sign.x.min() - 0.2, sign.x.max() + 0.2\nymin, ymax = sign.y.min() - 0.2, sign.y.max() + 0.2\n\nfig, ax = plt.subplots(figsize=(6,6))\n\n\n# Animation Function\n\ndef animation_frame(f):\n    \"\"\"\n    Draw a single frame of the sign animation.\n    \"\"\"\n    \n    # Filter dataframe for current frame\n    frame = sign[sign.frame == f]\n\n    ax.clear()\n\n    # Plot each body part separately\n    for t in ['face','left_hand','right_hand','pose']:\n        part = frame[frame.type == t]\n\n        if len(part) > 0:\n            ax.plot(part.x, part.y, '.')\n\n    # Keep consistent axis limits across frames\n    ax.set_xlim(xmin, xmax)\n    ax.set_ylim(ymin, ymax)\n    ax.set_title(sign_name)\n\n\n# Get sorted frame numbers\nframes = sorted(sign.frame.unique())\n\n# Create animation\nanimation = FuncAnimation(fig, animation_frame, frames=frames, interval=50)\n\n# Render animation in notebook\nHTML(animation.to_jshtml())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:15:57.237587Z","iopub.execute_input":"2026-02-20T13:15:57.237892Z","iopub.status.idle":"2026-02-20T13:15:58.907457Z","shell.execute_reply.started":"2026-02-20T13:15:57.237862Z","shell.execute_reply":"2026-02-20T13:15:58.906224Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Enable JS-based animation rendering inside notebook\nplt.rcParams[\"animation.html\"] = \"jshtml\"\n\n\n\n# Load Competition Data\n\n# Auto-detect dataset directory inside Kaggle input\nDATA_DIR = f\"/kaggle/input/{os.listdir('/kaggle/input')[0]}\"\n\n# Load training metadata and set sequence_id as index\ntrain_data = pd.read_csv(f\"{DATA_DIR}/train.csv\")\ntrain_data = train_data.set_index(\"sequence_id\")\n\n# Canvas size for rendering frames\nheight = 800\nwidth = 800\n\n\n# Helper Functions\n\ndef get_random_sequence_id(train_data):\n    \"\"\"Return a random sequence_id from the dataset.\"\"\"\n    return random.choice(train_data.index.tolist())\n\n\ndef read_landmark_data_by_id(sequence_id, train_data):\n    \"\"\"Load parquet landmark data for a given sequence_id.\"\"\"\n    path = train_data.loc[sequence_id][\"path\"]\n    df = pd.read_parquet(f\"{DATA_DIR}/{path}\")\n    \n    # Flip Y-axis for better visualization\n    df[\"y\"] *= -1\n    return df\n\n\ndef create_frame(data, frame_id, height=800, width=800):\n    \"\"\"Render a single frame as an image (numpy array).\"\"\"\n    \n    frame = data[data.frame == frame_id]\n\n    fig, ax = plt.subplots(figsize=(4,4))\n    \n    # Keep consistent axis limits across frames\n    ax.set_xlim(data.x.min()-0.2, data.x.max()+0.2)\n    ax.set_ylim(data.y.min()-0.2, data.y.max()+0.2)\n    ax.axis(\"off\")\n\n    # Draw landmarks as colored points\n    for t, color in zip(\n        ['face','left_hand','right_hand','pose'],\n        ['gray','red','blue','green']\n    ):\n        part = frame[frame.type == t]\n        if len(part) > 0:\n            ax.scatter(part.x, part.y, s=10, c=color)\n\n    # Convert matplotlib figure to image array\n    fig.canvas.draw()\n    image = np.asarray(fig.canvas.buffer_rgba())\n    plt.close(fig)\n\n    return image\n\n\ndef create_frames(sequence_id, train_data, height=800, width=800):\n    \"\"\"Generate all frames for a given sequence.\"\"\"\n    \n    data = read_landmark_data_by_id(sequence_id, train_data)\n    frame_ids = sorted(data['frame'].unique())\n\n    images = [\n        create_frame(data, fid, height, width)\n        for fid in frame_ids\n    ]\n\n    return np.array(images)\n\n\ndef create_animation(images, fig, ax):\n    \"\"\"Create matplotlib animation from image frames.\"\"\"\n    \n    ax.axis('off')\n    ims = []\n\n    for img in images:\n        im = ax.imshow(img, animated=True)\n        ims.append([im])\n\n    anim = mpl.animation.ArtistAnimation(\n        fig,\n        ims,\n        interval=80,\n        blit=True,\n        repeat_delay=1000\n    )\n\n    return anim\n\n\ndef play_animation(sequence_id, train_data, height, width, figsize=(4,4)):\n    \"\"\"Render and display animation for a given sequence.\"\"\"\n    \n    frames = create_frames(sequence_id, train_data, height, width)\n    sign = train_data.loc[sequence_id]['sign']\n\n    fig, ax = plt.subplots(figsize=figsize)\n    anim = create_animation(frames, fig, ax)\n\n    ax.set_title(f\"Sign: {sign}\")\n\n    display(HTML(anim.to_jshtml()))\n    plt.close()\n\n\n# Run Animation\n\n# Select random sequence\nsequence_id = get_random_sequence_id(train_data)\n\nprint(\"Sequence ID:\", sequence_id)\nprint(\"Sign:\", train_data.loc[sequence_id][\"sign\"])\n\n# Play animation\nplay_animation(sequence_id, train_data, height, width)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:15:58.910125Z","iopub.execute_input":"2026-02-20T13:15:58.910445Z","iopub.status.idle":"2026-02-20T13:16:05.890492Z","shell.execute_reply.started":"2026-02-20T13:15:58.910417Z","shell.execute_reply":"2026-02-20T13:16:05.889398Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Enable JS animation rendering\nplt.rcParams[\"animation.html\"] = \"jshtml\"\n\n\n# Load Data\n\n# Auto-detect dataset directory\nDATA_DIR = f\"/kaggle/input/{os.listdir('/kaggle/input')[0]}\"\n\n# Load training metadata\ntrain_df = pd.read_csv(f\"{DATA_DIR}/train.csv\")\n\n# Select two signs to compare\nsign_left  = \"no\"\nsign_right = \"yes\"\n\n\n# Helper Functions\n\ndef get_sequence_id(sign_name):\n    \"\"\"Return sequence_id and parquet path for a given sign.\"\"\"\n    sample = train_df[train_df[\"sign\"] == sign_name]\n    \n    if len(sample) == 0:\n        raise ValueError(f\"Sign '{sign_name}' not found\")\n    \n    return sample.iloc[0][\"sequence_id\"], sample.iloc[0][\"path\"]\n\n\ndef load_landmarks(path):\n    \"\"\"Load landmark parquet file and flip Y-axis.\"\"\"\n    df = pd.read_parquet(f\"{DATA_DIR}/{path}\")\n    df[\"y\"] *= -1\n    return df\n\n\n# Get sequence info\nseq1, path1 = get_sequence_id(sign_left)\nseq2, path2 = get_sequence_id(sign_right)\n\n# Load landmark data\ndata1 = load_landmarks(path1)\ndata2 = load_landmarks(path2)\n\n\n# Create Frame (Side-by-Side)\n\ndef create_frame(d1, d2, fid1, fid2):\n    \"\"\"Render one frame from each sign side-by-side.\"\"\"\n    \n    fig, axes = plt.subplots(1, 2, figsize=(6,3))\n\n    for ax, data, fid, title in zip(\n        axes,\n        [d1, d2],\n        [fid1, fid2],\n        [sign_left, sign_right]\n    ):\n        frame = data[data.frame == fid]\n\n        ax.axis(\"off\")\n        ax.set_xlim(data.x.min()-0.2, data.x.max()+0.2)\n        ax.set_ylim(data.y.min()-0.2, data.y.max()+0.2)\n\n        # Draw landmarks as colored points\n        for t, color in zip(\n            ['face','left_hand','right_hand','pose'],\n            ['gray','red','blue','green']\n        ):\n            part = frame[frame.type == t]\n            if len(part) > 0:\n                ax.scatter(part.x, part.y, s=8, c=color)\n\n        ax.set_title(title)\n\n    # Convert matplotlib figure to image array\n    fig.canvas.draw()\n    img = np.asarray(fig.canvas.buffer_rgba())\n    plt.close(fig)\n\n    return img\n\n\n# Get sorted frame IDs\nframes1 = sorted(data1.frame.unique())\nframes2 = sorted(data2.frame.unique())\n\n# Match shortest sequence length\nmax_len = min(len(frames1), len(frames2))\n\n# Generate frame images\nimages = [\n    create_frame(data1, data2, frames1[i], frames2[i])\n    for i in range(max_len)\n]\n\n\n# Create Animation\n\nfig, ax = plt.subplots()\nax.axis(\"off\")\n\nims = []\nfor img in images:\n    im = ax.imshow(img, animated=True)\n    ims.append([im])\n\nanim = mpl.animation.ArtistAnimation(\n    fig,\n    ims,\n    interval=80,\n    blit=True,\n    repeat_delay=1000\n)\n\ndisplay(HTML(anim.to_jshtml()))\nplt.close()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-20T13:16:05.892148Z","iopub.execute_input":"2026-02-20T13:16:05.892466Z","iopub.status.idle":"2026-02-20T13:16:07.042337Z","shell.execute_reply.started":"2026-02-20T13:16:05.892436Z","shell.execute_reply":"2026-02-20T13:16:07.041263Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Observations on Selected Signs\n\nIn this section, we analyze the motion patterns of the following signs:\n\n- **thankyou**\n- **no**\n- **yes**\n- **find**\n\nThe observations are based on landmark movement across frames (hands, face, and pose).\n\n---\n\n### 🔹 1. Sign: *thankyou*\n\n**Motion Characteristics:**\n- The movement typically starts near the chin or mouth area.\n- The dominant hand moves outward away from the face.\n- The face remains mostly stable.\n- Motion is primarily concentrated in the right hand.\n\n**Data Insights:**\n- Noticeable displacement in the hand landmarks along the x-axis.\n- Minimal variation in body pose.\n- Clear directional movement pattern.\n\n**Modeling Insight:**\n- Temporal modeling is essential.\n- A single frame is insufficient for classification.\n- Motion-based features (velocity or displacement) could improve performance.\n\n---\n\n### 🔹 2. Sign: *no*\n\n**Motion Characteristics:**\n- Very small and compact movement.\n- Motion is mostly in finger joints.\n- Limited spatial displacement.\n\n**Data Insights:**\n- Low variance in landmark coordinates.\n- Subtle changes between consecutive frames.\n- Minimal body movement.\n\n**Modeling Insight:**\n- Harder to classify due to small motion amplitude.\n- Requires high-resolution temporal modeling.\n- Fine-grained hand features are important.\n\n---\n\n### 🔹 3. Sign: *yes*\n\n**Motion Characteristics:**\n- Repetitive vertical hand motion.\n- Clear oscillation pattern.\n- Moderate amplitude movement.\n\n**Data Insights:**\n- Strong variation along the y-axis.\n- Periodic motion pattern.\n- Slight shoulder involvement in pose landmarks.\n\n**Modeling Insight:**\n- Frequency-based motion features could help.\n- Easier to distinguish than \"no\" due to larger movement.\n\n---\n\n### 🔹 4. Sign: *pencil*\n\n**Motion Characteristics:**\n- More complex movement involving both hands.\n- Larger spatial coverage.\n- Coordinated multi-joint motion.\n\n**Data Insights:**\n- Higher variance in both x and y axes.\n- Strong interaction between left and right hands.\n- More pose engagement.\n\n**Modeling Insight:**\n- Requires modeling spatial + temporal relationships.\n- Transformer-based architectures may capture this better than simple RNNs.\n\n---\n\n##  Comparative Summary\n\n| Sign       | Motion Size | Complexity | Temporal Importance |\n|------------|------------|------------|---------------------|\n| thankyou   | Medium     | Moderate   | High                |\n| no         | Small      | Low        | Very High           |\n| yes        | Medium     | Repetitive | High                |\n| find       | Large      | Complex    | Very High           |\n\n---\n\n###  Key Takeaways\n\n- Hand landmarks are the most informative features.\n- Small-motion signs (e.g., *no*) are more challenging to classify.\n- Temporal modeling is critical for all signs.\n- Static frame classification is insufficient for accurate recognition.","metadata":{}},{"cell_type":"markdown","source":"# Final Remarks\n<br>\n\n<a id=\"ASL\"></a><br><b style=\"text-decoration: underline; font-family: Verdana; font-size: 120%; text-transform: uppercase;\">American Sign Language (ASL)</b>\n\nAmerican Sign Language (ASL) is a complete and natural language used primarily by the Deaf community in North America.  \nIt has its own unique grammar, syntax, and vocabulary, and it is completely separate from English.\n\nASL is a visual-gestural language, meaning it relies on:\n\n- Hand movements  \n- Facial expressions  \n- Body language  \n\nASL developed more than 200 years ago, influenced by local sign systems and French Sign Language (LSF).  \nEarly exposure to language (signed or spoken) is critical for proper cognitive, social, and linguistic development.\n\n<br>\n\n<b><sub>Cute Animated ASL GIF ...</sub></b>\n\n<img src=\"https://media0.giphy.com/media/1xVfziksXMFBhe0weN/200w.gif?cid=82a1493b2jwtrnketklxa6jseau5xwv0ry3yxuasxufy7xtt&rid=200w.gif&ct=g\">\n\n<br>\n\n---\n\n<br>\n\n<a id=\"mediapipe\"></a><br><b style=\"text-decoration: underline; font-family: Verdana; font-size: 120%; text-transform: uppercase;\">MediaPipe Holistic Solution</b>\n\nMediaPipe Holistic is an open-source framework that performs real-time detection and tracking of:\n\n- Body pose  \n- Face landmarks  \n- Hand landmarks  \n\nIt provides a unified topology of 540+ keypoints using multiple coordinated machine learning models.  \nThe pipeline is optimized for mobile and desktop devices and enables applications such as:\n\n- Sign language recognition  \n- Gesture control  \n- Sports analytics  \n- Augmented reality systems  \n\n<br>\n\n<b><sub>Example of MediaPipe Holistic</sub></b>\n\n<img src=\"https://mediapipe.dev/images/mobile/holistic_sports_and_gestures_example.gif\">\n\n<br>\n\n<b><sub>MediaPipe Landmarks for Hands</sub></b>\n\n<img src=\"https://mediapipe.dev/images/mobile/hand_landmarks.png\">\n\n<br>\n\n---\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2026-02-18T15:33:27.296372Z","iopub.execute_input":"2026-02-18T15:33:27.296916Z","iopub.status.idle":"2026-02-18T15:33:27.343981Z","shell.execute_reply.started":"2026-02-18T15:33:27.296879Z","shell.execute_reply":"2026-02-18T15:33:27.342964Z"}}},{"cell_type":"markdown","source":"##### Thanks ","metadata":{}}]}