{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":130287,"databundleVersionId":15633993}],"dockerImageVersionId":31287,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"e2c2fea9","cell_type":"markdown","source":"# Motion-S Visual Baseline: TF-IDF + KNN Retrieval\n\nThis notebook is a detailed, visual, and submission-ready baseline for the **Motion-S: Hierarchical Text-to-Motion Generation for Sign Language** competition.\n\nThe core idea is intentionally simple:\n\n> If a test sentence/gloss is textually similar to a training sentence/gloss, reuse the training motion-token sequence as the prediction.\n\nThis is a **retrieval baseline**, not a generative neural model. It is useful because it is fast, stable, easy to debug, and produces a valid `submission.csv`.\n\n## What this notebook does\n\n| Stage | Purpose | Output |\n|---|---|---|\n| 1. Load data | Read train/test/sample submission files | `train`, `test`, `sample_sub` |\n| 2. Build text | Combine `gloss` and `sentence` into searchable text | `train_text`, `test_text` |\n| 3. Validate tokens | Remove unusable train rows | `train_good` |\n| 4. TF-IDF | Convert text into sparse character n-gram vectors | `X_train`, `X_test` |\n| 5. KNN retrieval | Find the closest training row for each test row | `idx`, `dist` |\n| 6. Token transfer | Copy and normalize token layers | `pred` |\n| 7. Validation | Check length, range, and layer alignment | validation report |\n| 8. Save | Write `submission.csv` | `/kaggle/working/submission.csv` |\n","metadata":{}},{"id":"cbae5a22","cell_type":"markdown","source":"## Visual Pipeline Overview\n\n```text\n┌────────────────────┐\n│ train.csv / test.csv│\n└─────────┬──────────┘\n          │\n          ▼\n┌────────────────────┐\n│ Text normalization  │\n│ gloss + sentence    │\n└─────────┬──────────┘\n          │\n          ▼\n┌────────────────────┐\n│ TF-IDF Vectorizer   │\n│ char_wb 3-6 grams   │\n└─────────┬──────────┘\n          │\n          ▼\n┌────────────────────┐\n│ NearestNeighbors    │\n│ cosine distance     │\n└─────────┬──────────┘\n          │\n          ▼\n┌────────────────────┐\n│ Copy motion tokens  │\n│ from nearest train  │\n└─────────┬──────────┘\n          │\n          ▼\n┌────────────────────┐\n│ Validate submission │\n│ length/range/layers │\n└─────────┬──────────┘\n          │\n          ▼\n┌────────────────────┐\n│ submission.csv      │\n└────────────────────┘\n```\n","metadata":{}},{"id":"cfefa934","cell_type":"code","source":"# ============================================================\n# CELL 1: Environment Setup\n# ============================================================\n# Purpose:\n# - Import required libraries.\n# - Fix random seeds for reproducibility.\n# - Define competition paths.\n#\n# Notes:\n# - No kagglehub.login() is required inside Kaggle notebooks.\n# - The notebook assumes the competition dataset is attached as input.\n\nimport os\nimport re\nimport gc\nimport random\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom IPython.display import display, Markdown, HTML\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.neighbors import NearestNeighbors\n\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\nINPUT_DIR = Path(\"/kaggle/input/motion-s-hierarchical-text-to-motion-generation-for-sign-language\")\nTRAIN_CSV  = INPUT_DIR / \"train.csv\"\nTEST_CSV   = INPUT_DIR / \"test.csv\"\nSAMPLE_SUB = INPUT_DIR / \"sample_submission.csv\"\nOUT_PATH   = Path(\"/kaggle/working/submission.csv\")\n\nprint(\"Input directory:\", INPUT_DIR)\nprint(\"Output path:\", OUT_PATH)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:18.961169Z","iopub.execute_input":"2026-05-10T01:23:18.962150Z","iopub.status.idle":"2026-05-10T01:23:18.969243Z","shell.execute_reply.started":"2026-05-10T01:23:18.962122Z","shell.execute_reply":"2026-05-10T01:23:18.968442Z"}},"outputs":[],"execution_count":null},{"id":"365f49e1","cell_type":"code","source":"# ============================================================\n# CELL 2: Helper Display Functions\n# ============================================================\n# Purpose:\n# - Make notebook output easier to read.\n# - Display important steps as visual cards.\n\ndef show_step(step, title, detail):\n    html = f\"\"\"\n    <div style='border-left:6px solid #4A90E2; padding:12px 16px; margin:12px 0; background:#f5f9ff;'>\n        <div style='font-size:18px; font-weight:700;'>STEP {step}: {title}</div>\n        <div style='font-size:14px; color:#333; margin-top:4px;'>{detail}</div>\n    </div>\n    \"\"\"\n    display(HTML(html))\n\ndef show_table(df, title=None, n=5):\n    if title:\n        display(Markdown(f\"### {title}\"))\n    display(df.head(n))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:18.970816Z","iopub.execute_input":"2026-05-10T01:23:18.971152Z","iopub.status.idle":"2026-05-10T01:23:18.991734Z","shell.execute_reply.started":"2026-05-10T01:23:18.971116Z","shell.execute_reply":"2026-05-10T01:23:18.991162Z"}},"outputs":[],"execution_count":null},{"id":"298ee93c","cell_type":"code","source":"# ============================================================\n# CELL 3: Load Competition Data\n# ============================================================\n\nshow_step(1, \"Load Data\", \"Read train, test, and sample submission files from Kaggle input.\")\ntrain = pd.read_csv(TRAIN_CSV)\ntest = pd.read_csv(TEST_CSV)\nsample_sub = pd.read_csv(SAMPLE_SUB)\nsummary = pd.DataFrame({\"dataset\":[\"train\",\"test\",\"sample_submission\"],\"rows\":[len(train),len(test),len(sample_sub)],\"columns\":[train.shape[1],test.shape[1],sample_sub.shape[1]]})\ndisplay(summary)\nprint(\"Train columns:\", train.columns.tolist())\nprint(\"Test columns:\", test.columns.tolist())\nprint(\"Sample submission columns:\", sample_sub.columns.tolist())\nshow_table(train, \"Train preview\")\nshow_table(test, \"Test preview\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:18.992607Z","iopub.execute_input":"2026-05-10T01:23:18.992899Z","iopub.status.idle":"2026-05-10T01:23:19.419007Z","shell.execute_reply.started":"2026-05-10T01:23:18.992879Z","shell.execute_reply":"2026-05-10T01:23:19.418403Z"}},"outputs":[],"execution_count":null},{"id":"d56baf6f","cell_type":"code","source":"# ============================================================\n# CELL 4: Text Normalization and Search Text Construction\n# ============================================================\n\nshow_step(2, \"Build Search Text\", \"Normalize gloss/sentence and combine them into retrieval text. Gloss is repeated to give it higher retrieval weight.\")\n_ws_re = re.compile(r\"\\s+\")\ndef norm_text(x):\n    if not isinstance(x, str): return \"\"\n    return _ws_re.sub(\" \", x.strip())\ndef build_text(df):\n    g = df[\"gloss\"].map(norm_text) if \"gloss\" in df.columns else pd.Series([\"\"]*len(df), index=df.index)\n    s = df[\"sentence\"].map(norm_text) if \"sentence\" in df.columns else pd.Series([\"\"]*len(df), index=df.index)\n    return (g + \" || \" + g + \" || \" + s).fillna(\"\")\ntrain_text = build_text(train)\ntest_text = build_text(test)\ndisplay(pd.DataFrame({\"train_text_example\": train_text.head(5)}))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:19.419896Z","iopub.execute_input":"2026-05-10T01:23:19.420174Z","iopub.status.idle":"2026-05-10T01:23:19.510937Z","shell.execute_reply.started":"2026-05-10T01:23:19.420145Z","shell.execute_reply":"2026-05-10T01:23:19.510288Z"}},"outputs":[],"execution_count":null},{"id":"d15a0dce","cell_type":"code","source":"# ============================================================\n# CELL 5: Visualize Text Length Distribution\n# ============================================================\n\nshow_step(3, \"Text Diagnostics\", \"Compare text length distributions for train and test. Very short text can weaken retrieval quality.\")\ntrain_text_len = train_text.str.len(); test_text_len = test_text.str.len()\nlength_summary = pd.DataFrame({\"dataset\":[\"train\",\"test\"],\"min\":[train_text_len.min(),test_text_len.min()],\"median\":[train_text_len.median(),test_text_len.median()],\"mean\":[train_text_len.mean(),test_text_len.mean()],\"max\":[train_text_len.max(),test_text_len.max()]})\ndisplay(length_summary)\nplt.figure(figsize=(10,5)); plt.hist(train_text_len, bins=50, alpha=0.6, label=\"train\"); plt.hist(test_text_len, bins=50, alpha=0.6, label=\"test\")\nplt.title(\"Search Text Length Distribution\"); plt.xlabel(\"Number of characters\"); plt.ylabel(\"Count\"); plt.legend(); plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:19.512910Z","iopub.execute_input":"2026-05-10T01:23:19.513214Z","iopub.status.idle":"2026-05-10T01:23:19.891332Z","shell.execute_reply.started":"2026-05-10T01:23:19.513192Z","shell.execute_reply":"2026-05-10T01:23:19.890552Z"}},"outputs":[],"execution_count":null},{"id":"d83b0240","cell_type":"code","source":"# ============================================================\n# CELL 6: Token Utility Functions\n# ============================================================\n\nshow_step(4, \"Token Utilities\", \"Define safe token parsing and length-normalization functions.\")\nTOKEN_COLS = [\"base_tokens\", \"residual_1\", \"residual_2\", \"residual_3\", \"residual_4\", \"residual_5\"]\nMIN_LEN, MAX_LEN = 40, 800\ndef parse_tokens(tok_str):\n    if not isinstance(tok_str, str): return []\n    tok_str = tok_str.strip()\n    if tok_str == \"\": return []\n    try: return [int(x) for x in tok_str.split()]\n    except ValueError: return []\ndef tokens_to_str(tokens): return \" \".join(map(str, tokens))\ndef enforce_len(tokens, min_len=MIN_LEN, max_len=MAX_LEN):\n    if len(tokens) == 0: return [random.randint(0, 511) for _ in range(min_len)]\n    if len(tokens) < min_len:\n        reps = (min_len + len(tokens) - 1) // len(tokens)\n        tokens = (tokens * reps)[:min_len]\n    if len(tokens) > max_len: tokens = tokens[:max_len]\n    return tokens\nprint(\"Token columns:\", TOKEN_COLS); print(\"Length range:\", MIN_LEN, \"to\", MAX_LEN)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:19.892197Z","iopub.execute_input":"2026-05-10T01:23:19.892517Z","iopub.status.idle":"2026-05-10T01:23:19.901067Z","shell.execute_reply.started":"2026-05-10T01:23:19.892495Z","shell.execute_reply":"2026-05-10T01:23:19.900356Z"}},"outputs":[],"execution_count":null},{"id":"3d9a8f79","cell_type":"code","source":"# ============================================================\n# CELL 7: Validate Training Token Rows\n# ============================================================\n\nshow_step(5, \"Training Token Validation\", \"Filter out train rows with invalid token values or inconsistent layer lengths.\")\ngood_mask = np.ones(len(train), dtype=bool); train_lens = []; invalid_reasons = []\nfor i, row in train.iterrows():\n    lens=[]; ok=True; reason=\"ok\"\n    for c in TOKEN_COLS:\n        t = parse_tokens(row.get(c, \"\"))\n        if len(t) == 0: ok=False; reason=f\"empty_{c}\"; break\n        if any((x < 0 or x > 511) for x in t): ok=False; reason=f\"out_of_range_{c}\"; break\n        lens.append(len(t))\n    if ok and len(set(lens)) != 1: ok=False; reason=\"layer_length_mismatch\"\n    if ok: train_lens.append(lens[0])\n    else: good_mask[i] = False; invalid_reasons.append(reason)\ntrain_good = train.loc[good_mask].reset_index(drop=True)\ntrain_text_good = train_text.loc[good_mask].reset_index(drop=True)\ndisplay(pd.DataFrame({\"metric\":[\"train rows\",\"usable rows\",\"removed rows\",\"usable ratio\"],\"value\":[len(train),len(train_good),len(train)-len(train_good),len(train_good)/max(len(train),1)]}))\nif invalid_reasons: display(pd.Series(invalid_reasons).value_counts().rename_axis(\"reason\").reset_index(name=\"count\"))\nelse: print(\"No invalid token rows found.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:19.902011Z","iopub.execute_input":"2026-05-10T01:23:19.902310Z","iopub.status.idle":"2026-05-10T01:23:22.571945Z","shell.execute_reply.started":"2026-05-10T01:23:19.902257Z","shell.execute_reply":"2026-05-10T01:23:22.571135Z"}},"outputs":[],"execution_count":null},{"id":"1f60afab","cell_type":"code","source":"# ============================================================\n# CELL 8: Visualize Motion Token Lengths\n# ============================================================\n\nshow_step(6, \"Token Length Diagnostics\", \"Visualize usable training token sequence lengths.\")\nif len(train_lens) > 0:\n    lens_series = pd.Series(train_lens, name=\"token_length\")\n    display(lens_series.describe().to_frame().T)\n    plt.figure(figsize=(10,5)); plt.hist(lens_series, bins=50, alpha=0.8)\n    plt.axvline(MIN_LEN, linestyle=\"--\", label=\"MIN_LEN\"); plt.axvline(MAX_LEN, linestyle=\"--\", label=\"MAX_LEN\")\n    plt.title(\"Training Motion Token Length Distribution\"); plt.xlabel(\"Token sequence length\"); plt.ylabel(\"Count\"); plt.legend(); plt.show()\nelse: print(\"No valid training token lengths available.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:22.572834Z","iopub.execute_input":"2026-05-10T01:23:22.573101Z","iopub.status.idle":"2026-05-10T01:23:22.780174Z","shell.execute_reply.started":"2026-05-10T01:23:22.573078Z","shell.execute_reply":"2026-05-10T01:23:22.779462Z"}},"outputs":[],"execution_count":null},{"id":"9db395e0","cell_type":"code","source":"# ============================================================\n# CELL 9: TF-IDF Feature Extraction\n# ============================================================\n\nshow_step(7, \"TF-IDF Vectorization\", \"Convert normalized text into sparse character n-gram feature vectors.\")\nvectorizer = TfidfVectorizer(lowercase=True, analyzer=\"char_wb\", ngram_range=(3,6), min_df=2, max_features=250000)\nX_train = vectorizer.fit_transform(train_text_good)\nX_test = vectorizer.transform(test_text)\ndisplay(pd.DataFrame({\"matrix\":[\"X_train\",\"X_test\"],\"rows\":[X_train.shape[0],X_test.shape[0]],\"features\":[X_train.shape[1],X_test.shape[1]],\"non_zero_values\":[X_train.nnz,X_test.nnz]}))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:22.781245Z","iopub.execute_input":"2026-05-10T01:23:22.781563Z","iopub.status.idle":"2026-05-10T01:23:24.407326Z","shell.execute_reply.started":"2026-05-10T01:23:22.781541Z","shell.execute_reply":"2026-05-10T01:23:24.406420Z"}},"outputs":[],"execution_count":null},{"id":"ff8305c2","cell_type":"code","source":"# ============================================================\n# CELL 10: Nearest Neighbor Retrieval\n# ============================================================\n\nshow_step(8, \"Nearest Neighbor Retrieval\", \"Find the closest training text for every test text using cosine distance.\")\nnn = NearestNeighbors(n_neighbors=1, metric=\"cosine\", algorithm=\"brute\")\nnn.fit(X_train)\ndist, idx = nn.kneighbors(X_test, return_distance=True)\nidx = idx.reshape(-1); dist = dist.reshape(-1)\ndisplay(pd.DataFrame({\"test_id\": test[\"id\"].values[:10] if \"id\" in test.columns else np.arange(min(10,len(test))),\"nearest_train_index\": idx[:10],\"cosine_distance\": dist[:10],\"test_text\": test_text.iloc[:10].values,\"nearest_train_text\": train_text_good.iloc[idx[:10]].values}))\nplt.figure(figsize=(10,5)); plt.hist(dist, bins=50, alpha=0.8)\nplt.title(\"Nearest Neighbor Cosine Distance Distribution\"); plt.xlabel(\"Cosine distance: lower is better\"); plt.ylabel(\"Number of test rows\"); plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:24.408378Z","iopub.execute_input":"2026-05-10T01:23:24.408655Z","iopub.status.idle":"2026-05-10T01:23:26.830512Z","shell.execute_reply.started":"2026-05-10T01:23:24.408626Z","shell.execute_reply":"2026-05-10T01:23:26.829734Z"}},"outputs":[],"execution_count":null},{"id":"4a5fd4a3","cell_type":"code","source":"# ============================================================\n# CELL 11: Token Prediction by Retrieval Transfer\n# ============================================================\n\nshow_step(9, \"Token Transfer\", \"Copy motion tokens from the nearest training sample and normalize layer lengths.\")\nid_col = \"id\" if \"id\" in test.columns else sample_sub.columns[0]\npred = pd.DataFrame({\"id\": test[id_col].values})\nfor c in TOKEN_COLS: pred[c] = \"\"\ntrain_tokens_cache = {c: [parse_tokens(x) for x in train_good[c].astype(str).tolist()] for c in TOKEN_COLS}\nfor j in range(len(test)):\n    k = int(idx[j])\n    layers = [train_tokens_cache[c][k] for c in TOKEN_COLS]\n    L = min(len(t) for t in layers)\n    if L <= 0: L = MIN_LEN\n    base = enforce_len(layers[0][:L], MIN_LEN, MAX_LEN)\n    L2 = len(base)\n    fixed_layers = [base]\n    for li in range(1, len(TOKEN_COLS)):\n        t = layers[li][:L]\n        t = enforce_len(t, L2, L2)\n        fixed_layers.append(t)\n    for c, tokens in zip(TOKEN_COLS, fixed_layers): pred.at[j, c] = tokens_to_str(tokens)\nshow_table(pred, \"Prediction preview\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:26.831476Z","iopub.execute_input":"2026-05-10T01:23:26.831829Z","iopub.status.idle":"2026-05-10T01:23:29.054218Z","shell.execute_reply.started":"2026-05-10T01:23:26.831806Z","shell.execute_reply":"2026-05-10T01:23:29.053579Z"}},"outputs":[],"execution_count":null},{"id":"f744ae22","cell_type":"code","source":"# ============================================================\n# CELL 12: Submission Validation\n# ============================================================\n\nshow_step(10, \"Submission Validation\", \"Check length, token range, and layer consistency for predicted rows.\")\ndef validate_row(r):\n    lens=[]\n    for c in TOKEN_COLS:\n        t = parse_tokens(r[c])\n        if len(t) < MIN_LEN or len(t) > MAX_LEN: return False\n        if any((x < 0 or x > 511) for x in t): return False\n        lens.append(len(t))\n    return len(set(lens)) == 1\ncheck_n = min(200, len(pred)); ok = sum(int(validate_row(pred.iloc[i])) for i in range(check_n))\ndisplay(pd.DataFrame({\"checked_rows\":[check_n],\"valid_rows\":[ok],\"valid_ratio\":[ok/max(check_n,1)]}))\npred_lens = pred[\"base_tokens\"].map(lambda x: len(parse_tokens(x)))\ndisplay(pred_lens.describe().to_frame(name=\"predicted_base_length\").T)\nplt.figure(figsize=(10,5)); plt.hist(pred_lens, bins=50, alpha=0.8)\nplt.title(\"Predicted Token Length Distribution\"); plt.xlabel(\"Token sequence length\"); plt.ylabel(\"Count\"); plt.show()\nassert ok == check_n, \"Some prediction rows failed validation. Please inspect pred before submission.\"\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:29.055250Z","iopub.execute_input":"2026-05-10T01:23:29.055563Z","iopub.status.idle":"2026-05-10T01:23:29.350800Z","shell.execute_reply.started":"2026-05-10T01:23:29.055531Z","shell.execute_reply":"2026-05-10T01:23:29.350101Z"}},"outputs":[],"execution_count":null},{"id":"bc691392","cell_type":"code","source":"# ============================================================\n# CELL 13: Save submission.csv\n# ============================================================\n\nshow_step(11, \"Save Submission\", \"Write the final submission file and show a final format checklist.\")\nexpected_cols = sample_sub.columns.tolist()\nif set(expected_cols).issubset(set(pred.columns)):\n    pred = pred[expected_cols]\nelse:\n    print(\"Warning: sample submission columns are not a subset of pred columns. Keeping pred column order.\")\npred.to_csv(OUT_PATH, index=False)\ndisplay(pd.DataFrame({\"item\":[\"Output path\",\"Rows\",\"Columns\",\"File exists\"],\"value\":[str(OUT_PATH), pred.shape[0], pred.shape[1], OUT_PATH.exists()]}))\ndisplay(pred.head())\nprint(\"Saved:\", OUT_PATH)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-10T01:23:29.351732Z","iopub.execute_input":"2026-05-10T01:23:29.352048Z","iopub.status.idle":"2026-05-10T01:23:29.573333Z","shell.execute_reply.started":"2026-05-10T01:23:29.352026Z","shell.execute_reply":"2026-05-10T01:23:29.572514Z"}},"outputs":[],"execution_count":null},{"id":"f74d601a","cell_type":"markdown","source":"## Improvement Ideas\n\n| Idea | Why it may help | Difficulty |\n|---|---|---|\n| Use top-k neighbors instead of only top-1 | Similar samples can be blended or selected more robustly | Medium |\n| Predict sequence length separately | The copied neighbor length may not match the ideal motion length | Medium |\n| Weight `gloss` and `sentence` differently | Gloss may be more aligned with sign structure | Easy |\n| Add word-level TF-IDF features | Character n-grams may miss semantic similarity | Easy |\n| Use sentence embeddings | Better semantic matching than TF-IDF | Medium/Hard |\n| Retrieval reranking | Use additional criteria after initial TF-IDF search | Medium |\n\n## Final Summary\n\nThis notebook creates a valid retrieval-based Motion-S submission by building searchable text, vectorizing it with TF-IDF, retrieving the nearest training example, transferring six motion-token layers, validating the result, and saving `submission.csv`.\n","metadata":{}}]}