{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":130287,"databundleVersionId":15633993}],"dockerImageVersionId":31286,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1: Environment Setup & Loading Data 🛠️","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os \nimport warnings\n\n# Ignore warning messages for cleaner output\nwarnings.filterwarnings(\"ignore\")\n\n# Define the dataset directory path\nDATA_DIR = \"/kaggle/input/competitions/motion-s-hierarchical-text-to-motion-generation-for-sign-language\"\n\n# Load the main CSV files\ntrain_df = pd.read_csv(os.path.join(DATA_DIR, \"train.csv\"))\ntest_df = pd.read_csv(os.path.join(DATA_DIR, \"test.csv\"))\nsample_submission = pd.read_csv(os.path.join(DATA_DIR, \"sample_submission.csv\"))\n\n# Print the dimensions of the datasets\nprint(f\"Train Dataset Shape: {train_df.shape}\")\nprint(f\"Test Dataset Shape: {test_df.shape}\")\nprint(f\"Sample Submission Shape: {sample_submission.shape}\\n\")\n\n# Display the first\ndisplay(train_df.head(3))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:14.476922Z","iopub.execute_input":"2026-03-21T11:12:14.477201Z","iopub.status.idle":"2026-03-21T11:12:14.909175Z","shell.execute_reply.started":"2026-03-21T11:12:14.477179Z","shell.execute_reply":"2026-03-21T11:12:14.908468Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 1\n\n1. **Target Variables:** The data our model needs to predict consists of six columns in our table: `base_tokens` and `residual_1` through `residual_5`. These tokens are stored as space-separated strings. Before moving to the modeling phase, we will need to convert these into sequences of integer tensors.\n\n2. **Input Features:** We have two columns with natural language structure: `sentence` and `gloss`, which follows Sign Grammar. Since sign language has its own unique syntax, the `gloss` column will be a much stronger signal (feature) for our model in predicting movements.\n\n3. **Dimensions:** We have approximately 12.5K training examples and 3K test examples.","metadata":{}},{"cell_type":"markdown","source":"# 2: Data Integrity & Sequence Length Analysis 🕵️‍♂️","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Check for missing values in the train dataset\nprint(\"Missing Values:\\n\", train_df.isnull().sum())\nprint(\"-\" * 50)\n\n# Calculate sequence lengths by splitting the base_tokens string\n# We assume the data is clean, but let's calculate the length of tokens for each row.\ntrain_df[\"seq_len\"] = train_df[\"base_tokens\"].apply(lambda x: len(str(x).split()))\n\n# Display basic statistical metrics for sequence lengths\nprint(\"Sequence Length Statistics:\")\nprint(train_df[\"seq_len\"].describe())\nprint(\"-\" * 50)\n\n# Verify if all 6 layers have the exact same length for the first sample\nsample_row = train_df.iloc[0]\ntoken_columns = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\nlengths = [len(str(sample_row[col]).split()) for col in token_columns]\nprint(f\"Lengths of the 6 layers for ID {sample_row[\"id\"]}: {lengths}\")\n\n# Plot the distribution of sequence lengths\nplt.figure(figsize=(10, 5))\nsns.histplot(train_df[\"seq_len\"], bins=50, kde=True, color=\"mediumpurple\")\nplt.title('Distribution of Token Sequence Lengths', fontsize=14, fontweight='bold')\nplt.xlabel('Sequence Length', fontsize=12)\nplt.ylabel('Frequency', fontsize=12)\n\n# Add vertical lines for competition constraints\nplt.axvline(x=40, color='red', linestyle='--', linewidth=2, label='Min Limit (40)')\nplt.axvline(x=800, color='red', linestyle='--', linewidth=2, label='Max Limit (800)')\n\nplt.legend()\nplt.grid(True, alpha=0.3)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:14.910494Z","iopub.execute_input":"2026-03-21T11:12:14.910810Z","iopub.status.idle":"2026-03-21T11:12:16.455478Z","shell.execute_reply.started":"2026-03-21T11:12:14.910789Z","shell.execute_reply":"2026-03-21T11:12:16.454669Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 2\n\n1.  **Missing Values:** Only 4 out of 12,467 rows have missing token data. Since this is a negligible amount relative to the total dataset, the safest approach is to delete these 4 rows directly, rather than dealing with complex methods like imputation.\n\n2.  **Outliers & Constraint Violations:** The competition rules clearly stated that sequences must be between 40 and 800 tokens. However, looking at the statistics (Min: 1, Max: 1853) and the histogram you sent, we see outliers both below 40 and above 800. If we train our model with data outside these limits, we risk getting 0 points on the test set!\n\n3.  **Layer Consistency:** In the first example we examined (ID: 1000648), we saw that the length of all layers was 81. This confirms that the base_tokens and the 5 residual layers are in a parallel and synchronous structure.","metadata":{}},{"cell_type":"markdown","source":"# 3: Data Cleaning & Filtering 🧹","metadata":{}},{"cell_type":"code","source":"# Drop rows with any missing values\ntrain_clean = train_df.dropna().copy()\n\n# Recalculate sequence lengths just to be safe after dropna\ntrain_clean[\"seq_len\"] = train_clean[\"base_tokens\"].apply(lambda x: len(str(x).split()))\n\n# Filter the dataset based on competition length constraints\n# Constraint: Length must be between 40 and 800\nvalid_length_mask = (train_clean[\"seq_len\"] >= 40) & (train_clean[\"seq_len\"] <= 800)\ntrain_clean = train_clean[valid_length_mask]\n\n# Reset index for the cleaned dataframe\ntrain_clean = train_clean.reset_index(drop=True)\n\n# Print the cleaning results\nprint(f\"Original Data Size: {len(train_df)}\")\nprint(f\"Cleaned Data Size: {len(train_clean)}\")\nprint(f\"Removed Samples: {len(train_df) - len(train_clean)}\")\n\n# Show the new min and max lengths to verify\nprint(f\"New Min Length: {train_clean['seq_len'].min()}\")\nprint(f\"New Max Length: {train_clean['seq_len'].max()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:16.456385Z","iopub.execute_input":"2026-03-21T11:12:16.456725Z","iopub.status.idle":"2026-03-21T11:12:16.547338Z","shell.execute_reply.started":"2026-03-21T11:12:16.456703Z","shell.execute_reply":"2026-03-21T11:12:16.546320Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 3\n\nWe have successfully eliminated 94 \"problematic\" (rule-violating or incomplete) rows. Our new limits are between 40 and 781. This was a critical step because we have protected ourselves from the nightmare of \"Submission Error\" or \"0 Points\" due to format errors in Kaggle competitions. Our data is now clean and 100% compliant with the competition rules.\n\nNow it's time to speak the language of machine learning models! Deep learning models understand numbers and tensors, not strings. In our dataset, `base_tokens` and the 5 residual layers are currently in a string format separated by spaces (e.g., \"379 295 376...\"). To perform mathematical operations on them, we must convert them into actual integer lists.","metadata":{}},{"cell_type":"markdown","source":"# 4: Target Variable Transformation 🔢","metadata":{}},{"cell_type":"code","source":"# Define the columns that contain the target tokens\ntoken_cols = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n\n# Function to convert space-separated string to a list of integers\ndef string_to_int_list(token_string):\n    # Split the string by spaces and convert each element to int\n    return [int(x) for x in str(token_string).split()]\n\n# Apply the conversion to all target token columns \nfor col in token_cols:\n    train_clean[col] = train_clean[col].apply(string_to_int_list)\n\n# Verify the transformation by checking the first sample\nfirst_sample_base = train_clean['base_tokens'].iloc[0]\n\nprint(f\"Data type of the sequence: {type(first_sample_base)}\")\nprint(f\"Data type of the first token: {type(first_sample_base[0])}\")\nprint(f\"First 5 tokens of the first sample: {first_sample_base[:5]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:16.549086Z","iopub.execute_input":"2026-03-21T11:12:16.549398Z","iopub.status.idle":"2026-03-21T11:12:18.193853Z","shell.execute_reply.started":"2026-03-21T11:12:16.549376Z","shell.execute_reply":"2026-03-21T11:12:18.193119Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 4\n\nGreat! Our target variables are now in a format ready for a deep learning model. We have successfully navigated the string manipulation phase. The list we see in the output `[379, 295, 376...]` represents the motion codes that our model ultimately needs to produce.\n\nNow that we have taken care of the `y` (target) side of the equation, we can focus on the `X` (input) side. There was a critical tip in the competition description: **Glossification**. Using text translated into sign language grammar (gloss) instead of a natural English sentence will lighten the model's load in learning word order.\n\nHowever, just like our target tokens, our model cannot understand these gloss texts directly. We also need to convert them into numerical vectors (Tokenization).","metadata":{}},{"cell_type":"markdown","source":"# 5: Text Tokenization & Input Length Analysis 🌠","metadata":{}},{"cell_type":"code","source":"from transformers import AutoTokenizer\n\n# Define the pre-trained text model name\nTEXT_MODEL_NAME = \"bert-base-uncased\"\n\n# Load the tokenizer from Hugging Faca\ntokenizer = AutoTokenizer.from_pretrained(TEXT_MODEL_NAME)\n\n# Calculate token lengths for the \"bloss\" column\n# We use encode() to get the actual token IDs including special tokens like [CLS] and [SEP]\ntrain_clean['gloss_token_len'] = train_clean['gloss'].apply(lambda x: len(tokenizer.encode(str(x))))\n\n# Print statistics of the gloss token lengths\nprint(\"Gloss Token Length Statistics:\")\nprint(train_clean[\"gloss_token_len\"].describe())\nprint(\"-\" * 50)\n\n# Calculate percentiles to decide on a safe max_length for padding\np90 = train_clean['gloss_token_len'].quantile(0.90)\np95 = train_clean['gloss_token_len'].quantile(0.95)\np99 = train_clean['gloss_token_len'].quantile(0.99)\n\nprint(f\"90th Percentile: {p90}\")\nprint(f\"95th Percentile: {p95}\")\nprint(f\"99th Percentile: {p99}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:18.194730Z","iopub.execute_input":"2026-03-21T11:12:18.195167Z","iopub.status.idle":"2026-03-21T11:12:41.528304Z","shell.execute_reply.started":"2026-03-21T11:12:18.195142Z","shell.execute_reply":"2026-03-21T11:12:41.527572Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 5\n\nAlthough the maximum token length was 97, 99% of the data was 17 tokens or less. If we were to pad the entire dataset to the longest sample of 97, the model would waste time and GPU memory calculating those empty spaces (\"0\" pad tokens) at every training step.\n\nBy looking at these statistics, we can set a maximum length (MAX_TEXT_LEN) for our input texts to a safe limit of 32 (lengths that are multiples of 8, like 32, are significantly more efficient for GPU Tensor Cores).\n\nAdditionally, our motion sequences ranged from 40 to 800. To teach the model where a sequence should end, we will need to define special tokens outside our existing vocabulary (0-511): an End-of-Sequence token (<EOS>: 513) and a Padding token (<PAD>: 512).","metadata":{}},{"cell_type":"markdown","source":"# 6: Custom PyTorch Dataset 🗂️","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import Dataset, DataLoader\n\n# Define maximum text length and special motion tokens \nMAX_TEXT_LEN = 32\nMOTION_PAD_IDX = 512 # Padding token outside the 0-511 VAE vocab\nMOTION_EOS_IDX = 513 # End of sequence token\n\nclass SignMotionDataset(Dataset):\n    def __init__(self, df, tokenizer, max_text_len):\n        # Initialize the dataset with dataframe, tokenizer and max length\n        self.df = df\n        self.tokenizer = tokenizer\n        self.max_text_len = max_text_len\n        self.token_cols = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n        \n    def __len__(self):\n        # Return the total number of samples\n        return len(self.df)\n    \n    def __getitem__(self, idx):\n        # Get the row at the given index\n        row = self.df.iloc[idx]\n        \n        # 1. Process Text Input\n        gloss_text = str(row['gloss'])\n        text_encoding = self.tokenizer(\n            gloss_text,\n            max_length=self.max_text_len,\n            padding='max_length',\n            truncation=True,\n            return_tensors=\"pt\" # Return PyTorch tensors\n        )\n        \n        # Remove the extra batch dimension added by tokenizer\n        input_ids = text_encoding['input_ids'].squeeze(0)\n        attention_mask = text_encoding['attention_mask'].squeeze(0)\n        \n        # 2. Process Motion Target\n        motion_layers = []\n        for col in self.token_cols:\n            # Append EOS token to indicate the end of the generated sequence\n            layer_tokens = row[col] + [MOTION_EOS_IDX]\n            motion_layers.append(layer_tokens)\n            \n        # Convert the list of lists into a 2D PyTorch tensor of shape (6, seq_len + 1)\n        motion_tensor = torch.tensor(motion_layers, dtype=torch.long)\n        \n        return {\n            \"input_ids\": input_ids,\n            \"attention_mask\": attention_mask,\n            \"motion_tokens\": motion_tensor,\n            \"id\": row['id']\n        }\n\n# Instantiate the dataset object\ntrain_dataset = SignMotionDataset(train_clean, tokenizer, MAX_TEXT_LEN)\n\n# Fetch the first sample to verify the pipeline\nsample_data = train_dataset[0]\n\nprint(\"--- First Sample PyTorch Data ---\")\nprint(f\"Text Input IDs Shape: {sample_data['input_ids'].shape}\")\nprint(f\"Text Attention Mask Shape: {sample_data['attention_mask'].shape}\")\nprint(f\"Motion Tokens Tensor Shape: {sample_data['motion_tokens'].shape}\")\nprint(f\"First 5 tokens of Base Layer: {sample_data['motion_tokens'][0][:5]}\")\nprint(f\"Last token of Base Layer: {sample_data['motion_tokens'][0][-1].item()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:41.529325Z","iopub.execute_input":"2026-03-21T11:12:41.530411Z","iopub.status.idle":"2026-03-21T11:12:41.593286Z","shell.execute_reply.started":"2026-03-21T11:12:41.530341Z","shell.execute_reply":"2026-03-21T11:12:41.592365Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 6\n\nOur PyTorch Dataset class is working perfectly. Our text inputs (`input_ids` and `attention_mask`) are fixed exactly to the specified length of 32. More importantly, our motion tensor has the shape `(6, 82)`, and its last element has successfully become the EOS token (513) as we defined (Original length 81 + 1 EOS = 82).\n\nHowever, we are facing a critical Kaggle problem here: **Variable Length Sequences**.\n\nWhile our first example has a length of 82, the next example could be 150 or 45. When the PyTorch DataLoader attempts to form a \"batch\" from this data, it will crash because the matrix dimensions will not match. Padding the entire dataset to the longest sequence (782) would cause a massive waste of memory (RAM/VRAM).\n\n**Solution: Dynamic Padding.** We will only find the longest sequence within the current batch and pad the others accordingly. To implement this, we need to write a custom `collate_fn`.","metadata":{}},{"cell_type":"markdown","source":"# 7: Dynamic Padding & DataLoader 🗜️","metadata":{}},{"cell_type":"code","source":"from torch.nn.utils.rnn import pad_sequence\n\n# Custom collate function for dynamic padding\ndef collate_fn(batch):\n    # Stack text inputs since they are already padded to MAX_TEXT_LEN\n    input_ids = torch.stack([item['input_ids'] for item in batch])\n    attention_mask = torch.stack([item['attention_mask'] for item in batch])\n    ids = [item['id'] for item in batch]\n\n    # Get all motion tensors in the batch\n    # Each is shape (6, seq_len)\n    motion_tensors = [item[\"motion_tokens\"] for item in batch]\n\n    # Find the maximum sequence length in this specific batch\n    max_len = max([t.shape[1] for t in motion_tensors])\n    batch_size = len(batch)\n\n    # Create an empty tensor filled with PAD tokens\n    # Shape: (batch_size, 6 layers, max_len)\n    padded_motions = torch.full((batch_size, 6, max_len), MOTION_PAD_IDX, dtype=torch.long)\n\n    # Copy the original tokens into the padded tensor\n    for i, tensor in enumerate(motion_tensors):\n        seq_len = tensor.shape[1]\n        padded_motions[i, :, :seq_len] = tensor\n\n    return {\n        \"input_ids\": input_ids,\n        \"attention_mask\": attention_mask,\n        \"motion_tokens\": padded_motions,\n        \"id\": ids\n    }\n\n# Set hyperparameters for the DataLoader\nBATCH_SIZE = 16 # Adjust based on GPU memory\n\n# Create the DataLoader for the training set\ntrain_loader = DataLoader(\n    train_dataset,\n    batch_size=BATCH_SIZE,\n    shuffle=True, # Shuffle for training\n    collate_fn=collate_fn,\n    drop_last=True # Drop the last incomplete batch\n)\n\n# Fetch a single batch to verify\nfor batch in train_loader:\n    print(\"--- Batch Tensor Shapes  ---\")\n    print(f\"Batch Text IDs): {batch['input_ids'].shape}\")\n    print(f\"Batch Attention Mask: {batch['attention_mask'].shape}\")\n    print(f\"Batch Motion Tokens: {batch['motion_tokens'].shape}\")\n    break # We only need one batch","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:41.594369Z","iopub.execute_input":"2026-03-21T11:12:41.594757Z","iopub.status.idle":"2026-03-21T11:12:41.634604Z","shell.execute_reply.started":"2026-03-21T11:12:41.594733Z","shell.execute_reply":"2026-03-21T11:12:41.633837Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 7\n\nThe longest sequence in this specific batch was 741 tokens long. Instead of blindly padding the entire dataset to 800 (our maximum constraint), our PyTorch data loader performed dynamic padding only up to 741, the actual length required for this batch. This freed up a massive amount of VRAM!","metadata":{}},{"cell_type":"markdown","source":"# 8: Building the Text-to-Motion Model 🧠","metadata":{}},{"cell_type":"code","source":"import torch.nn as nn\nfrom transformers import AutoModel\n\n# Define the custom Seq2Seq model for Text-to-Motion generation\nclass SignMotionGenerator(nn.Module):\n    def __init__(self, text_model_name=\"bert-base-uncased\", vocab_size=514, hidden_size=768, num_layers=4, nhead=8):\n        super().__init__()\n        \n        # 1. Text Encoder: Pre-trained BERT\n        self.text_encoder = AutoModel.from_pretrained(text_model_name)\n        \n        # Freeze BERT parameters to save memory and speed up baseline training\n        for param in self.text_encoder.parameters():\n            param.requires_grad = False\n            \n        # 2. Motion Embedding: Embeds the 0-513 tokens\n        self.motion_emb = nn.Embedding(vocab_size, hidden_size)\n        \n        # Positional Encoding for the decoder sequence\n        self.pos_encoder = nn.Embedding(1000, hidden_size) \n        \n        # 3. Transformer Decoder\n        # batch_first=True makes tensors (batch, seq, feature) instead of (seq, batch, feature)\n        decoder_layer = nn.TransformerDecoderLayer(d_model=hidden_size, nhead=nhead, batch_first=True)\n        self.decoder = nn.TransformerDecoder(decoder_layer, num_layers=num_layers)\n        \n        # 4. Output Heads: 6 separate linear layers for the 6 RVQ levels\n        self.heads = nn.ModuleList([nn.Linear(hidden_size, vocab_size) for _ in range(6)])\n        \n    def forward(self, input_ids, attention_mask, motion_tokens):\n        # -- ENCODER --\n        # Get text embeddings from BERT\n        text_outputs = self.text_encoder(input_ids=input_ids, attention_mask=attention_mask)\n        memory = text_outputs.last_hidden_state # Shape: (batch, text_len, hidden_size)\n        \n        # -- DECODER INPUT PREPARATION --\n        # motion_tokens shape is (batch, 6, seq_len)\n        # Transpose to (batch, seq_len, 6) for easier processing per time step\n        motions_t = motion_tokens.permute(0, 2, 1) \n        \n        # Embed all 6 tokens and sum them up to get a single vector per time step\n        # Shape becomes: (batch, seq_len, hidden_size)\n        emb = self.motion_emb(motions_t).sum(dim=2)\n        \n        # Add positional encoding\n        seq_len = emb.size(1)\n        positions = torch.arange(0, seq_len, device=emb.device).unsqueeze(0)\n        emb = emb + self.pos_encoder(positions)\n        \n        # Causal Mask: Prevents the decoder from looking at future tokens\n        tgt_mask = nn.Transformer.generate_square_subsequent_mask(seq_len).to(emb.device)\n        \n        # -- DECODER --\n        # Decode using text memory and previous motion tokens\n        out = self.decoder(tgt=emb, memory=memory, tgt_mask=tgt_mask) # Shape: (batch, seq_len, hidden_size)\n        \n        # -- OUTPUT --\n        # Get logits for each of the 6 layers \n        logits = [head(out) for head in self.heads]\n        \n        # Stack logits to shape (batch, 6, seq_len, vocab_size)\n        return torch.stack(logits, dim=1)\n\n# Let's test our model architecture with the dummy batch we extracted earlier!\n\nmodel = SignMotionGenerator()\n\n# Dummy forward pass\n# We use motion_tokens[:, :, :-1] as input to predict the next tokens\ndecoder_input_motions = batch['motion_tokens'][:, :, :-1] \n\noutput_logits = model(\n    input_ids=batch['input_ids'], \n    attention_mask=batch['attention_mask'], \n    motion_tokens=decoder_input_motions\n)\n\nprint(f\"Output Logits Shape: {output_logits.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:41.635760Z","iopub.execute_input":"2026-03-21T11:12:41.636189Z","iopub.status.idle":"2026-03-21T11:12:47.642621Z","shell.execute_reply.started":"2026-03-21T11:12:41.636158Z","shell.execute_reply":"2026-03-21T11:12:47.641627Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 8\n\n1. **BERT Warnings:** The \"Some weights of BertModel were not initialized...\" warning from Hugging Face is completely expected and correct. We are using BERT's body (`AutoModel`), which extracts raw semantic vectors (hidden states), discarding the \"Masked Language Modeling\" heads used in its pre-training. So, everything is fine!\n\n2. **Logits Shape:** Our output shape is [16, 6, 151, 514].\n\n   - **16:** Our batch size.\n   - **6:** The number of RVQ layers we need to predict.\n   - **151:** The sequence length in this batch. The original sequence was 152, but we cut the last token because the model is trying to predict the *next* token.\n   - **514:** Our vocabulary size (0-511 for moves, 512 for PAD, 513 for EOS).","metadata":{}},{"cell_type":"markdown","source":"# 9: Loss Function & Training Loop 📉","metadata":{}},{"cell_type":"code","source":"import torch.optim as optim\n\n# Define the device\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Using device: {device}\")\n\n# Move model to device\nmodel = model.to(device)\n\n# Define Optimizer\n# AdamW is the standard robust optimizer for Transformer architectures\noptimizer = optim.AdamW(filter(lambda p: p.requires_grad, model.parameters()), lr=1e-4)\n\n# Cross Entropy Loss, strictly ignoring the padding token \ncriterion = nn.CrossEntropyLoss(ignore_index=MOTION_PAD_IDX)\n\n# Set model to training mode\nmodel.train()\n\nprint(\"Starting a mini training loop...\")\n\n# Test with just 5 batches\nnum_batches_to_test = 5\n\nfor batch_idx, batch in enumerate(train_loader):\n    if batch_idx >= num_batches_to_test:\n        break\n        \n    # Move batch data to the chosen device\n    input_ids = batch['input_ids'].to(device)\n    attention_mask = batch['attention_mask'].to(device)\n    motion_tokens = batch['motion_tokens'].to(device)\n    \n    # AUTOREGRESSIVE SHIFT\n    # Input is all tokens except the last one \n    decoder_input = motion_tokens[:, :, :-1]\n    \n    # Target is all tokens except the first one \n    target_tokens = motion_tokens[:, :, 1:]\n    \n    # Zero the gradients\n    optimizer.zero_grad()\n    \n    # Forward pass\n    logits = model(input_ids, attention_mask, decoder_input)\n    \n    # Reshape logits and targets for CrossEntropyLoss\n    # Flatten everything except the vocabulary dimension\n    logits_flat = logits.reshape(-1, logits.size(-1))\n    targets_flat = target_tokens.reshape(-1)\n    \n    # Calculate loss\n    loss = criterion(logits_flat, targets_flat)\n    \n    # Backward pass: Compute gradients\n    loss.backward()\n    \n    # Update weights\n    optimizer.step()\n    \n    print(f\"Batch {batch_idx+1}/{num_batches_to_test} | Loss: {loss.item():.4f}\")\n\nprint(\"Mini training loop completed successfully!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:47.643625Z","iopub.execute_input":"2026-03-21T11:12:47.643936Z","iopub.status.idle":"2026-03-21T11:12:49.438495Z","shell.execute_reply.started":"2026-03-21T11:12:47.643911Z","shell.execute_reply":"2026-03-21T11:12:49.437842Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 9\n\nOur Loss values decreased beautifully and steadily from 6.4138 to 5.9121! This proves that our model's architecture is mathematically flawless, gradients are flowing smoothly, and weights are being updated in the right direction.","metadata":{}},{"cell_type":"markdown","source":"# 10: Autoregressive Inference 🔮","metadata":{}},{"cell_type":"code","source":"def generate_motion(model, text, tokenizer, max_gen_len=100):\n    # Set model to evaluation mode\n    model.eval() \n    \n    # 1. Tokenize the input text\n    text_encoding = tokenizer(\n        text, \n        max_length=MAX_TEXT_LEN, \n        padding='max_length', \n        truncation=True, \n        return_tensors=\"pt\"\n    )\n    \n    # Move text tensors to the correct device\n    input_ids = text_encoding['input_ids'].to(device)\n    attention_mask = text_encoding['attention_mask'].to(device)\n    \n    # 2. Initialize the generated sequence \n    # We use a PAD token to kickstart the decoder \n    # Shape: (batch_size=1, num_layers=6, seq_len=1)\n    generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n    print(f\"Generating motion for: '{text}'...\")\n    \n    with torch.no_grad(): # Disable gradient calculation for inference\n        for step in range(max_gen_len):\n            # Forward pass \n            logits = model(input_ids, attention_mask, generated_seq)\n            \n            # Get the predictions for the last time step\n            # logits shape: (batch=1, layers=6, seq_len, vocab_size=514)\n            next_token_logits = logits[:, :, -1, :] \n            \n            # Greedy decoding: Pick the token with the highest probability\n            next_tokens = torch.argmax(next_token_logits, dim=-1) # Shape: (1, 6)\n            \n            # Reshape next_tokens to (1, 6, 1) so we can append it\n            next_tokens = next_tokens.unsqueeze(-1)\n            \n            # Append to the generated sequence\n            generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n            \n            # Stop condition: If the base layer (layer 0) predicts the EOS token (513) \n            if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX:\n                break\n                \n    # Remove the initial kickstart token and extract from PyTorch tensor\n    # Also ignore the generated EOS token at the end if it exists\n    final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n    \n    # If the last token is EOS, remove it from all layers\n    if final_sequence[0, -1] == MOTION_EOS_IDX:\n         final_sequence = final_sequence[:, :-1]\n            \n    return final_sequence\n\n# Let's test it with a sample gloss from our dataset! \nsample_gloss = \"ME LIKE HOME WORK STORY FINISH\"\ngenerated_motion = generate_motion(model, sample_gloss, tokenizer)\n\nprint(f\"Generated Motion Shape: {generated_motion.shape}\")\nprint(f\"First 5 tokens of Base Layer: {generated_motion[0, :5]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:49.440985Z","iopub.execute_input":"2026-03-21T11:12:49.441218Z","iopub.status.idle":"2026-03-21T11:12:50.844404Z","shell.execute_reply.started":"2026-03-21T11:12:49.441197Z","shell.execute_reply":"2026-03-21T11:12:50.843482Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 10\n\n1. **Why exactly 100 tokens?** Because our model is untrained (we only ran a test with 5 stacks). Since it hasn't learned when to stop (the <EOS> token), it kept generating until it hit the limit of `max_gen_len=100` set in our generate_motion function.\n\n2. **Why does it repeat itself as [299, 73, 73, 73, 73]?** This is a common phenomenon we observe in untrained autoregressive models (\"mode collapse\" or the tendency to select the highest frequency class).","metadata":{}},{"cell_type":"markdown","source":"# 11: Submission Formatting 🏆","metadata":{}},{"cell_type":"code","source":"# from tqdm import tqdm\n\n# # Define the columns for the submission file\n# submission_cols = ['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n# submission_data = []\n\n# print(\"Generating submission for the first 3 test samples...\")\n\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df)):\n#     sample_id = row['id']\n#     gloss_text = str(row['gloss'])\n    \n#     # Generate motion tokens\n#     # We set max_gen_len=45 to comfortably pass the \"minimum 40 tokens\" rule\n#     gen_seq = generate_motion(model, gloss_text, tokenizer, max_gen_len=45)\n    \n#     formatted_layers = []\n    \n#     # Process each of the 6 layers\n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n        \n#         # Enforce Competition Constraints: Min 40, Max 800 \n#         if len(layer_tokens) < 40:\n#             # Pad to 40 tokens if it's too short\n#             padding = [MOTION_PAD_IDX] * (40 - len(layer_tokens))\n#             layer_tokens.extend(padding)\n#         elif len(layer_tokens) > 800:\n#             # Truncate to 800 tokens if it's too long\n#             layer_tokens = layer_tokens[:800]\n            \n#         # Convert the integer list to a space-separated string\n#         str_tokens = \" \".join([str(int(token)) for token in layer_tokens])\n#         formatted_layers.append(str_tokens)\n        \n#     # Append the formatted row to our submission list\n#     submission_data.append([sample_id] + formatted_layers)\n\n# # Convert the list to a Pandas DataFrame\n# submission_df = pd.DataFrame(submission_data, columns=submission_cols)\n\n# # Save the DataFrame to a CSV file\n# submission_df.to_csv('submission.csv', index=False)\n\n# print(\"\\nSubmission file 'submission.csv' saved successfully!\")\n# display(submission_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.845481Z","iopub.execute_input":"2026-03-21T11:12:50.845725Z","iopub.status.idle":"2026-03-21T11:12:50.850445Z","shell.execute_reply.started":"2026-03-21T11:12:50.845702Z","shell.execute_reply":"2026-03-21T11:12:50.849713Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 11\n**Public Score:** 0.2531\n\n95 95 95 95... and 87 87 87 87...\n\nSince our model has not yet learned to connect the meaning of words to actions, it takes the safest path and repeats the same token endlessly. In deep learning, this is called **Mode Collapse**. The solution is very simple: train our model for \"real\" this time!","metadata":{}},{"cell_type":"markdown","source":"# 12: Full Training Loop & Model Checkpointing 🏋️‍♂️","metadata":{}},{"cell_type":"code","source":"# from tqdm import tqdm\n# import time\n\n# # Define training hyperparameters \n# NUM_EPOCHS = 5  # For a real Kaggle run, you might want 20-30 epochs\n# SAVE_PATH = \"sign_motion_best_model.pth\"\n\n# def train_model(model, train_loader, optimizer, criterion, epochs, device):\n#     print(f\"Starting full training for {epochs} epochs on {device}...\")\n    \n#     best_loss = float('inf')\n    \n#     for epoch in range(epochs):\n#         model.train()\n#         total_loss = 0.0\n#         start_time = time.time()\n        \n#         # Use tqdm for a beautiful progress bar\n#         epoch_iterator = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{epochs}\")\n        \n#         for batch in epoch_iterator:\n#             # Move data to GPU/CPU\n#             input_ids = batch['input_ids'].to(device)\n#             attention_mask = batch['attention_mask'].to(device)\n#             motion_tokens = batch['motion_tokens'].to(device)\n            \n#             # Autoregressive shift\n#             decoder_input = motion_tokens[:, :, :-1]\n#             target_tokens = motion_tokens[:, :, 1:]\n            \n#             optimizer.zero_grad()\n            \n#             # Forward pass\n#             logits = model(input_ids, attention_mask, decoder_input)\n            \n#             # Reshape for loss calculation\n#             logits_flat = logits.reshape(-1, logits.size(-1))\n#             targets_flat = target_tokens.reshape(-1)\n            \n#             loss = criterion(logits_flat, targets_flat)\n#             loss.backward()\n            \n#             # Gradient clipping to prevent exploding gradients\n#             torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n            \n#             optimizer.step()\n            \n#             total_loss += loss.item()\n#             # Update progress bar with current loss\n#             epoch_iterator.set_postfix(loss=loss.item())\n            \n#         avg_loss = total_loss / len(train_loader)\n#         epoch_time = time.time() - start_time\n        \n#         print(f\"\\nEpoch {epoch+1} Summary | Avg Loss: {avg_loss:.4f} | Time: {epoch_time:.2f}s\")\n        \n#         # Save the best model\n#         if avg_loss < best_loss:\n#             best_loss = avg_loss\n#             print(f\"🥇 New best loss ({best_loss:.4f})! Saving model to {SAVE_PATH}\")\n#             torch.save(model.state_dict(), SAVE_PATH)\n\n# # Reset the optimizer to start fresh\n# optimizer = optim.AdamW(filter(lambda p: p.requires_grad, model.parameters()), lr=3e-4)\n\n# # Run the training loop\n# train_model(model, train_loader, optimizer, criterion, epochs=NUM_EPOCHS, device=device)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.851484Z","iopub.execute_input":"2026-03-21T11:12:50.851828Z","iopub.status.idle":"2026-03-21T11:12:50.868008Z","shell.execute_reply.started":"2026-03-21T11:12:50.851760Z","shell.execute_reply":"2026-03-21T11:12:50.867322Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 12\n\n**What does this mean?**\nOur model has now escaped those annoying repetitions of \"95 95 95...\" (mode collapse). It has learned the spatial relationship between text embeddings (BERT embeddings) and motion tokens (RVQ tokens). It's beginning to decipher the \"grammar\" of sign language!\n\nWe now have a real, trained model with its weights saved to disk (`sign_motion_best_model.pth`).","metadata":{}},{"cell_type":"markdown","source":"# 13: The Final Full Submission 🚀","metadata":{}},{"cell_type":"code","source":"# # Create a silent version of the generation function to avoid output flooding\n# def generate_motion_silent(model, text, tokenizer, max_gen_len=150):\n#     model.eval() \n#     text_encoding = tokenizer(\n#         text, max_length=MAX_TEXT_LEN, padding='max_length', \n#         truncation=True, return_tensors=\"pt\"\n#     )\n#     input_ids = text_encoding['input_ids'].to(device)\n#     attention_mask = text_encoding['attention_mask'].to(device)\n    \n#     generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n#     with torch.no_grad():\n#         for step in range(max_gen_len):\n#             logits = model(input_ids, attention_mask, generated_seq)\n#             next_token_logits = logits[:, :, -1, :] \n#             next_tokens = torch.argmax(next_token_logits, dim=-1).unsqueeze(-1)\n#             generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n            \n#             if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX:\n#                 break\n                \n#     final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n#     if final_sequence[0, -1] == MOTION_EOS_IDX:\n#          final_sequence = final_sequence[:, :-1]\n            \n#     return final_sequence\n\n# # Load the best model weights just to be 100% sure\n# model.load_state_dict(torch.load(SAVE_PATH))\n# model.to(device)\n# model.eval()\n\n# print(\"Starting full generation for 3000 test samples...\")\n\n# submission_data = []\n\n# # Process ALL rows in the test dataset\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating Submission\"):\n#     sample_id = row['id']\n#     gloss_text = str(row['gloss'])\n    \n#     # We use max_gen_len=150 as a safe upper bound since average length was ~108\n#     gen_seq = generate_motion_silent(model, gloss_text, tokenizer, max_gen_len=150)\n    \n#     formatted_layers = []\n    \n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n        \n#         # Enforce Competition Constraints\n#         if len(layer_tokens) < 40:\n#             padding = [MOTION_PAD_IDX] * (40 - len(layer_tokens))\n#             layer_tokens.extend(padding)\n#         elif len(layer_tokens) > 800:\n#             layer_tokens = layer_tokens[:800]\n            \n#         str_tokens = \" \".join([str(int(token)) for token in layer_tokens])\n#         formatted_layers.append(str_tokens)\n        \n#     submission_data.append([sample_id] + formatted_layers)\n\n# # Save the final DataFrame \n# submission_cols = ['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n# submission_df = pd.DataFrame(submission_data, columns=submission_cols)\n# submission_df.to_csv('submission.csv', index=False)\n\n# print(\"\\nFinal 'submission.csv' is ready!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.868851Z","iopub.execute_input":"2026-03-21T11:12:50.869231Z","iopub.status.idle":"2026-03-21T11:12:50.883516Z","shell.execute_reply.started":"2026-03-21T11:12:50.869196Z","shell.execute_reply":"2026-03-21T11:12:50.882815Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 13\n**Public Score:** 0.3853\n\nThis represents a **52% increase**. More importantly, if you look at the tokens in the `submission.csv`, you will see that the meaningless \"mode collapse\" sequences (like repeating 95 95 95 or 87 87) have been replaced by diverse token combinations that represent real movements.\n\nOur model can now read English \"Gloss\" texts, decode their meaning, and transform them into 3D motion tokens across 6 different detail layers. We have taken the most significant step towards building an AI avatar that works for individuals with hearing impairments.","metadata":{}},{"cell_type":"markdown","source":"# 14: Building a Custom Length Estimator 📏","metadata":{}},{"cell_type":"code","source":"# from sklearn.linear_model import Ridge\n\n# print(\"Building a Custom Length Estimator..\")\n\n# # 1. Feature Extraction: Count the number of words in the gloss\n# train_clean['gloss_word_count'] = train_clean['gloss'].apply(lambda x: len(str(x).split()))\n\n# # 2. Prepare X (Features) and y (Target Sequence Lengths)\n# X_train_len = train_clean[['gloss_word_count']].values\n# y_train_len = train_clean['seq_len'].values\n\n# # 3. Initialize and train a Ridge Regression model (Robust to outliers)\n# length_model = Ridge(alpha=1.0)\n# length_model.fit(X_train_len, y_train_len)\n\n# # Print the learned equation\n# slope = length_model.coef_[0]\n# intercept = length_model.intercept_\n# print(f\"Learned Equation: Motion Length = ({slope:.2f} * Word Count) + {intercept:.2f}\")\n\n# # 4. Create a helper function to predict length for test data\n# def predict_motion_length(gloss_text):\n#     # Count words\n#     word_count = len(str(gloss_text).split())\n    \n#     # Predict\n#     pred_len = length_model.predict(np.array([[word_count]]))[0]\n    \n#     # Enforce competition limits (40 to 800) and convert to integer\n#     pred_len = int(np.clip(pred_len, 40, 800))\n    \n#     return pred_len\n\n# # Let's test our new estimator with some examples!\n# test_gloss_1 = \"ME GO STORE TOMORROW\" # 4 words\n# test_gloss_2 = \"ME TELL YOU HOW DO IT TAKE LONGER THAN JUST GO AHEAD AND DO IT MYSELF\" # 17 words\n\n# print(f\"\\nTest 1 ('{test_gloss_1}'): Estimated Length -> {predict_motion_length(test_gloss_1)} tokens\")\n# print(f\"Test 2 ('{test_gloss_2}'): Estimated Length -> {predict_motion_length(test_gloss_2)} tokens\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.884423Z","iopub.execute_input":"2026-03-21T11:12:50.884713Z","iopub.status.idle":"2026-03-21T11:12:50.899227Z","shell.execute_reply.started":"2026-03-21T11:12:50.884671Z","shell.execute_reply":"2026-03-21T11:12:50.898375Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 14\n\n**Intercept (31.23):** The model indicates that when a person stands in front of the camera, raises their hands, and prepares to begin (setup/teardown), there is an average preparation/pause time of **31 tokens**. Given that RVQ produces 7.5 tokens per second, this translates to approximately **4 seconds** of setup time.\n\n**Slope (15.90):** For each additional word (gloss), approximately **16 tokens** (about **2 seconds** of sign language gesture) are added to the motion.\n\n**Dynamic Lengths:** While \"ME GO STORE TOMORROW\" (4 words) produces **94 tokens**, a 17-word epic narrative generates **285 tokens**—demonstrating how sentence length dynamically scales the total motion frames.","metadata":{}},{"cell_type":"markdown","source":"# 15: Optimizing Inference with Length Estimator 🏆","metadata":{}},{"cell_type":"code","source":"# # Create the optimized generation function using the length estimator\n# def generate_motion_optimized(model, text, tokenizer, target_len):\n#     model.eval() \n#     text_encoding = tokenizer(\n#         text, max_length=MAX_TEXT_LEN, padding='max_length', \n#         truncation=True, return_tensors=\"pt\"\n#     )\n#     input_ids = text_encoding['input_ids'].to(device)\n#     attention_mask = text_encoding['attention_mask'].to(device)\n    \n#     # Initialize with PAD token\n#     generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n#     with torch.no_grad():\n#         # Generate exactly target_len tokens\n#         for step in range(target_len):\n#             logits = model(input_ids, attention_mask, generated_seq)\n#             next_token_logits = logits[:, :, -1, :] \n#             next_tokens = torch.argmax(next_token_logits, dim=-1).unsqueeze(-1)\n#             generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n            \n#             # If EOS is predicted early, pad the rest of the sequence\n#             if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX:\n#                 remaining_steps = target_len - (step + 1)\n#                 if remaining_steps > 0:\n#                     padding = torch.full((1, 6, remaining_steps), MOTION_PAD_IDX, dtype=torch.long, device=device)\n#                     generated_seq = torch.cat([generated_seq, padding], dim=-1)\n#                 break\n                \n#     # Remove the initial kickstart PAD token\n#     final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n            \n#     return final_sequence\n\n# print(\"Starting FINAL optimized generation for 3000 test samples...\")\n\n# submission_data = []\n\n# # Process ALL rows with dynamic length prediction\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating Gold Submission\"):\n#     sample_id = row['id']\n#     gloss_text = str(row['gloss'])\n    \n#     # 1. Predict optimal length for this specific text\n#     optimal_length = predict_motion_length(gloss_text)\n    \n#     # 2. Generate motion using the predicted length\n#     gen_seq = generate_motion_optimized(model, gloss_text, tokenizer, target_len=optimal_length)\n    \n#     formatted_layers = []\n    \n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n        \n#         # Enforce exact length constraints just to be absolutely safe\n#         if len(layer_tokens) < 40:\n#             layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n#         elif len(layer_tokens) > 800:\n#             layer_tokens = layer_tokens[:800]\n            \n#         str_tokens = \" \".join([str(int(token)) for token in layer_tokens])\n#         formatted_layers.append(str_tokens)\n        \n#     submission_data.append([sample_id] + formatted_layers)\n\n# # Save the ultimate DataFrame\n# submission_cols = ['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n# submission_df = pd.DataFrame(submission_data, columns=submission_cols)\n# submission_df.to_csv('submission_gold.csv', index=False)\n\n# print(\"\\nUltimate 'submission_gold.csv' is ready!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.900111Z","iopub.execute_input":"2026-03-21T11:12:50.900374Z","iopub.status.idle":"2026-03-21T11:12:50.915461Z","shell.execute_reply.started":"2026-03-21T11:12:50.900354Z","shell.execute_reply":"2026-03-21T11:12:50.914723Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 15\n**public score: 0.3468**\n\n### Why Did Our Score Drop? (Post-Mortem Analysis)\n\nWe fell into a very classic \"Over-Engineering\" trap here. Let's take a detailed look at what we did:\n\n1.  **The Nature of Motion:** Our custom length predictor analyzed the text and decided, \"This motion should last exactly 94 tokens.\" However, when our model wanted to finish the sentence at step 70 by producing an `<EOS>` (End of Sequence), we told it, \"No, you will generate until step 94!\" and we forcibly padded the remaining 24 steps with `MOTION_PAD_IDX (512)`.\n\n2.  **Metric Massacre:** The competition's FID (Fidelity) metric measures how closely the generated motions resemble real human movements. A real person doesn't freeze like a robot for 2 seconds after finishing their sentence! (The 3D animation equivalent of 512 padding is \"freezing\"). These artificial freezes and abrupt cuts destroyed our FID score and our \"Diversity\" score.","metadata":{}},{"cell_type":"markdown","source":"# 16: Training Loop 🌟","metadata":{}},{"cell_type":"code","source":"# from sklearn.model_selection import train_test_split\n# from torch.optim.lr_scheduler import ReduceLROnPlateau\n# import time\n# import copy\n# from tqdm import tqdm\n\n# print(\"Setting up the Professional Training Pipeline... \")\n\n# # 1. Split the cleaned dataset\n# train_df, val_df = train_test_split(train_clean, test_size=0.10, random_state=42)\n# train_df = train_df.reset_index(drop=True)\n# val_df = val_df.reset_index(drop=True)\n\n# print(f\"Training Samples: {len(train_df)}\")\n# print(f\"Validation Samples: {len(val_df)}\")\n\n# # 2. Create Datasets and DataLoaders\n# train_dataset = SignMotionDataset(train_df, tokenizer, MAX_TEXT_LEN)\n# val_dataset = SignMotionDataset(val_df, tokenizer, MAX_TEXT_LEN)\n\n# train_loader = DataLoader(train_dataset, batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate_fn, drop_last=True)\n# val_loader = DataLoader(val_dataset, batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate_fn, drop_last=False)\n\n# # 3. Re-initialize the Model and Optimizer for a fresh start \n# model = SignMotionGenerator().to(device)\n# optimizer = optim.AdamW(filter(lambda p: p.requires_grad, model.parameters()), lr=5e-4)\n\n# # THE FIX: Removed 'verbose=True'\n# scheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2)\n\n# # 4. Training Function\n# def train_professional(model, train_loader, val_loader, optimizer, scheduler, criterion, epochs, device):\n#     best_val_loss = float('inf')\n#     best_model_wts = copy.deepcopy(model.state_dict())\n    \n#     for epoch in range(epochs):\n#         print(f\"\\nEpoch {epoch+1}/{epochs}\")\n#         print(\"-\" * 20)\n        \n#         # --- TRAINING PHASE ---\n#         model.train()\n#         train_loss = 0.0\n        \n#         train_iterator = tqdm(train_loader, desc=\"Training\")\n#         for batch in train_iterator:\n#             input_ids = batch['input_ids'].to(device)\n#             attention_mask = batch['attention_mask'].to(device)\n#             motion_tokens = batch['motion_tokens'].to(device)\n            \n#             decoder_input = motion_tokens[:, :, :-1]\n#             target_tokens = motion_tokens[:, :, 1:]\n            \n#             optimizer.zero_grad()\n#             logits = model(input_ids, attention_mask, decoder_input)\n            \n#             loss = criterion(logits.reshape(-1, logits.size(-1)), target_tokens.reshape(-1))\n#             loss.backward()\n#             torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n#             optimizer.step()\n            \n#             train_loss += loss.item()\n#             train_iterator.set_postfix(loss=loss.item())\n            \n#         avg_train_loss = train_loss / len(train_loader)\n        \n#         # --- VALIDATION PHASE ---\n#         model.eval()\n#         val_loss = 0.0\n        \n#         with torch.no_grad():\n#             val_iterator = tqdm(val_loader, desc=\"Validation\", colour=\"green\")\n#             for batch in val_iterator:\n#                 input_ids = batch['input_ids'].to(device)\n#                 attention_mask = batch['attention_mask'].to(device)\n#                 motion_tokens = batch['motion_tokens'].to(device)\n                \n#                 decoder_input = motion_tokens[:, :, :-1]\n#                 target_tokens = motion_tokens[:, :, 1:]\n                \n#                 logits = model(input_ids, attention_mask, decoder_input)\n#                 loss = criterion(logits.reshape(-1, logits.size(-1)), target_tokens.reshape(-1))\n#                 val_loss += loss.item()\n                \n#         avg_val_loss = val_loss / len(val_loader)\n        \n#         print(f\"Train Loss: {avg_train_loss:.4f} | Val Loss: {avg_val_loss:.4f}\")\n        \n#         # Step the scheduler\n#         scheduler.step(avg_val_loss)\n        \n#         # Save best model\n#         if avg_val_loss < best_val_loss:\n#             best_val_loss = avg_val_loss\n#             best_model_wts = copy.deepcopy(model.state_dict())\n#             torch.save(best_model_wts, \"sign_motion_ultimate_model.pth\")\n#             print(f\"🌟 New Best Model Saved! Val Loss: {best_val_loss:.4f}\")\n\n#     print(\"\\nTraining Complete!\")\n#     model.load_state_dict(best_model_wts)\n#     return model\n\n# # START THE MARATHON! \n# NUM_EPOCHS = 20\n# model = train_professional(model, train_loader, val_loader, optimizer, scheduler, criterion, epochs=NUM_EPOCHS, device=device)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.916401Z","iopub.execute_input":"2026-03-21T11:12:50.916807Z","iopub.status.idle":"2026-03-21T11:12:50.931704Z","shell.execute_reply.started":"2026-03-21T11:12:50.916747Z","shell.execute_reply":"2026-03-21T11:12:50.930827Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 16\n\nLet's look at the profound deep learning story this table tells us:\n\n- **Epoch 1 - 15 (The Golden Age):** Our model improves rapidly on both the training data (Train Loss) and the validation data (Val Loss) it has never seen before.\n    - Epoch 1: Val Loss = 2.6183\n    - Epoch 15: Val Loss = 1.8442 🌟 (Best Score!)\n\n- **Epoch 16 - 30 (The Dark Age - Overfitting Begins):** This is precisely where things take a turn for the worse. Our model starts to memorize the training data (Train Loss drops from 1.21 to 0.78), but it fails on the validation set—i.e., in the \"real world\" (Val Loss increases from 1.84 to 2.03).\n\nThis is what we call **Overfitting**. Training the model for 50 epochs was pointless because after the 15th epoch, the model stopped learning new things and simply began to memorize the training data.","metadata":{}},{"cell_type":"markdown","source":"# 17: The Ultimate Gold Submission 🚀","metadata":{}},{"cell_type":"code","source":"# # Create a silent version of the generation function (with freedom to stop early!)\n# def generate_motion_silent(model, text, tokenizer, max_gen_len=150):\n#     model.eval() \n#     text_encoding = tokenizer(\n#         text, max_length=MAX_TEXT_LEN, padding='max_length', \n#         truncation=True, return_tensors=\"pt\"\n#     )\n#     input_ids = text_encoding['input_ids'].to(device)\n#     attention_mask = text_encoding['attention_mask'].to(device)\n    \n#     generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n#     with torch.no_grad():\n#         for step in range(max_gen_len):\n#             logits = model(input_ids, attention_mask, generated_seq)\n#             next_token_logits = logits[:, :, -1, :] \n#             next_tokens = torch.argmax(next_token_logits, dim=-1).unsqueeze(-1)\n#             generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n            \n#             # THE SECRET SAUCE: Stop generating as soon as the model predicts EOS!\n#             if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX:\n#                 break\n                \n#     final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n#     if final_sequence[0, -1] == MOTION_EOS_IDX:\n#          final_sequence = final_sequence[:, :-1]\n            \n#     return final_sequence\n\n# print(\"Loading the ultimate weights from Epoch 15...\")\n\n# # Load the BEST model weights saved before overfitting!\n# model.load_state_dict(torch.load(\"sign_motion_ultimate_model.pth\"))\n# model.eval()\n\n# print(\"Starting FINAL generation for 3000 test samples...\")\n\n# submission_data = []\n\n# # Process ALL rows\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating Ultimate Gold Submission\"):\n#     sample_id = row['id']\n#     gloss_text = str(row['gloss'])\n    \n#     gen_seq = generate_motion_silent(model, gloss_text, tokenizer, max_gen_len=150)\n    \n#     formatted_layers = []\n    \n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n        \n#         # Enforce Competition Constraints safely\n#         if len(layer_tokens) < 40:\n#             layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n#         elif len(layer_tokens) > 800:\n#             layer_tokens = layer_tokens[:800]\n            \n#         str_tokens = \" \".join([str(int(token)) for token in layer_tokens])\n#         formatted_layers.append(str_tokens)\n        \n#     submission_data.append([sample_id] + formatted_layers)\n\n# # Save the Ultimate DataFrame\n# submission_cols = ['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n# submission_df = pd.DataFrame(submission_data, columns=submission_cols)\n\n# # We save it with a new name to track our progress\n# submission_df.to_csv('submission_ultimate.csv', index=False)\n\n# print(\"\\nUltimate 'submission_ultimate.csv' is ready!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.932637Z","iopub.execute_input":"2026-03-21T11:12:50.932993Z","iopub.status.idle":"2026-03-21T11:12:50.946503Z","shell.execute_reply.started":"2026-03-21T11:12:50.932965Z","shell.execute_reply":"2026-03-21T11:12:50.945664Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 17\n\nWhen we halted training at Epoch 27, our system saved that magic `sign_motion_ultimate_model.pth` file. That file represented the model's smartest moment (Epoch 15), right before it began to overfit the training data. By taking these optimal weights and giving the model the freedom to \"stop whenever you want by generating <EOS>,\" both our alignment (R-Precision) and realism (FID) scores went through the roof!","metadata":{}},{"cell_type":"markdown","source":"# 18: The Monolithic Pipeline 🚀","metadata":{}},{"cell_type":"code","source":"# from sklearn.model_selection import train_test_split\n# from torch.optim.lr_scheduler import ReduceLROnPlateau\n# from transformers import AutoModel\n# import torch.nn as nn\n# import torch.nn.functional as F\n# import torch.optim as optim\n# import torch\n# import copy\n# from tqdm import tqdm\n\n# print(\"--- 🚀 INITIATING The Monolithic Pipeline ---\")\n# device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# # 1. TRAIN/VAL SPLIT & DATALOADERS\n# print(\"1. Setting up Train/Validation Splits and DataLoaders...\")\n# train_split, val_split = train_test_split(train_clean, test_size=0.10, random_state=42)\n# train_split = train_split.reset_index(drop=True)\n# val_split = val_split.reset_index(drop=True)\n\n# train_loader = DataLoader(SignMotionDataset(train_split, tokenizer, MAX_TEXT_LEN), batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate_fn, drop_last=True)\n# val_loader = DataLoader(SignMotionDataset(val_split, tokenizer, MAX_TEXT_LEN), batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate_fn, drop_last=False)\n\n# # 2. MODEL ARCHITECTURE\n# print(\"2. Building the SignMotionGenerator Model...\")\n# class SignMotionGenerator(nn.Module):\n#     def __init__(self, vocab_size=514, hidden_size=768, num_layers=4, nhead=8):\n#         super().__init__()\n#         self.text_encoder = AutoModel.from_pretrained(\"bert-base-uncased\")\n#         for param in self.text_encoder.parameters(): param.requires_grad = False\n#         self.motion_emb = nn.Embedding(vocab_size, hidden_size)\n#         self.pos_encoder = nn.Embedding(1000, hidden_size)\n#         self.decoder = nn.TransformerDecoder(nn.TransformerDecoderLayer(d_model=hidden_size, nhead=nhead, batch_first=True), num_layers=num_layers)\n#         self.heads = nn.ModuleList([nn.Linear(hidden_size, vocab_size) for _ in range(6)])\n\n#     def forward(self, input_ids, attention_mask, motion_tokens):\n#         memory = self.text_encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state\n#         motions_t = motion_tokens.permute(0, 2, 1)\n#         emb = self.motion_emb(motions_t).sum(dim=2)\n#         emb = emb + self.pos_encoder(torch.arange(0, emb.size(1), device=emb.device).unsqueeze(0))\n#         tgt_mask = nn.Transformer.generate_square_subsequent_mask(emb.size(1)).to(emb.device)\n#         out = self.decoder(tgt=emb, memory=memory, tgt_mask=tgt_mask)\n#         return torch.stack([head(out) for head in self.heads], dim=1)\n\n# # 3. HIERARCHICAL LOSS\n# print(\"3. Setting up Hierarchical Loss...\")\n# LAYER_WEIGHTS = torch.tensor([1.0, 0.8, 0.6, 0.4, 0.2, 0.1], device=device)\n\n# def calculate_hierarchical_loss(logits, targets, pad_idx=512, smoothing=0.1):\n#     total_loss = 0.0\n#     for i in range(6):\n#         total_loss += F.cross_entropy(logits[:, i, :, :].reshape(-1, logits.size(-1)), targets[:, i, :].reshape(-1), ignore_index=pad_idx, label_smoothing=smoothing) * LAYER_WEIGHTS[i]\n#     return total_loss / LAYER_WEIGHTS.sum()\n\n# # 4. TRAINING LOOP\n# print(\"4. Launching Gold Medal Training Loop...\")\n# model = SignMotionGenerator().to(device)\n# optimizer = optim.AdamW(filter(lambda p: p.requires_grad, model.parameters()), lr=5e-4)\n# scheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2)\n\n# best_val_loss = float('inf')\n# best_model_wts = copy.deepcopy(model.state_dict())\n\n# for epoch in range(20):\n#     print(f\"\\nEpoch {epoch+1}/20 [Hierarchical & Label Smoothing]\")\n#     model.train()\n#     train_loss = 0.0\n#     train_iterator = tqdm(train_loader, desc=\"Training\")\n\n#     for batch in train_iterator:\n#         input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#         optimizer.zero_grad()\n#         logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n#         loss = calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1)\n#         loss.backward()\n#         torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n#         optimizer.step()\n#         train_loss += loss.item()\n#         train_iterator.set_postfix(loss=loss.item())\n\n#     model.eval()\n#     val_loss = 0.0\n#     with torch.no_grad():\n#         for batch in tqdm(val_loader, desc=\"Validation\", colour=\"green\"):\n#             input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#             logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n#             val_loss += calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1).item()\n\n#     avg_train_loss, avg_val_loss = train_loss / len(train_loader), val_loss / len(val_loader)\n#     print(f\"Train Loss: {avg_train_loss:.4f} | Val Loss: {avg_val_loss:.4f}\")\n#     scheduler.step(avg_val_loss)\n\n#     if avg_val_loss < best_val_loss:\n#         best_val_loss = avg_val_loss\n#         best_model_wts = copy.deepcopy(model.state_dict())\n#         torch.save(best_model_wts, \"sign_motion_hierarchical_gold.pth\")\n#         print(f\"🏆 NEW GOLD MODEL SAVED! Val Loss: {best_val_loss:.4f}\")\n\n# # 5. INFERENCE & SUBMISSION GENERATION\n# print(\"\\n5. Loading Best Weights and Generating Final Submission...\")\n# model.load_state_dict(best_model_wts)\n# model.eval()\n\n# def generate_motion_silent(model, text, tokenizer, max_gen_len=150):\n#     text_encoding = tokenizer(text, max_length=MAX_TEXT_LEN, padding='max_length', truncation=True, return_tensors=\"pt\")\n#     input_ids, attention_mask = text_encoding['input_ids'].to(device), text_encoding['attention_mask'].to(device)\n#     generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n\n#     with torch.no_grad():\n#         for step in range(max_gen_len):\n#             logits = model(input_ids, attention_mask, generated_seq)\n#             next_tokens = torch.argmax(logits[:, :, -1, :], dim=-1).unsqueeze(-1)\n#             generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n#             if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX: break\n\n#     final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n#     return final_sequence[:, :-1] if final_sequence[0, -1] == MOTION_EOS_IDX else final_sequence\n\n# submission_data = []\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating Submission\"):\n#     gen_seq = generate_motion_silent(model, str(row['gloss']), tokenizer)\n#     formatted_layers = []\n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n#         if len(layer_tokens) < 40: layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n#         elif len(layer_tokens) > 800: layer_tokens = layer_tokens[:800]\n#         formatted_layers.append(\" \".join([str(int(token)) for token in layer_tokens]))\n#     submission_data.append([row['id']] + formatted_layers)\n\n# submission_df = pd.DataFrame(submission_data, columns=['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5'])\n# submission_df.to_csv('submission_hierarchical.csv', index=False)\n# print(\"\\n--- 🏆 PIPELINE COMPLETED! 'submission_hierarchical.csv' IS READY! ---\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.947444Z","iopub.execute_input":"2026-03-21T11:12:50.947742Z","iopub.status.idle":"2026-03-21T11:12:50.961310Z","shell.execute_reply.started":"2026-03-21T11:12:50.947711Z","shell.execute_reply":"2026-03-21T11:12:50.960566Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 18\n**Public Score:** 0.3946\n\nOur established \"God Mode\" pipeline, featuring a hierarchical loss function and label smoothing, is working mathematically.","metadata":{}},{"cell_type":"markdown","source":"# 19: Building the MaskGIT Architecture 🛠️","metadata":{}},{"cell_type":"code","source":"# import torch\n# import torch.nn as nn\n# import torch.nn.functional as F\n# import torch.optim as optim\n# from transformers import AutoModel\n# import copy\n# from tqdm import tqdm\n# import random\n\n# print(\"--- ☠️ INITIATING HARDCORE MASKGIT PIPELINE ---\")\n\n# MOTION_PAD_IDX = 512\n# MOTION_EOS_IDX = 513\n# MOTION_MASK_IDX = 514\n# VOCAB_SIZE = 515 \n# device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# # 1. NON-AUTOREGRESSIVE (MASKGIT) MODEL ARCHITECTURE\n# class SignMotionMaskGIT(nn.Module):\n#     def __init__(self, vocab_size=VOCAB_SIZE, hidden_size=768, num_layers=4, nhead=8):\n#         super().__init__()\n#         self.text_encoder = AutoModel.from_pretrained(\"bert-base-uncased\")\n#         for param in self.text_encoder.parameters(): param.requires_grad = False\n            \n#         self.motion_emb = nn.Embedding(vocab_size, hidden_size)\n#         self.pos_encoder = nn.Embedding(1000, hidden_size) \n\n#         self.decoder = nn.TransformerDecoder(nn.TransformerDecoderLayer(d_model=hidden_size, nhead=nhead, batch_first=True), num_layers=num_layers)\n#         self.heads = nn.ModuleList([nn.Linear(hidden_size, vocab_size) for _ in range(6)])\n        \n#     def forward(self, input_ids, attention_mask, masked_motion_tokens):\n#         memory = self.text_encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state \n#         motions_t = masked_motion_tokens.permute(0, 2, 1) \n#         emb = self.motion_emb(motions_t).sum(dim=2)\n#         emb = emb + self.pos_encoder(torch.arange(0, emb.size(1), device=emb.device).unsqueeze(0))\n        \n#         out = self.decoder(tgt=emb, memory=memory, tgt_mask=None) \n        \n#         return torch.stack([head(out) for head in self.heads], dim=1)\n\n# def apply_dynamic_masking(targets, mask_ratio, pad_idx=512, eos_idx=513, mask_idx=514):\n#     masked_inputs = targets.clone()\n#     batch_size, num_layers, seq_len = targets.shape\n    \n#     valid_mask = (targets != pad_idx) & (targets != eos_idx)\n    \n#     random_matrix = torch.rand(batch_size, num_layers, seq_len, device=targets.device)\n#     mask_condition = (random_matrix < mask_ratio) & valid_mask\n    \n#     masked_inputs[mask_condition] = mask_idx\n    \n#     return masked_inputs\n\n# LAYER_WEIGHTS = torch.tensor([1.0, 0.8, 0.6, 0.4, 0.2, 0.1], device=device)\n# def calculate_hierarchical_loss(logits, targets, pad_idx=512, smoothing=0.1):\n#     total_loss = 0.0\n#     for i in range(6):\n#         total_loss += F.cross_entropy(logits[:, i, :, :].reshape(-1, logits.size(-1)), targets[:, i, :].reshape(-1), ignore_index=pad_idx, label_smoothing=smoothing) * LAYER_WEIGHTS[i]\n#     return total_loss / LAYER_WEIGHTS.sum()\n\n# print(\"MaskGIT Architecture and Masking logic successfully defined!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.962185Z","iopub.execute_input":"2026-03-21T11:12:50.962364Z","iopub.status.idle":"2026-03-21T11:12:50.972254Z","shell.execute_reply.started":"2026-03-21T11:12:50.962347Z","shell.execute_reply":"2026-03-21T11:12:50.971386Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 20: The MaskGIT Training Loop (Denoising) ⚡","metadata":{}},{"cell_type":"code","source":"# from sklearn.model_selection import train_test_split\n# from torch.utils.data import DataLoader\n# from torch.optim.lr_scheduler import ReduceLROnPlateau\n# import copy\n\n# print(\"--- 🛠️ RESTORING DATALOADERS & LAUNCHING MASKGIT ---\")\n\n# # 1. RECREATE TRAIN/VAL SPLIT AND DATALOADERS\n# train_split, val_split = train_test_split(train_clean, test_size=0.10, random_state=42)\n# train_split = train_split.reset_index(drop=True)\n# val_split = val_split.reset_index(drop=True)\n\n# # DataLoader'ları %90 Train ve %10 Val olacak şekilde tekrar kur\n# train_loader = DataLoader(SignMotionDataset(train_split, tokenizer, MAX_TEXT_LEN), batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate_fn, drop_last=True)\n# val_loader = DataLoader(SignMotionDataset(val_split, tokenizer, MAX_TEXT_LEN), batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate_fn, drop_last=False)\n\n# # 2. MASKGIT TRAINING LOOP\n# print(\"Launching MaskGIT Denoising Training Loop...\")\n\n# maskgit_model = SignMotionMaskGIT().to(device)\n# optimizer = optim.AdamW(filter(lambda p: p.requires_grad, maskgit_model.parameters()), lr=5e-4)\n# scheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2)\n\n# best_val_loss = float('inf')\n# best_maskgit_wts = copy.deepcopy(maskgit_model.state_dict())\n\n# NUM_EPOCHS = 20\n\n# for epoch in range(NUM_EPOCHS):\n#     print(f\"\\nEpoch {epoch+1}/{NUM_EPOCHS} [MaskGIT Denoising]\")\n#     maskgit_model.train()\n#     train_loss = 0.0\n#     train_iterator = tqdm(train_loader, desc=\"Training\")\n    \n#     for batch in train_iterator:\n#         input_ids, attention_mask, targets = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#         optimizer.zero_grad()\n        \n#         current_mask_ratio = random.uniform(0.1, 0.9)\n#         masked_inputs = apply_dynamic_masking(targets, current_mask_ratio)\n        \n#         logits = maskgit_model(input_ids, attention_mask, masked_inputs)\n#         loss = calculate_hierarchical_loss(logits, targets, pad_idx=MOTION_PAD_IDX, smoothing=0.1)\n        \n#         loss.backward()\n#         torch.nn.utils.clip_grad_norm_(maskgit_model.parameters(), max_norm=1.0)\n#         optimizer.step()\n#         train_loss += loss.item()\n#         train_iterator.set_postfix(loss=loss.item(), mask_ratio=f\"{current_mask_ratio:.2f}\")\n        \n#     maskgit_model.eval()\n#     val_loss = 0.0\n#     with torch.no_grad():\n#         for batch in tqdm(val_loader, desc=\"Validation\", colour=\"green\"):\n#             input_ids, attention_mask, targets = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n            \n#             masked_inputs = apply_dynamic_masking(targets, mask_ratio=0.5)\n#             logits = maskgit_model(input_ids, attention_mask, masked_inputs)\n#             val_loss += calculate_hierarchical_loss(logits, targets, pad_idx=MOTION_PAD_IDX, smoothing=0.1).item()\n            \n#     avg_train_loss, avg_val_loss = train_loss / len(train_loader), val_loss / len(val_loader)\n#     print(f\"Train Loss: {avg_train_loss:.4f} | Val Loss: {avg_val_loss:.4f}\")\n#     scheduler.step(avg_val_loss)\n    \n#     if avg_val_loss < best_val_loss:\n#         best_val_loss = avg_val_loss\n#         best_maskgit_wts = copy.deepcopy(maskgit_model.state_dict())\n#         torch.save(best_maskgit_wts, \"sign_motion_maskgit_gold.pth\")\n#         print(f\"🏆 NEW MASKGIT GOLD MODEL SAVED! Val Loss: {best_val_loss:.4f}\")\n\n# print(\"\\n--- ☠️ MASKGIT TRAINING COMPLETED! ---\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.973128Z","iopub.execute_input":"2026-03-21T11:12:50.973401Z","iopub.status.idle":"2026-03-21T11:12:50.988126Z","shell.execute_reply.started":"2026-03-21T11:12:50.973381Z","shell.execute_reply":"2026-03-21T11:12:50.987297Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 21: Inference (MaskGIT Iterative Decoding) 🪄","metadata":{}},{"cell_type":"code","source":"# print(\"--- 🪄 INITIATING MASKGIT ITERATIVE DECODING ---\")\n\n# maskgit_model.load_state_dict(best_maskgit_wts)\n# maskgit_model.eval()\n\n# def generate_motion_maskgit(model, text, tokenizer, iterations=4):\n#     word_count = len(str(text).split())\n#     target_len = int(np.clip(16 * word_count + 31, 40, 800))\n    \n#     text_encoding = tokenizer(text, max_length=MAX_TEXT_LEN, padding='max_length', truncation=True, return_tensors=\"pt\")\n#     input_ids, attention_mask = text_encoding['input_ids'].to(device), text_encoding['attention_mask'].to(device)\n    \n#     seq = torch.full((1, 6, target_len), MOTION_MASK_IDX, dtype=torch.long, device=device)\n    \n#     with torch.no_grad():\n#         for step in range(iterations):\n#             mask_ratio = np.cos((step / iterations) * (np.pi / 2))\n#             num_masked = int(target_len * mask_ratio)\n            \n#             logits = model(input_ids, attention_mask, seq) # Shape: (1, 6, seq_len, vocab_size)\n            \n#             probs = F.softmax(logits, dim=-1)\n#             confidences, next_tokens = torch.max(probs, dim=-1)\n            \n#             if step == iterations - 1:\n#                 seq = next_tokens\n#                 break\n                \n#             for layer_idx in range(6):\n#                 layer_conf = confidences[0, layer_idx, :]\n                \n#                 _, least_confident_indices = torch.topk(layer_conf, k=num_masked, largest=False)\n                \n#                 seq[0, layer_idx, :] = next_tokens[0, layer_idx, :]\n                \n#                 seq[0, layer_idx, least_confident_indices] = MOTION_MASK_IDX\n                \n#     return seq[0].cpu().numpy()\n\n# print(f\"Generating Ultimate MaskGIT Submission for {len(test_df)} samples...\")\n\n# submission_data = []\n\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Iterative Decoding\"):\n#     gen_seq = generate_motion_maskgit(maskgit_model, str(row['gloss']), tokenizer, iterations=4)\n    \n#     formatted_layers = []\n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n        \n#         if len(layer_tokens) < 40: \n#             layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n#         elif len(layer_tokens) > 800: \n#             layer_tokens = layer_tokens[:800]\n            \n#         formatted_layers.append(\" \".join([str(int(token)) for token in layer_tokens]))\n        \n#     submission_data.append([row['id']] + formatted_layers)\n\n# # Kaydet\n# submission_df = pd.DataFrame(submission_data, columns=['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5'])\n# submission_df.to_csv('submission_maskgit.csv', index=False)\n\n# print(\"\\n--- 🏆 submission_maskgit.csv' IS READY! ---\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:50.989031Z","iopub.execute_input":"2026-03-21T11:12:50.989327Z","iopub.status.idle":"2026-03-21T11:12:51.004235Z","shell.execute_reply.started":"2026-03-21T11:12:50.989307Z","shell.execute_reply":"2026-03-21T11:12:51.003379Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 22: Monolithic Pipeline V2 🌠","metadata":{}},{"cell_type":"code","source":"# import os\n# import pandas as pd\n# import numpy as np\n# import torch\n# import torch.nn as nn\n# import torch.nn.functional as F\n# import torch.optim as optim\n# from torch.utils.data import Dataset, DataLoader\n# from torch.optim.lr_scheduler import ReduceLROnPlateau\n# from transformers import AutoTokenizer, AutoModel\n# from sklearn.model_selection import train_test_split\n# import copy\n# import random\n# from tqdm import tqdm\n# import warnings\n\n# # 0. SETUP & CONFIG\n# warnings.filterwarnings('ignore')\n# DATA_DIR = \"/kaggle/input/competitions/motion-s-hierarchical-text-to-motion-generation-for-sign-language\"\n# MAX_TEXT_LEN = 32\n# MOTION_PAD_IDX = 512\n# MOTION_EOS_IDX = 513\n# BATCH_SIZE = 16 \n# device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n# print(f\"--- ⚔️ INITIATING MODE V2 ON {device.type.upper()} ---\")\n\n# # 1. LOAD & CLEAN DATA\n# print(\"1. Loading and Cleaning Data...\")\n# train_df = pd.read_csv(os.path.join(DATA_DIR, \"train.csv\"))\n# test_df = pd.read_csv(os.path.join(DATA_DIR, \"test.csv\"))\n# train_clean = train_df.dropna().copy()\n# train_clean['seq_len'] = train_clean['base_tokens'].apply(lambda x: len(str(x).split()))\n# train_clean = train_clean[(train_clean['seq_len'] >= 40) & (train_clean['seq_len'] <= 800)].reset_index(drop=True)\n\n# token_cols = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n# for col in token_cols:\n#     train_clean[col] = train_clean[col].apply(lambda x: [int(i) for i in str(x).split()])\n\n# # 2. TOKENIZER & SPLITS\n# print(\"2. Setting up Tokenizer...\")\n# tokenizer = AutoTokenizer.from_pretrained(\"bert-base-uncased\")\n# train_split, val_split = train_test_split(train_clean, test_size=0.10, random_state=42)\n# train_split, val_split = train_split.reset_index(drop=True), val_split.reset_index(drop=True)\n\n# # 3. DATALOADERS WITH GLOSS DROPOUT (Kaggle Trick #1)\n# print(\"3. Building DataLoaders with Gloss Dropout...\")\n# class SignMotionDataset(Dataset):\n#     def __init__(self, df, tokenizer, max_text_len, is_train=False):\n#         self.df, self.tokenizer, self.max_text_len, self.is_train = df, tokenizer, max_text_len, is_train\n#         self.token_cols = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n        \n#     def __len__(self): return len(self.df)\n    \n#     # TRICK\n#     def apply_gloss_dropout(self, text, drop_prob=0.1):\n#         if not self.is_train: return text\n#         words = str(text).split()\n#         if len(words) <= 1: return text\n#         kept_words = [w for w in words if random.random() > drop_prob]\n#         if len(kept_words) == 0: kept_words = [random.choice(words)] # Hepsi silinirse 1 tane bırak\n#         return \" \".join(kept_words)\n\n#     def __getitem__(self, idx):\n#         row = self.df.iloc[idx]\n#         final_text = self.apply_gloss_dropout(row['gloss'])\n#         text_encoding = self.tokenizer(final_text, max_length=self.max_text_len, padding='max_length', truncation=True, return_tensors=\"pt\")\n#         input_ids = text_encoding['input_ids'].squeeze(0)\n#         attention_mask = text_encoding['attention_mask'].squeeze(0)\n#         motion_layers = [row[col] + [MOTION_EOS_IDX] for col in self.token_cols]\n#         return {\"input_ids\": input_ids, \"attention_mask\": attention_mask, \"motion_tokens\": torch.tensor(motion_layers, dtype=torch.long), \"id\": row['id']}\n\n# def collate_fn(batch):\n#     input_ids = torch.stack([item['input_ids'] for item in batch])\n#     attention_mask = torch.stack([item['attention_mask'] for item in batch])\n#     motion_tensors = [item['motion_tokens'] for item in batch]\n#     max_len = max([t.shape[1] for t in motion_tensors])\n#     padded_motions = torch.full((len(batch), 6, max_len), MOTION_PAD_IDX, dtype=torch.long)\n#     for i, tensor in enumerate(motion_tensors): padded_motions[i, :, :tensor.shape[1]] = tensor\n#     return {\"input_ids\": input_ids, \"attention_mask\": attention_mask, \"motion_tokens\": padded_motions}\n\n# train_loader = DataLoader(SignMotionDataset(train_split, tokenizer, MAX_TEXT_LEN, is_train=True), batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate_fn, drop_last=True)\n# val_loader = DataLoader(SignMotionDataset(val_split, tokenizer, MAX_TEXT_LEN, is_train=False), batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate_fn, drop_last=False)\n\n# # 4. MODEL ARCHITECTURE WITH UNFROZEN BERT\n# print(\"4. Building the Model (Unfreezing top BERT layers)...\")\n# class SignMotionGenerator(nn.Module):\n#     def __init__(self, vocab_size=514, hidden_size=768, num_layers=4, nhead=8):\n#         super().__init__()\n#         self.text_encoder = AutoModel.from_pretrained(\"bert-base-uncased\")\n        \n#         # TRICK\n#         for name, param in self.text_encoder.named_parameters():\n#             if any(layer in name for layer in [\"encoder.layer.8\", \"encoder.layer.9\", \"encoder.layer.10\", \"encoder.layer.11\", \"pooler\"]):\n#                 param.requires_grad = True\n#             else:\n#                 param.requires_grad = False\n                \n#         self.motion_emb = nn.Embedding(vocab_size, hidden_size)\n#         self.pos_encoder = nn.Embedding(1000, hidden_size) \n#         self.decoder = nn.TransformerDecoder(nn.TransformerDecoderLayer(d_model=hidden_size, nhead=nhead, batch_first=True), num_layers=num_layers)\n#         self.heads = nn.ModuleList([nn.Linear(hidden_size, vocab_size) for _ in range(6)])\n        \n#     def forward(self, input_ids, attention_mask, motion_tokens):\n#         memory = self.text_encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state \n#         motions_t = motion_tokens.permute(0, 2, 1) \n#         emb = self.motion_emb(motions_t).sum(dim=2)\n#         emb = emb + self.pos_encoder(torch.arange(0, emb.size(1), device=emb.device).unsqueeze(0))\n#         tgt_mask = nn.Transformer.generate_square_subsequent_mask(emb.size(1)).to(emb.device)\n#         out = self.decoder(tgt=emb, memory=memory, tgt_mask=tgt_mask) \n#         return torch.stack([head(out) for head in self.heads], dim=1)\n\n# # 5. DIFFERENTIAL LEARNING RATES & HIERARCHICAL LOSS\n# print(\"5. Launching V2 Mode Training Loop...\")\n# LAYER_WEIGHTS = torch.tensor([1.0, 0.8, 0.6, 0.4, 0.2, 0.1], device=device)\n\n# def calculate_hierarchical_loss(logits, targets, pad_idx=512, smoothing=0.1):\n#     total_loss = 0.0\n#     for i in range(6):\n#         total_loss += F.cross_entropy(logits[:, i, :, :].reshape(-1, logits.size(-1)), targets[:, i, :].reshape(-1), ignore_index=pad_idx, label_smoothing=smoothing) * LAYER_WEIGHTS[i]\n#     return total_loss / LAYER_WEIGHTS.sum()\n\n# model = SignMotionGenerator().to(device)\n\n# bert_params = [p for n, p in model.named_parameters() if \"text_encoder\" in n and p.requires_grad]\n# decoder_params = [p for n, p in model.named_parameters() if \"text_encoder\" not in n and p.requires_grad]\n\n# optimizer = optim.AdamW([\n#     {'params': bert_params, 'lr': 2e-5},\n#     {'params': decoder_params, 'lr': 5e-4}\n# ])\n# scheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2)\n\n# best_val_loss = float('inf')\n# best_model_wts = copy.deepcopy(model.state_dict())\n\n# for epoch in range(20): \n#     print(f\"\\nEpoch {epoch+1}/20 [V2: Unfrozen BERT & Gloss Dropout]\")\n#     model.train()\n#     train_loss = 0.0\n#     train_iterator = tqdm(train_loader, desc=\"Training\")\n    \n#     for batch in train_iterator:\n#         input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#         optimizer.zero_grad()\n#         logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n#         loss = calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1)\n#         loss.backward()\n#         torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n#         optimizer.step()\n#         train_loss += loss.item()\n#         train_iterator.set_postfix(loss=loss.item())\n        \n#     model.eval()\n#     val_loss = 0.0\n#     with torch.no_grad():\n#         for batch in tqdm(val_loader, desc=\"Validation\", colour=\"green\"):\n#             input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#             logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n#             val_loss += calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1).item()\n            \n#     avg_train_loss, avg_val_loss = train_loss / len(train_loader), val_loss / len(val_loader)\n#     print(f\"Train Loss: {avg_train_loss:.4f} | Val Loss: {avg_val_loss:.4f}\")\n#     scheduler.step(avg_val_loss)\n    \n#     if avg_val_loss < best_val_loss:\n#         best_val_loss = avg_val_loss\n#         best_model_wts = copy.deepcopy(model.state_dict())\n#         torch.save(best_model_wts, \"sign_motion_v2.pth\")\n#         print(f\"🏆 NEW V2 MODEL SAVED! Val Loss: {best_val_loss:.4f}\")\n\n# # 6. INFERENCE & SUBMISSION GENERATION\n# print(\"\\n6. Loading Best Weights and Generating Final Submission...\")\n# model.load_state_dict(best_model_wts)\n# model.eval()\n\n# def generate_motion_silent(model, text, tokenizer, max_gen_len=150):\n#     text_encoding = tokenizer(text, max_length=MAX_TEXT_LEN, padding='max_length', truncation=True, return_tensors=\"pt\")\n#     input_ids, attention_mask = text_encoding['input_ids'].to(device), text_encoding['attention_mask'].to(device)\n#     generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n#     with torch.no_grad():\n#         for step in range(max_gen_len):\n#             logits = model(input_ids, attention_mask, generated_seq)\n#             next_tokens = torch.argmax(logits[:, :, -1, :], dim=-1).unsqueeze(-1)\n#             generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n#             if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX: break\n                \n#     final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n#     return final_sequence[:, :-1] if final_sequence[0, -1] == MOTION_EOS_IDX else final_sequence\n\n# submission_data = []\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating V2 Submission\"):\n#     gen_seq = generate_motion_silent(model, str(row['gloss']), tokenizer)\n#     formatted_layers = []\n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n#         if len(layer_tokens) < 40: layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n#         elif len(layer_tokens) > 800: layer_tokens = layer_tokens[:800]\n#         formatted_layers.append(\" \".join([str(int(token)) for token in layer_tokens]))\n#     submission_data.append([row['id']] + formatted_layers)\n\n# submission_df = pd.DataFrame(submission_data, columns=['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5'])\n# submission_df.to_csv('submission_v2.csv', index=False)\n# print(\"\\n--- 🏆 V2 PIPELINE COMPLETED! 'submission_v2.csv' IS READY! ---\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:51.005358Z","iopub.execute_input":"2026-03-21T11:12:51.005839Z","iopub.status.idle":"2026-03-21T11:12:51.020804Z","shell.execute_reply.started":"2026-03-21T11:12:51.005799Z","shell.execute_reply":"2026-03-21T11:12:51.019893Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 22\n**public score: 0.4159**\n\nIncreasing from 0.39 to 0.41 is ten times harder than increasing from 0.20 to 0.30.\n\n### Why Did We Achieve This Leap?\n\n*   **BERT Unfreeze:** Previously, BERT was like just an \"English Dictionary.\" When we unlocked it, BERT learned to directly match English words with 3D sign language movements.\n*   **Gloss Dropout:** The model is no longer blindly dependent on words. Even if a word is omitted, it focuses on the overall context of the sentence and generates movements accordingly.","metadata":{}},{"cell_type":"markdown","source":"# Step 23: Ensemble Pipline 👑","metadata":{}},{"cell_type":"code","source":"# import os\n# import pandas as pd\n# import numpy as np\n# import torch\n# import torch.nn as nn\n# import torch.nn.functional as F\n# import torch.optim as optim\n# from torch.utils.data import Dataset, DataLoader\n# from torch.optim.lr_scheduler import CosineAnnealingLR\n# from transformers import AutoModel\n# from sklearn.model_selection import train_test_split\n# import copy\n# import random\n# from tqdm import tqdm\n# import warnings\n\n# # 0. SETUP \n# warnings.filterwarnings('ignore')\n# MAX_TEXT_LEN = 32\n# MOTION_PAD_IDX = 512\n# MOTION_EOS_IDX = 513\n# BATCH_SIZE = 16\n# device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n# print(f\"--- 👑 INITIATING STEP 23: THE ENSEMBLE ON {device.type.upper()} ---\")\n\n# # 1. TRAIN/VAL SPLIT\n# print(\"1. Setting up Train/Val Splits...\")\n# train_split, val_split = train_test_split(train_clean, test_size=0.10, random_state=42)\n# train_split, val_split = train_split.reset_index(drop=True), val_split.reset_index(drop=True)\n\n# # 2. DATALOADERS\n# print(\"2. Building DataLoaders with Gloss Dropout...\")\n# class SignMotionDataset(Dataset):\n#     def __init__(self, df, tokenizer, max_text_len, is_train=False):\n#         self.df, self.tokenizer, self.max_text_len, self.is_train = df, tokenizer, max_text_len, is_train\n#         self.token_cols = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n#     def __len__(self): return len(self.df)\n    \n#     def apply_gloss_dropout(self, text, drop_prob=0.1):\n#         if not self.is_train: return text\n#         words = str(text).split()\n#         if len(words) <= 1: return text\n#         kept_words = [w for w in words if random.random() > drop_prob]\n#         if len(kept_words) == 0: kept_words = [random.choice(words)]\n#         return \" \".join(kept_words)\n        \n#     def __getitem__(self, idx):\n#         row = self.df.iloc[idx]\n#         final_text = self.apply_gloss_dropout(row['gloss'])\n#         text_encoding = self.tokenizer(final_text, max_length=self.max_text_len, padding='max_length', truncation=True, return_tensors=\"pt\")\n#         input_ids = text_encoding['input_ids'].squeeze(0)\n#         attention_mask = text_encoding['attention_mask'].squeeze(0)\n#         motion_layers = [row[col] + [MOTION_EOS_IDX] for col in self.token_cols]\n#         return {\"input_ids\": input_ids, \"attention_mask\": attention_mask, \"motion_tokens\": torch.tensor(motion_layers, dtype=torch.long), \"id\": row['id']}\n\n# def collate_fn(batch):\n#     input_ids = torch.stack([item['input_ids'] for item in batch])\n#     attention_mask = torch.stack([item['attention_mask'] for item in batch])\n#     motion_tensors = [item['motion_tokens'] for item in batch]\n#     max_len = max([t.shape[1] for t in motion_tensors])\n#     padded_motions = torch.full((len(batch), 6, max_len), MOTION_PAD_IDX, dtype=torch.long)\n#     for i, tensor in enumerate(motion_tensors): padded_motions[i, :, :tensor.shape[1]] = tensor\n#     return {\"input_ids\": input_ids, \"attention_mask\": attention_mask, \"motion_tokens\": padded_motions}\n\n# train_loader = DataLoader(SignMotionDataset(train_split, tokenizer, MAX_TEXT_LEN, is_train=True), batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate_fn, drop_last=True)\n# val_loader = DataLoader(SignMotionDataset(val_split, tokenizer, MAX_TEXT_LEN, is_train=False), batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate_fn, drop_last=False)\n\n# # 3. MODEL ARCHITECTURE\n# print(\"3. Building the Model Architecture...\")\n# class SignMotionGenerator(nn.Module):\n#     def __init__(self, vocab_size=514, hidden_size=768, num_layers=4, nhead=8):\n#         super().__init__()\n#         self.text_encoder = AutoModel.from_pretrained(\"bert-base-uncased\")\n#         for name, param in self.text_encoder.named_parameters():\n#             # BERT'in son 4 katmanını eğitime açıyoruz\n#             if any(layer in name for layer in [\"encoder.layer.8\", \"encoder.layer.9\", \"encoder.layer.10\", \"encoder.layer.11\", \"pooler\"]):\n#                 param.requires_grad = True\n#             else:\n#                 param.requires_grad = False\n#         self.motion_emb = nn.Embedding(vocab_size, hidden_size)\n#         self.pos_encoder = nn.Embedding(1000, hidden_size) \n#         self.decoder = nn.TransformerDecoder(nn.TransformerDecoderLayer(d_model=hidden_size, nhead=nhead, batch_first=True), num_layers=num_layers)\n#         self.heads = nn.ModuleList([nn.Linear(hidden_size, vocab_size) for _ in range(6)])\n        \n#     def forward(self, input_ids, attention_mask, motion_tokens):\n#         memory = self.text_encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state \n#         motions_t = motion_tokens.permute(0, 2, 1) \n#         emb = self.motion_emb(motions_t).sum(dim=2)\n#         emb = emb + self.pos_encoder(torch.arange(0, emb.size(1), device=emb.device).unsqueeze(0))\n#         tgt_mask = nn.Transformer.generate_square_subsequent_mask(emb.size(1)).to(emb.device)\n#         out = self.decoder(tgt=emb, memory=memory, tgt_mask=tgt_mask) \n#         return torch.stack([head(out) for head in self.heads], dim=1)\n\n# # 4. HIERARCHICAL LOSS\n# print(\"4. Defining Hierarchical Loss...\")\n# LAYER_WEIGHTS = torch.tensor([1.0, 0.8, 0.6, 0.4, 0.2, 0.1], device=device)\n# def calculate_hierarchical_loss(logits, targets, pad_idx=512, smoothing=0.1):\n#     total_loss = 0.0\n#     for i in range(6):\n#         total_loss += F.cross_entropy(logits[:, i, :, :].reshape(-1, logits.size(-1)), targets[:, i, :].reshape(-1), ignore_index=pad_idx, label_smoothing=smoothing) * LAYER_WEIGHTS[i]\n#     return total_loss / LAYER_WEIGHTS.sum()\n\n# # 5. ENSEMBLE TRAINING\n# NUM_MODELS = 3\n# SEEDS = [42, 100, 2026]\n# trained_models = []\n\n# for m_idx in range(NUM_MODELS):\n#     current_seed = SEEDS[m_idx]\n#     print(f\"\\n{'='*40}\\n🚀 TRAINING MODEL {m_idx+1}/{NUM_MODELS} (Seed: {current_seed})\\n{'='*40}\")\n    \n#     torch.manual_seed(current_seed)\n#     random.seed(current_seed)\n#     np.random.seed(current_seed)\n    \n#     model = SignMotionGenerator().to(device)\n#     bert_params = [p for n, p in model.named_parameters() if \"text_encoder\" in n and p.requires_grad]\n#     decoder_params = [p for n, p in model.named_parameters() if \"text_encoder\" not in n and p.requires_grad]\n#     optimizer = optim.AdamW([{'params': bert_params, 'lr': 2e-5}, {'params': decoder_params, 'lr': 5e-4}])\n#     scheduler = CosineAnnealingLR(optimizer, T_max=15, eta_min=1e-6)\n    \n#     best_val_loss = float('inf')\n#     best_wts = copy.deepcopy(model.state_dict())\n    \n#     for epoch in range(15): \n#         model.train()\n#         train_loss = 0.0\n#         for batch in tqdm(train_loader, desc=f\"M{m_idx+1} - Epoch {epoch+1} Train\", leave=False):\n#             input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#             optimizer.zero_grad()\n#             logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n#             loss = calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1)\n#             loss.backward()\n#             torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n#             optimizer.step()\n#             train_loss += loss.item()\n            \n#         model.eval()\n#         val_loss = 0.0\n#         with torch.no_grad():\n#             for batch in val_loader:\n#                 input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n#                 logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n#                 val_loss += calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1).item()\n                \n#         avg_train_loss, avg_val_loss = train_loss / len(train_loader), val_loss / len(val_loader)\n#         print(f\"Epoch {epoch+1} | Train: {avg_train_loss:.4f} | Val: {avg_val_loss:.4f}\")\n#         scheduler.step()\n        \n#         if avg_val_loss < best_val_loss:\n#             best_val_loss = avg_val_loss\n#             best_wts = copy.deepcopy(model.state_dict())\n            \n#     # En iyi modeli listeye ekle\n#     model.load_state_dict(best_wts)\n#     model.eval()\n#     trained_models.append(model)\n#     print(f\"Model {m_idx+1} completed! Best Val Loss: {best_val_loss:.4f}\")\n\n# # 6. ENSEMBLE INFERENCE (Ortak Akıl Tahmini)\n# print(\"\\n--- 🧠 LAUNCHING COLLECTIVE INTELLIGENCE (ENSEMBLE INFERENCE) ---\")\n# def generate_motion_ensemble(models, text, tokenizer, max_gen_len=150):\n#     text_encoding = tokenizer(text, max_length=MAX_TEXT_LEN, padding='max_length', truncation=True, return_tensors=\"pt\")\n#     input_ids, attention_mask = text_encoding['input_ids'].to(device), text_encoding['attention_mask'].to(device)\n#     generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n#     with torch.no_grad():\n#         for step in range(max_gen_len):\n#             ensemble_logits = 0\n#             for model in models:\n#                 logits = model(input_ids, attention_mask, generated_seq)\n#                 ensemble_logits += logits[:, :, -1, :] \n            \n#             avg_logits = ensemble_logits / len(models) # Modellerin tahminlerinin ortalaması\n#             next_tokens = torch.argmax(avg_logits, dim=-1).unsqueeze(-1)\n#             generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n#             if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX: break\n                \n#     final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n#     return final_sequence[:, :-1] if final_sequence[0, -1] == MOTION_EOS_IDX else final_sequence\n\n# submission_data = []\n# for _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating Ensemble Submission\"):\n#     gen_seq = generate_motion_ensemble(trained_models, str(row['gloss']), tokenizer)\n#     formatted_layers = []\n#     for layer_idx in range(6):\n#         layer_tokens = list(gen_seq[layer_idx, :])\n#         if len(layer_tokens) < 40: layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n#         elif len(layer_tokens) > 800: layer_tokens = layer_tokens[:800]\n#         formatted_layers.append(\" \".join([str(int(token)) for token in layer_tokens]))\n#     submission_data.append([row['id']] + formatted_layers)\n\n# # Kaydet\n# submission_df = pd.DataFrame(submission_data, columns=['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5'])\n# submission_df.to_csv('submission_ensemble_gold.csv', index=False)\n# print(\"\\n--- 🏆 ENSEMBLE PIPELINE COMPLETED! 'submission_ensemble_gold.csv' IS READY! ---\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:51.021957Z","iopub.execute_input":"2026-03-21T11:12:51.022270Z","iopub.status.idle":"2026-03-21T11:12:51.038014Z","shell.execute_reply.started":"2026-03-21T11:12:51.022246Z","shell.execute_reply":"2026-03-21T11:12:51.037266Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Insight 23\n**public score: 0.3904**\n\n### Auto-Regressive (AR) Ensemble Collapse\n\nIn classification or regression problems, averaging three models usually works. However, our model is autoregressive—it generates a time series.\n\nThink of it this way: We have three different drivers (Model 1, 2, and 3) driving the same car. The road splits ahead.\n\n- **Model 1** says: \"Turn right.\"\n- **Model 2** says: \"Turn left.\"\n\nBecause we are averaging the logits (probabilities), we hold the steering wheel \"straight,\" and the car crashes into the wall!\n\nIf one of the models had been left alone, it would have finished the route perfectly. However, by averaging each other's probabilities at every step, they disrupted the original sequence flow. Additionally, the fact that all their loss values remained around 2.24 indicates that instead of learning different things, the models all got stuck in the exact same local minima.","metadata":{}},{"cell_type":"markdown","source":"# 24: SWA + Stochastic Decoding 🦚","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.optim.lr_scheduler import ReduceLROnPlateau\nfrom transformers import AutoModel\nfrom sklearn.model_selection import train_test_split\nimport copy\nimport random\nfrom tqdm import tqdm\nimport warnings\n\n# 0. SETUP\nwarnings.filterwarnings('ignore')\nMAX_TEXT_LEN = 32\nMOTION_PAD_IDX = 512\nMOTION_EOS_IDX = 513\nBATCH_SIZE = 8\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"--- ⚔️ INITIATING MODE V3 (SWA SUPER MODEL) ON {device.type.upper()} ---\")\n\n# 1. TRAIN/VAL SPLIT \nprint(\"1. Setting up Train/Val Splits...\")\ntrain_split, val_split = train_test_split(train_clean, test_size=0.10, random_state=42)\ntrain_split, val_split = train_split.reset_index(drop=True), val_split.reset_index(drop=True)\n\n# 2. DATALOADERS (Gloss Dropout Aktif)\nclass SignMotionDataset(Dataset):\n    def __init__(self, df, tokenizer, max_text_len, is_train=False):\n        self.df, self.tokenizer, self.max_text_len, self.is_train = df, tokenizer, max_text_len, is_train\n        self.token_cols = ['base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']\n    def __len__(self): return len(self.df)\n    \n    def apply_gloss_dropout(self, text, drop_prob=0.1):\n        if not self.is_train: return text\n        words = str(text).split()\n        if len(words) <= 1: return text\n        kept_words = [w for w in words if random.random() > drop_prob]\n        if len(kept_words) == 0: kept_words = [random.choice(words)]\n        return \" \".join(kept_words)\n        \n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        final_text = self.apply_gloss_dropout(row['gloss'])\n        text_encoding = self.tokenizer(final_text, max_length=self.max_text_len, padding='max_length', truncation=True, return_tensors=\"pt\")\n        return {\"input_ids\": text_encoding['input_ids'].squeeze(0), \n                \"attention_mask\": text_encoding['attention_mask'].squeeze(0), \n                \"motion_tokens\": torch.tensor([row[col] + [MOTION_EOS_IDX] for col in self.token_cols], dtype=torch.long), \n                \"id\": row['id']}\n\ndef collate_fn(batch):\n    input_ids = torch.stack([item['input_ids'] for item in batch])\n    attention_mask = torch.stack([item['attention_mask'] for item in batch])\n    motion_tensors = [item['motion_tokens'] for item in batch]\n    max_len = max([t.shape[1] for t in motion_tensors])\n    padded_motions = torch.full((len(batch), 6, max_len), MOTION_PAD_IDX, dtype=torch.long)\n    for i, tensor in enumerate(motion_tensors): padded_motions[i, :, :tensor.shape[1]] = tensor\n    return {\"input_ids\": input_ids, \"attention_mask\": attention_mask, \"motion_tokens\": padded_motions}\n\ntrain_loader = DataLoader(SignMotionDataset(train_split, tokenizer, MAX_TEXT_LEN, is_train=True), batch_size=BATCH_SIZE, shuffle=True, collate_fn=collate_fn, drop_last=True)\nval_loader = DataLoader(SignMotionDataset(val_split, tokenizer, MAX_TEXT_LEN, is_train=False), batch_size=BATCH_SIZE, shuffle=False, collate_fn=collate_fn, drop_last=False)\n\n# 3. MODEL ARCHITECTURE (Unfrozen BERT)\nclass SignMotionGenerator(nn.Module):\n    def __init__(self, vocab_size=514, hidden_size=768, num_layers=4, nhead=8):\n        super().__init__()\n        self.text_encoder = AutoModel.from_pretrained(\"bert-base-uncased\")\n        for name, param in self.text_encoder.named_parameters():\n            if any(layer in name for layer in [\"encoder.layer.8\", \"encoder.layer.9\", \"encoder.layer.10\", \"encoder.layer.11\", \"pooler\"]):\n                param.requires_grad = True\n            else: param.requires_grad = False\n        self.motion_emb = nn.Embedding(vocab_size, hidden_size)\n        self.pos_encoder = nn.Embedding(1000, hidden_size) \n        self.decoder = nn.TransformerDecoder(nn.TransformerDecoderLayer(d_model=hidden_size, nhead=nhead, batch_first=True), num_layers=num_layers)\n        self.heads = nn.ModuleList([nn.Linear(hidden_size, vocab_size) for _ in range(6)])\n        \n    def forward(self, input_ids, attention_mask, motion_tokens):\n        memory = self.text_encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state \n        motions_t = motion_tokens.permute(0, 2, 1) \n        emb = self.motion_emb(motions_t).sum(dim=2) + self.pos_encoder(torch.arange(0, motion_tokens.size(2), device=motion_tokens.device).unsqueeze(0))\n        tgt_mask = nn.Transformer.generate_square_subsequent_mask(emb.size(1)).to(emb.device)\n        return torch.stack([head(self.decoder(tgt=emb, memory=memory, tgt_mask=tgt_mask)) for head in self.heads], dim=1)\n\n# 4. HIERARCHICAL LOSS\nLAYER_WEIGHTS = torch.tensor([1.0, 0.8, 0.6, 0.4, 0.2, 0.1], device=device)\ndef calculate_hierarchical_loss(logits, targets, pad_idx=512, smoothing=0.1):\n    total_loss = sum(F.cross_entropy(logits[:, i, :, :].reshape(-1, logits.size(-1)), targets[:, i, :].reshape(-1), ignore_index=pad_idx, label_smoothing=smoothing) * LAYER_WEIGHTS[i] for i in range(6))\n    return total_loss / LAYER_WEIGHTS.sum()\n\n# 5. SWA TRAINING LOOP\nprint(\"5. Launching SWA Training Loop...\")\nmodel = SignMotionGenerator().to(device)\noptimizer = optim.AdamW([\n    {'params': [p for n, p in model.named_parameters() if \"text_encoder\" in n and p.requires_grad], 'lr': 2e-5},\n    {'params': [p for n, p in model.named_parameters() if \"text_encoder\" not in n and p.requires_grad], 'lr': 5e-4}\n])\nscheduler = ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2)\n\nNUM_EPOCHS = 20\ncheckpoint_weights = [] \n\nfor epoch in range(NUM_EPOCHS): \n    print(f\"\\nEpoch {epoch+1}/{NUM_EPOCHS} [Mode V3]\")\n    model.train()\n    train_loss = 0.0\n    for batch in tqdm(train_loader, desc=\"Training\", leave=False):\n        input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n        optimizer.zero_grad()\n        logits = model(input_ids, attention_mask, motion_tokens[:, :, :-1])\n        loss = calculate_hierarchical_loss(logits, motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1)\n        loss.backward()\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n        optimizer.step()\n        train_loss += loss.item()\n        \n    model.eval()\n    val_loss = 0.0\n    with torch.no_grad():\n        for batch in val_loader:\n            input_ids, attention_mask, motion_tokens = batch['input_ids'].to(device), batch['attention_mask'].to(device), batch['motion_tokens'].to(device)\n            val_loss += calculate_hierarchical_loss(model(input_ids, attention_mask, motion_tokens[:, :, :-1]), motion_tokens[:, :, 1:], pad_idx=MOTION_PAD_IDX, smoothing=0.1).item()\n            \n    print(f\"Train Loss: {train_loss/len(train_loader):.4f} | Val Loss: {val_loss/len(val_loader):.4f}\")\n    scheduler.step(val_loss/len(val_loader))\n    \n    # GRANDMASTER SWA\n    if epoch >= NUM_EPOCHS - 5:\n        checkpoint_weights.append(copy.deepcopy(model.state_dict()))\n        print(f\"💾 Checkpoint saved for SWA (Model {len(checkpoint_weights)}/5)\")\n\n# 6. CREATE THE SUPER MODEL\nprint(\"\\n--- 🧠 CREATING SWA SUPER MODEL ---\")\nswa_state_dict = checkpoint_weights[0]\nfor key in swa_state_dict.keys():\n    for i in range(1, len(checkpoint_weights)):\n        swa_state_dict[key] += checkpoint_weights[i][key]\n    swa_state_dict[key] = swa_state_dict[key] / float(len(checkpoint_weights))\n\nmodel.load_state_dict(swa_state_dict)\ntorch.save(model.state_dict(), \"sign_motion_swa_super.pth\")\nprint(\"SWA Super Model created and loaded!\")\n\n# 7. GENERATION & SUBMISSION\nprint(\"\\n7. Generating Final Submission...\")\nmodel.eval()\ndef generate_motion_silent(model, text, tokenizer, max_gen_len=150):\n    text_encoding = tokenizer(text, max_length=MAX_TEXT_LEN, padding='max_length', truncation=True, return_tensors=\"pt\")\n    input_ids, attention_mask = text_encoding['input_ids'].to(device), text_encoding['attention_mask'].to(device)\n    generated_seq = torch.full((1, 6, 1), MOTION_PAD_IDX, dtype=torch.long, device=device)\n    \n    with torch.no_grad():\n        for step in range(max_gen_len):\n            logits = model(input_ids, attention_mask, generated_seq)\n            next_tokens = torch.argmax(logits[:, :, -1, :], dim=-1).unsqueeze(-1)\n            generated_seq = torch.cat([generated_seq, next_tokens], dim=-1)\n            if next_tokens[0, 0, 0].item() == MOTION_EOS_IDX: break\n                \n    final_sequence = generated_seq[0, :, 1:].cpu().numpy()\n    return final_sequence[:, :-1] if final_sequence[0, -1] == MOTION_EOS_IDX else final_sequence\n\nsubmission_data = []\nfor _, row in tqdm(test_df.iterrows(), total=len(test_df), desc=\"Generating SWA Submission\"):\n    gen_seq = generate_motion_silent(model, str(row['gloss']), tokenizer)\n    formatted_layers = []\n    for layer_idx in range(6):\n        layer_tokens = list(gen_seq[layer_idx, :])\n        if len(layer_tokens) < 40: layer_tokens.extend([MOTION_PAD_IDX] * (40 - len(layer_tokens)))\n        elif len(layer_tokens) > 800: layer_tokens = layer_tokens[:800]\n        formatted_layers.append(\" \".join([str(int(token)) for token in layer_tokens]))\n    submission_data.append([row['id']] + formatted_layers)\n\npd.DataFrame(submission_data, columns=['id', 'base_tokens', 'residual_1', 'residual_2', 'residual_3', 'residual_4', 'residual_5']).to_csv('submission_swa.csv', index=False)\nprint(\"\\n--- 🏆 V3 PIPELINE COMPLETED! 'submission_swa.csv' IS READY! ---\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-21T11:12:51.039106Z","iopub.execute_input":"2026-03-21T11:12:51.039469Z"}},"outputs":[],"execution_count":null}]}