{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[]},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":117416,"databundleVersionId":14067006,"sourceType":"competition"},{"sourceId":13417376,"sourceType":"datasetVersion","datasetId":8515827}],"dockerImageVersionId":31154,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":2582.745677,"end_time":"2025-10-08T16:39:31.900291","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2025-10-08T15:56:29.154614","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"e2e7cbb8-1f49-4428-80a0-737940145135","cell_type":"markdown","source":"# AFML Hackathon Part 2\n\nHey guys, congratulations on making it this far. We hope you've had some fun with part 1 of this hackathon. Part 2 of this hackathon relies on the output of model  `test-part1.csv`. What you have denoised are actually weights of another model that you will be using for this part.\n\nHere's some code to get you started for part 2. You will have to load the model from the csv, understand the architecture of the model given here and train it to meet the objective of this challenge.","metadata":{"id":"e2e7cbb8-1f49-4428-80a0-737940145135"}},{"id":"6736392d-b625-430f-9294-84bc2398387a","cell_type":"markdown","source":"## The model\n\nThis is an LSTM based encoder-decoder model, an architecture that was used for translation purposes before the Transoformer came along. The input and output vocab sizes are fixed based on your datasets and so are the embedding sizes and hidden dimensions.","metadata":{"id":"6736392d-b625-430f-9294-84bc2398387a"}},{"id":"WgA6nBe5pOKf","cell_type":"markdown","source":"Before you begin, make sure you have:\n- Your denoised weights: `<TeamID>_denoised_test_part1.csv` (from Part 1)\n- Your training data: `train_part2.csv` (in your team folder)\n- Your test phrases: `test_part2.csv` (in your team folder)\n\nUpdate all file paths in this notebook to match your team ID!","metadata":{"id":"WgA6nBe5pOKf"}},{"id":"8420b82f-f638-4729-a54d-ddfd55695b1a","cell_type":"code","source":"import gc\nimport os\nos.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'expandable_segments:True'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T18:02:16.431521Z","iopub.execute_input":"2025-10-17T18:02:16.432386Z","iopub.status.idle":"2025-10-17T18:02:16.436198Z","shell.execute_reply.started":"2025-10-17T18:02:16.432360Z","shell.execute_reply":"2025-10-17T18:02:16.435270Z"}},"outputs":[],"execution_count":null},{"id":"47ff5c51-7313-49ec-8a84-053d52776e9d","cell_type":"code","source":"import torch\n# Clear memory\ntorch.cuda.empty_cache()\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T18:02:16.437311Z","iopub.execute_input":"2025-10-17T18:02:16.437663Z","iopub.status.idle":"2025-10-17T18:02:21.000076Z","shell.execute_reply.started":"2025-10-17T18:02:16.437634Z","shell.execute_reply":"2025-10-17T18:02:20.999334Z"}},"outputs":[],"execution_count":null},{"id":"9ea5586e-4c9f-4b51-93bf-b0155e52f94d","cell_type":"code","source":"import torch\nimport torch.nn as nn\n\nclass CharLSTMTranslator(nn.Module):\n    def __init__(self, input_vocab_size, output_vocab_size, emb_size=64, hidden_size=128, num_layers=1, max_len=512):\n        super().__init__()\n        self.src_embedding = nn.Embedding(input_vocab_size, emb_size, padding_idx=0)\n        self.tgt_embedding = nn.Embedding(output_vocab_size, emb_size, padding_idx=0)\n\n        self.pos_embedding = nn.Embedding(max_len, emb_size)\n\n        self.encoder = nn.LSTM(emb_size, hidden_size, num_layers, batch_first=True)\n        self.decoder = nn.LSTM(emb_size, hidden_size, num_layers, batch_first=True)\n        self.fc = nn.Linear(hidden_size, output_vocab_size)\n\n    def forward(self, src, tgt):\n        batch_size, seq_len = src.size()\n\n        num_pos = self.pos_embedding.num_embeddings\n        pos_idx = torch.arange(seq_len, device=src.device) % num_pos\n        pos_idx = pos_idx.unsqueeze(0).repeat(batch_size, 1)\n        pos_embedded = self.pos_embedding(pos_idx)\n        embedded_src = self.src_embedding(src) + pos_embedded\n\n        _, (hidden, cell) = self.encoder(embedded_src)\n\n        embedded_tgt = self.tgt_embedding(tgt)\n        outputs, _ = self.decoder(embedded_tgt, (hidden, cell))\n\n        logits = self.fc(outputs)\n        return logits","metadata":{"execution":{"iopub.status.busy":"2025-10-17T18:02:21.000813Z","iopub.execute_input":"2025-10-17T18:02:21.001151Z","iopub.status.idle":"2025-10-17T18:02:21.007746Z","shell.execute_reply.started":"2025-10-17T18:02:21.001134Z","shell.execute_reply":"2025-10-17T18:02:21.007055Z"},"id":"9ea5586e-4c9f-4b51-93bf-b0155e52f94d","trusted":true},"outputs":[],"execution_count":null},{"id":"8bff82ed-9864-4613-b713-214918e6fcac","cell_type":"markdown","source":"## Load the csv","metadata":{"id":"8bff82ed-9864-4613-b713-214918e6fcac"}},{"id":"5jwiubkVpTgK","cell_type":"markdown","source":"Make sure to update the file path above to your denoised output:\n- File should be named: `<TeamID>_denoised_test_part1.csv`\n- Example: If your team is 55_AFMLTEAM, use `\"55_denoised_test_part1.csv\"`","metadata":{"id":"5jwiubkVpTgK"}},{"id":"ed738e97-6280-40d5-a032-f07985cd8779","cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(\"/kaggle/input/05-denoised-test-part1/05_denoised_test_part1.csv\").to_numpy()","metadata":{"execution":{"iopub.status.busy":"2025-10-17T18:02:21.009410Z","iopub.execute_input":"2025-10-17T18:02:21.009659Z","iopub.status.idle":"2025-10-17T18:02:21.474690Z","shell.execute_reply.started":"2025-10-17T18:02:21.009644Z","shell.execute_reply":"2025-10-17T18:02:21.474144Z"},"id":"ed738e97-6280-40d5-a032-f07985cd8779","trusted":true},"outputs":[],"execution_count":null},{"id":"2fe8440e-220b-4a28-8583-0140bb0b023a","cell_type":"markdown","source":"## Load model weights from the csv\n\nDon't worry, you don't have to do anything here","metadata":{"id":"2fe8440e-220b-4a28-8583-0140bb0b023a"}},{"id":"d4a89b88-1343-4a72-aea2-226501b041d1","cell_type":"code","source":"def load_model_from_matrix(model, weights_matrix, original_len):\n    weights_matrix = torch.tensor(weights_matrix)\n    flat_weights = weights_matrix.reshape(-1)[:original_len]\n\n    offset = 0\n    for p in model.parameters():\n        numel = p.numel()\n        new_data = flat_weights[offset : offset + numel].view_as(p)\n        p.data.copy_(new_data)\n        offset += numel\n\n    print(\"Model weights successfully restored!\")\n    return model\n\nmodel = CharLSTMTranslator(input_vocab_size=73, output_vocab_size=96)\nmodel = load_model_from_matrix(model, df, 254624)","metadata":{"id":"d4a89b88-1343-4a72-aea2-226501b041d1","trusted":true,"execution":{"iopub.status.busy":"2025-10-17T18:02:21.475407Z","iopub.execute_input":"2025-10-17T18:02:21.475588Z","iopub.status.idle":"2025-10-17T18:02:21.574084Z","shell.execute_reply.started":"2025-10-17T18:02:21.475574Z","shell.execute_reply":"2025-10-17T18:02:21.573398Z"}},"outputs":[],"execution_count":null},{"id":"f99492b0-267f-4763-af64-ef66e5cd466e","cell_type":"markdown","source":"## Working with the training data\n\nSince the model has been pre-trained a little bit, it's best to use the vocabulary and padding scheme that has been used in the pretraining. Go through this code to understand how the data is being loaded. You will still have to do some work of your own to actually structure the data for the model.","metadata":{"id":"f99492b0-267f-4763-af64-ef66e5cd466e"}},{"id":"af829385","cell_type":"code","source":"import pandas as pd\nfrom collections import Counter\nimport torch\nfrom torch.nn.utils.rnn import pad_sequence\nfrom torch.utils.data import Dataset, DataLoader\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\n\n# Load training data (fixed path)\ndf = pd.read_csv(\"/kaggle/input/afml-assignment-1-ec-campus/Team 5/train_part2.csv\")\nencoded_texts = df['encoded_text'].tolist()\nenglish_texts = df['text'].tolist()\n\n# Character sets\nall_encoded_chars = set(''.join(encoded_texts))\nall_english_chars = set(''.join(english_texts))\n\n# Special tokens\nSPECIAL_TOKENS = {\n    'PAD': 0,\n    'SOS': 1,\n    'EOS': 2,\n    'UNK': 3,\n}\n\n# Deterministic ordering for reproducibility\nencoded_chars_sorted = sorted(list(all_encoded_chars))\nenglish_chars_sorted = sorted(list(all_english_chars))\n\n# Build vocabs with space for specials at front\nencoded_vocab = {c: i + len(SPECIAL_TOKENS) for i, c in enumerate(encoded_chars_sorted)}\nenglish_vocab = {c: i + len(SPECIAL_TOKENS) for i, c in enumerate(english_chars_sorted)}\n\n# Add specials\nfor tok, idx in SPECIAL_TOKENS.items():\n    encoded_vocab[f'<{tok}>'] = idx\n    english_vocab[f'<{tok}>'] = idx\n\n# Reverse mapping\nrev_english_vocab = {idx: ch for ch, idx in english_vocab.items()}\n\n# Helpers\ndef text_to_seq_single(text, vocab):\n    return [vocab.get(ch, SPECIAL_TOKENS['UNK']) for ch in text]\n\ndef with_sos_eos(seq):\n    return [SPECIAL_TOKENS['SOS']] + seq + [SPECIAL_TOKENS['EOS']]\n\n# Encode sequences\nencoded_seqs = [text_to_seq_single(t, encoded_vocab) for t in encoded_texts]\nenglish_seqs = [text_to_seq_single(t, english_vocab) for t in english_texts]\n\n# Target input/output for teacher forcing\ntgt_input_seqs = [with_sos_eos(s)[:-1] for s in english_seqs]\ntgt_output_seqs = [with_sos_eos(s)[1:] for s in english_seqs]\n\ninput_vocab_size = max(encoded_vocab.values()) + 1\noutput_vocab_size = max(english_vocab.values()) + 1\nprint('Input vocab size:', input_vocab_size)\nprint('Output vocab size:', output_vocab_size)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2025-10-17T18:02:21.574876Z","iopub.execute_input":"2025-10-17T18:02:21.575199Z","iopub.status.idle":"2025-10-17T18:03:23.785564Z","shell.execute_reply.started":"2025-10-17T18:02:21.575180Z","shell.execute_reply":"2025-10-17T18:03:23.784899Z"},"id":"af829385","papermill":{"duration":50.635053,"end_time":"2025-10-08T15:57:24.973176","exception":false,"start_time":"2025-10-08T15:56:34.338123","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"1fb7b700","cell_type":"code","source":"len(all_encoded_chars), len(all_english_chars)","metadata":{"execution":{"iopub.status.busy":"2025-10-17T18:03:23.786367Z","iopub.execute_input":"2025-10-17T18:03:23.786663Z","iopub.status.idle":"2025-10-17T18:03:23.791667Z","shell.execute_reply.started":"2025-10-17T18:03:23.786646Z","shell.execute_reply":"2025-10-17T18:03:23.790916Z"},"id":"1fb7b700","outputId":"b2a572be-07b5-4026-b3fa-8c10dc91af60","papermill":{"duration":0.009686,"end_time":"2025-10-08T15:57:24.995293","exception":false,"start_time":"2025-10-08T15:57:24.985607","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"XikpFf4CqJEb","cell_type":"markdown","source":"**Instructions for use:**\n1. Keep all existing markdown cells in the notebook as they are\n2. Add these new markdown cells in the suggested positions\n","metadata":{"id":"XikpFf4CqJEb"}},{"id":"mzZKsv-3qJw6","cell_type":"code","source":"# Dataset, Training, and Inference\nimport math\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n\n# Ensure model is on device\nmodel = model.to(device)\n\n# Dataset and DataLoader\nclass Seq2SeqDataset(Dataset):\n    def __init__(self, src_seqs, tgt_in_seqs, tgt_out_seqs):\n        self.src = src_seqs\n        self.tgt_in = tgt_in_seqs\n        self.tgt_out = tgt_out_seqs\n    def __len__(self):\n        return len(self.src)\n    def __getitem__(self, idx):\n        return (\n            torch.tensor(self.src[idx], dtype=torch.long),\n            torch.tensor(self.tgt_in[idx], dtype=torch.long),\n            torch.tensor(self.tgt_out[idx], dtype=torch.long),\n        )\n\nPAD_IDX = SPECIAL_TOKENS['PAD']\n\ndef collate_fn(batch):\n    src, tgt_in, tgt_out = zip(*batch)\n    src_pad = pad_sequence(src, batch_first=True, padding_value=PAD_IDX)\n    tgt_in_pad = pad_sequence(tgt_in, batch_first=True, padding_value=PAD_IDX)\n    tgt_out_pad = pad_sequence(tgt_out, batch_first=True, padding_value=PAD_IDX)\n    return src_pad, tgt_in_pad, tgt_out_pad\n\ntrain_ds = Seq2SeqDataset(encoded_seqs, tgt_input_seqs, tgt_output_seqs)\ntrain_loader = DataLoader(train_ds, batch_size=16, shuffle=True, collate_fn=collate_fn)\n\n# Training setup\ncriterion = nn.CrossEntropyLoss(ignore_index=PAD_IDX)\noptimizer = optim.Adam(model.parameters(), lr=1e-3)\n\ndef train_epoch():\n    model.train()\n    total_loss = 0.0\n    total_tokens = 0\n    for src, tgt_in, tgt_out in train_loader:\n        src = src.to(device)\n        tgt_in = tgt_in.to(device)\n        tgt_out = tgt_out.to(device)\n\n        optimizer.zero_grad()\n        logits = model(src, tgt_in)  # [B, T, V]\n        B, T, V = logits.size()\n        loss = criterion(logits.reshape(B*T, V), tgt_out.reshape(B*T))\n        loss.backward()\n        nn.utils.clip_grad_norm_(model.parameters(), 1.0)\n        optimizer.step()\n\n        with torch.no_grad():\n            mask = (tgt_out != PAD_IDX)\n            num = mask.sum().item()\n            total_tokens += num\n            total_loss += loss.item() * (B*T)\n\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n    ppl = math.exp(total_loss / max(1, total_tokens))\n    return total_loss / max(1, total_tokens), ppl\n\n# Run a few epochs (adjust as needed)\nEPOCHS = 3\nfor epoch in range(1, EPOCHS+1):\n    avg_loss, ppl = train_epoch()\n    print(f\"Epoch {epoch}: avg_token_loss={avg_loss:.4f} ppl={ppl:.2f}\")\n\n# Greedy decoding\nEOS_IDX = SPECIAL_TOKENS['EOS']\nSOS_IDX = SPECIAL_TOKENS['SOS']\n\ndef greedy_decode_single(src_seq, max_len=256):\n    model.eval()\n    with torch.no_grad():\n        src_tensor = torch.tensor(src_seq, dtype=torch.long, device=device).unsqueeze(0)\n        B, S = src_tensor.size()\n        # Positional embedding for encoder\n        num_pos = model.pos_embedding.num_embeddings\n        pos_idx = (torch.arange(S, device=device) % num_pos).unsqueeze(0).repeat(B, 1)\n        pos_embedded = model.pos_embedding(pos_idx)\n        embedded_src = model.src_embedding(src_tensor) + pos_embedded\n        _, (hidden, cell) = model.encoder(embedded_src)\n\n        # Start with SOS\n        cur = torch.tensor([[SOS_IDX]], dtype=torch.long, device=device)\n        outputs = []\n        for _ in range(max_len):\n            emb = model.tgt_embedding(cur)  # [1,1,E]\n            out, (hidden, cell) = model.decoder(emb, (hidden, cell))\n            logit = model.fc(out[:, -1, :])  # [1, V]\n            next_token = torch.argmax(logit, dim=-1)  # [1]\n            idx = next_token.item()\n            if idx == EOS_IDX:\n                break\n            outputs.append(idx)\n            cur = next_token.unsqueeze(0)  # shape [1,1]\n        # Map to string\n        chars = [rev_english_vocab.get(i, '?') for i in outputs]\n        return ''.join(chars)\n\n# Translate test set and save\ntry:\n    with open('test_part2.txt', 'r', encoding='utf-8') as f:\n        test_lines = [line.strip('\\n') for line in f.readlines() if len(line.strip()) > 0]\nexcept FileNotFoundError:\n    # Fallback name if needed\n    with open('test_part2.csv', 'r', encoding='utf-8') as f:\n        import csv\n        reader = csv.DictReader(f)\n        test_lines = [row['encoded_text'] for row in reader]\n\nencoded_test = [[encoded_vocab.get(ch, SPECIAL_TOKENS['UNK']) for ch in line] for line in test_lines]\n\ntranslations = [greedy_decode_single(seq) for seq in encoded_test]\n\nout_path = 'translations_part2.txt'\nwith open(out_path, 'w', encoding='utf-8') as f:\n    for line, trans in zip(test_lines, translations):\n        f.write(f\"{line}\\t{trans}\\n\")\nprint(f\"Saved translations to {out_path}\")\n","metadata":{"id":"mzZKsv-3qJw6","trusted":true,"execution":{"iopub.status.busy":"2025-10-17T18:03:23.792331Z","iopub.execute_input":"2025-10-17T18:03:23.792502Z","iopub.status.idle":"2025-10-17T18:03:31.654839Z","shell.execute_reply.started":"2025-10-17T18:03:23.792488Z","shell.execute_reply":"2025-10-17T18:03:31.653784Z"}},"outputs":[],"execution_count":null},{"id":"26fb7dd2-32ac-490e-8707-d8b8cee51d29","cell_type":"code","source":"# Dataset, Training, and Inference with Multi-GPU Support (Memory Optimized)\nimport math\n\n# Clear any existing memory\nif torch.cuda.is_available():\n    torch.cuda.empty_cache()\n    gc.collect()\n\n# Check available GPUs\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nn_gpus = torch.cuda.device_count() if torch.cuda.is_available() else 0\n\nprint(f\"{'='*60}\")\nprint(f\"GPU Configuration:\")\nprint(f\"{'='*60}\")\nprint(f\"Available GPUs: {n_gpus}\")\nif n_gpus > 0:\n    for i in range(n_gpus):\n        print(f\"  GPU {i}: {torch.cuda.get_device_name(i)}\")\n        print(f\"    Memory: {torch.cuda.get_device_properties(i).total_memory / 1e9:.2f} GB\")\nprint(f\"{'='*60}\\n\")\n\n# Move model to device\nmodel = model.to(device)\n\n# Use conservative batch size due to long sequences\nif n_gpus > 1:\n    print(f\"🚀 Using DataParallel with {n_gpus} GPUs\")\n    model = nn.DataParallel(model)\n    # Very conservative batch size due to long sequences\n    batch_size = 8 * n_gpus  # 8 per GPU\n    print(f\"Batch size: {batch_size} (per-GPU: {batch_size // n_gpus})\")\nelse:\n    batch_size = 4\n    print(f\"Using single GPU with batch size: {batch_size}\")\n\nprint()\n\n# Dataset and DataLoader\nclass Seq2SeqDataset(Dataset):\n    def __init__(self, src_seqs, tgt_in_seqs, tgt_out_seqs):\n        self.src = src_seqs\n        self.tgt_in = tgt_in_seqs\n        self.tgt_out = tgt_out_seqs\n    \n    def __len__(self):\n        return len(self.src)\n    \n    def __getitem__(self, idx):\n        return (\n            torch.tensor(self.src[idx], dtype=torch.long),\n            torch.tensor(self.tgt_in[idx], dtype=torch.long),\n            torch.tensor(self.tgt_out[idx], dtype=torch.long),\n        )\n\nPAD_IDX = SPECIAL_TOKENS['PAD']\n\ndef collate_fn(batch):\n    src, tgt_in, tgt_out = zip(*batch)\n    src_pad = pad_sequence(src, batch_first=True, padding_value=PAD_IDX)\n    tgt_in_pad = pad_sequence(tgt_in, batch_first=True, padding_value=PAD_IDX)\n    tgt_out_pad = pad_sequence(tgt_out, batch_first=True, padding_value=PAD_IDX)\n    return src_pad, tgt_in_pad, tgt_out_pad\n\n# Create dataset and dataloader\ntrain_ds = Seq2SeqDataset(encoded_seqs, tgt_input_seqs, tgt_output_seqs)\ntrain_loader = DataLoader(\n    train_ds, \n    batch_size=batch_size, \n    shuffle=True, \n    collate_fn=collate_fn,\n    num_workers=2,\n    pin_memory=True\n)\n\nprint(f\"Dataset size: {len(train_ds)} samples\")\nprint(f\"Number of batches: {len(train_loader)}\\n\")\n\n# Training setup with gradient accumulation\ncriterion = nn.CrossEntropyLoss(ignore_index=PAD_IDX)\noptimizer = optim.Adam(model.parameters(), lr=1e-3)\n\n# Gradient accumulation to simulate larger batch size\naccumulation_steps = 3  # Effective batch size = batch_size * accumulation_steps\n\ndef train_epoch(epoch_num):\n    model.train()\n    total_loss = 0.0\n    total_tokens = 0\n    batch_count = 0\n    \n    optimizer.zero_grad()  # Initialize gradients\n    \n    for batch_idx, (src, tgt_in, tgt_out) in enumerate(train_loader):\n        # Move to device\n        src = src.to(device, non_blocking=True)\n        tgt_in = tgt_in.to(device, non_blocking=True)\n        tgt_out = tgt_out.to(device, non_blocking=True)\n\n        # Forward pass\n        logits = model(src, tgt_in)\n        B, T, V = logits.size()\n        \n        # Calculate loss and normalize by accumulation steps\n        loss = criterion(logits.reshape(B*T, V), tgt_out.reshape(B*T))\n        loss = loss / accumulation_steps\n        \n        # Backward pass\n        loss.backward()\n        \n        # Update weights every accumulation_steps\n        if (batch_idx + 1) % accumulation_steps == 0:\n            nn.utils.clip_grad_norm_(model.parameters(), 1.0)\n            optimizer.step()\n            optimizer.zero_grad()\n        \n        # Track metrics\n        with torch.no_grad():\n            mask = (tgt_out != PAD_IDX)\n            num = mask.sum().item()\n            total_tokens += num\n            # Multiply back by accumulation_steps for true loss\n            total_loss += loss.item() * accumulation_steps * (B*T)\n            batch_count += 1\n        \n        # Progress indicator\n        if (batch_idx + 1) % 50 == 0:\n            avg_loss_so_far = total_loss / max(1, total_tokens)\n            ppl_so_far = math.exp(min(avg_loss_so_far, 100))\n            print(f\"  Epoch {epoch_num} - Batch {batch_idx + 1}/{len(train_loader)} | \"\n                  f\"Loss: {avg_loss_so_far:.4f} | PPL: {ppl_so_far:.2f}\", end='\\r')\n        \n        # Clear cache periodically\n        if (batch_idx + 1) % 100 == 0:\n            if torch.cuda.is_available():\n                torch.cuda.empty_cache()\n    \n    # Final optimizer step if there are remaining gradients\n    if (batch_idx + 1) % accumulation_steps != 0:\n        nn.utils.clip_grad_norm_(model.parameters(), 1.0)\n        optimizer.step()\n        optimizer.zero_grad()\n    \n    # Calculate final metrics\n    avg_loss = total_loss / max(1, total_tokens)\n    ppl = math.exp(min(avg_loss, 100))\n    \n    return avg_loss, ppl\n\n# Training loop\nEPOCHS = 5\nprint(f\"{'='*60}\")\nprint(f\"Starting Training - {EPOCHS} epochs\")\nprint(f\"Effective batch size: {batch_size * accumulation_steps} (via gradient accumulation)\")\nprint(f\"{'='*60}\\n\")\n\nbest_loss = float('inf')\n\nfor epoch in range(1, EPOCHS + 1):\n    print(f\"Epoch {epoch}/{EPOCHS}\")\n    avg_loss, ppl = train_epoch(epoch)\n    print(f\"\\n  Epoch {epoch}: avg_token_loss={avg_loss:.4f} | perplexity={ppl:.2f}\")\n    \n    if avg_loss < best_loss:\n        best_loss = avg_loss\n        print(f\"  ✓ New best loss: {best_loss:.4f}\")\n    \n    # Clear cache after each epoch\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n        gc.collect()\n    print()\n\nprint(f\"{'='*60}\")\nprint(f\"Training completed! Best loss: {best_loss:.4f}\")\nprint(f\"{'='*60}\\n\")\n\n# Inference - Greedy decoding\nEOS_IDX = SPECIAL_TOKENS['EOS']\nSOS_IDX = SPECIAL_TOKENS['SOS']\n\ndef greedy_decode_single(src_seq, max_len=256):\n    # Access the actual model if wrapped in DataParallel\n    actual_model = model.module if isinstance(model, nn.DataParallel) else model\n    actual_model.eval()\n    \n    with torch.no_grad():\n        src_tensor = torch.tensor(src_seq, dtype=torch.long, device=device).unsqueeze(0)\n        B, S = src_tensor.size()\n        \n        # Positional embedding for encoder\n        num_pos = actual_model.pos_embedding.num_embeddings\n        pos_idx = (torch.arange(S, device=device) % num_pos).unsqueeze(0).repeat(B, 1)\n        pos_embedded = actual_model.pos_embedding(pos_idx)\n        embedded_src = actual_model.src_embedding(src_tensor) + pos_embedded\n        _, (hidden, cell) = actual_model.encoder(embedded_src)\n\n        # Start with SOS\n        cur = torch.tensor([[SOS_IDX]], dtype=torch.long, device=device)\n        outputs = []\n        \n        for _ in range(max_len):\n            emb = actual_model.tgt_embedding(cur)\n            out, (hidden, cell) = actual_model.decoder(emb, (hidden, cell))\n            logit = actual_model.fc(out[:, -1, :])\n            next_token = torch.argmax(logit, dim=-1)\n            idx = next_token.item()\n            \n            if idx == EOS_IDX:\n                break\n            \n            outputs.append(idx)\n            cur = next_token.unsqueeze(0)\n        \n        # Map to string\n        chars = [rev_english_vocab.get(i, '?') for i in outputs]\n        return ''.join(chars)\n\n# Load and translate test set\nprint(f\"{'='*60}\")\nprint(\"Loading test data and generating translations...\")\nprint(f\"{'='*60}\\n\")\n\ntest_file_path = \"/kaggle/input/afml-assignment-1-ec-campus/Team 5/test_part2.csv\"\n\ntry:\n    test_df = pd.read_csv(test_file_path)\n    test_lines = test_df['encoded_text'].tolist()\n    print(f\"Loaded {len(test_lines)} test samples from CSV\")\nexcept Exception as e:\n    print(f\"Error reading CSV: {e}\")\n    try:\n        with open(test_file_path.replace('.csv', '.txt'), 'r', encoding='utf-8') as f:\n            test_lines = [line.strip() for line in f.readlines() if line.strip()]\n        print(f\"Loaded {len(test_lines)} test samples from TXT\")\n    except:\n        print(\"Could not load test file. Please check the path.\")\n        test_lines = []\n\nif test_lines:\n    # Encode test data\n    encoded_test = [[encoded_vocab.get(ch, SPECIAL_TOKENS['UNK']) for ch in line] for line in test_lines]\n    \n    # Generate translations with progress indicator\n    translations = []\n    print(f\"Generating translations...\")\n    for i, seq in enumerate(encoded_test):\n        trans = greedy_decode_single(seq)\n        translations.append(trans)\n        if (i + 1) % 50 == 0:\n            print(f\"  Translated {i + 1}/{len(encoded_test)} samples...\", end='\\r')\n        # Clear cache periodically during inference\n        if (i + 1) % 200 == 0 and torch.cuda.is_available():\n            torch.cuda.empty_cache()\n    \n    print(f\"  Translated {len(translations)}/{len(encoded_test)} samples - Complete!   \\n\")\n    \n    # Save translations\n    out_path = '/kaggle/working/05_translations_part2.txt'\n    with open(out_path, 'w', encoding='utf-8') as f:\n        for line, trans in zip(test_lines, translations):\n            f.write(f\"{line}\\t{trans}\\n\")\n    \n    print(f\"✅ Saved translations to: {out_path}\")\n    \n    # Show sample translations\n    print(f\"\\n{'='*60}\")\n    print(\"Sample Translations (first 5):\")\n    print(f\"{'='*60}\")\n    for i in range(min(5, len(translations))):\n        enc_preview = test_lines[i][:50] + ('...' if len(test_lines[i]) > 50 else '')\n        dec_preview = translations[i][:50] + ('...' if len(translations[i]) > 50 else '')\n        print(f\"\\n{i+1}. Encoded: {enc_preview}\")\n        print(f\"   Decoded: {dec_preview}\")\n    print(f\"\\n{'='*60}\")\nelse:\n    print(\"❌ No test data found!\")\n\nprint(\"\\n🎉 Pipeline complete!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T18:03:43.947374Z","iopub.execute_input":"2025-10-17T18:03:43.947658Z"}},"outputs":[],"execution_count":null}]}