{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"background-color:#1B79FD; color:#19180F; font-size:30px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> XLM(Cross-lingual language model)</div> \n<div style=\"background-color:#74A9EC; color:#19180F; font-size:20px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> Architecture Overview </div>\n<br>\nIt is a powerful transformer based architecture designed primarily for multilingual natural language processing tasks. It is built on original transformers architecture. The unique building blocks are defined below :\n\n1. Transformer Encoder : XLM utilizes a transformer encoder as its core building block. This encoder is composed of stacked self attention layers and feed forward neural networks. The self attention mechanism allows the model to capture the dependencies between different words in a sentence while feed forward networks provide non linear transformations to the input.\n\n2. Bilingual Training : One of the distinctive aspects of XLM is its training process. It leverages parallel data from multiple languages during training, which helps it learn shared representations across languages. This bilingual training enables the model to transfer knowledge b/w different languages and improve its cross lingual capabilities.\n\n3. Masked Language Model(MLM): Similar to other transformer architectures, XLM employs a masked language model objective during pretraining. This objective involves randomly masking certain tokens in the input and training the model to predict the original tokens based on the context. This task encourages the model to understand the relationships between words and learn meaningful representations.\n\n4. Cross-lingual Language Model(CLM): In addn. to MLM objective, XLM introduces a cross-lingual language model objective. This objective involves training the model to predict words in one language given the context of another language. by using this objective, XLM learns to transfer the knowledge across langs and develop a shared representation space.\n\n5. Language discriminator : XLM includes a language discriminator module to further enhance its cross lingual capabilities. The language discriminator is trained to predict the language of a given text. By jointly training the language discriminator and cross lang model, XLM learns to understand and differentiate b/w multiple languages in a much better way.\n\n6. Unsupervised Training : XLM utilizes unsupervised training methods, which meants that it can learn from large amounts of monolingual and parallel data without requiring explicitly large language annotations. This approach makes it useful for low resource languages where labeled data is scarce.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:#74A9EC; color:#19180F; font-size:20px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> Architecture Diagram </div>\n\n[View bigger diagram](https://drive.google.com/file/d/1qIdKfkDaj3I_Kv-4t119T8iroi_BYOwi/view?usp=sharing)","metadata":{}},{"cell_type":"code","source":"from IPython.display import display\nfrom PIL import Image\n\nimage_path = '/kaggle/input/notebook-images/XLM.jpg'\n\nimage = Image.open(image_path)\n\ndisplay(image)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T22:55:51.851412Z","iopub.execute_input":"2023-06-08T22:55:51.851844Z","iopub.status.idle":"2023-06-08T22:55:52.726505Z","shell.execute_reply.started":"2023-06-08T22:55:51.851791Z","shell.execute_reply":"2023-06-08T22:55:52.725573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install sacremoses --quiet","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:43:56.623076Z","iopub.execute_input":"2023-06-08T21:43:56.623436Z","iopub.status.idle":"2023-06-08T21:44:09.847290Z","shell.execute_reply.started":"2023-06-08T21:43:56.623406Z","shell.execute_reply":"2023-06-08T21:44:09.845916Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:20px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> Importing modules and loading the tokenizer </div>","metadata":{}},{"cell_type":"code","source":"\nimport torch\nimport pandas as pd\nfrom torch.utils.data import DataLoader\nfrom transformers import XLMTokenizer, XLMForSequenceClassification, AdamW\nfrom sklearn.metrics import roc_auc_score\n\n# Load the XLM tokenizer\ntokenizer = XLMTokenizer.from_pretrained('xlm-mlm-100-1280')","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:44:09.856770Z","iopub.execute_input":"2023-06-08T21:44:09.857476Z","iopub.status.idle":"2023-06-08T21:44:19.809613Z","shell.execute_reply.started":"2023-06-08T21:44:09.857430Z","shell.execute_reply":"2023-06-08T21:44:19.808437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> This code snippet loads the training and validation datasets from CSV files. It then tokenizes the comments in the datasets using a tokenizer. The tokenizer function takes the comment texts as input, performs truncation and padding, and returns the encoded representations of the comments as train_encodings and val_encodings respectively. </div>","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train-processed-seqlen128.csv')\nval_df = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation-processed-seqlen128.csv')\n\ntrain_encodings = tokenizer(train_df.comment_text.tolist(), truncation=True, padding=True)\nval_encodings = tokenizer(val_df.comment_text.tolist(), truncation=True, padding=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:44:19.811264Z","iopub.execute_input":"2023-06-08T21:44:19.812074Z","iopub.status.idle":"2023-06-08T21:54:18.846354Z","shell.execute_reply.started":"2023-06-08T21:44:19.812035Z","shell.execute_reply":"2023-06-08T21:54:18.845182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> This code snippet defines a custom PyTorch dataset class called `ToxicCommentDataset` for the training and validation data. The class takes two arguments, `encodings` and `labels`, which are the encoded representations of the comments and their corresponding labels, respectively.<br>\n<br>\nThe `__getitem__` method returns a dictionary of tensors for each encoding key (such as `input_ids`, `attention_mask`, etc.) and the corresponding label tensor for a given index `idx`. The `__len__` method returns the length of the dataset, which is equal to the number of labels.<br> </div>","metadata":{}},{"cell_type":"code","source":"class ToxicCommentDataset(torch.utils.data.Dataset):\n    def __init__(self, encodings, labels):\n        self.encodings = encodings\n        self.labels = labels\n\n    def __getitem__(self, idx):\n        return {key: torch.tensor(val[idx]) for key, val in self.encodings.items()}, torch.tensor(self.labels[idx])\n\n    def __len__(self):\n        return len(self.labels)","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:54:18.851090Z","iopub.execute_input":"2023-06-08T21:54:18.852082Z","iopub.status.idle":"2023-06-08T21:54:18.861016Z","shell.execute_reply.started":"2023-06-08T21:54:18.852038Z","shell.execute_reply":"2023-06-08T21:54:18.859858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> These lines of code create instances of the `ToxicCommentDataset` class for the training and validation data. The `train_dataset` is initialized with `train_encodings` (the encoded representations of the training comments) and `train_df.toxic.values` (the corresponding toxic labels). Similarly, the `val_dataset` is initialized with `val_encodings` (the encoded representations of the validation comments) and `val_df.toxic.values` (the corresponding toxic labels).</div>","metadata":{}},{"cell_type":"code","source":"train_dataset = ToxicCommentDataset(train_encodings, train_df.toxic.values)\nval_dataset = ToxicCommentDataset(val_encodings, val_df.toxic.values)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:54:18.862522Z","iopub.execute_input":"2023-06-08T21:54:18.863699Z","iopub.status.idle":"2023-06-08T21:54:18.875057Z","shell.execute_reply.started":"2023-06-08T21:54:18.863657Z","shell.execute_reply":"2023-06-08T21:54:18.873980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> These lines of code create data loaders for the training and validation datasets. The `train_loader` is created using the `DataLoader` class with `train_dataset` as the dataset, a batch size of 4, and `shuffle=True` to randomize the order of samples in each batch during training. Similarly, the `val_loader` is created with `val_dataset` as the dataset, a batch size of 4, and `shuffle=False` to maintain the order of samples during validation. </div>","metadata":{}},{"cell_type":"code","source":"train_loader = DataLoader(train_dataset, batch_size=4, shuffle=True)\nval_loader = DataLoader(val_dataset, batch_size=4, shuffle=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:54:18.877077Z","iopub.execute_input":"2023-06-08T21:54:18.877526Z","iopub.status.idle":"2023-06-08T21:54:18.886578Z","shell.execute_reply.started":"2023-06-08T21:54:18.877490Z","shell.execute_reply":"2023-06-08T21:54:18.885441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> Performing sanity check of the dataloader </div>","metadata":{}},{"cell_type":"code","source":"for batch in train_loader:\n    print(batch)\n    break","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:54:18.888072Z","iopub.execute_input":"2023-06-08T21:54:18.889032Z","iopub.status.idle":"2023-06-08T21:54:18.950635Z","shell.execute_reply.started":"2023-06-08T21:54:18.888989Z","shell.execute_reply":"2023-06-08T21:54:18.949348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> The first line of code sets the device to use CUDA for GPU acceleration.<br>\n<br>\nThe second line of code loads the XLM model for sequence classification using the `XLMForSequenceClassification` class from the \"xlm-mlm-100-1280\" pretrained model. The `num_labels` parameter is set to 2, indicating the number of classes in the classification task.<br>\n<br>\nThe third line of code moves the model to the specified device (CUDA) using the `to` method.<br>\n<br>\nThe fourth line of code sets the model to training mode by invoking the `train` method.<br> </div>","metadata":{}},{"cell_type":"code","source":"device = torch.device(\"cuda\")\n\nmodel = XLMForSequenceClassification.from_pretrained('xlm-mlm-100-1280', num_labels=2)\nmodel.to(device)\nmodel.train()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:55:03.469184Z","iopub.execute_input":"2023-06-08T21:55:03.469582Z","iopub.status.idle":"2023-06-08T21:55:16.684036Z","shell.execute_reply.started":"2023-06-08T21:55:03.469549Z","shell.execute_reply":"2023-06-08T21:55:16.682909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"><br> The first line of code initializes the AdamW optimizer with the parameters of the `model`. The learning rate (`lr`) is set to 5e-5.\n<br>\nThe second line of code creates a step-based learning rate scheduler using the `StepLR` class from `torch.optim.lr_scheduler`. The scheduler is associated with the `optimizer` and has a step size of 1. After each step, the learning rate will be multiplied by the specified `gamma` value of 0.1.<br> </div>","metadata":{}},{"cell_type":"code","source":"optimizer = AdamW(model.parameters(), lr=5e-5)\nscheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=1, gamma=0.1)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T21:55:16.686231Z","iopub.execute_input":"2023-06-08T21:55:16.688151Z","iopub.status.idle":"2023-06-08T21:55:16.701241Z","shell.execute_reply.started":"2023-06-08T21:55:16.688107Z","shell.execute_reply":"2023-06-08T21:55:16.699649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> This code snippet performs a training loop for 2 epochs. You can increase the training for more epochs to obtain higher accuracy. <br>\n<br>\nWithin each epoch, it iterates over the batches in the `train_loader` using the `tqdm` function for progress tracking. <br>\n<br>\nFor each batch, it moves the input data (`input_ids`, `attention_mask`, `labels`) to the device (CUDA). <br>\n<br>\nThen, it zeros the gradients, performs a forward pass through the model, calculates the loss, and performs a backward pass and optimizer step to update the model parameters.<br>\n<br>\nAfter each batch, if the step number is divisible by 100, it prints the current step number and the corresponding loss.<br>\n<br>\nThen, it switches the model to evaluation mode (`model.eval()`) and performs inference on the validation dataset (`val_loader`) using a similar loop structure. It computes probabilities using softmax and selects the predicted class label. The predicted labels and true labels are accumulated to calculate the ROC AUC score for validation (`val_auc`).<br>\n<br>\nFinally, it prints the validation AUC score for each epoch, and the learning rate scheduler is updated with `scheduler.step()` at the end of each epoch to adjust the learning rate based on the predefined step size and gamma values.<br> </div>","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nfor epoch in range(2):\n    for step,batch in tqdm(enumerate(train_loader)):\n        input_ids = batch[0]['input_ids'].to(device)\n        attention_mask = batch[0]['attention_mask'].to(device)\n        labels = batch[1].to(device)\n\n        # Zero the gradients\n        optimizer.zero_grad()\n\n        # Forward pass\n        outputs = model(input_ids, attention_mask=attention_mask, labels=labels)\n\n        # Backward pass\n        loss = outputs.loss\n        \n        loss.backward()\n        optimizer.step()\n    if step%100==0:\n        print(\"Step-{},Loss-{}\".format(step,loss.item()))\n\n    model.eval()\n    val_predictions = []\n    val_labels = []\n    with torch.no_grad():\n        for step,batch in tqdm(enumerate(val_loader)):\n            # Move the batch to the device\n            input_ids = batch[0]['input_ids'].to(device)\n            attention_mask = batch[0]['attention_mask'].to(device)\n            labels = batch[1].to(device)\n\n            outputs = model(input_ids, attention_mask=attention_mask)\n            probabilities = torch.softmax(outputs.logits, dim=1)\n            _, predicted = torch.max(probabilities, dim=1)\n\n            val_predictions.extend(predicted.cpu().numpy())\n            val_labels.extend(labels.cpu().numpy())\n\n    val_auc = roc_auc_score(val_labels, val_predictions)\n\n    print(f'Epoch {epoch + 1}: Validation AUC = {val_auc}')\n\n    scheduler.step()\n#keyboard interrupt --> intentional since train time is higher, you may uncomment and train for more.","metadata":{"execution":{"iopub.status.busy":"2023-06-08T22:13:13.122533Z","iopub.execute_input":"2023-06-08T22:13:13.123782Z","iopub.status.idle":"2023-06-08T22:14:35.549311Z","shell.execute_reply.started":"2023-06-08T22:13:13.123728Z","shell.execute_reply":"2023-06-08T22:14:35.547658Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> <br>This code snippet loads the test dataset from a CSV file and tokenizes the comments in the test dataset using the same tokenizer as before. The resulting encodings are stored in `test_encodings`.<br>\n<br>\nA PyTorch dataset for the test data is created using the `ToxicCommentDataset` class with `test_encodings` as the encodings and a list of zeros with the length of the test dataset as the labels.<br>\n<br>\nA data loader for the test data is defined using the `DataLoader` class, with `test_dataset` as the dataset, a batch size of 32, and `shuffle=False` to maintain the order of samples.<br>\n<br>\nFinally, the model is switched to evaluation mode using `model.eval()`.<br> </div>","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')\ntest_encodings = tokenizer(test_df.content.tolist(), truncation=True, padding=True)\n\ntest_dataset = ToxicCommentDataset(test_encodings, [0] * len(test_df))\n\ntest_loader = DataLoader(test_dataset, batch_size=32, shuffle=False)\n\nmodel.eval()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-08T22:15:33.908480Z","iopub.execute_input":"2023-06-08T22:15:33.908940Z","iopub.status.idle":"2023-06-08T22:18:32.615920Z","shell.execute_reply.started":"2023-06-08T22:15:33.908901Z","shell.execute_reply":"2023-06-08T22:18:32.614838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> This code snippet creates an empty list called `predictions` to store the predicted classes.<br>\n<br>\nThen, it iterates over the test dataset (`test_loader`) using the `tqdm` function for progress tracking.<br>\n<br>\nFor each batch in the test dataset, it moves the input data (`input_ids`, `attention_mask`) to the device (CUDA).<br>\n<br>\nNext, it performs inference on the model by passing the input data through the model. It computes the probabilities using softmax and selects the predicted class label.\n<br>\nThe predicted classes are appended to the `predictions` list by extending it with the numpy array of predicted labels converted to CPU.<br>\n<br>\nAfter the loop completes, the `predictions` list will contain all the predicted classes for the test dataset.<br> </div>","metadata":{}},{"cell_type":"code","source":"predictions = []\n\nwith torch.no_grad():\n    for batch in tqdm(test_loader):\n        input_ids = batch[0]['input_ids'].to(device)\n        attention_mask = batch[0]['attention_mask'].to(device)\n\n        outputs = model(input_ids, attention_mask=attention_mask)\n        probabilities = torch.softmax(outputs.logits, dim=1)\n        _, predicted = torch.max(probabilities, dim=1)\n\n        predictions.extend(predicted.cpu().numpy())\n#keyboard interrupt --> intentional since inference time is higher, you may uncomment and infer.","metadata":{"execution":{"iopub.status.busy":"2023-06-08T22:23:16.826552Z","iopub.execute_input":"2023-06-08T22:23:16.827113Z","iopub.status.idle":"2023-06-08T22:30:25.827011Z","shell.execute_reply.started":"2023-06-08T22:23:16.827069Z","shell.execute_reply":"2023-06-08T22:30:25.825468Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#D4E8FF; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> This code snippet creates a DataFrame called `submission_df` with two columns: 'id' and 'toxic'. The 'id' column contains the corresponding IDs from the test dataset, and the 'toxic' column contains the predicted classes stored in the `predictions` list.<br>\n<br>\nThen, it saves the `submission_df` DataFrame to a CSV file named 'submission.csv' using the `to_csv` method. The `index=False` parameter ensures that the row indices are not included in the CSV file. <br></div>","metadata":{}},{"cell_type":"code","source":"# Create a dataframe with the predictions\nsubmission_df = pd.DataFrame({'id': test_df.id, 'toxic': predictions})\n\n# Save the dataframe to a CSV file\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-06-08T22:23:10.156470Z","iopub.status.idle":"2023-06-08T22:23:10.157972Z","shell.execute_reply.started":"2023-06-08T22:23:10.157625Z","shell.execute_reply":"2023-06-08T22:23:10.157661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df","metadata":{},"execution_count":null,"outputs":[]}]}