{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"background-color:#035FCA; color:#19180F; font-size:40px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> XLNet(eXtreme Lite Transformer) </div>\n<div style=\"background-color:#568FD1; color:#19180F; font-size:30px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> 📝 Architectural Overview.\n </div>\n<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\n- It is a state of the art llm developed by CMU and Google AI. It is based on transformer architecture and was designed to overcome the limitations of prev language models such as BERT and GPT2.<br>\n- The architecture of XLNet is similar to BERT with some key diff. The main diff being XLNet uses a permutation based training approach which allows it to model dependencies between all tokens in a sequence rather than just the tokens that come before the current token.<br>\n <br>\nThe block diagram of XLNet consists of three main components.\n<br>\n- Input embedding layer - This layer takes the input text and converts it into a vector representation that can be processed by the model. The input embedding layer uses a pretrained word embedding model to convert each word in the input text into a vector.<br>\n- Transformer encoder layers - The transformer encoder layers are the core buiding blocks of the model. They use self attn mechanism to process the input text and generate a contextualised representation of each word in the text. The transformer encoder layers are stacked on top of each other to create a deep neural network.<br>\n- Permutation based training - XLNet uses a permutation based training approach which allows it to model dependencies between all tokens in a sequence rather than just the tokens that come before the current token. This is achieved by randomly permuting the input seq during training and using a modified loss function that takes into account all possible permutations of the input sequence<br>\n<br>\nThe advantages of XLNet are<br>\n1. Improved modeling of dependencies. - It is able to model dependencies between all tokens in a sequence rather than just the tokens that come before it. It yields better and coherent text.<br>\n2. It performs better on Question-answering, sentiment analysis and text classification in comparision to bert and GPT2.<br>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:#568FD1; color:#19180F; font-size:30px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> 🏢 Architecture Diagram.\n </div>\n<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\n1. Input Layer:<br>\n   - The diagram starts with the \"Input Layer\" represented by the `input` node. It represents the input text that is fed into the XLNet model for processing.\n<br>\n2. Segment Embeddings:<br>\n   - The input text is passed through the \"Segment Embeddings\" layer, represented by the `segment` node. This layer assigns different embeddings to different segments or parts of the input text.\n<br>\n3. Position Embeddings:<br>\n   - The input text is also passed through the \"Position Embeddings\" layer, represented by the `position` node. This layer assigns embeddings based on the position or order of the words in the input text.\n<br>\n4. Attention Layer:<br>\n   - The segment embeddings and position embeddings are combined and fed into the \"Attention Layer,\" represented by the `attention` node. The attention layer performs self-attention, allowing the model to focus on different parts of the input text while considering the dependencies between words.\n<br>\n5. Feed-Forward Layer:<br>\n   - The output from the attention layer is passed through the \"Feed-Forward Layer,\" represented by the `feed_forward` node. This layer applies a neural network with multiple layers and nonlinear transformations to capture complex patterns in the data.\n<br>\n6. Residual Connections and Layer Normalization:<br>\n   - To facilitate better information flow and mitigate the vanishing gradient problem, \"Residual Connections\" are added between the attention layer and feed-forward layer. The \"Add & Layer Norm\" operations, represented by `add_norm_1` and `add_norm_2`, respectively, combine the output of the previous layer with its input and apply layer normalization.\n<br>\n7. Output Layer:<br>\n   - Finally, the output from the feed-forward layer passes through the \"Output Layer,\" represented by the `output` node, to generate the final output of the XLNet model.<br>\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import SVG, display\n\n# Load the SVG file and display it\nsvg_file = '/kaggle/input/notebook-images/xlnet.svg'\ndisplay(SVG(filename=svg_file))","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:17:06.946819Z","iopub.execute_input":"2023-06-09T01:17:06.947832Z","iopub.status.idle":"2023-06-09T01:17:06.967158Z","shell.execute_reply.started":"2023-06-09T01:17:06.947796Z","shell.execute_reply":"2023-06-09T01:17:06.966248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe import the required libraries including PyTorch, XLNetForSequenceClassification, XLNetTokenizer from the transformers package, and pandas for data handling.</div>\n\n","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nfrom torch.utils.data import DataLoader\nfrom transformers import XLNetForSequenceClassification, XLNetTokenizer\nfrom sklearn.metrics import f1_score\nimport pandas as pd\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T00:54:20.009229Z","iopub.execute_input":"2023-06-09T00:54:20.009801Z","iopub.status.idle":"2023-06-09T00:54:30.793979Z","shell.execute_reply.started":"2023-06-09T00:54:20.009768Z","shell.execute_reply":"2023-06-09T00:54:30.792785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe check if a GPU is available and set the device accordingly. This enables GPU acceleration if available.</div>","metadata":{}},{"cell_type":"code","source":"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T00:54:30.796360Z","iopub.execute_input":"2023-06-09T00:54:30.796749Z","iopub.status.idle":"2023-06-09T00:54:30.825169Z","shell.execute_reply.started":"2023-06-09T00:54:30.796717Z","shell.execute_reply":"2023-06-09T00:54:30.823743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe initialize the XLNet model for sequence classification and its corresponding tokenizer. We use the 'xlnet-base-cased' pre-trained model.</div>","metadata":{}},{"cell_type":"code","source":"# Define XLNet model and tokenizer\nmodel = XLNetForSequenceClassification.from_pretrained('xlnet-base-cased', num_labels=2)\ntokenizer = XLNetTokenizer.from_pretrained('xlnet-base-cased')\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T00:54:30.826713Z","iopub.execute_input":"2023-06-09T00:54:30.827656Z","iopub.status.idle":"2023-06-09T00:54:51.235407Z","shell.execute_reply.started":"2023-06-09T00:54:30.827619Z","shell.execute_reply":"2023-06-09T00:54:51.234422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe read the Quora Insincere Question Classification dataset from CSV files using pandas.</div>","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/train.csv')\ntest_df = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/test.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T00:54:51.237844Z","iopub.execute_input":"2023-06-09T00:54:51.238175Z","iopub.status.idle":"2023-06-09T00:54:56.690640Z","shell.execute_reply.started":"2023-06-09T00:54:51.238142Z","shell.execute_reply":"2023-06-09T00:54:56.689583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe tokenize and encode the question texts using the XLNet tokenizer. The encoded sequences include input IDs, attention masks, and the target labels for the training dataset.</div>","metadata":{}},{"cell_type":"code","source":"train_encoded = tokenizer.batch_encode_plus(train_df['question_text'].tolist(),\n                                            add_special_tokens=True,\n                                            padding='longest',\n                                            truncation=True,\n                                            return_tensors='pt', max_length=64) #increase max length to 512 if there are no memory restrictions","metadata":{"execution":{"iopub.status.busy":"2023-06-09T00:54:56.692054Z","iopub.execute_input":"2023-06-09T00:54:56.692423Z","iopub.status.idle":"2023-06-09T01:00:34.848509Z","shell.execute_reply.started":"2023-06-09T00:54:56.692390Z","shell.execute_reply":"2023-06-09T01:00:34.847510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_encoded = tokenizer.batch_encode_plus(test_df['question_text'].tolist(),\n                                           add_special_tokens=True,\n                                           padding='longest',\n                                           truncation=True,\n                                           return_tensors='pt',max_length=64)#increase max length to 512 if there are no memory restrictions\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:00:34.850057Z","iopub.execute_input":"2023-06-09T01:00:34.850598Z","iopub.status.idle":"2023-06-09T01:02:14.404155Z","shell.execute_reply.started":"2023-06-09T01:00:34.850558Z","shell.execute_reply":"2023-06-09T01:02:14.403131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe create data loaders for both the training and test datasets using the encoded sequences. </div>","metadata":{}},{"cell_type":"code","source":"train_dataset = torch.utils.data.TensorDataset(train_encoded['input_ids'],\n                                               train_encoded['attention_mask'],\n                                               torch.tensor(train_df['target'].tolist()))\ntest_dataset = torch.utils.data.TensorDataset(test_encoded['input_ids'], test_encoded['attention_mask'])\ntrain_loader = DataLoader(train_dataset, batch_size=16*4, shuffle=True)\ntest_loader = DataLoader(test_dataset, batch_size=16*4, shuffle=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:14.405816Z","iopub.execute_input":"2023-06-09T01:02:14.406177Z","iopub.status.idle":"2023-06-09T01:02:14.842095Z","shell.execute_reply.started":"2023-06-09T01:02:14.406144Z","shell.execute_reply":"2023-06-09T01:02:14.841027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nPerformed sanity check of the dataloaders</div","metadata":{}},{"cell_type":"code","source":"for batch in train_loader:\n    print(batch)\n    break","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:14.843789Z","iopub.execute_input":"2023-06-09T01:02:14.844459Z","iopub.status.idle":"2023-06-09T01:02:14.980077Z","shell.execute_reply.started":"2023-06-09T01:02:14.844421Z","shell.execute_reply":"2023-06-09T01:02:14.979123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe define the number of training epochs and the learning rate.</div","metadata":{}},{"cell_type":"code","source":"num_epochs = 1\nlearning_rate = 2e-5\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:14.981471Z","iopub.execute_input":"2023-06-09T01:02:14.982010Z","iopub.status.idle":"2023-06-09T01:02:14.986623Z","shell.execute_reply.started":"2023-06-09T01:02:14.981975Z","shell.execute_reply":"2023-06-09T01:02:14.985585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe specify the loss function (cross-entropy loss) and the optimizer (AdamW) for model training.</div>","metadata":{}},{"cell_type":"code","source":"criterion = nn.CrossEntropyLoss()\noptimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:14.990528Z","iopub.execute_input":"2023-06-09T01:02:14.991129Z","iopub.status.idle":"2023-06-09T01:02:14.998036Z","shell.execute_reply.started":"2023-06-09T01:02:14.991095Z","shell.execute_reply":"2023-06-09T01:02:14.997104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe moved the model to device and put it to train.</div>","metadata":{}},{"cell_type":"code","source":"model.to(device)\nmodel.train()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:14.999554Z","iopub.execute_input":"2023-06-09T01:02:14.999973Z","iopub.status.idle":"2023-06-09T01:02:19.599220Z","shell.execute_reply.started":"2023-06-09T01:02:14.999936Z","shell.execute_reply":"2023-06-09T01:02:19.598199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe train the XLNet model by iterating over the training data loader. The model is set to train mode, and the optimizer is zeroed before computing and backpropagating the loss.</div>","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\n\nfor epoch in range(num_epochs):\n    total_loss = 0\n    for step, (inputs, masks, targets) in enumerate(tqdm(train_loader)):\n        inputs, masks, targets = inputs.to(device), masks.to(device), targets.to(device)\n\n        optimizer.zero_grad()\n\n        outputs = model(input_ids=inputs, attention_mask=masks)[0]\n        loss = criterion(outputs, targets)\n        loss.backward()\n        optimizer.step()\n        if step%1000==0:\n            print(\"Step - {}, Loss - {}\".format(step, loss.item()))\n            break\n\n        total_loss += loss.item()\n\n    print(f\"Epoch {epoch + 1}/{num_epochs}, Loss: {total_loss}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:19.601363Z","iopub.execute_input":"2023-06-09T01:02:19.602179Z","iopub.status.idle":"2023-06-09T01:02:21.517637Z","shell.execute_reply.started":"2023-06-09T01:02:19.602141Z","shell.execute_reply":"2023-06-09T01:02:21.516698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe intentionally trained for few steps of epoch 0, you should train more !</div>","metadata":{}},{"cell_type":"code","source":"model.eval()\npredictions = []\n","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:21.519124Z","iopub.execute_input":"2023-06-09T01:02:21.519719Z","iopub.status.idle":"2023-06-09T01:02:21.526725Z","shell.execute_reply.started":"2023-06-09T01:02:21.519683Z","shell.execute_reply":"2023-06-09T01:02:21.525684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nWe switch the model to evaluation mode and iterate over the test data loader to make predictions on the test dataset.</div>","metadata":{}},{"cell_type":"code","source":"with torch.no_grad():\n    for inputs, masks in tqdm(test_loader):\n        inputs, masks = inputs.to(device), masks.to(device)\n        outputs = model(input_ids=inputs, attention_mask=masks)[0]\n        _, predicted_labels = torch.max(outputs, 1)\n        predictions.extend(predicted_labels.tolist())","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:02:21.528201Z","iopub.execute_input":"2023-06-09T01:02:21.528574Z","iopub.status.idle":"2023-06-09T01:17:05.797147Z","shell.execute_reply.started":"2023-06-09T01:02:21.528541Z","shell.execute_reply":"2023-06-09T01:17:05.796142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nsubmission.csv generated</div>","metadata":{}},{"cell_type":"code","source":"submission_df = pd.DataFrame({'qid': test_df['qid'], 'prediction': predictions})\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:17:05.798749Z","iopub.execute_input":"2023-06-09T01:17:05.799738Z","iopub.status.idle":"2023-06-09T01:17:06.945504Z","shell.execute_reply.started":"2023-06-09T01:17:05.799698Z","shell.execute_reply":"2023-06-09T01:17:06.944534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df","metadata":{"execution":{"iopub.status.busy":"2023-06-09T01:18:25.460758Z","iopub.execute_input":"2023-06-09T01:18:25.461139Z","iopub.status.idle":"2023-06-09T01:18:25.479460Z","shell.execute_reply.started":"2023-06-09T01:18:25.461107Z","shell.execute_reply":"2023-06-09T01:18:25.478330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}