{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"background-color:#035FCA; color:#19180F; font-size:40px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> CTRL(Conditional Transformer Language Model) </div>\n\n<div style=\"background-color:#568FD1; color:#19180F; font-size:30px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">🔧Architecture Overview⚙️\n </div>\n<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nIt is a language model architecture designed to generate human like text based on a given control code or prompt. It combines the power of transformers with the ability to condition the model's output based on specific instructions.<br>\n\nThe architecture of CTRL consists of various building blocks which are explained below<br>\n\n1. Input Encoding : The input text is first tokenized. These tokens capture the meaning & structure of the text. Each token is then converted into a high dimensional vector representation.<br>\n\n2. Transformer Encoder - The encoded input tokens are passed through multiple layers of transformer encoders. Each encoder layer consists of two sub layers - A multi headed self attn mechanism and a position wise feedfwd neural network. The self attn mechanism helps the model capture dependencies between diff words in the i/p whereas the feed forward network applies non linear transformations to each position in the sequence.<br>\n\n3. Control code embedding : CTRL introduces an additional control code embedding to allow conditioning the model's output on specific instructions. The control code is appenedd to the input tokens and has its own learned embedding representations. This enables fine grained control over the generated text.<br>\n\n4. Transformer Decoder : The output of the control code embedding is passed through a series of transformer decoder layers. Each decoder layer also consists of self attn and feed forward sub layers. The decoder layers refine the representation of the control code and generate the final output tokens.<br>\n\n5. Output decoding - The decoder's o/p tokens are decoded to generate the final text. The decoding process involves converting the token representations back into text.<br>\n\n\n<b>Architectural differences w.r.t ELECTRA, ALBERT, BERT, RoBERTa, GPT2, XLNet and T5</b><br>\n\n- Controlled Generation - CTRL is specifically designed for controlled text generation where a control code or prompt guides the output. It allows users to specify the desired style, topic or other attributes of the generated text.<br>\n\n- Bidirectionality - Unlike models like GPT2 and XLNet, which are unidirectional llms, CTRL is bidirectional i.e It can utilize both the left and right contexts of a token to generate its representations. This allows the model to have a better understanding of the dependencies within the text like Bidirectional LSTM does.<br>\n\n- Control code conditioning - It introduces control code embedding which is a unique component in its architecture. It enables the model to condition its output based on the given control code making it highly flexible for various language generation tasks.<br>\n\n- Diff pretraining objective- UNlike ELECTRA and BERT which employs MLM as their pretraining objective, CTRL employs a new objective called \"relevance ranking\" which involves compairing different parts of the input text to identify the most relevant content.<br>\n\n- Domain adaptation - CTRL supports domain adaptation, allowing fine tuning of the model on specific domains or datasets. This improves the model's performance and adaptability for particular applications.<br>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:#568FD1; color:#19180F; font-size:30px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\"> 🏢 Architecture Diagram.\n </div>\n<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\n\n1. Input: This is the text that is entered into the CTRL Transformer architecture.<br>\n\n2. Tokenization: The tokenizer performs the tokenization procedure on the supplied text. Tokenization divides the input text into individual tokens, which are the basic units of input for the language model.<br>\n\n3. Encoder: The essential component of the CTRL Transformer design is the encoder. It is made up of many encoder blocks, each of which is responsible for processing a piece of the input text. Four encoder blocks are presented in pairs in this figure.<br>\n\n4. Encoder Blocks: The encoder blocks take the tokenized input text and conduct different actions on it in order to capture the tokens' contextual information. Each encoder block has a number of components that are shared by the pairings.<br>\n\n5. Encoder Components: The following are the major components found in each encoder block:<br>\n   - Self-Attention: This component enables the model to focus on various segments of the input sequence while taking into account token connections.<br>\n   - Multi-Head Attention: It expands self-attention by attending to and mixing distinct representations (heads) of the information.<br>\n   - Feed-Forward Network: This component processes attended representations and extracts higher-level features using a collection of fully linked layers.<br>\n   - Residual Connection: It improves information flow by connecting the feed-forward network's output to the encoder block's input.<br>\n   - Layer Normalisation: This component normalises the encoder block's output, assisting in the stabilisation of the training process.<br>\n\n6. Connections: The connections between the parts show how information and data move across the encoder blocks. The data transmission direction is shown by the arrows.<br>\n\n7. Output: The processed representations go to the last layer normalisation component after passing via the encoder blocks, where they produce the output text.<br>\n\nThe illustration shows the process by which the input text is tokenized, then processed by the encoder blocks with their corresponding components, and ultimately turned into the output text.<br>\n</div>\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import SVG, display\n\n# Load the SVG file and display it\nsvg_file = '/kaggle/input/notebook-images/ctrl.svg'\ndisplay(SVG(filename=svg_file))","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:26:57.188194Z","iopub.execute_input":"2023-06-10T22:26:57.188514Z","iopub.status.idle":"2023-06-10T22:26:57.221634Z","shell.execute_reply.started":"2023-06-10T22:26:57.188483Z","shell.execute_reply":"2023-06-10T22:26:57.220820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThis code snippet imports the necessary libraries.</div>","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import DataLoader\nfrom torchvision import transforms\nfrom PIL import Image\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom transformers import CTRLConfig, CTRLLMHeadModel, CTRLTokenizer, CTRLForSequenceClassification, AdamW","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:26:57.222912Z","iopub.execute_input":"2023-06-10T22:26:57.223165Z","iopub.status.idle":"2023-06-10T22:27:10.925031Z","shell.execute_reply.started":"2023-06-10T22:26:57.223143Z","shell.execute_reply":"2023-06-10T22:27:10.924111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nDefining the path to train.csv and train folder</div>","metadata":{}},{"cell_type":"code","source":"train_csv_path = '/kaggle/input/landmark-recognition-2021/train.csv'\ntrain_folder_path = '/kaggle/input/landmark-recognition-2021/train/'\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:10.926956Z","iopub.execute_input":"2023-06-10T22:27:10.927285Z","iopub.status.idle":"2023-06-10T22:27:10.932110Z","shell.execute_reply.started":"2023-06-10T22:27:10.927253Z","shell.execute_reply":"2023-06-10T22:27:10.930778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThe code snippet reads a CSV file (`train_csv_path`) using pandas' `read_csv` function and stores the data in a DataFrame called `train_df`. The DataFrame represents a dataset containing information about images, including their IDs and corresponding image paths.<br>\n\nThe next line modifies the `image_path` column in the `train_df` DataFrame. It concatenates the `train_folder_path` (the base directory for the training images) with the subdirectories based on the ID of each image. The `train_df['id'].str[0]` extracts the first character of the ID, `train_df['id'].str[1]` extracts the second character, `train_df['id'].str[2]` extracts the third character, and `train_df['id']` represents the full ID. These components are concatenated using '/' as separators, and the '.jpg' extension is added at the end to form the complete image path for each image in the dataset.<br></div>\n","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(train_csv_path)\n# Modify image_path column to match the directory structure\ntrain_df['image_path'] = train_folder_path + train_df['id'].str[0] + '/' + train_df['id'].str[1] + '/' + train_df['id'].str[2] + '/' + train_df['id'] + '.jpg'\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:10.933590Z","iopub.execute_input":"2023-06-10T22:27:10.934240Z","iopub.status.idle":"2023-06-10T22:27:17.572987Z","shell.execute_reply.started":"2023-06-10T22:27:10.934209Z","shell.execute_reply":"2023-06-10T22:27:17.572050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nPerforming sanity check of the dataframe's image_path column</div>","metadata":{}},{"cell_type":"code","source":"train_df['image_path'][0]","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:17.575755Z","iopub.execute_input":"2023-06-10T22:27:17.576120Z","iopub.status.idle":"2023-06-10T22:27:17.583141Z","shell.execute_reply.started":"2023-06-10T22:27:17.576088Z","shell.execute_reply":"2023-06-10T22:27:17.582104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nSplitting the modified dataframe into train and val data</div>","metadata":{}},{"cell_type":"code","source":"train_data, val_data = train_test_split(train_df, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:17.584846Z","iopub.execute_input":"2023-06-10T22:27:17.585229Z","iopub.status.idle":"2023-06-10T22:27:18.074033Z","shell.execute_reply.started":"2023-06-10T22:27:17.585194Z","shell.execute_reply":"2023-06-10T22:27:18.073034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThe code snippet defines a series of image transformations using `torchvision.transforms.Compose`. These transformations are applied to each image in order to preprocess them before feeding them into a neural network model for training or inference.<br>\n\nThe transformations specified in `image_transforms` are as follows:<br>\n\n1. `transforms.Resize((224, 224))`: Resizes the input image to a fixed size of 224x224 pixels. This is a common size used in many computer vision models.<br>\n\n2. `transforms.ToTensor()`: Converts the image from PIL Image format to a PyTorch tensor. This allows the image to be processed by PyTorch and passed through neural networks.<br>\n\n3. `transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])`: Normalizes the tensor image by subtracting the mean values `[0.485, 0.456, 0.406]` from each channel and dividing by the standard deviation values `[0.229, 0.224, 0.225]`. Normalization helps in standardizing the pixel values across images and improves model performance.<br></div>\n","metadata":{}},{"cell_type":"code","source":"image_transforms = transforms.Compose([\n    transforms.Resize((224, 224)),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])\n])","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:18.075468Z","iopub.execute_input":"2023-06-10T22:27:18.076683Z","iopub.status.idle":"2023-06-10T22:27:18.082611Z","shell.execute_reply.started":"2023-06-10T22:27:18.076646Z","shell.execute_reply":"2023-06-10T22:27:18.081682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThe code snippet defines a custom dataset class called `LandmarkDataset`. This class inherits from `torch.utils.data.Dataset`, which is a PyTorch base class for creating custom datasets.<br>\n\nThe `LandmarkDataset` class has the following methods:<br>\n\n1. `__init__(self, data, transform=None)`: The initialization method takes two parameters: `data` and `transform`. The `data` parameter represents the dataset, typically a DataFrame containing information about the images. The `transform` parameter is an optional parameter that represents the image transformations to be applied to each image. These transformations are defined earlier using `transforms.Compose`.<br>\n\n2. `__len__(self)`: This method returns the length of the dataset, i.e., the total number of samples in the dataset.<br>\n\n3. `__getitem__(self, index)`: This method retrieves an item from the dataset at the specified `index`. It first retrieves the image path corresponding to the given index from the dataset. Then, it opens the image using `PIL.Image.open` and converts it to the RGB mode using `.convert('RGB')`. This ensures that the image has three channels (red, green, and blue) required by most deep learning models.<br>\n\n   If a transformation is specified (`self.transform is not None`), it applies the transformation to the image using `self.transform(image)`. The transformed image is then returned.<br>\n\n   Finally, the method returns the image as the output.<br></div>\n","metadata":{}},{"cell_type":"code","source":"class LandmarkDataset(torch.utils.data.Dataset):\n    def __init__(self, data, transform=None):\n        self.data = data\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.data)\n\n    def __getitem__(self, index):\n        image_path = self.data.iloc[index]['image_path']\n        image = Image.open(image_path).convert('RGB')\n\n        if self.transform is not None:\n            image = self.transform(image)\n\n        return image\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:18.084300Z","iopub.execute_input":"2023-06-10T22:27:18.084687Z","iopub.status.idle":"2023-06-10T22:27:18.093095Z","shell.execute_reply.started":"2023-06-10T22:27:18.084628Z","shell.execute_reply":"2023-06-10T22:27:18.092191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nCreating train and val dataset using the train, val splits of the data.</div>","metadata":{}},{"cell_type":"code","source":"# Create DataLoader objects for training and validation data\ntrain_dataset = LandmarkDataset(train_data, transform=image_transforms)\nval_dataset = LandmarkDataset(val_data, transform=image_transforms)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:18.094436Z","iopub.execute_input":"2023-06-10T22:27:18.095054Z","iopub.status.idle":"2023-06-10T22:27:18.103347Z","shell.execute_reply.started":"2023-06-10T22:27:18.095022Z","shell.execute_reply":"2023-06-10T22:27:18.102385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nCreating dataloaders using the train and val dataset</div>","metadata":{}},{"cell_type":"code","source":"train_loader = DataLoader(train_dataset, batch_size=1, shuffle=True)\nval_loader = DataLoader(val_dataset, batch_size=1)","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:18.104885Z","iopub.execute_input":"2023-06-10T22:27:18.105252Z","iopub.status.idle":"2023-06-10T22:27:18.114351Z","shell.execute_reply.started":"2023-06-10T22:27:18.105223Z","shell.execute_reply":"2023-06-10T22:27:18.113459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nAltering the configuration file of CTRL as per our need.</div>","metadata":{}},{"cell_type":"code","source":"config = CTRLConfig.from_pretrained('ctrl')\nconfig.num_labels = 1  # Number of output labels (landmark or not)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:18.118348Z","iopub.execute_input":"2023-06-10T22:27:18.118672Z","iopub.status.idle":"2023-06-10T22:27:19.289538Z","shell.execute_reply.started":"2023-06-10T22:27:18.118642Z","shell.execute_reply":"2023-06-10T22:27:19.288626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\n\nThe code snippet defines a custom model called `CustomCTRLModel` that extends the `CTRLForSequenceClassification` class from the Transformers library.<br>\n\n1. `class CustomCTRLModel(CTRLForSequenceClassification)`: This line defines the `CustomCTRLModel` class that inherits from `CTRLForSequenceClassification`. This allows you to customize and extend the functionality of the base model.<br>\n\n2. `def __init__(self, config)`: This is the constructor method of the custom model. It takes a `config` object as an argument, which contains the configuration parameters for the model. Inside the constructor, the `super().__init__(config)` line calls the constructor of the base class to initialize the model using the provided configuration.<br>\n\n3. `self.image_embedding = nn.Linear(image_feature_size, config.hidden_size)`: This line creates a linear layer (`nn.Linear`) called `image_embedding`. It maps the `image_feature_size` to the `hidden_size` specified in the model's configuration.<br>\n\n4. `self.fusion = nn.Linear(config.hidden_size+1, config.hidden_size)`: This line creates another linear layer called `fusion`. It takes as input the concatenation of the `hidden_size` of the text outputs and the size of the image embeddings plus 1 (to account for the additional dimension introduced by `unsqueeze`). The output size is set to `hidden_size`.<br>\n\n5. `def forward(self, input_ids, inputs_embeds=None, image_embeds=None, **kwargs)`: This method defines the forward pass of the model. It takes input IDs, input embeddings, and image embeddings as arguments. The `**kwargs` parameter allows for additional keyword arguments that may be passed to the base model's forward method.<br>\n\n6. `text_outputs = super().forward(input_ids=input_ids, inputs_embeds=inputs_embeds, **kwargs)`: This line calls the `forward` method of the base class (`CTRLForSequenceClassification`) with the provided arguments. It computes the text outputs of the model.<br>\n\n7. `if image_embeds is not None:`: This conditional statement checks if image embeddings are provided as input.<br>\n\n8. `image_outputs = self.image_embedding(image_embeds)`: If image embeddings are provided, this line passes the image embeddings through the `image_embedding` linear layer to obtain `image_outputs`.<br>\n\n9. `combined_outputs = torch.cat((text_outputs[0], image_outputs.unsqueeze(0)), dim=1)`: This line concatenates the text outputs and the image outputs along the second dimension (`dim=1`). It uses `torch.cat` to concatenate the tensors. The image outputs are unsqueezed to add an additional dimension to match the shape of the text outputs.<br>\n\n10. `fused_outputs = self.fusion(combined_outputs)`: This line passes the concatenated outputs through the `fusion` linear layer to obtain the final fused outputs.<br>\n\n11. `return fused_outputs, text_outputs[1:]`: This line returns the fused outputs and the remaining text outputs (excluding the first element, which is the pooled output) as a tuple.<br>\n\n12. `else: return fused_outputs`: If image embeddings are not provided, this line simply returns the fused outputs.<br>\n\nThis custom model allows for the fusion of text and image features by concatenating their outputs and passing them through a linear layer for further processing.<br></div>","metadata":{}},{"cell_type":"code","source":"import torch.nn as nn\nimage_feature_size = 2048\nclass CustomCTRLModel(CTRLForSequenceClassification):\n    def __init__(self, config):\n        super().__init__(config)\n        self.image_embedding = nn.Linear(image_feature_size, config.hidden_size)\n        self.fusion = nn.Linear(config.hidden_size+1, config.hidden_size)\n\n    def forward(self, input_ids, inputs_embeds=None, image_embeds=None, **kwargs):\n        text_outputs = super().forward(input_ids=input_ids, inputs_embeds=inputs_embeds, **kwargs)\n\n        if image_embeds is not None:\n            image_outputs = self.image_embedding(image_embeds)\n            #rint(text_outputs[0].shape, image_outputs.shape)\n            combined_outputs = torch.cat((text_outputs[0], image_outputs.unsqueeze(0)), dim=1)\n            fused_outputs = self.fusion(combined_outputs)\n            return fused_outputs, text_outputs[1:]\n        else:\n            return fused_outputs\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:19.290879Z","iopub.execute_input":"2023-06-10T22:27:19.291764Z","iopub.status.idle":"2023-06-10T22:27:19.301074Z","shell.execute_reply.started":"2023-06-10T22:27:19.291732Z","shell.execute_reply":"2023-06-10T22:27:19.299872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nInitialising the tokenizer of CTRL</div>","metadata":{}},{"cell_type":"code","source":"tokenizer = CTRLTokenizer.from_pretrained('ctrl')","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:19.302348Z","iopub.execute_input":"2023-06-10T22:27:19.302850Z","iopub.status.idle":"2023-06-10T22:27:26.497899Z","shell.execute_reply.started":"2023-06-10T22:27:19.302819Z","shell.execute_reply":"2023-06-10T22:27:26.496917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThe  code snippet demonstrates how to utilize multiple GPUs for training by using DataParallel in PyTorch.<br>\n\n1. `model = CustomCTRLModel(config=config)`: This line creates an instance of the `CustomCTRLModel` by passing the configuration object `config` to its constructor. This assumes that you have already defined the model architecture and imported the necessary modules.<br>\n\n2. `if torch.cuda.device_count() > 1:`: This conditional statement checks if there are multiple GPUs available for training. The `torch.cuda.device_count()` function returns the number of available GPUs.<br>\n\n3. `model = DataParallel(model)`: If multiple GPUs are available, this line wraps the model with `DataParallel`. The `DataParallel` class is responsible for distributing the input batches across multiple GPUs and aggregating the gradients during backpropagation.<br>\n\n4. `device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")`: This line selects the device for training. If a GPU is available, it sets the device to CUDA (`\"cuda\"`), otherwise, it sets it to CPU (`\"cpu\"`).<br>\n\n5. `model.to(device)`: This line moves the model to the selected device. By calling the `to` method on the model and passing the device as an argument, the model's parameters and buffers are transferred to the specified device.<br>\n\nAfter executing this code snippet, the `model` is ready to be trained using either a single GPU or multiple GPUs, depending on the availability in the kaggle kernel.<br></div>","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nfrom torch.nn.parallel import DataParallel\n\n# Assuming you have already created your model\nmodel = CustomCTRLModel(config=config)\n\n# Check if multiple GPUs are available\nif torch.cuda.device_count() > 1:\n    # Create DataParallel model\n    model = DataParallel(model)\n\n# Move the model to the GPU(s)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel.to(device)","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:26.499313Z","iopub.execute_input":"2023-06-10T22:27:26.499887Z","iopub.status.idle":"2023-06-10T22:27:56.787085Z","shell.execute_reply.started":"2023-06-10T22:27:26.499853Z","shell.execute_reply":"2023-06-10T22:27:56.786155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nIt creates an instance of the AdamW optimizer by passing the model parameters (`model.parameters()`) and the learning rate (`lr=1e-5`) to its constructor. The `model.parameters()` function returns an iterator over all the trainable parameters of the model.<br>\n\nThe AdamW optimizer is a variant of the Adam optimizer that incorporates weight decay regularization. It is commonly used in deep learning for optimizing neural networks.</div>\n","metadata":{}},{"cell_type":"code","source":"optimizer = AdamW(model.parameters(), lr=1e-5)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:56.790408Z","iopub.execute_input":"2023-06-10T22:27:56.791352Z","iopub.status.idle":"2023-06-10T22:27:56.813093Z","shell.execute_reply.started":"2023-06-10T22:27:56.791319Z","shell.execute_reply":"2023-06-10T22:27:56.812311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nChecking the length of train_loader and val_loader to figure out the number of steps that will be needed to complete one epoch as per declared batch_size</div>","metadata":{}},{"cell_type":"code","source":"len(train_loader), len(val_loader)","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:56.817037Z","iopub.execute_input":"2023-06-10T22:27:56.819177Z","iopub.status.idle":"2023-06-10T22:27:56.829024Z","shell.execute_reply.started":"2023-06-10T22:27:56.819145Z","shell.execute_reply":"2023-06-10T22:27:56.827895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThe code snippet defines an image encoder using ResNet-50. The purpose of the image encoder is to extract image features from input images.\n1. `class ImageEncoder:`: This defines a custom class called `ImageEncoder` for the image encoder.<br>\n\n2. `def __init__(self):`: This is the constructor method of the `ImageEncoder` class. It initializes the ResNet-50 model (`self.resnet`) with pretrained weights and sets its fully connected layer (`self.resnet.fc`) to a `torch.nn.Identity()` layer, effectively removing the fully connected layer.<br>\n\n3. `self.transform = transforms.Compose([...])`: This defines a sequence of image transformations to be applied to the input images. It includes resizing the images to a fixed size of 224x224 pixels and normalizing the image pixels using mean and standard deviation values commonly used for pre-trained models.<br>\n\n4. `def __call__(self, images):`: This is the callable method of the `ImageEncoder` class. It takes input images as input and performs the image encoding process.<br>\n\n5. `images = self.transform(images)`: This applies the defined transformations to the input images.<br>\n\n6. `features = self.resnet(images)`: This passes the preprocessed images through the ResNet-50 model to obtain the image features. The output is a tensor of shape `(batch_size, num_features, 1, 1)`.<br>\n\n7. `return features.squeeze()`: This squeezes the tensor to remove the dimensions of size 1, resulting in a tensor of shape `(batch_size, num_features)`.<br>\n\nFinally, an instance of the `ImageEncoder` class is created as `image_encoder`, which can be used to extract image features by calling it with input images.<br></div>","metadata":{}},{"cell_type":"code","source":"import torch\nimport torchvision.models as models\nimport torchvision.transforms as transforms\n\n\n# Define the image encoder using ResNet-50\nclass ImageEncoder:\n    def __init__(self):\n        self.resnet = models.resnet50(pretrained=True).to(device)\n        self.resnet.fc = torch.nn.Identity()  # Remove the fully connected layer\n\n        self.transform = transforms.Compose([\n            transforms.Resize((224, 224)),\n            transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])\n        ])\n\n    def __call__(self, images):\n        images = self.transform(images)\n        features = self.resnet(images)\n        return features.squeeze()\n\n# Create an instance of the image encoder\nimage_encoder = ImageEncoder()\n\n","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:27:56.833493Z","iopub.execute_input":"2023-06-10T22:27:56.834616Z","iopub.status.idle":"2023-06-10T22:27:58.045214Z","shell.execute_reply.started":"2023-06-10T22:27:56.834586Z","shell.execute_reply":"2023-06-10T22:27:58.044244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\n\nThe code defines a training loop for a model using landmark images and text prompts. The steps that are executed in the code are -<br>\n\n1. `from tqdm import tqdm`: This imports the `tqdm` library, which provides a progress bar for iterations.<br>\n\n2. `num_epochs = 1`: This sets the number of training epochs.<br>\n\n3. `best_val_loss = float('inf')`: This initializes a variable to keep track of the best validation loss.<br>\n\n4. The code enters a loop that iterates over the specified number of epochs.<br>\n\n5. `model.train()`: This sets the model to training mode, enabling gradient computation and parameter updates.<br>\n\n6. `train_loss = 0.0`: This initializes the training loss to 0.<br>\n\n7. The loop iterates over the training data using `train_loader`, which presumably loads batches of images.<br>\n\n8. `images = images.to(device)`: This moves the input images to the appropriate device (e.g., GPU).<br>\n\n9. `input_ids = tokenizer.encode(...)`: This encodes a text prompt, such as \"Which landmark is depicted in this image?\", using a tokenizer. The resulting input_ids tensor is moved to the device.<br>\n\n10. `image_features = image_encoder(images)`: This extracts image features from the input images using the `image_encoder` object, which is an instance of the `ImageEncoder` class defined earlier.<br>\n\n11. `predictions = model(input_ids=input_ids, image_embeds=image_features)[0][0]`: This passes the input_ids and image features through the model to obtain landmark predictions. The predictions tensor is extracted for further processing.<br>\n\n12. `targets = torch.ones_like(predictions)`: This creates a tensor of the same shape as predictions, filled with ones. It assumes that all landmarks in the training data are present.<br>\n\n13. `loss = F.binary_cross_entropy_with_logits(predictions, targets)`: This computes the binary cross-entropy loss between the predictions and targets.<br>\n\n14. `optimizer.zero_grad()`: This zeroes the gradients of the model parameters.<br>\n\n15. `loss.backward()`: This performs backpropagation by computing gradients of the loss with respect to the model parameters.<br>\n\n16. `optimizer.step()`: This updates the model parameters based on the computed gradients.<br>\n\n17. The training loss is updated by adding the loss multiplied by the number of images in the batch.<br>\n\n18. After the inner training loop, the training loss is divided by the size of the training dataset to obtain the average training loss.<br>\n\n19. The code then enters the validation phase.<br>\n\n20. `model.eval()`: This sets the model to evaluation mode, disabling gradient computation and parameter updates.<br>\n\n21. `val_loss = 0.0`: This initializes the validation loss to 0.<br>\n\n22. The loop iterates over the validation data using `val_loader`, which presumably loads batches of validation images.<br>\n\n23. Similar to the training loop, the input images are moved to the appropriate device, and image features are extracted using the `image_encoder`.<br>\n\n24. The model is used to make predictions on the input_ids and image features, and the predictions tensor is extracted.<br>\n\n25. `targets = torch.zeros_like(predictions)`: This creates a tensor of the same shape as predictions, filled with zeros. It assumes that no landmarks are present in the validation data.<br>\n\n26. The validation loss is computed using binary cross-entropy with logits.<br>\n\n27. The validation loss is updated by adding the loss multiplied by the number of images in the batch.<br>\n\n28. After the inner validation loop, the validation loss is divided by the size of the validation dataset to obtain the average validation loss.<br>\n\n29. If the current validation loss is lower than the previous best validation loss, the model's state_dict is saved to a file called '\n\nbest_model.pth'. This allows later loading of the model with the best performance.<br>\n\n30. The training and validation loss for the current epoch are printed.<br>\n\n</div>","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nimport torch.nn.functional as F\nnum_epochs = 1\nbest_val_loss = float('inf')\n\nfor epoch in range(num_epochs):\n    model.train()\n    train_loss = 0.0\n\n    for step, images in tqdm(enumerate(train_loader)):\n        images = images.to(device)\n\n        # Generate landmark predictions using CTRL model\n        input_ids = tokenizer.encode(\"Which landmark is depicted in this image?\", add_special_tokens=True, return_tensors=\"pt\").to(device)\n        image_features = image_encoder(images)  # Extract image features using a pre-trained CNN\n        image_features = image_features.to(device)\n        input_ids = input_ids.to(device)\n        predictions = model(input_ids=input_ids, image_embeds=image_features)[0][0]\n\n        # Compute loss and perform backpropagation\n        targets = torch.ones_like(predictions)  # Assume all landmarks in training data\n        loss = F.binary_cross_entropy_with_logits(predictions, targets)\n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        if step % 1 == 0:\n            print(\"Step-{}, Loss-{}\".format(step, loss.item()))\n            break\n\n        train_loss += loss.item() * images.size(0)\n\n    train_loss /= len(train_loader.dataset)\n    break\n    # Validation\n    model.eval()\n    val_loss = 0.0\n\n    with torch.no_grad():\n        for images in val_loader:\n            images = images.to(device)\n\n            input_ids = tokenizer.encode(\"Which landmark is depicted in this image?\", add_special_tokens=True, return_tensors=\"pt\").to(device)\n            image_features = image_encoder(images)  # Extract image features using a pre-trained CNN\n            image_features = image_features.to(device)\n            input_ids = input_ids.to(device)\n            predictions = model(input_ids=input_ids, image_embeds=image_features)[0][0]\n\n            targets = torch.zeros_like(predictions)\n            loss = F.binary_cross_entropy_with_logits(predictions, targets)\n            val_loss += loss.item() * images.size(0)\n\n    val_loss /= len(val_loader.dataset)\n\n    # Save the model with the best validation loss\n    if val_loss < best_val_loss:\n        torch.save(model.module.state_dict(), 'best_model.pth')  # Save the module's state_dict for DataParallel model\n        best_val_loss = val_loss\n\n    print(f'Epoch {epoch+1}/{num_epochs}, Train Loss: {train_loss:.4f}, Val Loss: {val_loss:.4f}')","metadata":{"execution":{"iopub.status.busy":"2023-06-10T22:29:12.585505Z","iopub.execute_input":"2023-06-10T22:29:12.586174Z","iopub.status.idle":"2023-06-10T22:29:13.095265Z","shell.execute_reply.started":"2023-06-10T22:29:12.586143Z","shell.execute_reply":"2023-06-10T22:29:13.093759Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nThe code is used for generating predictions on the test data using the trained model. The execution of the code can be broken down into the following steps<br>\n\n1. `test_folder_path`: This variable specifies the path to the test folder containing the image files.<br>\n\n2. `test_files = os.listdir(test_folder_path)`: This retrieves the list of image files in the test folder.<br>\n\n3. `predictions_final = []`: This initializes an empty list to store the final predictions.<br>\n\n4. `model.eval()`: This sets the model to evaluation mode.<br>\n\n5. The code enters a loop that iterates over the files in the test folder using `os.walk(test_folder_path)`.<br>\n\n6. Within the loop, the code checks if the file has the '.jpg' extension to ensure it's an image file.<br>\n\n7. `image_path = os.path.join(root, file)`: This creates the full path to the image file.<br>\n\n8. `image = Image.open(image_path).convert('RGB')`: This opens the image file and converts it to RGB mode using the PIL library.<br>\n\n9. `image = image_transforms(image).unsqueeze(0).to(device)`: This applies the image transformations defined earlier (`image_transforms`) to preprocess the image for model input. It then unsqueezes the image tensor to add a batch dimension and moves it to the appropriate device.<br>\n\n10. `image_features = image_encoder(image)`: This extracts image features from the preprocessed image using the `image_encoder` object defined earlier.<br>\n\n11. `input_ids = tokenizer.encode(...)`: This encodes the text prompt, \"Which landmark is depicted in this image?\", using the tokenizer. The resulting input_ids tensor is moved to the device.<br>\n\n12. `batch_size = image.size(0)`: This retrieves the batch size, which is 1 in this case.<br>\n\n13. `input_ids = input_ids.expand(batch_size, -1)`: This expands the input_ids tensor to match the batch size, effectively duplicating the input_ids for each image in the batch.<br>\n\n14. `predictions = model(input_ids=input_ids, image_embeds=image_features)[0][0]`: This passes the input_ids and image features through the model to obtain landmark predictions. The predictions tensor is extracted for further processing.<br>\n\n15. `probs = torch.sigmoid(predictions)`: This applies the sigmoid function to the predictions to obtain probabilities.<br>\n\n16. `confidence_scores = probs.tolist()`: This converts the probability tensor to a Python list.<br>\n\n17. `image_id = os.path.splitext(file)[0]`: This retrieves the image ID by removing the file extension from the file name.<br>\n\n18. `prediction = f\"{image_id},\"`: This initializes the prediction string with the image ID.<br>\n\n19. If there are confidence scores available:<br>\n\n    a. `landmark_id = confidence_scores.index(max(confidence_scores))`: This finds the index of the maximum confidence score, representing the predicted landmark ID.<br>\n    \n    b. `confidence_score = max(confidence_scores)`: This retrieves the maximum confidence score.<br>\n    \n    c. `prediction += f\"{landmark_id} {confidence_score}\"`: This appends the predicted landmark ID and confidence score to the prediction string.<br>\n    \n20. `predictions_final.append(prediction)`: This adds the prediction string to the `predictions_final` list.<br>\n\nThe code iterates over all the image files in the test folder, generates predictions using the trained model, and stores the predictions in the `predictions_final` list.<br></div>","metadata":{}},{"cell_type":"code","source":"import os\n\ntest_folder_path = '/kaggle/input/landmark-recognition-2021/test/'\n\n# Get the list of image files in the test folder\ntest_files = os.listdir(test_folder_path)\n\npredictions_final = []\n\n#model.load_state_dict(torch.load('best_model.pth'))\nmodel.eval()\n\nwith torch.no_grad():\n    for root, dirs, files in tqdm(os.walk(test_folder_path)):\n        for file in files:\n            if file.endswith('.jpg'):\n                image_path = os.path.join(root, file)\n                image = Image.open(image_path).convert('RGB')\n                image = image_transforms(image).unsqueeze(0).to(device)\n                # Generate image features using the image encoder\n                image_features = image_encoder(image)\n                # Generate landmark predictions using CTRL model\n                input_ids = tokenizer.encode(\"Which landmark is depicted in this image?\", add_special_tokens=True, return_tensors=\"pt\").to(device)\n                batch_size = image.size(0)\n                input_ids = input_ids.expand(batch_size, -1)  # Expand input_ids to match the batch size\n\n                predictions = model(input_ids=input_ids, image_embeds=image_features)[0][0]\n                \n                probs = torch.sigmoid(predictions)\n                confidence_scores = probs.tolist()\n                # Add the prediction to the list\n                image_id = os.path.splitext(file)[0]\n                prediction = f\"{image_id},\"\n                if confidence_scores:\n                    landmark_id = confidence_scores.index(max(confidence_scores))\n                    confidence_score = max(confidence_scores)\n                    prediction += f\"{landmark_id} {confidence_score}\"\n                predictions_final.append(prediction)\n\n\n#intentional keyboard interrupt as the inference time is high!","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-06-10T22:29:14.716015Z","iopub.execute_input":"2023-06-10T22:29:14.716368Z","iopub.status.idle":"2023-06-10T22:29:34.611311Z","shell.execute_reply.started":"2023-06-10T22:29:14.716339Z","shell.execute_reply":"2023-06-10T22:29:34.609608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:#A7C6F5; color:#19180F; font-size:15px; font-family:Verdana; padding:10px; border: 5px solid #19180F;\">\nSanity check of the predictions_final df,Desired output which is ready to be written to submission.csv</div>","metadata":{}},{"cell_type":"code","source":"predictions_final","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-06-10T22:29:56.956369Z","iopub.execute_input":"2023-06-10T22:29:56.957061Z","iopub.status.idle":"2023-06-10T22:29:56.972261Z","shell.execute_reply.started":"2023-06-10T22:29:56.957027Z","shell.execute_reply":"2023-06-10T22:29:56.971075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}