{"metadata":{"kernelspec":{"display_name":"base","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.10"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31153,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Problem description\nThe goal of this project is to develop an image classification model to detect cancerous cells in histopathologic samples. This is a computer vision problem applied to medicine, aiming to assist in early cancer detection through automated image analysis.\n\n##### Project Objectives\n\n- Train and compare different convolutional neural network (CNN) architectures for binary classification.\n\n- Evaluate model performance using AUC and other metrics like Accuracy.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport copy\nimport os\nimport torch\nfrom PIL import Image\nfrom torch.utils.data import Dataset\nimport torchvision.transforms as transforms\nfrom torch.utils.data import random_split\nfrom torch.optim.lr_scheduler import ReduceLROnPlateau\nimport torch.nn as nn\n%matplotlib inline","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data description\nThe data is donwloaded from https://www.kaggle.com/competitions/histopathologic-cancer-detection/data competition. It is composed by the following files and folders:\n- Test folder: This contains the images to classify for the challenge\n- Train folder: This contains the images to train the model\n- sample_submission.csv: Is the file as example of the sumbission that we need to upload\n- tran_labels.csv: The file containing the label for each image in the training folder\n\nLet's explore each of them:","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv('histo_data/train_labels.csv')\nprint(train_labels.head())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Train shape:')\nprint(train_labels.shape)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The train dataset is composed for 220K samples, let's see if we have duplicated id's","metadata":{}},{"cell_type":"code","source":"len(train_labels['id'].unique())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Now let's explore the test folder\nfrom pathlib import Path\n\nfolder_path = Path(\"histo_data/test\")\n\ntest_ids = []\n\nfor f in folder_path.iterdir():\n    if f.is_file():\n        test_ids.append(f.name)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Test shape:')\nprint(len(test_ids))\nprint('Test unique ids (Check for duplicates)')\nprint(len(np.unique(test_ids)))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_submission = pd.read_csv('histo_data/sample_submission.csv')\nsample_submission.head()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Sample submission shape:')\nprint(sample_submission.shape)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can see that the sample submission and the test data is the same lenght, which makes sense.","metadata":{}},{"cell_type":"code","source":"test_ids[0:10]","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Also, as we can see the format of images is tif, let's explore few of them:","metadata":{}},{"cell_type":"code","source":"image_path = f'histo_data/test/{test_ids[0]}'\nimg = Image.open(image_path)\nplt.imshow(img)\nplt.axis('off')\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Image shape:')\nnp.array(img).shape","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The images are 96x96 and 3 channels","metadata":{}},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"First, let's see the target distribution, in order to see if we have an imbalanced problem.","metadata":{}},{"cell_type":"code","source":"train_labels['label'].value_counts() / len(train_labels)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(5,3))\nsns.countplot(data=train_labels, x='label', palette=\"Set2\")\n\nplt.xlabel(\"Cancer\")\nplt.ylabel(\"Frequency\")\nplt.title(\"Distriuction of the target variable (Cancer)\")\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The target is a little bit imbalanced as we have fewer cancer images, so we need to take care of this in the processing step","metadata":{}},{"cell_type":"markdown","source":"Now I'll take a look at the images between cancer vs non-cancer.","metadata":{}},{"cell_type":"code","source":"cancer_sample_images = train_labels[train_labels['label'] == 1].sample(9)['id'].values\nnon_cancer_sample_images = train_labels[train_labels['label'] == 0].sample(9)['id'].values","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(8, 8))\n\nfor i, img_name in enumerate(cancer_sample_images):\n    img_path = f'histo_data/train/{img_name}.tif'\n    img = Image.open(img_path)\n    \n    plt.subplot(3, 3, i+1)  # 3x3 grid\n    plt.imshow(img)\n    plt.axis('off')\n\nplt.suptitle('Sample of Cancer images', fontsize=16)\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(8, 8))\n\nfor i, img_name in enumerate(non_cancer_sample_images):\n    img_path = f'histo_data/train/{img_name}.tif'\n    img = Image.open(img_path)\n    \n    plt.subplot(3, 3, i+1)  # 3x3 grid\n    plt.imshow(img)\n    plt.axis('off')\n\nplt.suptitle('Sample of Benign images')\nplt.tight_layout()\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It's interesting to see that the cancer images looks that the cells follow a more irregular patron compared to the benign ones, and they are bigger.","metadata":{}},{"cell_type":"markdown","source":"Based on the EDA I'll do the following:\n- Sample the data to balance the classes\n- Normalize the images to have a more stable traing process\n- Use only a fraction of training samples, to speed up the training\n- Use data augmentations techniques, this is typicaly done in computer vision problems","metadata":{}},{"cell_type":"markdown","source":"### Data preparation\nIn this section I'll partition the data, use the data augmentation techniques, sample it, and so on.","metadata":{}},{"cell_type":"code","source":"train_labels['label'].value_counts()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Firts, let's sample the data to reduce it and balancing it.\nsample_size = 80000\nrandom_state = 102\ntrain_labels_1 = train_labels[train_labels['label'] == 1].sample(\n    int(sample_size/2), random_state=random_state)\ntrain_labels_0 = train_labels[train_labels['label'] == 0].sample(\n    int(sample_size/2), random_state=random_state)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels_sample = pd.concat([train_labels_1, train_labels_0]).reset_index(drop=True)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels_sample['label'].value_counts()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have a balanced data to train, let's prepare it.\n\nIn pytorch you need to create a class to load your data and prepare it for training","metadata":{}},{"cell_type":"code","source":"from torch.utils.data import Dataset, DataLoader, random_split","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class TrainData(Dataset):\n    def __init__(self, train_labels_sample, img_dir, transform=None):\n        '''\n        train_labels_sample: The training samples id's and their label.\n        img_dir: the path in which the images are located.\n        transform: The transformation to be applied to the images. (Resize, augmentation, normalization)\n        '''\n        self.data = train_labels_sample\n        self.img_dir = img_dir\n        self.transform = transform\n\n    def __len__(self):\n        ''' \n        This is the typical template in Pytorch, \n        it needs a __len__ method and __getitem__, to obtain the input and label\n        based on index.\n        '''\n        return len(self.data)\n\n    def __getitem__(self, idx):\n        img_name = self.data.iloc[idx, 0] + \".tif\"\n        label = int(self.data.iloc[idx, 1])\n        img_path = os.path.join(self.img_dir, img_name)\n\n        image = Image.open(img_path)\n\n        if self.transform:\n            image = self.transform(image)\n\n        return image, label\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can learn more about image transfornations in Pytorch in this link:\nhttps://docs.pytorch.org/vision/0.22/auto_examples/transforms/plot_transforms_illustrations.html#sphx-glr-auto-examples-transforms-plot-transforms-illustrations-py","metadata":{}},{"cell_type":"code","source":"# 64x64\ntrain_transforms = transforms.Compose([\n    transforms.Resize((96,96)), # Resize the image to ensure that all have 96x96 shape\n    transforms.RandomHorizontalFlip(), # Performs horizonal flip in image randomly\n    transforms.RandomVerticalFlip(), # Performs vertical flip in image randomly\n    transforms.RandomRotation(20), # Rotates the image randomly between (-20, 20) degrees\n    transforms.ColorJitter(\n        brightness=0.2,\n        contrast=0.2, \n        saturation=0.2, hue=0.05), # transform randomly changes the brightness, contrast, saturation, hue\n    transforms.ToTensor(),\n    transforms.Normalize([0.5,0.5,0.5], [0.5,0.5,0.5])\n])\n\n# This is for validation/testing data, as it doesn't require to do the part \n# of data augmentation\nval_transforms = transforms.Compose([\n    transforms.Resize((96,96)),\n    transforms.ToTensor(),\n    transforms.Normalize([0.5,0.5,0.5], [0.5,0.5,0.5])\n])","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset = TrainData(train_labels_sample, img_dir=\"histo_data/train\", transform=train_transforms)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Train, validation split","metadata":{}},{"cell_type":"code","source":"train_size = int(0.8 * len(dataset))\nval_size = len(dataset) - train_size\ntrain_dataset, val_dataset = random_split(dataset, [train_size, val_size])\n\n# For the validation data only use the parte that normalizes and resize the images, not the augmentation\nval_dataset.dataset.transform = val_transforms","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Train data size:')\nprint(len(train_dataset))\nprint('Validation data size:')\nprint(len(val_dataset))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# This is one of the hyperparameters, it controls the batch size in which the data will be partitioned\nbatch_size = 32\n# Now, let's pass each dataset to a DataLoader Pytorch class, this is in charge\n# of suffling, creating the batch, and can do many other things.\ntrain_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=True)\nval_loader = DataLoader(val_dataset, batch_size=batch_size, shuffle=False)\n\n\n\nimages, labels = next(iter(train_loader))\nprint(f\"Train batch images: {images.shape}, labels: {labels.shape}\")\n\nimages, labels = next(iter(val_loader))\nprint(f\"Validation batch images: {images.shape}, labels: {labels.shape}\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Model architecture","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.nn.functional as F","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Firts, I'll use this simple model architecture in order to see how it performs.\nThe architecture is a typical one, except that it has normalization in each layer\nand max pooling.\n\nIt's simpler than VGNet, as it applies pooling after ecah Convolution, not two.\n","metadata":{}},{"cell_type":"code","source":"class SimpleCNN(nn.Module):\n    def __init__(self):\n        super(SimpleCNN, self).__init__()\n        # Convolution 1\n        self.conv1 = nn.Conv2d(in_channels=3, out_channels=32, kernel_size=3, padding=1)\n        # 32 is the number of filters, and kernel size the filter size\n        self.bn1 = nn.BatchNorm2d(32)\n        \n        # Convolution 2\n        self.conv2 = nn.Conv2d(32, 64, 3, padding=1)\n        self.bn2 = nn.BatchNorm2d(64)\n        \n        # Convolution 3\n        self.conv3 = nn.Conv2d(64, 128, 3, padding=1)\n        self.bn3 = nn.BatchNorm2d(128)\n        \n        # Pooling\n        self.pool = nn.MaxPool2d(2, 2)\n        \n        # Fully connected\n        self.fc1 = nn.Linear(128 * 12 * 12, 256)  # 96x96 -> 3 poolings => 12x12\n        self.dropout = nn.Dropout(0.5)\n        self.fc2 = nn.Linear(256, 1)  # 1 for binary classification\n\n    def forward(self, x):\n        x = self.pool(F.relu(self.bn1(self.conv1(x))))  # 96x96 -> 48x48\n        x = self.pool(F.relu(self.bn2(self.conv2(x))))  # 48x48 -> 24x24\n        x = self.pool(F.relu(self.bn3(self.conv3(x))))  # 24x24 -> 12x12\n        x = x.view(-1, 128 * 12 * 12)\n        x = F.relu(self.fc1(x))\n        x = self.dropout(x)\n        x = self.fc2(x)\n        return x  # Whithout sigmoid becase we'll use BCEWithLogitsLoss\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, let's try a more complicated architure, in which I'll try 2 convolutions, then\na MaxPooling and a Dropout for trying to reduxe overfitting. And this will be repeated 3 times, so, 3 blocks and finally a fully conected layer.\n\nI'll use ReLU as activation function for all the layers, as it works well in this type of problems.\nand 3x3 filters (kernel size).","metadata":{}},{"cell_type":"code","source":"class CNNAdvanced(nn.Module):\n    def __init__(self):\n        super(CNNAdvanced, self).__init__()\n\n        # First block\n        self.block1 = nn.Sequential(\n            nn.Conv2d(in_channels=3, out_channels=32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(32, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(32, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2, 2),\n            nn.Dropout(0.4)\n        )\n\n        # Second block\n        self.block2 = nn.Sequential(\n            nn.Conv2d(32, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2, 2),\n            nn.Dropout(0.4)\n        )\n\n        # Third block\n        self.block3 = nn.Sequential(\n            nn.Conv2d(64, 128, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(128, 128, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(128, 128, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2, 2),\n            nn.Dropout(0.4)\n        )\n\n        # Fully connected layer\n        # height and weight size after 3 pools: 12x12\n        # that's because we have 3 blocks, and each block has a MaxPool of (2,2): 96 -> 48 -> 24 -> 12\n        self.flatten_dim = 128 * 12 * 12\n        # And note that 128 is the depth size of the 3 block\n        self.fc = nn.Sequential(\n            nn.Linear(self.flatten_dim, 256),\n            nn.ReLU(),\n            nn.Dropout(0.3),\n            nn.Linear(256, 1)  # 1 for binary classification\n        )\n\n    def forward(self, x):\n        x = self.block1(x)\n        x = self.block2(x)\n        x = self.block3(x)\n        x = x.view(-1, self.flatten_dim)  # flatten\n        x = self.fc(x)\n        return x # Whithout sigmoid becase we'll use BCEWithLogitsLoss","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nsimple_model = SimpleCNN().to(device)\nprint('Simple model architecture:')\nprint(simple_model)\nprint('*'*100)\n\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Advanced model architecture:')\nadvanced_model = CNNAdvanced().to(device)\nprint(advanced_model)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Optimizer and Loss function","metadata":{}},{"cell_type":"code","source":"import torch.optim as optim\n# For binarry classifier, and as I mentioned earlier, this function applies sigmoid while evaluationg the Loss.\nloss_function_simple_model = nn.BCEWithLogitsLoss()\nloss_function_advanced_model= nn.BCEWithLogitsLoss() \n\n\n# I'll use the Adam optimizer because is good optimization method.\noptimizer_simple_model = optim.Adam(simple_model.parameters(), lr=0.0001)\noptimizer_advanced_model = optim.Adam(advanced_model.parameters(), lr=0.0001)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### Scheduler\nThis is for adjusting the learning reate after each epoch, in order to obtain better results, as it can happen that the parameters are arriving to an optimal value and the learning rate needs to be reduced.","metadata":{}},{"cell_type":"code","source":"scheduler_simple_model = optim.lr_scheduler.ReduceLROnPlateau(\n    optimizer_simple_model, mode='max', factor=0.5, patience=2, min_lr=1e-5\n)\n\nscheduler_advanced_model = optim.lr_scheduler.ReduceLROnPlateau(\n    optimizer_advanced_model, mode='max', factor=0.5, patience=2, min_lr=1e-5\n)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Training\nIn order to see the performance after each epoch I'll use the following metrics:\n- Accuracy\n- AUC: This is a good metric for binary classification models, as it does't depend on threshold, so it's more robust.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\n\ndef training_model(num_epochs, optimizer, model, loss_function, scheduler, model_save_path):\n    # This is the value to store the best validation AUC and save the best model\n    best_val_auc = -1.0\n    # Iterate all the data in each epoch\n    for epoch in range(num_epochs):\n        model.train()\n        running_loss = 0.0\n        # In this part, it iterates over each batch\n        for images, labels in train_loader:\n            images, labels = images.to(device), labels.float().unsqueeze(1).to(device)\n\n            # Set gradients to zero, because Pytorch acumulates the previous ones.\n            optimizer.zero_grad()\n            outputs = model(images)\n            # calculate the loss\n            loss = loss_function(outputs, labels)\n            # Calculate the gradients respect to the loss\n            loss.backward()\n            # Finally update the parameters based on the gradients calculated with loss.backward()\n            optimizer.step()\n\n            # Calculate the loss\n            running_loss += loss.item() * images.size(0)\n        \n        epoch_loss = running_loss / len(train_loader.dataset)\n        print(f\"Epoch {epoch+1}/{num_epochs}, Training Loss: {epoch_loss:.4f}\")\n        \n        # Validation\n        model.eval()\n        val_loss = 0.0\n        correct = 0\n        # here I'll save the labels and prediction for each batch in the validation set\n        all_labels = []\n        all_probs = []\n\n        with torch.no_grad():\n            for images, labels in val_loader:\n                images, labels = images.to(device), labels.float().unsqueeze(1).to(device)\n                outputs = model(images)\n                loss = loss_function(outputs, labels)\n                val_loss += loss.item() * images.size(0)\n                \n                probs = torch.sigmoid(outputs)\n                preds = probs >= 0.5\n                correct += (preds.float() == labels).sum().item()\n                \n                all_labels.append(labels.cpu())\n                all_probs.append(probs.cpu())\n\n        val_loss = val_loss / len(val_loader.dataset)\n        val_acc = correct / len(val_loader.dataset)\n\n        # Concat for all the batches\n        all_labels = torch.cat(all_labels).numpy()\n        all_probs = torch.cat(all_probs).numpy()\n        val_auc = roc_auc_score(all_labels, all_probs)\n\n        print(f\"Validation Loss: {val_loss:.4f}, Accuracy: {val_acc:.4f}, AUC: {val_auc:.4f}\")\n        if val_auc > best_val_auc:\n            best_val_auc = val_auc\n            # Save the best model parameters\n            torch.save(model.state_dict(), model_save_path)\n            print(f\"Best model saved with validation AUC: {best_val_auc:.4f}\")\n\n        # Update the learning rate with the scheduler\n        scheduler.step(val_auc) # Based on AUC, maximize the AUC\n\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, lets train the first simple model","metadata":{}},{"cell_type":"code","source":"training_model(\n    num_epochs=15, optimizer=optimizer_simple_model,\n    model=simple_model, loss_function=loss_function_simple_model,\n    scheduler=scheduler_simple_model,\n    model_save_path='models/best_simple_cnn_model.pth')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's load the best simple model","metadata":{}},{"cell_type":"code","source":"simple_model = SimpleCNN().to(device)\nsimple_model.load_state_dict(torch.load(\"models/best_simple_cnn_model.pth\"))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's plot the train and validation AUC","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\nimport matplotlib.pyplot as plt","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def calculate_probability_of_cancer(model, data_loader):\n    '''\n    Model: The model to calculate the probability\n    data_loader: The data batches, train or validation\n    '''\n    all_labels = []\n    all_probs = []\n    with torch.no_grad():\n        for images, labels in data_loader:\n            images, labels = images.to(device), labels.float().unsqueeze(1).to(device)\n            outputs = model(images)\n            # This is necessesary because the model outputs the raw after \n            # the conected layer, no the probability\n            probs = torch.sigmoid(outputs)\n            all_labels.append(labels.cpu())\n            all_probs.append(probs.cpu())\n    all_labels = torch.cat(all_labels).numpy()\n    all_probs = torch.cat(all_probs).numpy()\n    return all_labels, all_probs","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels, train_probs = calculate_probability_of_cancer(\n    simple_model, train_loader)\nval_labels, val_probs = calculate_probability_of_cancer(\n    simple_model, val_loader)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_fpr, train_tpr, _ = roc_curve(train_labels, train_probs)\nval_fpr, val_tpr, _ = roc_curve(val_labels, val_probs)\n\ntrain_auc = auc(train_fpr, train_tpr)\nval_auc = auc(val_fpr, val_tpr)\n\n# Plot the AUC\nplt.figure(figsize=(8,6))\nplt.plot(train_fpr, train_tpr, label=f'Train AUC = {train_auc:.4f})')\nplt.plot(val_fpr, val_tpr, label=f'Validation AUC = {val_auc:.4f})')\nplt.plot([0,1], [0,1], 'k--')  # diagonal\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Simple model ROC Curves')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For me looks a good model as it doesn't overfit in the training data compared to the validation set, and also has a great AUC.\n\nBut let's check with the advanced model archirecture.","metadata":{}},{"cell_type":"markdown","source":"Training the advanced model architecture","metadata":{}},{"cell_type":"code","source":"training_model(\n    num_epochs=15, optimizer=optimizer_advanced_model,\n    model=advanced_model, loss_function=loss_function_advanced_model,\n    scheduler=scheduler_advanced_model,\n    model_save_path='models/best_advanced_cnn_model.pth')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Load the best advanced model","metadata":{}},{"cell_type":"code","source":"advanced_model = CNNAdvanced().to(device)\nadvanced_model.load_state_dict(torch.load(\"models/best_advanced_cnn_model.pth\"))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels, train_probs = calculate_probability_of_cancer(\n    advanced_model, train_loader)\nval_labels, val_probs = calculate_probability_of_cancer(\n    advanced_model, val_loader)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_fpr, train_tpr, _ = roc_curve(train_labels, train_probs)\nval_fpr, val_tpr, _ = roc_curve(val_labels, val_probs)\n\ntrain_auc = auc(train_fpr, train_tpr)\nval_auc = auc(val_fpr, val_tpr)\n\n# Plot the AUC\nplt.figure(figsize=(8,6))\nplt.plot(train_fpr, train_tpr, label=f'Tran AUC = {train_auc:.4f})')\nplt.plot(val_fpr, val_tpr, label=f'Validation AUC = {val_auc:.4f})')\nplt.plot([0,1], [0,1], 'k--')  # diagonal\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Advanced model ROC Curves')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This model looks better in the following terms:\n- More validation AUC\n- Less discrepancy between train/val AUC\n\nSo, I'll choose the advanced architecture to do parameter tuning.","metadata":{}},{"cell_type":"markdown","source":"# Hyperparameter tuning\nHere, I'll use the simpler model architecture and tune the following hyperparameters:\n- Dropout propotion\n- Learning rate\n- Depth in each convolution layer\n- Units in the fully conected layer, the last one to prediction\n\n\nI'll be using optuna is an open-source Python library designed for automatic hyperparameter optimization.\nYou can learn more about Optuna here: \nhttps://optuna.readthedocs.io/en/stable/","metadata":{}},{"cell_type":"code","source":"import optuna","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def objective(trial):\n\n    # Hyperparameter space\n    first_filters = trial.suggest_categorical('first_filters', [16, 32, 64])\n    second_filters = trial.suggest_categorical('second_filters', [32, 64, 128])\n    third_filters = trial.suggest_categorical('third_filters', [64, 128, 256])\n    fc1_units = trial.suggest_categorical('fc1_units', [128, 256, 512])\n    dropout_rate = trial.suggest_float('dropout_rate', 0.3, 0.6)\n    dropout_rate_fc_layer = trial.suggest_float('dropout_rate_fc_layer', 0.3, 0.6)\n    lr = trial.suggest_float('lr', 1e-4, 1e-2, log=True)\n\n    # Model definition based on hyperparameters\n    class TunedCNN(nn.Module):\n        def __init__(self):\n            super(TunedCNN, self).__init__()\n\n            # First block\n            self.block1 = nn.Sequential(\n                nn.Conv2d(in_channels=3, out_channels=first_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.Conv2d(first_filters, first_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.Conv2d(first_filters, first_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.MaxPool2d(2, 2),\n                nn.Dropout(dropout_rate)\n            )\n\n            # Second block\n            self.block2 = nn.Sequential(\n                nn.Conv2d(first_filters, second_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.Conv2d(second_filters, second_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.Conv2d(second_filters, second_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.MaxPool2d(2, 2),\n                nn.Dropout(dropout_rate)\n            )\n\n            # Third block\n            self.block3 = nn.Sequential(\n                nn.Conv2d(second_filters, third_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.Conv2d(third_filters, third_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.Conv2d(third_filters, third_filters, kernel_size=3, padding=1),\n                nn.ReLU(),\n                nn.MaxPool2d(2, 2),\n                nn.Dropout(dropout_rate)\n            )\n\n            # Fully connected layer\n            # height and weight size after 3 pools: 12x12\n            # that's because we have 3 blocks, and each block has a MaxPool of (2,2): 96 -> 48 -> 24 -> 12\n            self.flatten_dim = third_filters * 12 * 12\n            self.fc = nn.Sequential(\n                nn.Linear(self.flatten_dim, fc1_units),\n                nn.ReLU(),\n                nn.Dropout(dropout_rate_fc_layer),\n                nn.Linear(fc1_units, 1)  # 1 for binary classification\n            )\n\n        def forward(self, x):\n            x = self.block1(x)\n            x = self.block2(x)\n            x = self.block3(x)\n            x = x.view(-1, self.flatten_dim)  # flatten\n            x = self.fc(x)\n            return x # Whithout sigmoid becase we'll use BCEWithLogitsLoss\n    \n    # Define the model, loss and optimizer\n    model = TunedCNN().to(device)\n    optimizer = optim.Adam(model.parameters(), lr=lr)\n    loss_fn = nn.BCEWithLogitsLoss()\n\n    # Training the model with only 3 epochs\n    for epoch in range(3):\n        model.train()\n        for images, labels in train_loader:\n            images, labels = images.to(device), labels.float().unsqueeze(1).to(device)\n            optimizer.zero_grad()\n            outputs = model(images)\n            loss = loss_fn(outputs, labels)\n            loss.backward()\n            optimizer.step()\n\n    # Validation set\n    # Here we calculate the AUC in validation after the 3 epochs\n    model.eval()\n    all_labels, all_probs = [], []\n    with torch.no_grad():\n        for images, labels in val_loader:\n            images, labels = images.to(device), labels.float().unsqueeze(1).to(device)\n            outputs = model(images)\n            probs = torch.sigmoid(outputs)\n            all_labels.append(labels.cpu())\n            all_probs.append(probs.cpu())\n\n    all_labels = torch.cat(all_labels).numpy()\n    all_probs = torch.cat(all_probs).numpy()\n    val_auc = roc_auc_score(all_labels, all_probs)\n\n    return val_auc  # Maximize AUC\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Use optuna for hyperparameter tuning\nstudy = optuna.create_study(direction='maximize')\nstudy.optimize(objective, n_trials=20)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can see that there are hyperparameters combinations that doesn't work well, so the valid AUC is 0.5","metadata":{}},{"cell_type":"code","source":"print(\"Best hyperparameters:\", study.best_params)\nprint(\"Best AUC:\", study.best_value)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have the best hyperparameters, let's define the model architecture with those and train it with more epochs.","metadata":{}},{"cell_type":"code","source":"class TunedCNN(nn.Module):\n    def __init__(self):\n        super(TunedCNN, self).__init__()\n\n        # First block\n        self.block1 = nn.Sequential(\n            nn.Conv2d(in_channels=3, out_channels=64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2, 2),\n            nn.Dropout(0.5106408629195538)\n        )\n\n        # Second block\n        self.block2 = nn.Sequential(\n            nn.Conv2d(64, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(32, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(32, 32, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2, 2),\n            nn.Dropout(0.5106408629195538)\n        )\n\n        # Third block\n        self.block3 = nn.Sequential(\n            nn.Conv2d(32, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.Conv2d(64, 64, kernel_size=3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2, 2),\n            nn.Dropout(0.5106408629195538)\n        )\n\n        # Fully connected layer\n        # height and weight size after 3 pools: 12x12\n        # that's because we have 3 blocks, and each block has a MaxPool of (2,2): 96 -> 48 -> 24 -> 12\n        self.flatten_dim = 64 * 12 * 12\n        # And note that 128 is the depth size of the 3 block\n        self.fc = nn.Sequential(\n            nn.Linear(self.flatten_dim, 256),\n            nn.ReLU(),\n            nn.Dropout(0.4776451070371182),\n            nn.Linear(256, 1)  # 1 for binary classification\n        )\n\n    def forward(self, x):\n        x = self.block1(x)\n        x = self.block2(x)\n        x = self.block3(x)\n        x = x.view(-1, self.flatten_dim)  # flatten\n        x = self.fc(x)\n        return x # Whithout sigmoid becase we'll use BCEWithLogitsLoss\n# Define the model, loss and optimizer\nmodel_tunned = TunedCNN().to(device)\noptimizer = optim.Adam(model_tunned.parameters(), lr=0.00034503462288880974)\nloss_fn = nn.BCEWithLogitsLoss()\nscheduler = optim.lr_scheduler.ReduceLROnPlateau(\n    optimizer, mode='max', factor=0.5, patience=2, min_lr=1e-5\n)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"training_model(\n    num_epochs=15, optimizer=optimizer,\n    model=model_tunned, loss_function=loss_fn,\n    scheduler=scheduler,\n    model_save_path='models/best_tuned_cnn_model.pth')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's load the best tuned model","metadata":{}},{"cell_type":"code","source":"model_tunned = TunedCNN().to(device)\nmodel_tunned.load_state_dict(torch.load(\"models/best_tuned_cnn_model.pth\"))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels, train_probs = calculate_probability_of_cancer(\n    model_tunned, train_loader)\nval_labels, val_probs = calculate_probability_of_cancer(\n    model_tunned, val_loader)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_fpr, train_tpr, _ = roc_curve(train_labels, train_probs)\nval_fpr, val_tpr, _ = roc_curve(val_labels, val_probs)\n\ntrain_auc = auc(train_fpr, train_tpr)\nval_auc = auc(val_fpr, val_tpr)\n\n# Plot the AUC\nplt.figure(figsize=(8,6))\nplt.plot(train_fpr, train_tpr, label=f'Train AUC = {train_auc:.4f})')\nplt.plot(val_fpr, val_tpr, label=f'Validation AUC = {val_auc:.4f})')\nplt.plot([0,1], [0,1], 'k--')  # diagonal\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Tuned model ROC Curves')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Test data for leaderboard\nNow, let's predict the probability for the test data","metadata":{}},{"cell_type":"code","source":"from torch.utils.data import Dataset","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"First, we need to upload the data and prepare to ingest it to the model","metadata":{}},{"cell_type":"code","source":"\n# we need to create a custom class with the same template of Training data, but this is simpler.\nclass TestDataset(Dataset):\n    def __init__(self, folder_path, transform=None):\n        self.folder_path = folder_path\n        self.image_files = os.listdir(folder_path)\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.image_files)\n\n    def __getitem__(self, idx):\n        img_name = self.image_files[idx]\n        img_path = os.path.join(self.folder_path, img_name)\n        image = Image.open(img_path)\n        if self.transform:\n            image = self.transform(image)\n        return image, img_name\n\n# The tranfsomation for the test, is equal to the validation set.\ntest_transforms = transforms.Compose([\n    transforms.Resize((96,96)),\n    transforms.ToTensor(),\n    transforms.Normalize([0.5,0.5,0.5], [0.5,0.5,0.5])\n])\n\ntest_dataset = TestDataset(\"histo_data/test\", transform=test_transforms)\ntest_loader = DataLoader(test_dataset, batch_size=64, shuffle=False)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, evaluate with each model","metadata":{}},{"cell_type":"code","source":"def obtain_test_predictions(model):\n    model.eval()\n    predictions = []\n\n    with torch.no_grad():\n        # Iterate over the batches\n        for images, file_names in test_loader:\n            images = images.to(device)\n            outputs = model(images)\n            probs = torch.sigmoid(outputs).cpu().numpy()  # Apply the sigmoid for convert to probability\n\n            # save each probability sample in the batch\n            for fname, prob in zip(file_names, probs):\n                predictions.append((fname, prob[0]))\n    return predictions","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### Simple model","metadata":{}},{"cell_type":"code","source":"predictions_simple_model = obtain_test_predictions(simple_model)\nsubmission_simple_model = pd.DataFrame(predictions_simple_model, columns=['id', 'label'])\n\nsubmission_simple_model['id'] = submission_simple_model['id'].str.replace('.tif', '')\nsubmission_simple_model.to_csv(\"submission_simple_model.csv\", index=False)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### Advanced model","metadata":{}},{"cell_type":"code","source":"predictions_advanced_model = obtain_test_predictions(advanced_model)\nsubmission_advanced_model = pd.DataFrame(predictions_advanced_model, columns=['id', 'label'])\n\nsubmission_advanced_model['id'] = submission_advanced_model['id'].str.replace('.tif', '')\nsubmission_advanced_model.to_csv(\"submission_advanced_model.csv\", index=False)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Tuned model","metadata":{}},{"cell_type":"code","source":"predictions_tunned_model = obtain_test_predictions(model_tunned)\nsubmission_tunned_model = pd.DataFrame(predictions_tunned_model, columns=['id', 'label'])\n\nsubmission_tunned_model['id'] = submission_tunned_model['id'].str.replace('.tif', '')\nsubmission_tunned_model.to_csv(\"submission_tunned_model.csv\", index=False)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After passing the test set in LeaderBoard at Kaggle I obtained the following public and private scores:\n- Simple model:\n    - Public score: 0.9296\n    - Private score: 0.8934\n- Advanced model:\n    - Public score: 0.9463\n    - Private score: 0.9181\n- Tuned model:\n    - Public score: 0.9503\n    - Private score: 0.9105","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n- Is better using an advanced architecture for this problem than a simpler one\n- The model with hyperparameter tuning is better in public score but worse in private score\n- The best model in leaderboard is: The advanced model but without hyperparameter tuning\n- I didn't showd this, but at the beginning I started resizing the images from 96x96 to 64x64, and the training time is almost the same, but the performance is better using 96x96. I think that using more resolution allows the model to learn finer details in images.\n","metadata":{}},{"cell_type":"markdown","source":"# Next steps:\n- Optimize the architecture, trying to adjust the trial function in Optuna to change the architecture\n- I used 80,000 samples for train/validation, so, the next step is trying with more data\n- Try more epochs in the hyperparameter optimization, rather than 3\n- Try more n_trials, I used 20\n","metadata":{}}]}