{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CNN Cancer Detection\n\nThis noteboook will attempt to implement a CNN model for the Histopathologic Cancer Detection competition. Our dataset, provided by Kaggle, is based on the PCam dataset which contains pre-split and pre-labelled images. The kaggle dataset specifically is also already de-duplicated.\nIn terms of the data labels: a positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. For training, our labels are available at train_labels.csv and the images themselves in ./train/, and the submission output should follow a similar structure for the images sitting in ./test/\n\n","metadata":{}},{"cell_type":"markdown","source":"### Exploratory Data Analysis\n#### Load labels & basic counts\n\nWe'll start by just trying to see \"what does our data even look like?\", and getting a rough sense of the balance of the labels we have.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os, pathlib\nimport matplotlib.pyplot as plt\n\nDATA_DIR = pathlib.Path('/kaggle/input/histopathologic-cancer-detection')\nTRAIN_DIR = os.path.join(DATA_DIR, 'train/')\nTEST_DIR = os.path.join(DATA_DIR, 'test/')\ndf = pd.read_csv(DATA_DIR/'train_labels.csv')\n\n# Getting a basic sample of the data\nprint(df.sample(3))\n\n# Checking for balance\nprint(df['label'].value_counts(normalize=True))\n\n# Example for label 0\nsample_ids_label_0 = df[df.label==0].sample(3).id.values\nfig, axes = plt.subplots(1,3)\nfor ax, img_id in zip(axes.flat, sample_ids_label_0):\n    img = plt.imread(TRAIN_DIR+f'{img_id}.tif')\n    ax.imshow(img)\nplt.suptitle(f'Examples for Label 0', y=0.75)\n\n# Example for label 1\nsample_ids_label_1 = df[df.label==1].sample(3).id.values\nfig, axes = plt.subplots(1,3)\nfor ax, img_id in zip(axes.flat, sample_ids_label_1):\n    img = plt.imread(TRAIN_DIR+f'{img_id}.tif')\n    ax.imshow(img)\nplt.suptitle(f'Examples for Label 1', y=0.75)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T10:17:50.596416Z","iopub.execute_input":"2025-05-25T10:17:50.597148Z","iopub.status.idle":"2025-05-25T10:17:51.469969Z","shell.execute_reply.started":"2025-05-25T10:17:50.597118Z","shell.execute_reply":"2025-05-25T10:17:51.469193Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It seems like our data is roughly 60/40 split between healthy and cancerous results: not perfectly balanced (as all things should be), but probably \"good enough for now\". We also see all the data is labelled with the expected 0s and 1s (not unexpected values). There doesn't seem to be much for us to clean up.\n\n### Data Preprocessing:\n\nBefore working on any model, we need to prepare our data. The df needs to be split into a training set and a validation set. This allows us to evaluate our model's performance on unseen data during development and tune hyperparameters. Both will come from the original \"train\" folder since it's the only labelled one.","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.model_selection import train_test_split\nfrom PIL import Image\nimport time\n\ndf['filename'] = df['id'] + '.tif'\ndf['label_str'] = df['label'].astype(str)\n\ntrain_subset_df, validation_subset_df = train_test_split(df, test_size=0.2, random_state=42)\n\nprint(f\"Training set size: {len(train_subset_df)}\")\nprint(f\"Validation set size: {len(validation_subset_df)}\")\n\n# Image parameters\nIMG_WIDTH, IMG_HEIGHT = 96, 96\nBATCH_SIZE = 512\n\nclass HistopathologyDataset(Dataset):\n    def __init__(self, dataframe, image_dir, transform=None, is_test=False):\n        self.dataframe = dataframe\n        self.image_dir = image_dir\n        self.transform = transform\n        self.is_test = is_test\n\n    def __len__(self):\n        return len(self.dataframe)\n\n    def __getitem__(self, idx):\n        img_name = os.path.join(self.image_dir, self.dataframe.iloc[idx]['id'] + '.tif')\n        image = Image.open(img_name).convert('RGB')\n        image = self.transform(image)\n\n        if self.is_test:\n            return image, self.dataframe.iloc[idx]['id']\n        else:\n            label = self.dataframe.iloc[idx]['label']\n            return image, torch.tensor(label, dtype=torch.float32)\n\n# Include augmentation for training\ntrain_transforms = transforms.Compose([\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.RandomRotation(20),\n    transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5])\n])\n\n# Only normalization for validation\nval_test_transforms = transforms.Compose([\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5])\n])\n\n# Create Datasets\ntrain_dataset = HistopathologyDataset(train_subset_df, TRAIN_DIR, transform=train_transforms)\nvalidation_dataset = HistopathologyDataset(validation_subset_df, TRAIN_DIR, transform=val_test_transforms)\n\n# Create DataLoaders\ntrain_loader = DataLoader(train_dataset, batch_size=BATCH_SIZE, shuffle=True, num_workers=4, pin_memory=True)\nvalidation_loader = DataLoader(validation_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=4, pin_memory=True)\nprint(\"\\nDataLoaders created.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T10:17:58.120862Z","iopub.execute_input":"2025-05-25T10:17:58.121397Z","iopub.status.idle":"2025-05-25T10:17:58.275701Z","shell.execute_reply.started":"2025-05-25T10:17:58.121371Z","shell.execute_reply":"2025-05-25T10:17:58.274952Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Training a Model\nNext we'll train a simple CNN model which we'll later iterate on to try and improve.\nWe'll use CNNs as they are particulaly well-suited for the task since they can efficiently extract features from images and learn composed hierarchies. Additionally, we'll use CNNs because this is a homework assignment about CNNs so this probably shouldn't surprise the precisely 3 people reading this to do the peer review (hey there).\nWe'll start by using 3 convolutional layers and a couple of fully connected layers. We'll use ReLU activation for most layers for reduced likelihood of vanishing gradiants, and sigmoid for the output for efficient binary classification.","metadata":{}},{"cell_type":"code","source":"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nclass BasicCNN(nn.Module):\n    def __init__(self):\n        super(BasicCNN, self).__init__()\n        # Input: (batch_size, 3, 96, 96)\n        self.conv1 = nn.Conv2d(in_channels=3, out_channels=32, kernel_size=3, padding=1)\n        self.relu1 = nn.ReLU()\n        self.pool1 = nn.MaxPool2d(kernel_size=2, stride=2)\n\n        self.conv2 = nn.Conv2d(in_channels=32, out_channels=32, kernel_size=3, padding=1)\n        self.relu2 = nn.ReLU()\n        self.pool2 = nn.MaxPool2d(kernel_size=2, stride=2)\n\n        self.conv3 = nn.Conv2d(in_channels=32, out_channels=32, kernel_size=3, padding=1)\n        self.relu3 = nn.ReLU()\n        self.pool3 = nn.MaxPool2d(kernel_size=2, stride=2)\n\n        self.flatten = nn.Flatten()\n        self.fc1 = nn.Linear(32 * 12 * 12, 128)\n        self.relu4 = nn.ReLU()\n        self.fc2 = nn.Linear(128, 1)\n        self.sigmoid = nn.Sigmoid()\n\n    def forward(self, x):\n        x = self.pool1(self.relu1(self.conv1(x)))\n        x = self.pool2(self.relu2(self.conv2(x)))\n        x = self.pool3(self.relu3(self.conv3(x)))\n        x = self.flatten(x)\n        x = self.relu4(self.fc1(x))\n        x = self.fc2(x)\n        x = self.sigmoid(x) # Output probability\n        return x\n\n# Instantiate the model and move to device\nmodel_basic_pt = BasicCNN().to(device)\nprint(\"\\nBasic Model:\")\nprint(model_basic_pt)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T10:18:02.450754Z","iopub.execute_input":"2025-05-25T10:18:02.451447Z","iopub.status.idle":"2025-05-25T10:18:02.467478Z","shell.execute_reply.started":"2025-05-25T10:18:02.451422Z","shell.execute_reply":"2025-05-25T10:18:02.466691Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Training Code\nNow that we have our basic architecture set up, let's train it and see what happens","metadata":{}},{"cell_type":"code","source":"criterion = nn.BCELoss()\nhistory_storage_pt = {}\n\ndef train_model(model, train_loader, validation_loader, criterion, optimizer, num_epochs=30):\n    train_losses, val_losses = [], []\n    train_aucs, val_aucs = [], []\n    train_accs, val_accs = [], []\n\n    for epoch in range(num_epochs):\n        print(f\"Start of epoch {epoch}\")\n        start_time = time.time()\n        model.train()\n        running_loss = 0.0\n        train_preds, train_targets = [], []\n\n        for inputs, labels in train_loader:\n            inputs, labels = inputs.to(device), labels.to(device).unsqueeze(1)\n\n            optimizer.zero_grad()\n            outputs = model(inputs)\n            loss = criterion(outputs, labels)\n            loss.backward()\n            optimizer.step()\n\n            running_loss += loss.item() * inputs.size(0)\n            train_preds.extend(outputs.detach().cpu().numpy())\n            train_targets.extend(labels.detach().cpu().numpy())\n\n        print(f\"Epoch {epoch} finished training loop, now calculating stats\")\n        \n        epoch_loss = running_loss / len(train_loader.dataset)\n        train_losses.append(epoch_loss)\n        \n        train_preds_flat = np.array(train_preds).flatten()\n        train_targets_flat = np.array(train_targets).flatten()\n        \n        current_train_auc = roc_auc_score(train_targets_flat, train_preds_flat)\n        train_aucs.append(current_train_auc)\n        \n        train_predicted_classes = (train_preds_flat > 0.5).astype(int)\n        current_train_acc = np.mean(train_predicted_classes == train_targets_flat)\n        train_accs.append(current_train_acc)\n\n        print(f\"Starting eval for epoch {epoch}\")\n        model.eval()\n        val_running_loss = 0.0\n        val_preds, val_targets = [], []\n        with torch.no_grad():\n            for inputs, labels in validation_loader:\n                inputs, labels = inputs.to(device), labels.to(device).unsqueeze(1)\n                outputs = model(inputs)\n                loss = criterion(outputs, labels)\n                val_running_loss += loss.item() * inputs.size(0)\n                val_preds.extend(outputs.cpu().numpy())\n                val_targets.extend(labels.cpu().numpy())\n\n        epoch_val_loss = val_running_loss / len(validation_loader.dataset)\n        val_losses.append(epoch_val_loss)\n\n        val_preds_flat = np.array(val_preds).flatten()\n        val_targets_flat = np.array(val_targets).flatten()\n        \n        current_val_auc = roc_auc_score(val_targets_flat, val_preds_flat)\n        val_aucs.append(current_val_auc)\n\n        val_predicted_classes = (val_preds_flat > 0.5).astype(int)\n        current_val_acc = np.mean(val_predicted_classes == val_targets_flat)\n        val_accs.append(current_val_acc)\n\n        epoch_duration = time.time() - start_time\n        print(f\"Epoch {epoch+1}/{num_epochs} | Time: {epoch_duration:.2f}s | \"\n              f\"Train Loss: {epoch_loss:.4f} | Train AUC: {current_train_auc:.4f} | Train Acc: {current_train_acc:.4f} | \"\n              f\"Val Loss: {epoch_val_loss:.4f} | Val AUC: {current_val_auc:.4f} | Val Acc: {current_val_acc:.4f}\")\n\n    history = {\n        'loss': train_losses, 'val_loss': val_losses,\n        'auc': train_aucs, 'val_auc': val_aucs,\n        'accuracy': train_accs, 'val_accuracy': val_accs\n    }\n    print(f\"Training complete. Final Validation AUC: {current_val_auc:.4f}\")\n    return model, history\n\nEPOCHS_PT = 10","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T11:10:31.716608Z","iopub.execute_input":"2025-05-25T11:10:31.717141Z","iopub.status.idle":"2025-05-25T11:10:31.727644Z","shell.execute_reply.started":"2025-05-25T11:10:31.717120Z","shell.execute_reply":"2025-05-25T11:10:31.726618Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\n--- Training Basic Model ---\")\nmodel_basic_pt = BasicCNN().to(device)\noptimizer_basic = optim.Adam(model_basic_pt.parameters(), lr=0.001)\n\n\nmodel_basic_pt, history_basic_pt = train_model(\n    model_basic_pt, train_loader, validation_loader, criterion, optimizer_basic, num_epochs=EPOCHS_PT\n)\nhistory_storage_pt['basic_pt'] = history_basic_pt\n\n# Final evaluation on validation set with the best model\nmodel_basic_pt.eval()\nfinal_val_preds, final_val_targets = [], []\nwith torch.no_grad():\n    for inputs, labels in validation_loader:\n        inputs, labels = inputs.to(device), labels.to(device).unsqueeze(1)\n        outputs = model_basic_pt(inputs)\n        final_val_preds.extend(outputs.cpu().numpy())\n        final_val_targets.extend(labels.cpu().numpy())\n\nfinal_val_auc_basic = roc_auc_score(np.array(final_val_targets).flatten(), np.array(final_val_preds).flatten())\nprint(f\"Basic Model - Final Validation AUC: {final_val_auc_basic:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T10:32:06.223599Z","iopub.execute_input":"2025-05-25T10:32:06.223884Z","iopub.status.idle":"2025-05-25T11:09:49.432023Z","shell.execute_reply.started":"2025-05-25T10:32:06.223863Z","shell.execute_reply":"2025-05-25T11:09:49.431197Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Checking out our performance\nWith the training done, let's visualise the results and see how the model performed.","metadata":{}},{"cell_type":"code","source":"def plot_training_history_pt(history_dict, model_name):\n    fig, axes = plt.subplots(1, 3, figsize=(22, 5))\n    fig.suptitle(f'Training History for {model_name}', fontsize=16)\n\n    # Plot Loss\n    axes[0].plot(history_dict.get('loss', []), label='Training Loss')\n    axes[0].plot(history_dict.get('val_loss', []), label='Validation Loss')\n    axes[0].set_title('Loss vs. Epochs')\n    axes[0].set_xlabel('Epochs')\n    axes[0].set_ylabel('Loss')\n    axes[0].legend()\n    axes[0].grid(True)\n\n    # Plot AUC\n    axes[1].plot(history_dict.get('auc', []), label='Training AUC')\n    axes[1].plot(history_dict.get('val_auc', []), label='Validation AUC')\n    axes[1].set_title('AUC vs. Epochs')\n    axes[1].set_xlabel('Epochs')\n    axes[1].set_ylabel('AUC')\n    axes[1].legend()\n    axes[1].grid(True)\n    \n    # Plot Accuracy\n    axes[2].plot(history_dict.get('accuracy', []), label='Training Accuracy')\n    axes[2].plot(history_dict.get('val_accuracy', []), label='Validation Accuracy')\n    axes[2].set_title('Accuracy vs. Epochs')\n    axes[2].set_xlabel('Epochs')\n    axes[2].set_ylabel('Accuracy')\n    axes[2].legend()\n    axes[2].grid(True)\n\n    plt.tight_layout(rect=[0, 0, 1, 0.95])\n    plt.show()\n\nplot_training_history_pt(history_storage_pt['basic_pt'], 'Basic Model')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T11:09:54.677609Z","iopub.execute_input":"2025-05-25T11:09:54.677892Z","iopub.status.idle":"2025-05-25T11:09:55.252585Z","shell.execute_reply.started":"2025-05-25T11:09:54.677870Z","shell.execute_reply":"2025-05-25T11:09:55.251746Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Improving the model\nTo try and improve our model, we can attempt a few additions:\n1. Increasing the size of our layers might help the model extract more features.\n2. Adding an additional convolutional layer might help the model extract more sophisticated features from the images.\n3. Adding regularisation can help with concerns of potential overfitting (especially since we're increasing the model size)\n4. We can add batch normalisation to improve the speed of convergence.","metadata":{}},{"cell_type":"code","source":"class EnhancedCNN(nn.Module):\n    def __init__(self, num_classes=1):\n        super(EnhancedCNN, self).__init__()\n        \n        self.conv1 = nn.Conv2d(3, 32, kernel_size=3, padding=1)\n        self.bn1 = nn.BatchNorm2d(32)\n        self.relu1 = nn.ReLU()\n        self.pool1 = nn.MaxPool2d(2, 2)\n\n        self.conv2 = nn.Conv2d(32, 64, kernel_size=3, padding=1)\n        self.bn2 = nn.BatchNorm2d(64)\n        self.relu2 = nn.ReLU()\n        self.pool2 = nn.MaxPool2d(2, 2)\n\n        self.conv3 = nn.Conv2d(64, 128, kernel_size=3, padding=1)\n        self.bn3 = nn.BatchNorm2d(128)\n        self.relu3 = nn.ReLU()\n        self.pool3 = nn.MaxPool2d(2, 2)\n\n        self.conv4 = nn.Conv2d(128, 256, kernel_size=3, padding=1)\n        self.bn4 = nn.BatchNorm2d(256)\n        self.relu4 = nn.ReLU()\n        self.pool4 = nn.MaxPool2d(2, 2)\n\n        self.flatten = nn.Flatten()\n        self.fc1 = nn.Linear(256 * 6 * 6, 256)\n        self.bn_fc = nn.BatchNorm1d(256)\n        self.relu5 = nn.ReLU()\n        self.dropout = nn.Dropout(0.5)\n        self.fc2 = nn.Linear(256, num_classes)\n        self.sigmoid = nn.Sigmoid()\n\n    def forward(self, x):\n        x = self.pool1(self.relu1(self.bn1(self.conv1(x))))\n        x = self.pool2(self.relu2(self.bn2(self.conv2(x))))\n        x = self.pool3(self.relu3(self.bn3(self.conv3(x))))\n        x = self.pool4(self.relu4(self.bn4(self.conv4(x))))\n        x = self.flatten(x)\n        x = self.relu5(self.bn_fc(self.fc1(x)))\n        x = self.dropout(x)\n        x = self.sigmoid(self.fc2(x))\n        return x\n\nmodel_enhanced_pt = EnhancedCNN().to(device)\nprint(\"\\nEnhanced Model:\")\nprint(model_enhanced_pt)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T11:10:12.052858Z","iopub.execute_input":"2025-05-25T11:10:12.053159Z","iopub.status.idle":"2025-05-25T11:10:12.095747Z","shell.execute_reply.started":"2025-05-25T11:10:12.053138Z","shell.execute_reply":"2025-05-25T11:10:12.095005Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Training and evaluating the enhanced model","metadata":{}},{"cell_type":"code","source":"optimizer_enhanced = optim.Adam(model_enhanced_pt.parameters(), lr=0.001)\n\nmodel_enhanced_pt, history_enhanced_pt = train_model(\n    model_enhanced_pt, train_loader, validation_loader, criterion, optimizer_enhanced, num_epochs=EPOCHS_PT\n)\nhistory_storage_pt['enhanced_pt'] = history_enhanced_pt\n\nmodel_enhanced_pt.eval()\nfinal_val_preds_enh, final_val_targets_enh = [], []\nwith torch.no_grad():\n    for inputs, labels in validation_loader:\n        inputs, labels = inputs.to(device), labels.to(device).unsqueeze(1)\n        outputs = model_enhanced_pt(inputs)\n        final_val_preds_enh.extend(outputs.cpu().numpy())\n        final_val_targets_enh.extend(labels.cpu().numpy())\n        \nfinal_val_auc_enhanced = roc_auc_score(np.array(final_val_targets_enh).flatten(), np.array(final_val_preds_enh).flatten())\nprint(f\"Enhanced Model - Final Validation AUC: {final_val_auc_enhanced:.4f}\")\n\nplot_training_history_pt(history_storage_pt['enhanced_pt'], 'Enhanced Model')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T11:10:43.562817Z","iopub.execute_input":"2025-05-25T11:10:43.563074Z","iopub.status.idle":"2025-05-25T11:48:40.088711Z","shell.execute_reply.started":"2025-05-25T11:10:43.563058Z","shell.execute_reply":"2025-05-25T11:48:40.087742Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Results Summary\nAs we can see from the results, our tuning improved the performance of the model:\n\n| Model        | Validation AUC | Validation Loss | Validation Accuracy |\n| ------------ | -------------- | --------------- | ------------------- |\n| Basic CNN    | 0.9625         | 0.2426          | 0.9015              |\n| Enhanced CNN | 0.9809         | 0.1803          | 0.9312              |\n\nWith that, we can attempt using it against the test data and submit our results to Kaggle for evaluation.","metadata":{}},{"cell_type":"code","source":"print(\"\\n--- Generating Predictions for Test Set ---\")\n\nmodel_enhanced_pt.eval()\n\ndf = pd.read_csv(DATA_DIR/'sample_submission.csv')\n\ntest_df = pd.DataFrame({'id': df['id']})\ntest_df['label'] = -1\n\n\ntest_dataset_pt = HistopathologyDataset(\n    dataframe=test_df,\n    image_dir=TEST_DIR,\n    transform=val_test_transforms,\n    is_test=True\n)\n\ntest_loader_pt = DataLoader(\n    test_dataset_pt,\n    batch_size=BATCH_SIZE,\n    shuffle=False,\n    num_workers=4,\n    pin_memory=True\n)\n\nprint(f\"Predicting on {len(test_dataset_pt)} test images using enhanced model...\")\nall_predictions_pt = []\nall_ids_pt = []\n\nwith torch.no_grad():\n    for images, ids_batch in test_loader_pt:\n        images = images.to(device)\n        outputs = model_enhanced_pt(images)\n        all_predictions_pt.extend(outputs.cpu().numpy().flatten())\n        all_ids_pt.extend(ids_batch)\n\npred_dict = {img_id: pred for img_id, pred in zip(all_ids_pt, all_predictions_pt)}\n\nfinal_ordered_predictions = [pred_dict.get(img_id, 0.5) for img_id in df['id']] # Default to 0.5 if an ID was missed\n\nsubmission_df_pt = pd.DataFrame({\n    'id': df['id'],\n    'label': final_ordered_predictions\n})\n\nsubmission_path_pt = 'submission.csv'\nsubmission_df_pt.to_csv(submission_path_pt, index=False)\nprint(f\"\\nSubmission file created: {submission_path_pt}\")\nprint(submission_df_pt.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-25T11:55:07.206634Z","iopub.execute_input":"2025-05-25T11:55:07.207361Z","iopub.status.idle":"2025-05-25T11:55:40.712403Z","shell.execute_reply.started":"2025-05-25T11:55:07.207332Z","shell.execute_reply":"2025-05-25T11:55:40.711554Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Conclusions and future improvements\nWe saw that CNNs do a decent job at binary classification for this dataset. As expected, giving the model a decent size in both width and depth can allow the model to learn more, while using techniques like dropouts help us avoid overfitting too quickly. At the same time, due to Kaggle resource constraints we only trained the model for 10 epoch and the graphs show there might be more room for the model to learn from this dataset. To avoid overfitting in longer epoch we could also apply additional transformations to our training data, potentially extracting a bit more from it.\nIn conclusion: while no one should trust this model in a medical setting, it's a good demonstration of how CNNs learn visual classification from labelled data.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}