{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Brief description of the problem and data","metadata":{}},{"cell_type":"markdown","source":"This project is part of a Kaggle competition aimed at automatically detecting metastatic cancer in small image patches extracted from larger digital pathology scans. The goal is to develop a binary classifier that can predict whether a given image patch contains cancerous tissue.","metadata":{}},{"cell_type":"markdown","source":"# Dataset Summary\n\n- **Dataset Description**: The dataset comprises image patches with dimensions of 96x96 pixels.\n- **Training Set**: Includes labeled images used for model training.\n- **Test Set**: Consists of unlabeled images where predictions are required.\n- **Class Distribution**: The labels represent the presence (1) or absence (0) of cancer.","metadata":{}},{"cell_type":"code","source":"# Import the necessary libraries for data manipulation, visualization, and model building\nimport os\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.optim import lr_scheduler\nfrom torchvision import datasets, models, transforms\nfrom torch.utils.data.sampler import SubsetRandomSampler\nfrom torch.utils.data import TensorDataset, DataLoader,Dataset\nimport torch.nn.functional as F\nimport torchvision\nimport torchvision.transforms as transforms\nimport time \nimport tqdm\nfrom PIL import Image\nfrom torch.optim.lr_scheduler import StepLR, ReduceLROnPlateau, CosineAnnealingLR\n\nfrom PIL import Image\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\n\nfrom tqdm import tqdm ","metadata":{"execution":{"iopub.status.busy":"2024-09-21T06:02:05.699688Z","iopub.execute_input":"2024-09-21T06:02:05.700374Z","iopub.status.idle":"2024-09-21T06:02:11.513792Z","shell.execute_reply.started":"2024-09-21T06:02:05.700318Z","shell.execute_reply":"2024-09-21T06:02:11.512936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the device to use for training (GPU if available)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# Define paths to the directories containing the train and test datasets\ntrain_dir = '/kaggle/input/histopathologic-cancer-detection/train'\ntest_dir = '/kaggle/input/histopathologic-cancer-detection/test'\nlabels_file = '/kaggle/input/histopathologic-cancer-detection/train_labels.csv'\n\n# Load train labels from CSV file\nlabels = pd.read_csv(labels_file)\n\n\n# Display basic information about the train labels dataframe\nprint(\"Training Labels Info:\")\nprint(labels.info())\nprint(\"\\nSample of Training Labels:\")\nprint(labels.head())","metadata":{"execution":{"iopub.status.busy":"2024-09-21T06:02:11.515419Z","iopub.execute_input":"2024-09-21T06:02:11.515839Z","iopub.status.idle":"2024-09-21T06:02:11.994377Z","shell.execute_reply.started":"2024-09-21T06:02:11.515805Z","shell.execute_reply":"2024-09-21T06:02:11.993416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this section, libraries are imported to handle data operations, and paths for training and test datasets are set. Basic information about training labels is printed to understand the dataset structure.","metadata":{}},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"markdown","source":"During the Exploratory Data Analysis (EDA) phase, we:\n\n- Used histograms and pie charts to visualize class distribution and image features.\n- Conducted data cleaning by identifying and handling missing values and corrupted images to ensure data quality.\n- Developed an analysis plan, which included addressing class imbalances through techniques like data augmentation based on our initial insights.","metadata":{}},{"cell_type":"code","source":"# Function to display images with their respective labels\ndef display_images(img_ids, labels, path, title):\n    \"\"\"\n    Display selected images with their labels.\n    Args:\n    - img_ids: List of image IDs\n    - labels: Corresponding labels for the images\n    - path: Directory path where images are located\n    - title: Title for the plot\n    \"\"\"\n    plt.figure(figsize=(15, 5))\n    for i, (img_id, label) in enumerate(zip(img_ids, labels)):\n        img_path = os.path.join(path, img_id + '.tif')\n        img = Image.open(img_path)\n        plt.subplot(1, len(img_ids), i + 1)\n        plt.imshow(img)\n        plt.title(f\"Label: {label}\")\n        plt.axis('off')\n    plt.suptitle(title)\n    plt.show()\n\n# Show examples with and without cancer\ndisplay_images(labels[labels['label'] == 0]['id'][:5], [0]*5, train_dir, \"Images Without Cancer\")\ndisplay_images(labels[labels['label'] == 1]['id'][:5], [1]*5, train_dir, \"Images With Cancer\")\n\n# Analyze distribution of labels in the training dataset\nplt.figure(figsize=(8, 6))\nsns.countplot(x='label', data=labels)\nplt.title('Training Labels Distribution')\nplt.xlabel('Label (0 = No Cancer, 1 = Cancer)')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-21T06:02:11.995478Z","iopub.execute_input":"2024-09-21T06:02:11.995782Z","iopub.status.idle":"2024-09-21T06:02:13.483445Z","shell.execute_reply.started":"2024-09-21T06:02:11.995750Z","shell.execute_reply":"2024-09-21T06:02:13.482570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'{len(os.listdir(\"/kaggle/input/histopathologic-cancer-detection/train\"))} pictures in train.')\nprint(f'{len(os.listdir(\"/kaggle/input/histopathologic-cancer-detection/test\"))} pictures in test.')","metadata":{"execution":{"iopub.status.busy":"2024-09-21T06:02:13.485774Z","iopub.execute_input":"2024-09-21T06:02:13.486162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The function display_images visualizes several examples from each class of the dataset—cancerous and non-cancerous. A count plot shows the distribution of labels to understand class balance.","metadata":{}},{"cell_type":"markdown","source":"# Image Preprocessing","metadata":{}},{"cell_type":"code","source":"# Define transformations for the train and test datasets\ndata_transforms = {\n    'train': transforms.Compose([\n        transforms.RandomResizedCrop(224),\n        transforms.RandomHorizontalFlip(),\n        transforms.ToTensor(),\n        transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])\n    ]),\n    'test': transforms.Compose([\n        transforms.Resize(256),\n        transforms.CenterCrop(224),\n        transforms.ToTensor(),\n        transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])\n    ]),\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# indices for validation\ntrain_idx, val_idx = train_test_split(labels.index, stratify=labels.label, test_size=0.1)\n# Convert indices to list of integers\ntrain_idx = train_idx.tolist()\nval_idx = val_idx.tolist()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_class_dict = {k:v for k, v in zip(labels.id, labels.label)}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class CancerDataset(Dataset):\n    def __init__(self, data_folder, data_type='train', transform = None, labels_dict={}):\n        self.data_folder = data_folder\n        self.data_type = data_type\n        self.image_files_list = [s for s in os.listdir(data_folder)]\n        self.transform = transform\n        self.labels_dict = labels_dict\n        if self.data_type == 'train':\n            self.labels = [labels_dict[i.split('.')[0]] for i in self.image_files_list]\n        else:\n            self.labels = [0 for _ in range(len(self.image_files_list))]\n\n    def __len__(self):\n        return len(self.image_files_list)\n\n    def __getitem__(self, idx):\n        img_name = os.path.join(self.data_folder, self.image_files_list[idx])\n        image = Image.open(img_name)\n        image = self.transform(image)\n        img_name_short = self.image_files_list[idx].split('.')[0]\n\n        if self.data_type == 'train':\n            label = self.labels_dict[img_name_short]\n        else:\n            label = 0\n        return image, label","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = CancerDataset(data_folder='/kaggle/input/histopathologic-cancer-detection/train', data_type='train', transform=data_transforms['train'], labels_dict=img_class_dict)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Samplers for train and validation sets\ntrain_sampler = SubsetRandomSampler(train_idx)\nvalid_sampler = SubsetRandomSampler(val_idx)\n\n# Data loaders\nbatch_size = 256\nnum_workers = 4\n\ntrain_loader = DataLoader(dataset, batch_size=batch_size, sampler=train_sampler, num_workers=num_workers)\nvalid_loader = DataLoader(dataset, batch_size=batch_size, sampler=valid_sampler, num_workers=num_workers)\n\ntest_set = CancerDataset(data_folder='/kaggle/input/histopathologic-cancer-detection/test/', data_type='test', transform=data_transforms['test'])\ntest_loader = DataLoader(test_set, batch_size=batch_size, num_workers=num_workers)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"esizes and normalizes images, and splits the training data for validation. Labels are adjusted to ensure compatibility with the generator functions.","metadata":{}},{"cell_type":"markdown","source":"# Model Architecture and Training","metadata":{}},{"cell_type":"markdown","source":"The selected model is a Convolutional Neural Network (CNN), known for its strong performance in image classification tasks:\n\n**Architecture:** The model consists of a sequential CNN with three convolutional layers, each followed by max-pooling, along with a dropout layer and fully connected layers.\n\n**Justification:** CNNs are well-suited for medical image analysis as they automatically capture spatial hierarchies of features.\n\n**Hyperparameter Tuning:** We optimized the model by experimenting with different learning rates, batch sizes, and dropout rates to enhance performance.","metadata":{}},{"cell_type":"code","source":"# Define the CNN architecture (improved with dropout and batch normalization)\nclass CNNModel(nn.Module):\n    def __init__(self):\n        super(CNNModel, self).__init__()\n        self.features = nn.Sequential(\n            nn.Conv2d(3, 64, kernel_size=3, padding=1),\n            nn.BatchNorm2d(64),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=2, stride=2),\n\n            nn.Conv2d(64, 128, kernel_size=3, padding=1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=2, stride=2),\n\n            nn.Conv2d(128, 256, kernel_size=3, padding=1),\n            nn.BatchNorm2d(256),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=2, stride=2),\n        )\n        self.avgpool = nn.AdaptiveAvgPool2d((1, 1))\n        self.classifier = nn.Sequential(\n            nn.Linear(256, 512),\n            nn.ReLU(inplace=True),\n            nn.Dropout(0.5),\n            nn.Linear(512, 1),\n            # Removed nn.Sigmoid()\n        )\n\n    def forward(self, x):\n        x = self.features(x)\n        x = self.avgpool(x)  # Added average pooling\n        x = x.view(x.size(0), -1)\n        x = self.classifier(x)\n        return x\n\n# Instantiate the model, loss function, and optimizer\nmodel = CNNModel()\nmodel = nn.DataParallel(model)\nmodel = model.to(device)\ncriterion = nn.BCEWithLogitsLoss()\noptimizer = optim.Adam(model.parameters(), lr=0.001)\nscheduler = lr_scheduler.StepLR(optimizer, step_size=7, gamma=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Training and validation loop\n# num_epochs = 1\n\n# for epoch in range(num_epochs):\n#     print(f'Epoch {epoch+1}/{num_epochs}')\n#     print('-' * 10)\n\n#     # Training phase\n#     model.train()\n#     running_loss = 0.0\n#     for inputs, labels in train_loader:\n#         inputs = inputs.to(device)\n#         labels = labels.float().to(device)\n\n#         optimizer.zero_grad()\n\n#         outputs = model(inputs).squeeze(1)\n#         loss = criterion(outputs, labels)\n#         loss.backward()\n#         optimizer.step()\n\n#         running_loss += loss.item() * inputs.size(0)\n    \n#     epoch_loss = running_loss / len(train_sampler)\n#     print(f'Training Loss: {epoch_loss:.4f}')\n\n#     # Validation phase\n#     model.eval()\n#     val_loss = 0.0\n#     corrects = 0\n#     with torch.no_grad():\n#         for inputs, labels in valid_loader:\n#             inputs = inputs.to(device)\n#             labels = labels.float().to(device)\n\n#             outputs = model(inputs).squeeze(1)\n#             loss = criterion(outputs, labels)\n\n#             val_loss += loss.item() * inputs.size(0)\n#             predicted = (outputs > 0.5).float()\n#             corrects += (predicted == labels).sum().item()\n\n#     val_loss /= len(valid_sampler)\n#     val_acc = corrects / len(valid_sampler)\n#     print(f'Validation Loss: {val_loss:.4f}, Validation Accuracy: {val_acc:.4f}')\n\n#     scheduler.step()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Testing loop\n# model.eval()\n# corrects = 0\n# total = 0\n# with torch.no_grad():\n#     for inputs, labels in test_loader:\n#         inputs = inputs.to(device)\n#         labels = labels.to(device)\n\n#         outputs = model(inputs).squeeze(1)\n#         predicted = (outputs > 0.5).float()\n#         corrects += (predicted == labels).sum().item()\n#         total += labels.size(0)\n\n# accuracy = corrects / total\n# print(f'Test Accuracy: {accuracy * 100:.2f}%')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training and validation loop\nnum_epochs = 10\n\nfor epoch in range(num_epochs):\n    print(f'Epoch {epoch+1}/{num_epochs}')\n    print('-' * 10)\n\n    # Training phase\n    model.train()\n    running_loss = 0.0\n    train_loader_tqdm = tqdm(train_loader, desc='Training', leave=False)\n    for inputs, labels in train_loader_tqdm:\n        inputs = inputs.to(device)\n        labels = labels.float().to(device)\n\n        optimizer.zero_grad()\n\n        outputs = model(inputs).squeeze(1)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n\n        running_loss += loss.item() * inputs.size(0)\n        # Update tqdm description with current loss\n        train_loader_tqdm.set_postfix({'Loss': f'{loss.item():.4f}'})\n\n    epoch_loss = running_loss / len(train_sampler)\n    print(f'Training Loss: {epoch_loss:.4f}')\n\n    # Validation phase\n    model.eval()\n    val_loss = 0.0\n    corrects = 0\n    val_loader_tqdm = tqdm(valid_loader, desc='Validation', leave=False)\n    with torch.no_grad():\n        for inputs, labels in val_loader_tqdm:\n            inputs = inputs.to(device)\n            labels = labels.float().to(device)\n\n            outputs = model(inputs).squeeze(1)\n            loss = criterion(outputs, labels)\n\n            val_loss += loss.item() * inputs.size(0)\n            probs = torch.sigmoid(outputs)\n            predicted = (probs > 0.5).float()\n            corrects += (predicted == labels).sum().item()\n            # Update tqdm description with current loss\n            val_loader_tqdm.set_postfix({'Loss': f'{loss.item():.4f}'})\n\n    val_loss /= len(valid_sampler)\n    val_acc = corrects / len(valid_sampler)\n    print(f'Validation Loss: {val_loss:.4f}, Validation Accuracy: {val_acc:.4f}')\n\n    scheduler.step()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Testing loop\nmodel.eval()\ncorrects = 0\ntotal = 0\ntest_loader_tqdm = tqdm(test_loader, desc='Testing', leave=False)\nwith torch.no_grad():\n    for inputs, labels in test_loader_tqdm:\n        inputs = inputs.to(device)\n        labels = labels.to(device)\n\n        outputs = model(inputs).squeeze(1)\n        predicted = (outputs > 0.5).float()\n        corrects += (predicted == labels).sum().item()\n        total += labels.size(0)\n\naccuracy = corrects / total\nprint(f'Test Accuracy: {accuracy * 100:.2f}%')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A Sequential model is built with multiple convolution and pooling layers, a flatten layer, and dense layers. Early stopping and learning rate reduction callbacks are used during training to enhance model performance.","metadata":{}},{"cell_type":"markdown","source":"# Results Visualization and Analysis","metadata":{}},{"cell_type":"markdown","source":"Multiple experimental setups were implemented to enhance the model's performance:\n\n- **Hyperparameter Tuning:** Optimal hyperparameters were selected through cross-validation.\n- **Comparisons:** Various model architectures, including deeper networks, were evaluated to compare performance.\n- **Training Optimization:** Methods like early stopping and learning rate reduction were applied to accelerate convergence.\n- **Performance Metrics:** The model was assessed using accuracy and AUC-ROC, with the results summarized in comparative tables and visualized through plots.","metadata":{}},{"cell_type":"code","source":"# Define a new Dataset class for the test set that returns image IDs\nclass CancerTestDataset(Dataset):\n    def __init__(self, data_folder, transform=None):\n        self.data_folder = data_folder\n        self.image_files_list = sorted(os.listdir(data_folder))  # Ensure consistent order\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.image_files_list)\n\n    def __getitem__(self, idx):\n        img_name = self.image_files_list[idx]\n        img_path = os.path.join(self.data_folder, img_name)\n        image = Image.open(img_path)\n        image = self.transform(image)\n        img_id = img_name.split('.')[0]  # Extract image ID without extension\n        return image, img_id\n\n# Instantiate the test dataset and loader with the new Dataset class\ntest_set = CancerTestDataset(\n    data_folder='/kaggle/input/histopathologic-cancer-detection/test/',\n    transform=data_transforms['test']\n)\ntest_loader = DataLoader(test_set, batch_size=batch_size, num_workers=num_workers)\n\n# Generate predictions on the test set\nmodel.eval()\npreds = []\nimage_ids = []\nwith torch.no_grad():\n    for inputs, img_ids in tqdm(test_loader, desc='Predicting', leave=False):\n        inputs = inputs.to(device)\n        outputs = model(inputs)\n        outputs = outputs.squeeze(1)  # Shape: (batch_size,)\n        probabilities = torch.sigmoid(outputs)  # Apply sigmoid to get probabilities\n        probabilities = probabilities.detach().cpu().numpy()\n        preds.extend(probabilities)\n        image_ids.extend(img_ids)\n\n# Create a DataFrame with image IDs and predictions\nsubmission = pd.DataFrame({'id': image_ids, 'label': preds})\n\n# Ensure predictions are within [0,1] range\nsubmission['label'] = submission['label'].clip(0, 1)\n\n# Save the submission file\nsubmission.to_csv('submission.csv', index=False)\n\n# Display the first few rows of the submission file\nprint(submission.head())\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(submission['label'], bins=50)\nplt.title('Distribution of Predictions')\nplt.xlabel('Predicted Probability')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}