{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"},{"sourceId":12199268,"sourceType":"datasetVersion","datasetId":7684526},{"sourceId":12209779,"sourceType":"datasetVersion","datasetId":7691630}],"dockerImageVersionId":31040,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection – Mini‑Project\n\nAuthor: **Janmejay Buranpuri**  \nDate: 2025-06-17\n\n*Course mini‑project for binary classification of metastatic cancer in histopathology image patches (Kaggle competition).*  \n","metadata":{}},{"cell_type":"markdown","source":"### Problem Statement\n\nThe goal of this competition is to identify metastatic cancer in small image patches taken from larger digital pathology scans of lymph node sections. It is a **binary image classification** task, where each image patch is labeled as either containing metastatic tissue (`label=1`) or not (`label=0`).\n\n### Data Description\n\n- Images are 96x96 pixel RGB patches (PNG format).\n- There are over 220,000 labeled training images and 57,000 test images.\n- Each image has a unique ID. Labels are provided in `train_labels.csv` (columns: `id`, `label`).\n- The data is imbalanced: cancer-positive patches are less common.\n\n> For this mini-project, we will use a **subset** of the data to reduce training time.\n","metadata":{}},{"cell_type":"code","source":"# Imports\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nfrom PIL import Image\nfrom tqdm import tqdm\n\n# Data paths (change if needed)\nDATA_DIR = \"../input/histopathologic-cancer-detection\"\nTRAIN_IMG_DIR = os.path.join(DATA_DIR, \"train\")\nLABELS_PATH = os.path.join(DATA_DIR, \"train_labels.csv\")\n\n# Load labels\nlabels_df = pd.read_csv(LABELS_PATH)\nprint(f\"Total images: {len(labels_df)}\")\nlabels_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:37:44.585056Z","iopub.execute_input":"2025-06-18T18:37:44.585507Z","iopub.status.idle":"2025-06-18T18:37:44.876125Z","shell.execute_reply.started":"2025-06-18T18:37:44.585476Z","shell.execute_reply":"2025-06-18T18:37:44.875135Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check class distribution\nsns.countplot(x=\"label\", data=labels_df)\nplt.title(\"Label Distribution\")\nplt.show()\n\nprint(labels_df['label'].value_counts(normalize=True))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:37:44.877595Z","iopub.execute_input":"2025-06-18T18:37:44.878135Z","iopub.status.idle":"2025-06-18T18:37:45.057305Z","shell.execute_reply.started":"2025-06-18T18:37:44.878108Z","shell.execute_reply":"2025-06-18T18:37:45.056413Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Show random sample of images from each class\ndef plot_sample_images(df, img_dir, label, n=5):\n    ids = df[df.label==label].sample(n, random_state=1)['id'].values\n    plt.figure(figsize=(15,3))\n    for i, img_id in enumerate(ids):\n        img = Image.open(os.path.join(img_dir, img_id + \".tif\"))\n        plt.subplot(1, n, i+1)\n        plt.imshow(img)\n        plt.title(f\"Label: {label}\")\n        plt.axis('off')\n    plt.show()\n\nplot_sample_images(labels_df, TRAIN_IMG_DIR, label=0)\nplot_sample_images(labels_df, TRAIN_IMG_DIR, label=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:37:45.058522Z","iopub.execute_input":"2025-06-18T18:37:45.058900Z","iopub.status.idle":"2025-06-18T18:37:45.991741Z","shell.execute_reply.started":"2025-06-18T18:37:45.058869Z","shell.execute_reply":"2025-06-18T18:37:45.990625Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Observations**\n\n* The dataset is reasonably large for medical imaging (≈220k images).  \n* Class imbalance is manageable but data‑augmentation of the minority class can help.  \n","metadata":{}},{"cell_type":"markdown","source":"#### EDA Summary\n\n- The dataset is **imbalanced**: far more negative than positive samples.\n- Images are small (96x96, 3 channels).\n- Cancer-positive patches are visually harder to distinguish.\n\n### Data Cleaning\n\n- No missing values in labels.\n- All images referenced exist.\n\n**Plan:**  \nWe'll build a CNN model for classification. To reduce imbalance impact, we’ll use balanced sampling or class weights.  \nWe will use a small sample for training for speed.\n","metadata":{}},{"cell_type":"markdown","source":"### Model Choices\n\n- Baseline: Simple CNN (Conv2D layers + MaxPooling + Dense).\n- Comparison: Pretrained model (e.g., MobileNetV2 via transfer learning).\n- We'll use Keras, with data augmentation and early stopping.\n\n**Rationale:**  \n- CNNs are state-of-the-art for image tasks.\n- Transfer learning should improve performance even with less data.\n\n> For speed, we'll train on 5,000 negative and 5,000 positive samples.\n","metadata":{}},{"cell_type":"code","source":"# Sampling a balanced dataset\nN_SAMPLES = 5000  # For each class\n\npos_df = labels_df[labels_df.label==1].sample(N_SAMPLES, random_state=42)\nneg_df = labels_df[labels_df.label==0].sample(N_SAMPLES, random_state=42)\nsample_df = pd.concat([pos_df, neg_df]).sample(frac=1, random_state=1).reset_index(drop=True)\n\nprint(\"Sample dataset shape:\", sample_df.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:37:45.993671Z","iopub.execute_input":"2025-06-18T18:37:45.994048Z","iopub.status.idle":"2025-06-18T18:37:46.045255Z","shell.execute_reply.started":"2025-06-18T18:37:45.994024Z","shell.execute_reply":"2025-06-18T18:37:46.044246Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Image loader (fast)\nIMG_SIZE = 96\n\ndef load_images(df, img_dir, img_size=IMG_SIZE):\n    X = []\n    for img_id in tqdm(df['id']):\n        img = Image.open(os.path.join(img_dir, img_id + \".tif\")).resize((img_size, img_size))\n        X.append(np.array(img))\n    return np.array(X)\n\nX = load_images(sample_df, TRAIN_IMG_DIR)\nX = X.astype(\"float32\") / 255.0   # <--- THIS IS CRUCIAL!\ny = sample_df['label'].values\nprint(\"X shape:\", X.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:37:46.046116Z","iopub.execute_input":"2025-06-18T18:37:46.046627Z","iopub.status.idle":"2025-06-18T18:38:11.849138Z","shell.execute_reply.started":"2025-06-18T18:37:46.046596Z","shell.execute_reply":"2025-06-18T18:38:11.848112Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Train-validation split\nfrom sklearn.model_selection import train_test_split\n\nX_train, X_val, y_train, y_val = train_test_split(\n    X, y, test_size=0.2, random_state=42, stratify=y)\n\nprint(\"Train shape:\", X_train.shape, \"Val shape:\", X_val.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:38:11.850275Z","iopub.execute_input":"2025-06-18T18:38:11.850560Z","iopub.status.idle":"2025-06-18T18:38:12.213823Z","shell.execute_reply.started":"2025-06-18T18:38:11.850539Z","shell.execute_reply":"2025-06-18T18:38:12.212832Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\n\ndef get_simple_cnn(input_shape):\n    model = keras.Sequential([\n        layers.Input(shape=input_shape),\n        layers.Conv2D(32, 3, activation=\"relu\"),\n        layers.MaxPooling2D(),\n        layers.Conv2D(64, 3, activation=\"relu\"),\n        layers.MaxPooling2D(),\n        layers.Flatten(),\n        layers.Dense(64, activation=\"relu\"),\n        layers.Dropout(0.5),\n        layers.Dense(1, activation=\"sigmoid\")\n    ])\n    model.compile(optimizer=\"adam\", loss=\"binary_crossentropy\", metrics=[\"AUC\", \"accuracy\"])\n    return model\n\ncnn = get_simple_cnn((IMG_SIZE, IMG_SIZE, 3))\ncnn.summary()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:38:12.214971Z","iopub.execute_input":"2025-06-18T18:38:12.215321Z","iopub.status.idle":"2025-06-18T18:38:12.307984Z","shell.execute_reply.started":"2025-06-18T18:38:12.215293Z","shell.execute_reply":"2025-06-18T18:38:12.307009Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Data augmentation for training\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\ndatagen = ImageDataGenerator(horizontal_flip=True, vertical_flip=True, rotation_range=20)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:38:12.308940Z","iopub.execute_input":"2025-06-18T18:38:12.309203Z","iopub.status.idle":"2025-06-18T18:38:12.314900Z","shell.execute_reply.started":"2025-06-18T18:38:12.309184Z","shell.execute_reply":"2025-06-18T18:38:12.313595Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"BATCH_SIZE = 32\nEPOCHS = 10\n\n# Use EarlyStopping for efficiency\ncallback = keras.callbacks.EarlyStopping(monitor=\"val_auc\", patience=3, mode=\"max\", restore_best_weights=True)\n\nhistory = cnn.fit(\n    datagen.flow(X_train, y_train, batch_size=BATCH_SIZE),\n    validation_data=(X_val, y_val),\n    epochs=EPOCHS,\n    callbacks=[callback],\n    class_weight={0:1, 1:1.5}  # simple positive class weight\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T18:38:12.315884Z","iopub.execute_input":"2025-06-18T18:38:12.316349Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\nplt.subplot(1,2,1)\nplt.plot(history.history['loss'], label='train')\nplt.plot(history.history['val_loss'], label='val')\nplt.title(\"Loss\")\nplt.legend()\nplt.subplot(1,2,2)\nplt.plot(history.history['AUC'], label='train')      \nplt.plot(history.history['val_AUC'], label='val')   \nplt.title(\"AUC\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Validation performance\nval_preds = cnn.predict(X_val)\nfrom sklearn.metrics import roc_auc_score, accuracy_score, confusion_matrix\n\nauc = roc_auc_score(y_val, val_preds)\nacc = accuracy_score(y_val, (val_preds > 0.5).astype(int))\nprint(f\"Validation AUC: {auc:.4f} | Accuracy: {acc:.4f}\")\n\ncm = confusion_matrix(y_val, (val_preds > 0.5).astype(int))\nsns.heatmap(cm, annot=True, fmt=\"d\")\nplt.title(\"Validation Confusion Matrix\")\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tensorflow.keras.applications import MobileNetV2\n\n\ndef get_transfer_model(input_shape):\n    base = MobileNetV2(\n        weights=\"/kaggle/input/imagenet/mobilenet_v2_weights_tf_dim_ordering_tf_kernels_1.0_96_no_top.h5\",\n        include_top=False,\n        input_shape=input_shape\n    )\n    base.trainable = False  # freeze base\n    model = keras.Sequential([\n        base,\n        layers.GlobalAveragePooling2D(),\n        layers.Dense(64, activation=\"relu\"),\n        layers.Dropout(0.5),\n        layers.Dense(1, activation=\"sigmoid\")\n    ])\n    model.compile(optimizer=\"adam\", loss=\"binary_crossentropy\", metrics=[\"AUC\", \"accuracy\"])\n    return model\n\n\ntransfer_model = get_transfer_model((IMG_SIZE, IMG_SIZE, 3))\nhistory2 = transfer_model.fit(\n    datagen.flow(X_train, y_train, batch_size=BATCH_SIZE),\n    validation_data=(X_val, y_val),\n    epochs=EPOCHS,\n    callbacks=[callback],\n    class_weight={0:1, 1:1.5}\n)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compare AUCs\nval_preds2 = transfer_model.predict(X_val)\nauc2 = roc_auc_score(y_val, val_preds2)\nprint(f\"Transfer Model Validation AUC: {auc2:.4f}\")\n\nplt.plot(history2.history['val_AUC'], label=\"Transfer Model\")\nplt.plot(history.history['val_AUC'], label=\"Simple CNN\")\nplt.title(\"Validation AUC Comparison\")\nplt.legend()\nplt.show()\n\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import glob\n\n# 1. Load test image filenames\ntest_path = '../input/histopathologic-cancer-detection/test/'\ntest_files = glob.glob(test_path + '*.tif')\ntest_ids = [os.path.basename(x)[:-4] for x in test_files]\n\n# 2. Load and preprocess test images\ndef load_test_images(test_ids, test_path, img_size=IMG_SIZE):\n    X_test = []\n    for img_id in tqdm(test_ids):\n        img = Image.open(os.path.join(test_path, img_id + '.tif')).resize((img_size, img_size))\n        X_test.append(np.array(img))\n    X_test = np.array(X_test).astype(\"float32\") / 255.0\n    return X_test\n\nX_test = load_test_images(test_ids, test_path)\n\n# 3. Predict probabilities (use transfer_model or your best model)\ny_pred = transfer_model.predict(X_test, batch_size=32).flatten()\n\n# 4. Create the submission DataFrame\nsubmission = pd.DataFrame({'id': test_ids, 'label': y_pred})\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file saved as submission.csv\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-18T19:18:28.856709Z","iopub.execute_input":"2025-06-18T19:18:28.857101Z","iopub.status.idle":"2025-06-18T19:18:33.432124Z","shell.execute_reply.started":"2025-06-18T19:18:28.857073Z","shell.execute_reply":"2025-06-18T19:18:33.430723Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Results Summary\n\n- Simple CNN AUC\n- MobileNetV2 Transfer AUC\n\nTransfer learning yielded higher AUC and was faster to converge. Data augmentation and class weights both helped.\n","metadata":{}},{"cell_type":"markdown","source":"### Conclusions & Learnings\n\n- **Transfer learning** with MobileNetV2 gave the best results for this subset.\n- **Class imbalance** must be addressed (class weights, balanced sampling, or oversampling).\n- **Data augmentation** helped generalization.\n- **Limitations:** Only used a small subset and a few epochs for speed.\n- **Future improvements:** Use larger sample, fine-tune the base model, experiment with other architectures and regularization, more hyperparameter tuning.\n\n","metadata":{}}]}