{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"d9ce5925","cell_type":"markdown","source":"# Histopathologic Cancer Detection - Mini Project\n\n**Kaggle Competition**: [Histopathologic Cancer Detection](https://www.kaggle.com/c/histopathologic-cancer-detection)\n\n## Problem Description\nThis challenge is a binary image classification problem where the goal is to detect metastatic cancer in small image patches derived from larger pathology slides. Identifying cancerous cells is crucial for proper diagnosis and treatment planning.\n\n## Data Description\nThe dataset consists of:\n- **Images**: 96x96 color image patches in `.tif` format.\n- **Train Labels**: CSV file with two columns: `id` and `label` (1 for cancer, 0 for non-cancer).\n- The training set contains over 220,000 images.\n\n---","metadata":{}},{"id":"81ee503c","cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\nWe will start by inspecting the dataset structure, looking at the distribution of labels, and displaying a few example images to understand what the model will learn from.","metadata":{}},{"id":"09ab1e20","cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom pathlib import Path\nfrom PIL import Image\nimport os\nimport numpy as np\n\n# Load labels\ndf = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\ndf['label'].value_counts().plot(kind='bar', title='Label Distribution')\ndf['label'] = df['label'].astype(str)\nplt.show()\n\n# Show a few example images\nsample_ids = df.sample(6)['id'].values\nfig, axes = plt.subplots(2, 3, figsize=(10, 7))\nfor ax, img_id in zip(axes.flatten(), sample_ids):\n    img = Image.open(f'/kaggle/input/histopathologic-cancer-detection/train/{img_id}.tif')\n    ax.imshow(img)\n    ax.set_title(f'ID: {img_id}')\n    ax.axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T16:03:51.380982Z","iopub.execute_input":"2025-04-25T16:03:51.381415Z","iopub.status.idle":"2025-04-25T16:03:52.185131Z","shell.execute_reply.started":"2025-04-25T16:03:51.381392Z","shell.execute_reply":"2025-04-25T16:03:52.184472Z"}},"outputs":[],"execution_count":null},{"id":"62dc3b61-2c82-4b39-bc61-3d29f0d0ee2b","cell_type":"markdown","source":"## Pre-Processing","metadata":{}},{"id":"fc262b8a-ce4f-4bb3-8a29-79897e8cc973","cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nimage_dir = \"/kaggle/input/histopathologic-cancer-detection/train\"\ndf[\"id\"] = df[\"id\"].astype(str) + \".tif\"\ndf[\"path\"] = df[\"id\"].apply(lambda x: os.path.join(image_dir, x))\n\n# Parameters\nIMG_SIZE = 96\nBATCH_SIZE = 64\n\n# Data Generators\ndatagen = ImageDataGenerator(validation_split=0.2, rescale=1./255)\n\ntrain_generator = datagen.flow_from_dataframe(\n    dataframe=df,\n    directory='/kaggle/input/histopathologic-cancer-detection/train',\n    x_col='path',\n    y_col='label',\n    subset='training',\n    batch_size=BATCH_SIZE,\n    seed=42,\n    shuffle=True,\n    class_mode='binary',\n    target_size=(IMG_SIZE, IMG_SIZE),\n    color_mode='rgb'\n)\n\nval_generator = datagen.flow_from_dataframe(\n    dataframe=df,\n    directory='/kaggle/input/histopathologic-cancer-detection/train',\n    x_col='path',\n    y_col='label',\n    subset='validation',\n    batch_size=BATCH_SIZE,\n    seed=42,\n    shuffle=True,\n    class_mode='binary',\n    target_size=(IMG_SIZE, IMG_SIZE),\n    color_mode='rgb'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T16:03:52.186458Z","iopub.execute_input":"2025-04-25T16:03:52.186735Z","iopub.status.idle":"2025-04-25T16:24:16.605670Z","shell.execute_reply.started":"2025-04-25T16:03:52.186711Z","shell.execute_reply":"2025-04-25T16:24:16.604877Z"}},"outputs":[],"execution_count":null},{"id":"33c746e5","cell_type":"markdown","source":"## Model Architecture\nI will use Convolutional Neural Networks (CNNs) for this task as they are well-suited for image classification problems.\n\n### Baseline Model\nA simple CNN with a few Conv2D and MaxPooling2D layers followed by Dense layers.\n\n### Optimized Model\nI will use an identical model, and implement the RMSprop optimizer rather than adam.\n\nI will compare model performance mainly using validation accuracy.","metadata":{}},{"id":"a8491046","cell_type":"markdown","source":"## Model Training and Evaluation\nWe will use early stopping, learning rate scheduling, and model checkpointing during training to prevent overfitting and select the best model.","metadata":{}},{"id":"0b69938a","cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout, BatchNormalization\n\n# Simple CNN\nmodel = Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(IMG_SIZE, IMG_SIZE, 3)),\n    MaxPooling2D(2, 2),\n    Conv2D(64, (3, 3), activation='relu'),\n    MaxPooling2D(2, 2),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dropout(0.5),\n    Dense(1, activation='sigmoid')\n])\n\nmodel.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T16:24:16.606624Z","iopub.execute_input":"2025-04-25T16:24:16.606863Z","iopub.status.idle":"2025-04-25T16:24:18.731575Z","shell.execute_reply.started":"2025-04-25T16:24:16.606846Z","shell.execute_reply":"2025-04-25T16:24:18.731003Z"}},"outputs":[],"execution_count":null},{"id":"0ba5c430-3e34-4db2-81ce-047eebacde96","cell_type":"code","source":"from tensorflow.keras.callbacks import EarlyStopping, ModelCheckpoint\n\nearly_stop = EarlyStopping(patience=5, restore_best_weights=True)\ncheckpoint = ModelCheckpoint(\"baseline_model.keras\", save_best_only=True)\n\nbaseline_history = model.fit(\n    train_generator,\n    epochs=10,\n    validation_data=val_generator,\n    callbacks=[early_stop, checkpoint]\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T16:24:18.733137Z","iopub.execute_input":"2025-04-25T16:24:18.733648Z","iopub.status.idle":"2025-04-25T18:12:30.244364Z","shell.execute_reply.started":"2025-04-25T16:24:18.733630Z","shell.execute_reply":"2025-04-25T18:12:30.243708Z"}},"outputs":[],"execution_count":null},{"id":"c877b46c","cell_type":"markdown","source":"## Hyperparameter Tuning\nTo improve performance, we can experiment with:\n- Different optimizers (Adam, RMSprop, SGD)\n- Learning rates\n- Number of filters and dropout rates\n\nIn this case, I will try optimizing with RMSprop, and evaluate the models. Additional improvements may include fine-tuning learning rate, optimizer selection, and augmentations.","metadata":{}},{"id":"e020aae1","cell_type":"code","source":"from tensorflow.keras.optimizers import RMSprop, SGD\n\nmodel = Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(IMG_SIZE, IMG_SIZE, 3)),\n    MaxPooling2D(2, 2),\n    Conv2D(64, (3, 3), activation='relu'),\n    MaxPooling2D(2, 2),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dropout(0.5),\n    Dense(1, activation='sigmoid')\n])\n\ncheckpoint = ModelCheckpoint(\"optimized_model.keras\", save_best_only=True)\n# Update model with RMSprop\nprint(\"Compiling model\")\nmodel.compile(optimizer=RMSprop(learning_rate=0.0001), loss='binary_crossentropy', metrics=['accuracy'])\n\nprint(\"Training optimized model\")\noptimized_history = model.fit(\n    train_generator,\n    epochs=10,\n    validation_data=val_generator,\n    callbacks=[early_stop, checkpoint]\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T18:12:30.245239Z","iopub.execute_input":"2025-04-25T18:12:30.245566Z","iopub.status.idle":"2025-04-25T18:46:16.287748Z","shell.execute_reply.started":"2025-04-25T18:12:30.245545Z","shell.execute_reply":"2025-04-25T18:46:16.286957Z"}},"outputs":[],"execution_count":null},{"id":"f3a22b9c","cell_type":"markdown","source":"## Results Visualization\nWe visualize training and validation accuracy and loss to evaluate overfitting or underfitting.","metadata":{}},{"id":"361ee3cb","cell_type":"code","source":"plt.figure(figsize=(14, 5))\nplt.subplot(1, 2, 1)\nplt.plot(baseline_history.history['accuracy'], label='Train Accuracy')\nplt.plot(baseline_history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Baseline Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\nplt.subplot(1, 2, 2)\nplt.plot(optimized_history.history['accuracy'], label='Train Accuracy')\nplt.plot(optimized_history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Optimized Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T18:46:16.288782Z","iopub.execute_input":"2025-04-25T18:46:16.289031Z","iopub.status.idle":"2025-04-25T18:46:16.723967Z","shell.execute_reply.started":"2025-04-25T18:46:16.289012Z","shell.execute_reply":"2025-04-25T18:46:16.723151Z"}},"outputs":[],"execution_count":null},{"id":"c9caecbb","cell_type":"markdown","source":"## Final Results Table\nWe summarize results using accuracy and loss from the best performing model.","metadata":{}},{"id":"44be8037","cell_type":"code","source":"baseline_res = baseline_history.history\noptimized_res = optimized_history.history\n\nresults_df = pd.DataFrame({\n    'Metric': ['Train Accuracy', 'Validation Accuracy', 'Train Loss', 'Validation Loss'],\n    'Baseline': [baseline_res['accuracy'][-1], baseline_res['val_accuracy'][-1], baseline_res['loss'][-1], baseline_res['val_loss'][-1]],\n    'Optimized': [optimized_res['accuracy'][-1], optimized_res['val_accuracy'][-1], optimized_res['loss'][-1], optimized_res['val_loss'][-1]]\n})\ndisplay(results_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T18:46:16.724911Z","iopub.execute_input":"2025-04-25T18:46:16.725233Z","iopub.status.idle":"2025-04-25T18:46:16.775336Z","shell.execute_reply.started":"2025-04-25T18:46:16.725191Z","shell.execute_reply":"2025-04-25T18:46:16.774737Z"}},"outputs":[],"execution_count":null},{"id":"689fad15-3e72-4e46-9aa0-00f58d5d68f2","cell_type":"markdown","source":"## Conclusion\nWe explored a CNN-based approach to identify cancer in histopathologic images. The baseline model performs reasonably, even surpassing the RMSprop optimized model. I assume the reason for this is that Adam optimizer combines the momentum and RMSprop techniques and in turn provides a more balanced and efficient optimization process. Further improvements can include using advanced techniques such as:\n- Transfer learning\n- Better augmentation\n- Hyperparameter tuning\n- Adding more filters\n\nFurther exploration can involve using attention-based networks or ensemble models for better accuracy.","metadata":{}},{"id":"89a2f28c-7aea-4981-9cd5-fc0ca4f12ff4","cell_type":"markdown","source":"## Submission","metadata":{}},{"id":"919b50b7-d929-455d-ac32-b6fdbf1bbdaf","cell_type":"code","source":"# test data generator\ntest_datagen = ImageDataGenerator(rescale=1./255)\n\n# test file names DF\ntest_files = os.listdir('/kaggle/input/histopathologic-cancer-detection/test')\ntest_df = pd.DataFrame({\n    'id': [os.path.splitext(file)[0] for file in test_files],\n    'filename': test_files\n})\n\n# make test data gen\nsubmission_batch_size = BATCH_SIZE * 2 \n\ntest_generator = test_datagen.flow_from_dataframe(\n    dataframe=test_df,\n    directory='/kaggle/input/histopathologic-cancer-detection/test',\n    x_col='filename',\n    y_col=None,  # no labels for test data\n    target_size=(IMG_SIZE, IMG_SIZE),\n    batch_size=BATCH_SIZE,\n    class_mode=None,\n    shuffle=False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T18:46:16.776034Z","iopub.execute_input":"2025-04-25T18:46:16.776369Z","iopub.status.idle":"2025-04-25T18:49:59.894035Z","shell.execute_reply.started":"2025-04-25T18:46:16.776351Z","shell.execute_reply":"2025-04-25T18:49:59.893257Z"}},"outputs":[],"execution_count":null},{"id":"3943b061-9f8a-4eae-adb9-9a9641eda888","cell_type":"code","source":"# predict\nprint(\"\\nGenerating predictions for submission...\")\noptimized_model = tf.keras.models.load_model('/kaggle/working/baseline_model.keras')\npredictions = optimized_model.predict(\n    test_generator,\n    verbose=1\n)\npredicted_classes = (predictions > 0.5).astype(int).flatten()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T18:49:59.894962Z","iopub.execute_input":"2025-04-25T18:49:59.895229Z","iopub.status.idle":"2025-04-25T18:58:56.790582Z","shell.execute_reply.started":"2025-04-25T18:49:59.895191Z","shell.execute_reply":"2025-04-25T18:58:56.789866Z"}},"outputs":[],"execution_count":null},{"id":"bb6b11b9-b367-47d4-9f1f-6feaab98ea96","cell_type":"code","source":"# save\nsubmission_df = pd.DataFrame({\n    'id': test_df['id'][:len(predicted_classes)],\n    'label': predicted_classes\n})\n\nsubmission_path = 'submission.csv'\nsubmission_df.to_csv(submission_path, index=False)\nprint(f\"Submission saved to {submission_path}\")\nprint(f\"Sample of submission file:\\n{submission_df.head()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T18:58:56.792928Z","iopub.execute_input":"2025-04-25T18:58:56.793144Z","iopub.status.idle":"2025-04-25T18:58:56.907726Z","shell.execute_reply.started":"2025-04-25T18:58:56.793128Z","shell.execute_reply":"2025-04-25T18:58:56.906933Z"}},"outputs":[],"execution_count":null}]}