{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":""},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":10338,"databundleVersionId":862042},{"sourceType":"datasetVersion","sourceId":8626327,"datasetId":5164530,"databundleVersionId":8774821}],"dockerImageVersionId":31287,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🩺 Project #8: RSNA Pneumonia Detection (CNN vs ResNet50)\n\n**Project Objective:** Building a robust Computer Vision diagnostic engine to detect Pneumonia opacities in chest X-Rays. We will pit a deeply engineered Custom CNN against an industry-standard Transfer Learning giant (ResNet50) to see which architecture handles medical class imbalance better.\n\n---\n\n### 🏗️ Project Architecture: 10-Step Development Process\n\n1. **Defining the Project Goal:** Detecting Pneumonia in chest X-Rays (Binary Classification).\n2. **Data Reading and Exploratory Data Analysis (EDA):** Reading the `stage_2_train_labels.csv` file using Pandas and analyzing the healthy/sick class distribution to identify Class Imbalance.\n3. **Selecting Relevant Columns:** Isolating only the `patientId` and `Target` columns for the model's target mapping.\n4. **Categorical Transformations:** Cleaning duplicate records to ensure a single clear diagnosis per patient and converting the `Target` column to a string format for Keras image generators.\n5. **Data Manipulation:** Vectorizing the addition of the `.jpg` extension to patient IDs so the Keras engine can locate the files *(Optimized completely without the use of slow for-loops)*.\n6. **Feature Engineering:** Applying classic `1./255` pixel normalization for our Custom CNN, while preparing the specific `preprocess_input` function tailored for ResNet50's ImageNet weights.\n7. **Encoding & Data Flow:** Building two separate `flow_from_dataframe` generators to process over 26,000 images seamlessly without overloading system memory.\n8. **Data Splitting:** Dividing the dataset into 80% Training and 20% Validation sets.\n9. **Model Training (Fit):** Simultaneously training the \"ResNet50\" model alongside our Custom Deep CNN—which is deeply armored with \"Batch Normalization\" layers—on the exact same dataset.\n10. **Performance Audit:** Avoiding the deceptive \"Accuracy\" illusion by comparing **Recall** metrics (the true rate of detecting the rare disease) via Heatmaps (Confusion Matrices) to crown the champion architecture.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.metrics import confusion_matrix, classification_report, accuracy_score\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Flatten, Conv2D, MaxPooling2D, Dropout, BatchNormalization, InputLayer\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.applications.resnet50 import ResNet50, preprocess_input\nfrom tensorflow.keras.callbacks import EarlyStopping\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T17:35:51.623391Z","iopub.execute_input":"2026-04-20T17:35:51.623990Z","iopub.status.idle":"2026-04-20T17:36:19.991192Z","shell.execute_reply.started":"2026-04-20T17:35:51.623956Z","shell.execute_reply":"2026-04-20T17:36:19.990239Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 1: Project Objective ---\n# Goal: Binary Classification of Pneumonia (1: Positive, 0: Negative) \n# using Custom CNN and ResNet50 architectures.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:21:04.504397Z","iopub.execute_input":"2026-04-20T16:21:04.505186Z","iopub.status.idle":"2026-04-20T16:21:04.517199Z","shell.execute_reply.started":"2026-04-20T16:21:04.505162Z","shell.execute_reply":"2026-04-20T16:21:04.516406Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 2: Read and analyze the data (EDA) ---\ndf = pd.read_csv('/kaggle/input/competitions/rsna-pneumonia-detection-challenge/stage_2_train_labels.csv')\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:45.377036Z","iopub.execute_input":"2026-04-20T16:23:45.377358Z","iopub.status.idle":"2026-04-20T16:23:45.412824Z","shell.execute_reply.started":"2026-04-20T16:23:45.377331Z","shell.execute_reply":"2026-04-20T16:23:45.412059Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:45.901958Z","iopub.execute_input":"2026-04-20T16:23:45.902716Z","iopub.status.idle":"2026-04-20T16:23:45.913228Z","shell.execute_reply.started":"2026-04-20T16:23:45.902683Z","shell.execute_reply":"2026-04-20T16:23:45.912673Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:46.367722Z","iopub.execute_input":"2026-04-20T16:23:46.368399Z","iopub.status.idle":"2026-04-20T16:23:46.379849Z","shell.execute_reply.started":"2026-04-20T16:23:46.368372Z","shell.execute_reply":"2026-04-20T16:23:46.378977Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:48.805155Z","iopub.execute_input":"2026-04-20T16:23:48.805749Z","iopub.status.idle":"2026-04-20T16:23:48.814088Z","shell.execute_reply.started":"2026-04-20T16:23:48.805721Z","shell.execute_reply":"2026-04-20T16:23:48.813372Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.countplot(x=df['Target'])\nplt.title(\"Class Distribution\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:49.347724Z","iopub.execute_input":"2026-04-20T16:23:49.348545Z","iopub.status.idle":"2026-04-20T16:23:49.508860Z","shell.execute_reply.started":"2026-04-20T16:23:49.348506Z","shell.execute_reply":"2026-04-20T16:23:49.507877Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 3: Select suitable columns ---\n# Focusing on patientId for image mapping and Target for labels.\ndf = df[['patientId', 'Target']]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:51.893290Z","iopub.execute_input":"2026-04-20T16:23:51.893560Z","iopub.status.idle":"2026-04-20T16:23:51.898989Z","shell.execute_reply.started":"2026-04-20T16:23:51.893540Z","shell.execute_reply":"2026-04-20T16:23:51.898313Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 4: Convert categories ---\n# Dropping duplicate IDs (some patients have multiple boxes, but we need 1 label per image)\ndf = df.drop_duplicates(subset='patientId').reset_index(drop=True)\ndf['Target'] = df['Target'].astype(str)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:52.216200Z","iopub.execute_input":"2026-04-20T16:23:52.216792Z","iopub.status.idle":"2026-04-20T16:23:52.232446Z","shell.execute_reply.started":"2026-04-20T16:23:52.216764Z","shell.execute_reply":"2026-04-20T16:23:52.231765Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 5: Data manipulations ---\n# Vectorized filename creation. Keras needs JPG files.\n# Assuming a JPG-converted version of the dataset is used.\ndf['filename'] = df['patientId'] + '.jpg'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:23:52.536297Z","iopub.execute_input":"2026-04-20T16:23:52.536931Z","iopub.status.idle":"2026-04-20T16:23:52.543880Z","shell.execute_reply.started":"2026-04-20T16:23:52.536904Z","shell.execute_reply":"2026-04-20T16:23:52.543242Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 6: Feature Engineering & STEP 7: Encoding ---\n# Using specific generators for each model to handle preprocessing\ndatagen_cnn = ImageDataGenerator(rescale=1./255, validation_split=0.2)\ndatagen_resnet = ImageDataGenerator(preprocessing_function=preprocess_input, validation_split=0.2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:24:00.027199Z","iopub.execute_input":"2026-04-20T16:24:00.027533Z","iopub.status.idle":"2026-04-20T16:24:00.031895Z","shell.execute_reply.started":"2026-04-20T16:24:00.027507Z","shell.execute_reply":"2026-04-20T16:24:00.030940Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 8. Split data into X and y\n# Pointing to the dataset folder\nimg_dir = '/kaggle/input/datasets/chamina1ch/rsna-pneumonia-jpg/train'","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Generators for Custom CNN ---\ntrain_gen_cnn = datagen_cnn.flow_from_dataframe(\n    df,\n    directory=img_dir,\n    x_col='filename',\n    y_col='Target',\n    target_size=(224, 224),\n    class_mode='binary',\n    subset='training',\n    batch_size=32\n)\n\n# shuffle=False is critical for the confusion matrix\nval_gen_cnn = datagen_cnn.flow_from_dataframe(\n    df,\n    directory=img_dir,\n    x_col='filename',\n    y_col='Target',\n    target_size=(224, 224),\n    class_mode='binary',\n    subset='validation',\n    batch_size=32,\n    shuffle=False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:24:05.563131Z","iopub.execute_input":"2026-04-20T16:24:05.563725Z","iopub.status.idle":"2026-04-20T16:24:45.920739Z","shell.execute_reply.started":"2026-04-20T16:24:05.563700Z","shell.execute_reply":"2026-04-20T16:24:45.919972Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Generators for ResNet50 ---\ntrain_gen_res = datagen_resnet.flow_from_dataframe(\n    df,\n    directory=img_dir,\n    x_col='filename',\n    y_col='Target',\n    target_size=(224, 224),\n    class_mode='binary',\n    subset='training',\n    batch_size=32\n)\n\nval_gen_res = datagen_resnet.flow_from_dataframe(\n    df,\n    directory=img_dir,\n    x_col='filename',\n    y_col='Target',\n    target_size=(224, 224),\n    class_mode='binary',\n    subset='validation',\n    batch_size=32,\n    shuffle=False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:25:06.315204Z","iopub.execute_input":"2026-04-20T16:25:06.315727Z","iopub.status.idle":"2026-04-20T16:25:15.500448Z","shell.execute_reply.started":"2026-04-20T16:25:06.315698Z","shell.execute_reply":"2026-04-20T16:25:15.499394Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"early_stop = EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:26:46.794930Z","iopub.execute_input":"2026-04-20T16:26:46.795748Z","iopub.status.idle":"2026-04-20T16:26:46.799362Z","shell.execute_reply.started":"2026-04-20T16:26:46.795719Z","shell.execute_reply":"2026-04-20T16:26:46.798653Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 9: Model Execution (Fit-Predict) ---\n\n# --- MODEL A: Deep Custom CNN  ---\nmodel_cnn = Sequential([\n    InputLayer(input_shape=(224, 224, 3)),\n    Conv2D(32, (3, 3), activation='relu', padding='same'),\n    BatchNormalization(),\n    MaxPooling2D((2, 2)),\n    \n    Conv2D(64, (3, 3), activation='relu', padding='same'),\n    BatchNormalization(),\n    MaxPooling2D((2, 2)),\n    \n    Conv2D(128, (3, 3), activation='relu', padding='same'),\n    BatchNormalization(),\n    MaxPooling2D((2, 2)),\n    \n    Flatten(),\n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.5),\n    Dense(1, activation='sigmoid')\n])\n\nmodel_cnn.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\nprint(\"\\nTraining Custom CNN...\")\nmodel_cnn.fit(train_gen_cnn, validation_data=val_gen_cnn, epochs=8, callbacks=[early_stop])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:26:47.795123Z","iopub.execute_input":"2026-04-20T16:26:47.795737Z","iopub.status.idle":"2026-04-20T16:45:17.317251Z","shell.execute_reply.started":"2026-04-20T16:26:47.795710Z","shell.execute_reply":"2026-04-20T16:45:17.316655Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- MODEL B: ResNet50 Transfer Learning ---\nmodel_resnet = Sequential([\n    ResNet50(weights='imagenet', include_top=False, input_shape=(224, 224, 3)),\n    Flatten(),\n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.5),\n    Dense(1, activation='sigmoid')\n])\n\nmodel_resnet.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\nprint(\"\\nTraining ResNet50...\")\nmodel_resnet.fit(train_gen_res, validation_data=val_gen_res, epochs=8, callbacks=[early_stop])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T16:46:02.875024Z","iopub.execute_input":"2026-04-20T16:46:02.875337Z","iopub.status.idle":"2026-04-20T17:04:28.595827Z","shell.execute_reply.started":"2026-04-20T16:46:02.875312Z","shell.execute_reply":"2026-04-20T17:04:28.594981Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- STEP 10: Performance Audit ---\n\ndef evaluate_model(model, generator, name, color):\n    print(f\"\\n--- {name} Results ---\")\n    probs = model.predict(generator)\n    preds = np.round(probs).astype(int)\n    y_true = generator.classes\n    \n    print(f\"Accuracy: {accuracy_score(y_true, preds):.4f}\")\n    print(classification_report(y_true, preds))\n    \n    plt.figure(figsize=(5, 4))\n    sns.heatmap(confusion_matrix(y_true, preds), annot=True, fmt='d', cmap=color)\n    plt.title(f\"{name} Confusion Matrix\")\n    plt.show()\n\n# Audit both models\nevaluate_model(model_cnn, val_gen_cnn, \"Custom CNN\", \"Blues\")\nevaluate_model(model_resnet, val_gen_res, \"ResNet50\", \"Greens\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T17:05:02.735724Z","iopub.execute_input":"2026-04-20T17:05:02.736350Z","iopub.status.idle":"2026-04-20T17:06:03.620488Z","shell.execute_reply.started":"2026-04-20T17:05:02.736323Z","shell.execute_reply":"2026-04-20T17:06:03.619902Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- EXPORTING THE CHAMPION MODEL ---\n\n\nmodel_cnn.save('PneumoVision_CNN.h5')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-20T17:11:18.700470Z","iopub.execute_input":"2026-04-20T17:11:18.701127Z","iopub.status.idle":"2026-04-20T17:11:19.099670Z","shell.execute_reply.started":"2026-04-20T17:11:18.701100Z","shell.execute_reply":"2026-04-20T17:11:19.099026Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🏁 Project # 8 Final Audit: PneumoVision AI\n\n## 🚀 LIVE DEPLOYMENT\nThe victorious Custom CNN model has been serialized and successfully deployed as an interactive medical diagnostic web service. \n### 🔗 [LIVE ENGINE: PneumoVision AI on Hugging Face](https://huggingface.co/spaces/Ironside35/PneumoVision-AI)\n\n---\n\n## 🧠 Architectural Conclusion & Final Thoughts\n\n### 1. The Fall of the Giant (ResNet50)\nIn this execution, **ResNet50** fell violently back into the \"Dominant Class Trap.\" Despite its massive architecture, the model took the lazy route to maintain a 71% accuracy by predicting almost every X-ray as 'Healthy'. It successfully found only **14 out of 1077** actual pneumonia cases (Recall: 0.01)—a catastrophic failure for a medical diagnostic tool.\n\n### 2. The Custom CNN Victory\nOur engineered **Custom Deep CNN** proved its robustness. Thanks to the \"Batch Normalization\" layers maintaining gradient stability, it actively hunted for opacities and correctly identified **530 pneumonia cases** (Recall: 0.49). It completely bypassed the class imbalance trap that defeated ResNet50.\n\n### 3. Champion Model Selection\nBased on the superior **F1-Score (0.55 vs 0.03)** and exponentially higher Recall, the **Custom CNN** is crowned as the undisputed champion of this project. It mathematically proves that a custom-tailored architecture can outperform a blind, pre-trained giant in sensitive medical environments.\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}