{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"85346462","cell_type":"markdown","source":"# Fine-Tuned ResNet50 for Retinal Disease Classification — Reproduction Notebook (v2)\n\n**Base paper:** Rani & Gupta, *\"Automated Retinal Disease Classification Using Fine-Tuned ResNet50: A Deep Learning Approach for Early Diagnosis,\"* IITCEE 2025 (DOI: 10.1109/IITCEE64140.2025.10915328)\n\n**Reproduction status:** METHODOLOGICAL, not exact numerical.\n- Paper's dataset: 6-class disease-type classification (AMD, DR, DN, MH, ODC, Normal), private/unspecified source, no public link.\n- This notebook: APTOS 2019 Blindness Detection (Kaggle) — 5-class **DR severity grading** (0=No DR … 4=Proliferative DR).\n- Label space differs from the paper, so reported numbers are **not directly comparable**. State this explicitly in any report.\n\n**v2 changes from the first run (which hit 68.9% acc / 0.52 macro-F1):**\n1. Simplified head — GlobalAveragePooling instead of a bolted-on 3-layer Conv2D stack, which was too much randomly-initialized capacity for 2562 training images.\n2. Two-phase training — head-only warmup (frozen backbone), then fine-tune conv4+conv5 at a low LR.\n3. Ben Graham-style local color normalization — standard APTOS preprocessing trick, corrects lighting/contrast variance across cameras.\n4. Quadratic weighted kappa (QWK) tracked alongside accuracy/F1 — DR grading is ordinal, and this is the actual competition metric, more informative than raw accuracy for \"how far off\" misclassifications are.\n\n**Dataset:** Add `aptos2019-blindness-detection` via Kaggle \"Add Data\" before running. Check `/kaggle/input/` for the exact folder name — it varies (`aptos2019-blindness-detection` vs `competitions/aptos2019-blindness-detection`) depending on how it was added.\n","metadata":{}},{"id":"1066e33a","cell_type":"markdown","source":"## 1. Setup & Reproducibility","metadata":{}},{"id":"6ab4d603","cell_type":"code","source":"import os, random, json\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom pathlib import Path\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import (classification_report, confusion_matrix,\n                              roc_auc_score, roc_curve, f1_score,\n                              precision_score, recall_score, accuracy_score,\n                              cohen_kappa_score)\nfrom sklearn.utils.class_weight import compute_class_weight\n\nSEED = 42\nos.environ['PYTHONHASHSEED'] = str(SEED)\nrandom.seed(SEED)\nnp.random.seed(SEED)\ntf.random.set_seed(SEED)\n\nprint(\"TF version:\", tf.__version__)\nprint(\"GPUs:\", tf.config.list_physical_devices('GPU'))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:45:54.176255Z","iopub.execute_input":"2026-08-13T10:45:54.176509Z","iopub.status.idle":"2026-08-13T10:46:12.134626Z","shell.execute_reply.started":"2026-08-13T10:45:54.176474Z","shell.execute_reply":"2026-08-13T10:46:12.133733Z"}},"outputs":[],"execution_count":null},{"id":"4342f120","cell_type":"markdown","source":"## 2. Load Data","metadata":{}},{"id":"db9fe9cb","cell_type":"code","source":"# Check the actual mount path first if this errors — it varies by how the dataset was added\nDATA_DIR = Path(\"/kaggle/input/competitions/aptos2019-blindness-detection\")\nif not DATA_DIR.exists():\n    DATA_DIR = Path(\"/kaggle/input/competitions/aptos2019-blindness-detection\")\nassert DATA_DIR.exists(), f\"Dataset not found. Run !ls /kaggle/input/ to locate it.\"\n\nTRAIN_IMG_DIR = DATA_DIR / \"train_images\"\n\ntrain_df = pd.read_csv(DATA_DIR / \"train.csv\")\ntrain_df['filepath'] = train_df['id_code'].apply(lambda x: str(TRAIN_IMG_DIR / f\"{x}.png\"))\ntrain_df['diagnosis'] = train_df['diagnosis'].astype(int)\n\nCLASS_NAMES = {0: \"No DR\", 1: \"Mild\", 2: \"Moderate\", 3: \"Severe\", 4: \"Proliferative DR\"}\nNUM_CLASSES = 5\n\nprint(train_df.shape)\ntrain_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:12.136220Z","iopub.execute_input":"2026-08-13T10:46:12.136748Z","iopub.status.idle":"2026-08-13T10:46:12.200126Z","shell.execute_reply.started":"2026-08-13T10:46:12.136721Z","shell.execute_reply":"2026-08-13T10:46:12.199327Z"}},"outputs":[],"execution_count":null},{"id":"e952ed11","cell_type":"code","source":"# Sanity check: every listed image file actually exists\nmissing = train_df[~train_df['filepath'].apply(os.path.exists)]\nassert len(missing) == 0, f\"{len(missing)} missing image files\"\nprint(\"All files present.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:12.203253Z","iopub.execute_input":"2026-08-13T10:46:12.203540Z","iopub.status.idle":"2026-08-13T10:46:21.296717Z","shell.execute_reply.started":"2026-08-13T10:46:12.203517Z","shell.execute_reply":"2026-08-13T10:46:21.296034Z"}},"outputs":[],"execution_count":null},{"id":"d3a767c3","cell_type":"markdown","source":"## 3. EDA — Class Distribution & Samples","metadata":{}},{"id":"845ca431","cell_type":"code","source":"counts = train_df['diagnosis'].value_counts().sort_index()\nlabels = [CLASS_NAMES[i] for i in counts.index]\n\nplt.figure(figsize=(8,5))\nsns.barplot(x=labels, y=counts.values, hue=labels, palette=\"rocket\", legend=False)\nplt.title(\"Class distribution (APTOS 2019 train)\")\nplt.ylabel(\"count\")\nplt.xticks(rotation=20)\nplt.tight_layout()\nplt.savefig(\"/kaggle/working/class_distribution.png\", dpi=150)\nplt.show()\nprint(counts)\nprint(\"\\nImbalance ratio (max/min):\", counts.max() / counts.min())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:21.297652Z","iopub.execute_input":"2026-08-13T10:46:21.298210Z","iopub.status.idle":"2026-08-13T10:46:21.678340Z","shell.execute_reply.started":"2026-08-13T10:46:21.298185Z","shell.execute_reply":"2026-08-13T10:46:21.677556Z"}},"outputs":[],"execution_count":null},{"id":"438bd655","cell_type":"code","source":"fig, axes = plt.subplots(2, 4, figsize=(16, 8))\nsample = train_df.sample(8, random_state=SEED).reset_index(drop=True)\nfor i, ax in enumerate(axes.flat):\n    img = plt.imread(sample.loc[i, 'filepath'])\n    ax.imshow(img)\n    ax.set_title(CLASS_NAMES[sample.loc[i, 'diagnosis']])\n    ax.axis('off')\nplt.tight_layout()\nplt.savefig(\"/kaggle/working/sample_images.png\", dpi=150)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:21.680188Z","iopub.execute_input":"2026-08-13T10:46:21.680457Z","iopub.status.idle":"2026-08-13T10:46:32.234039Z","shell.execute_reply.started":"2026-08-13T10:46:21.680434Z","shell.execute_reply":"2026-08-13T10:46:32.232947Z"}},"outputs":[],"execution_count":null},{"id":"58b77617","cell_type":"markdown","source":"## 4. Train / Validation / Test Split\n\nAPTOS 2019 provides no patient/eye identifier in `train.csv`, so a strict patient-independent split\nisn't verifiable here — flagging this as a known limitation, not glossing over it.\nStratified split at the image level: 70% train / 15% val / 15% test, fixed seed.\n`test.csv` has no public labels, so the held-out test split below is used for final evaluation instead.\n","metadata":{}},{"id":"be8a302c","cell_type":"code","source":"train_val_df, test_df = train_test_split(\n    train_df, test_size=0.15, stratify=train_df['diagnosis'], random_state=SEED\n)\ntrain_split_df, val_df = train_test_split(\n    train_val_df, test_size=0.1765,  # 0.15/0.85 -> ~15% of total\n    stratify=train_val_df['diagnosis'], random_state=SEED\n)\n\nprint(\"Train:\", len(train_split_df), \"Val:\", len(val_df), \"Test:\", len(test_df))\nfor name, d in [(\"train\", train_split_df), (\"val\", val_df), (\"test\", test_df)]:\n    print(name, dict(d['diagnosis'].value_counts().sort_index()))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:32.235184Z","iopub.execute_input":"2026-08-13T10:46:32.235482Z","iopub.status.idle":"2026-08-13T10:46:32.254314Z","shell.execute_reply.started":"2026-08-13T10:46:32.235448Z","shell.execute_reply":"2026-08-13T10:46:32.253420Z"}},"outputs":[],"execution_count":null},{"id":"87cb1b11","cell_type":"markdown","source":"## 5. tf.data Pipeline with Ben Graham-Style Color Normalization\n\nAPTOS fundus images vary a lot in lighting/contrast across cameras and clinics. Local color\nnormalization (subtract a heavily-blurred version of the image, re-center at gray) flattens\nthat variance and is the single biggest known accuracy lever on this specific dataset.\n","metadata":{}},{"id":"83802ee3","cell_type":"code","source":"IMG_SIZE = (224, 224)\nBATCH_SIZE = 32\nAUTOTUNE = tf.data.AUTOTUNE\n\ndef ben_graham_normalize(img):\n    # img: float32 tensor, HWC, range [0,255]\n    blurred = tf.numpy_function(\n        lambda x: __import__('cv2').GaussianBlur(x, (0, 0), 10),\n        [img], tf.float32\n    )\n    blurred.set_shape(img.shape)\n    normalized = tf.clip_by_value(4.0 * img - 4.0 * blurred + 128.0, 0, 255)\n    return normalized\n\ndef load_image(path, label):\n    img = tf.io.read_file(path)\n    img = tf.image.decode_png(img, channels=3)\n    img = tf.image.resize(img, IMG_SIZE)\n    img = ben_graham_normalize(img)\n    img = tf.keras.applications.resnet50.preprocess_input(img)\n    label = tf.one_hot(label, NUM_CLASSES)\n    return img, label\n\ndata_augmentation = tf.keras.Sequential([\n    tf.keras.layers.RandomFlip(\"horizontal\"),\n    tf.keras.layers.RandomRotation(0.08),\n    tf.keras.layers.RandomZoom(0.15),\n    tf.keras.layers.RandomTranslation(0.15, 0.15),\n], name=\"augmentation\")\n\ndef make_dataset(df, training=False):\n    ds = tf.data.Dataset.from_tensor_slices((df['filepath'].values, df['diagnosis'].values))\n    if training:\n        ds = ds.shuffle(buffer_size=len(df), seed=SEED)\n    ds = ds.map(load_image, num_parallel_calls=AUTOTUNE)\n    if training:\n        ds = ds.map(lambda x, y: (data_augmentation(x, training=True), y), num_parallel_calls=AUTOTUNE)\n    ds = ds.batch(BATCH_SIZE).prefetch(AUTOTUNE)\n    return ds\n\ntrain_ds = make_dataset(train_split_df, training=True)\nval_ds = make_dataset(val_df, training=False)\ntest_ds = make_dataset(test_df, training=False)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:32.255243Z","iopub.execute_input":"2026-08-13T10:46:32.255511Z","iopub.status.idle":"2026-08-13T10:46:34.783365Z","shell.execute_reply.started":"2026-08-13T10:46:32.255475Z","shell.execute_reply":"2026-08-13T10:46:34.782694Z"}},"outputs":[],"execution_count":null},{"id":"841387e6","cell_type":"code","source":"# Visual check: does the normalization actually help? Compare raw vs normalized.\nraw = plt.imread(train_split_df.iloc[0]['filepath'])\nimg_norm, _ = load_image(train_split_df.iloc[0]['filepath'], train_split_df.iloc[0]['diagnosis'])\n# undo resnet preprocess_input roughly for display (mean-centered BGR) -> just show clipped raw normalized version instead\nimg_tf = tf.io.read_file(train_split_df.iloc[0]['filepath'])\nimg_tf = tf.image.decode_png(img_tf, channels=3)\nimg_tf = tf.image.resize(img_tf, IMG_SIZE)\nimg_display = ben_graham_normalize(img_tf).numpy().astype('uint8')\n\nfig, axes = plt.subplots(1, 2, figsize=(10, 5))\naxes[0].imshow(raw); axes[0].set_title(\"Raw\"); axes[0].axis('off')\naxes[1].imshow(img_display); axes[1].set_title(\"Ben Graham normalized\"); axes[1].axis('off')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:34.784348Z","iopub.execute_input":"2026-08-13T10:46:34.784672Z","iopub.status.idle":"2026-08-13T10:46:37.244412Z","shell.execute_reply.started":"2026-08-13T10:46:34.784638Z","shell.execute_reply":"2026-08-13T10:46:37.243441Z"}},"outputs":[],"execution_count":null},{"id":"8ffb1497","cell_type":"code","source":"# Class weights to counter DR-severity imbalance\nclass_weights_arr = compute_class_weight(\n    class_weight='balanced',\n    classes=np.arange(NUM_CLASSES),\n    y=train_split_df['diagnosis'].values\n)\nclass_weights = {i: w for i, w in enumerate(class_weights_arr)}\nprint(class_weights)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:37.245679Z","iopub.execute_input":"2026-08-13T10:46:37.246146Z","iopub.status.idle":"2026-08-13T10:46:37.252877Z","shell.execute_reply.started":"2026-08-13T10:46:37.246107Z","shell.execute_reply":"2026-08-13T10:46:37.252034Z"}},"outputs":[],"execution_count":null},{"id":"78f823fb","cell_type":"markdown","source":"## 6. Model — Fine-Tuned ResNet50 (v2: simplified head)\n\nChanges from v1:\n- Dropped the bolted-on 3-layer Conv2D stack (~10.7M randomly-initialized params — too much\n  fresh capacity for 2562 training images, was fighting the pretrained features).\n- GlobalAveragePooling2D straight off the ResNet50 backbone, standard transfer-learning head.\n- Backbone starts fully frozen; conv4+conv5 blocks unfrozen only in phase 2 (see training section).\n","metadata":{}},{"id":"a17ec9ed","cell_type":"code","source":"def build_model(num_classes=NUM_CLASSES):\n    base = tf.keras.applications.ResNet50(\n        include_top=False, weights='imagenet', input_shape=(224, 224, 3)\n    )\n    base.trainable = False  # fully frozen for phase 1\n\n    x = base.output\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dropout(0.5)(x)\n    x = tf.keras.layers.Dense(256, activation='relu')(x)\n    x = tf.keras.layers.Dropout(0.3)(x)\n    outputs = tf.keras.layers.Dense(num_classes, activation='softmax')(x)\n\n    model = tf.keras.Model(inputs=base.input, outputs=outputs)\n    return model, base\n\nmodel, backbone = build_model()\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n    loss='categorical_crossentropy',\n    metrics=['accuracy', tf.keras.metrics.AUC(name='auc', multi_label=True)]\n)\nmodel.summary()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:37.254181Z","iopub.execute_input":"2026-08-13T10:46:37.254585Z","iopub.status.idle":"2026-08-13T10:46:40.909216Z","shell.execute_reply.started":"2026-08-13T10:46:37.254547Z","shell.execute_reply":"2026-08-13T10:46:40.908564Z"}},"outputs":[],"execution_count":null},{"id":"b93fac54","cell_type":"markdown","source":"## 7. Two-Phase Training\n\n**Phase 1:** backbone fully frozen, train only the new head (fast, stabilizes the random-init layers before touching pretrained weights).\n**Phase 2:** unfreeze conv4_block* and conv5_block*, fine-tune the whole thing at a much lower LR so pretrained features aren't wrecked by large gradients from the still-adapting head.\n","metadata":{}},{"id":"ac1bf264","cell_type":"code","source":"os.makedirs(\"/kaggle/working/checkpoints\", exist_ok=True)\n\ndef make_callbacks(tag):\n    return [\n        tf.keras.callbacks.ModelCheckpoint(\n            f\"/kaggle/working/checkpoints/best_model_{tag}.keras\",\n            monitor='val_auc', mode='max', save_best_only=True, verbose=1\n        ),\n        tf.keras.callbacks.EarlyStopping(\n            monitor='val_auc', mode='max', patience=6, restore_best_weights=True, verbose=1\n        ),\n        tf.keras.callbacks.ReduceLROnPlateau(\n            monitor='val_loss', factor=0.5, patience=3, min_lr=1e-7, verbose=1\n        ),\n        tf.keras.callbacks.CSVLogger(f\"/kaggle/working/training_log_{tag}.csv\"),\n    ]\n\n# ---- Phase 1: head-only warmup ----\nPHASE1_EPOCHS = 6\n\nhistory1 = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=PHASE1_EPOCHS,\n    class_weight=class_weights,\n    callbacks=make_callbacks(\"phase1\")\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:46:40.910217Z","iopub.execute_input":"2026-08-13T10:46:40.910525Z","iopub.status.idle":"2026-08-13T10:59:05.564723Z","shell.execute_reply.started":"2026-08-13T10:46:40.910503Z","shell.execute_reply":"2026-08-13T10:59:05.563524Z"}},"outputs":[],"execution_count":null},{"id":"760f7f2c","cell_type":"code","source":"# ---- Phase 2: unfreeze conv4 + conv5, fine-tune at low LR ----\nfor layer in backbone.layers:\n    if layer.name.startswith('conv4_block') or layer.name.startswith('conv5_block'):\n        layer.trainable = True\n\ntrainable_count = sum(1 for l in backbone.layers if l.trainable)\nprint(f\"Unfrozen backbone layers: {trainable_count}\")\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-5),\n    loss='categorical_crossentropy',\n    metrics=['accuracy', tf.keras.metrics.AUC(name='auc', multi_label=True)]\n)\n\nPHASE2_EPOCHS = 20\n\nhistory2 = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=PHASE2_EPOCHS,\n    class_weight=class_weights,\n    callbacks=make_callbacks(\"phase2\")\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T10:59:05.565825Z","iopub.execute_input":"2026-08-13T10:59:05.566190Z","iopub.status.idle":"2026-08-13T11:38:09.875142Z","shell.execute_reply.started":"2026-08-13T10:59:05.566163Z","shell.execute_reply":"2026-08-13T11:38:09.874287Z"}},"outputs":[],"execution_count":null},{"id":"2d020457","cell_type":"markdown","source":"## 8. Training Curves (combined across both phases)","metadata":{}},{"id":"d809fed1","cell_type":"code","source":"hist1 = pd.DataFrame(history1.history)\nhist2 = pd.DataFrame(history2.history)\nhist1['phase'] = 1\nhist2['phase'] = 2\nhist_df = pd.concat([hist1, hist2], ignore_index=True)\nhist_df.to_csv(\"/kaggle/working/history_combined.csv\", index=False)\n\nphase_boundary = len(hist1)\n\nfig, axes = plt.subplots(1, 3, figsize=(18, 5))\nfor ax, metric in zip(axes, ['accuracy', 'loss', 'auc']):\n    ax.plot(hist_df[metric], label='train')\n    ax.plot(hist_df[f'val_{metric}'], label='val')\n    ax.axvline(phase_boundary - 0.5, color='gray', linestyle='--', alpha=0.5, label='phase boundary')\n    ax.set_title(metric)\n    ax.set_xlabel('epoch')\n    ax.legend()\nplt.tight_layout()\nplt.savefig(\"/kaggle/working/training_curves.png\", dpi=150)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:09.876316Z","iopub.execute_input":"2026-08-13T11:38:09.876697Z","iopub.status.idle":"2026-08-13T11:38:10.695925Z","shell.execute_reply.started":"2026-08-13T11:38:09.876673Z","shell.execute_reply":"2026-08-13T11:38:10.695314Z"}},"outputs":[],"execution_count":null},{"id":"0542e8c0","cell_type":"markdown","source":"## 9. Evaluation on Held-Out Test Set","metadata":{}},{"id":"0a0dce41","cell_type":"code","source":"best_model = tf.keras.models.load_model(\"/kaggle/working/checkpoints/best_model_phase2.keras\")\n\ny_true = test_df['diagnosis'].values\ny_pred_probs = best_model.predict(test_ds)\ny_pred = np.argmax(y_pred_probs, axis=1)\n\nacc = accuracy_score(y_true, y_pred)\nmacro_f1 = f1_score(y_true, y_pred, average='macro')\nweighted_f1 = f1_score(y_true, y_pred, average='weighted')\nmacro_prec = precision_score(y_true, y_pred, average='macro')\nmacro_rec = recall_score(y_true, y_pred, average='macro')\nqwk = cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\nprint(f\"Accuracy:                {acc:.4f}\")\nprint(f\"Macro F1:                {macro_f1:.4f}\")\nprint(f\"Weighted F1:             {weighted_f1:.4f}\")\nprint(f\"Macro Precision:         {macro_prec:.4f}\")\nprint(f\"Macro Recall:            {macro_rec:.4f}\")\nprint(f\"Quadratic Weighted Kappa: {qwk:.4f}  (ordinal agreement — the actual APTOS competition metric)\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:10.696984Z","iopub.execute_input":"2026-08-13T11:38:10.697423Z","iopub.status.idle":"2026-08-13T11:38:35.771757Z","shell.execute_reply.started":"2026-08-13T11:38:10.697400Z","shell.execute_reply":"2026-08-13T11:38:35.771151Z"}},"outputs":[],"execution_count":null},{"id":"c700aab0","cell_type":"code","source":"y_true_onehot = tf.one_hot(y_true, NUM_CLASSES).numpy()\nauc_ovr = roc_auc_score(y_true_onehot, y_pred_probs, average='macro', multi_class='ovr')\nprint(f\"Macro AUC (OvR): {auc_ovr:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:35.772594Z","iopub.execute_input":"2026-08-13T11:38:35.772909Z","iopub.status.idle":"2026-08-13T11:38:35.784920Z","shell.execute_reply.started":"2026-08-13T11:38:35.772885Z","shell.execute_reply":"2026-08-13T11:38:35.784340Z"}},"outputs":[],"execution_count":null},{"id":"0a76c112","cell_type":"code","source":"report = classification_report(\n    y_true, y_pred, target_names=[CLASS_NAMES[i] for i in range(NUM_CLASSES)],\n    output_dict=True\n)\nreport_df = pd.DataFrame(report).transpose()\nreport_df.to_csv(\"/kaggle/working/classification_report.csv\")\nreport_df\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:35.786123Z","iopub.execute_input":"2026-08-13T11:38:35.786437Z","iopub.status.idle":"2026-08-13T11:38:35.809225Z","shell.execute_reply.started":"2026-08-13T11:38:35.786415Z","shell.execute_reply":"2026-08-13T11:38:35.808357Z"}},"outputs":[],"execution_count":null},{"id":"ca4774ed","cell_type":"code","source":"cm = confusion_matrix(y_true, y_pred)\nplt.figure(figsize=(8, 6))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues',\n            xticklabels=[CLASS_NAMES[i] for i in range(NUM_CLASSES)],\n            yticklabels=[CLASS_NAMES[i] for i in range(NUM_CLASSES)])\nplt.title(f\"Confusion Matrix — Accuracy {acc*100:.2f}% | QWK {qwk:.3f}\")\nplt.xlabel(\"Predicted\")\nplt.ylabel(\"True\")\nplt.tight_layout()\nplt.savefig(\"/kaggle/working/confusion_matrix.png\", dpi=150)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:35.811615Z","iopub.execute_input":"2026-08-13T11:38:35.811878Z","iopub.status.idle":"2026-08-13T11:38:36.199735Z","shell.execute_reply.started":"2026-08-13T11:38:35.811858Z","shell.execute_reply":"2026-08-13T11:38:36.199007Z"}},"outputs":[],"execution_count":null},{"id":"3375ae99","cell_type":"code","source":"plt.figure(figsize=(8, 6))\nfor i in range(NUM_CLASSES):\n    fpr, tpr, _ = roc_curve(y_true_onehot[:, i], y_pred_probs[:, i])\n    class_auc = roc_auc_score(y_true_onehot[:, i], y_pred_probs[:, i])\n    plt.plot(fpr, tpr, label=f\"{CLASS_NAMES[i]} (AUC={class_auc:.3f})\")\nplt.plot([0, 1], [0, 1], 'k--', alpha=0.3)\nplt.xlabel(\"False Positive Rate\")\nplt.ylabel(\"True Positive Rate\")\nplt.title(\"Per-Class ROC Curves\")\nplt.legend()\nplt.tight_layout()\nplt.savefig(\"/kaggle/working/roc_curves.png\", dpi=150)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:36.200538Z","iopub.execute_input":"2026-08-13T11:38:36.200753Z","iopub.status.idle":"2026-08-13T11:38:36.595617Z","shell.execute_reply.started":"2026-08-13T11:38:36.200721Z","shell.execute_reply":"2026-08-13T11:38:36.594965Z"}},"outputs":[],"execution_count":null},{"id":"408df3c5","cell_type":"markdown","source":"## 10. Smoke Test — Single Forward Pass (pipeline sanity check)","metadata":{}},{"id":"cf43019b","cell_type":"code","source":"sample_path = test_df.iloc[0]['filepath']\nsample_label = test_df.iloc[0]['diagnosis']\nimg, _ = load_image(sample_path, sample_label)\nimg_batch = tf.expand_dims(img, 0)\npred = best_model.predict(img_batch, verbose=0)\npredicted_class = np.argmax(pred, axis=1)[0]\n\nprint(f\"True label: {CLASS_NAMES[sample_label]}\")\nprint(f\"Predicted:  {CLASS_NAMES[predicted_class]}\")\nprint(f\"Confidence: {pred[0][predicted_class]:.4f}\")\nassert pred.shape == (1, NUM_CLASSES), \"Output shape mismatch — pipeline broken\"\nprint(\"\\nSmoke test passed: pipeline runs end-to-end on a single image.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:36.596651Z","iopub.execute_input":"2026-08-13T11:38:36.596972Z","iopub.status.idle":"2026-08-13T11:38:41.091653Z","shell.execute_reply.started":"2026-08-13T11:38:36.596950Z","shell.execute_reply":"2026-08-13T11:38:41.090774Z"}},"outputs":[],"execution_count":null},{"id":"1fde7f16","cell_type":"markdown","source":"## 11. Summary & Reproducibility Notes","metadata":{}},{"id":"41de00c0","cell_type":"code","source":"summary = {\n    \"paper\": \"Rani & Gupta, IITCEE 2025, DOI 10.1109/IITCEE64140.2025.10915328\",\n    \"paper_task\": \"6-class retinal disease type classification (AMD, DR, DN, MH, ODC, Normal)\",\n    \"paper_dataset\": \"Not publicly released / unspecified source\",\n    \"this_run_dataset\": \"APTOS 2019 Blindness Detection (Kaggle, public)\",\n    \"this_run_task\": \"5-class DR severity grading (0-4)\",\n    \"reproduction_type\": \"Methodological (architecture reproduced), NOT numerical\",\n    \"architecture\": \"ResNet50 (ImageNet), phase1: fully frozen + GAP head; phase2: conv4+conv5 unfrozen, fine-tuned at 1e-5\",\n    \"preprocessing\": \"224x224 resize + Ben Graham local color normalization\",\n    \"phase1_epochs\": PHASE1_EPOCHS,\n    \"phase2_epochs\": PHASE2_EPOCHS,\n    \"seed\": SEED,\n    \"test_accuracy\": float(acc),\n    \"macro_f1\": float(macro_f1),\n    \"macro_auc_ovr\": float(auc_ovr),\n    \"quadratic_weighted_kappa\": float(qwk),\n    \"known_limitations\": [\n        \"No patient/eye ID in APTOS 2019 train.csv — patient-independence of split not verifiable\",\n        \"Label space differs from paper (DR severity vs disease type) — results not numerically comparable\",\n        \"Paper's exact learning rate, batch size, and layer-freezing boundary were not stated; reasonable defaults used\"\n    ]\n}\n\nwith open(\"/kaggle/working/run_summary.json\", \"w\") as f:\n    json.dump(summary, f, indent=2)\n\nprint(json.dumps(summary, indent=2))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-13T11:38:41.092760Z","iopub.execute_input":"2026-08-13T11:38:41.093265Z","iopub.status.idle":"2026-08-13T11:38:41.099909Z","shell.execute_reply.started":"2026-08-13T11:38:41.093238Z","shell.execute_reply":"2026-08-13T11:38:41.099017Z"}},"outputs":[],"execution_count":null}]}