{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cross-Dataset Generalization of Deep Learning Models for Diabetic Retinopathy Grading\n\n**Dissertation notebook** — trains a baseline EfficientNetB0 classifier on APTOS 2019 and\nevaluates zero-shot generalization on Messidor-2 and IDRiD.\n\n**Contents**\n1. Setup & Imports\n2. Data Loading\n3. Label & Shape Verification\n4. Data Cleaning — Deduplication\n5. Exploratory Data Analysis\n6. Train / Val / Test Split (APTOS)\n7. Image Loading into Arrays\n8. Class Weights\n9. Baseline Model — EfficientNetB0\n10. Training Curves\n11. External Dataset Preparation (Messidor-2, IDRiD)\n12. Evaluation — RQ1 (APTOS, in-distribution)\n13. Evaluation — RQ2 & RQ4 (Cross-dataset generalization)\n14. Results Summary\n    \n    14.1. Bootstrap Confidence Interval\n    \n    14.2. Ablation: Class Weighting vs No Class Weighting\n\n    14.3. Ablation: Frozen vs Fine-Tuned Layers\n\n    14.4. Multi-seed Stability Analysis\n    \n15. Mitigation Technique — Test-Time Augmentation (RQ3)\n\n    15.1. Repeated TTA Evaluation for Uncertainty\n    \n17. Appendix — Models Not Used","metadata":{"_uuid":"e3248c86-11ee-42ef-9409-3f7bebf91de9","_cell_guid":"aa627205-0101-4f15-a71d-25743cc1aadb","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"This notebook implements the full experimental pipeline for the dissertation:\ntraining an EfficientNetB0 baseline on APTOS 2019, evaluating zero-shot on\nMessidor-2 and IDRiD, testing a class-weighting and a frozen-layers ablation,\nrunning a 6-seed stability analysis, and evaluating Test-Time Augmentation (TTA)\nas a lightweight mitigation technique.\n\n**To reproduce:** run all cells sequentially from top to bottom. Random seeds\nare fixed (Section 1), though residual GPU-level nondeterminism means exact\nvalues may vary slightly between runs — see Section 14.2 for the documented range.","metadata":{}},{"cell_type":"markdown","source":"## 1. Setup & Imports","metadata":{"_uuid":"9aa9c083-9723-4f00-b8d5-ae231b936f6c","_cell_guid":"5d186b2b-face-4982-9290-3474323555b7","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport hashlib\nimport os\nimport warnings\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport tensorflow as tf\nfrom PIL import Image\nfrom tqdm import tqdm\nimport time\nimport sklearn\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, BatchNormalization, GlobalAveragePooling2D, Dense, Dropout\nfrom tensorflow.keras.models import Sequential, Model, load_model\nfrom tensorflow.keras.applications import EfficientNetB0, EfficientNetB7\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils.class_weight import compute_class_weight\nfrom sklearn.metrics import classification_report, cohen_kappa_score,confusion_matrix, recall_score, f1_score, accuracy_score\nfrom tensorflow.keras.callbacks import EarlyStopping, ModelCheckpoint, ReduceLROnPlateau\n\n\nwarnings.filterwarnings(\"ignore\")\nprint(\"GPU available:\", tf.config.list_physical_devices('GPU'))","metadata":{"_uuid":"848fff04-7593-4614-97b2-0e442a0d1bec","_cell_guid":"a3c89347-1b92-40e7-9ecb-770ff6e7747b","trusted":true,"execution":{"iopub.status.busy":"2026-09-06T23:19:53.876046Z","iopub.execute_input":"2026-09-06T23:19:53.876498Z","iopub.status.idle":"2026-09-06T23:20:09.215117Z","shell.execute_reply.started":"2026-09-06T23:19:53.876466Z","shell.execute_reply":"2026-09-06T23:20:09.214161Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f\"TensorFlow: {tf.__version__}\")\nprint(f\"scikit-learn: {sklearn.__version__}\")\nprint(f\"NumPy: {np.__version__}\")\nprint(f\"Pandas: {pd.__version__}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T23:20:09.216641Z","iopub.execute_input":"2026-09-06T23:20:09.217350Z","iopub.status.idle":"2026-09-06T23:20:09.223881Z","shell.execute_reply.started":"2026-09-06T23:20:09.217314Z","shell.execute_reply":"2026-09-06T23:20:09.222933Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tf.random.set_seed(42)\nnp.random.seed(42)\n\nos.environ['TF_DETERMINISTIC_OPS'] = '1'\nos.environ['TF_CUDNN_DETERMINISTIC'] = '1'\ntf.config.experimental.enable_op_determinism()","metadata":{"_uuid":"52978423-7062-4d19-b3b8-efac16393c35","_cell_guid":"63ff5a71-6e17-43bc-9774-93e606409c5c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.427188Z","iopub.execute_input":"2026-09-06T19:34:32.427743Z","iopub.status.idle":"2026-09-06T19:34:32.432696Z","shell.execute_reply.started":"2026-09-06T19:34:32.427718Z","shell.execute_reply":"2026-09-06T19:34:32.431542Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Data Loading","metadata":{"_uuid":"a68f9a42-a5ed-4565-98c4-95de722f0d14","_cell_guid":"ebc079e2-d532-48ff-b803-252dce3837ab","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"apt_df=pd.read_csv(\"/kaggle/input/competitions/aptos2019-blindness-detection/train.csv\")\nms_df=pd.read_csv(\"/kaggle/input/datasets/mariaherrerot/messidor2preprocess/messidor_data.csv\", usecols=['id_code', 'diagnosis'])\nidr_df=pd.read_csv(\"/kaggle/input/datasets/mariaherrerot/idrid-dataset/idrid_labels.csv\", usecols=['id_code', 'diagnosis'])","metadata":{"_uuid":"3bbeeeee-f0f4-4073-9a89-dfc4910b1621","_cell_guid":"b727c766-662a-46fb-ba34-6208816723c0","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.433629Z","iopub.execute_input":"2026-09-06T19:34:32.434139Z","iopub.status.idle":"2026-09-06T19:34:32.504408Z","shell.execute_reply.started":"2026-09-06T19:34:32.434100Z","shell.execute_reply":"2026-09-06T19:34:32.503592Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"apt_df.head()","metadata":{"_uuid":"5091a5f3-bb43-4fe3-9e36-867fe5b7f12c","_cell_guid":"a896bb42-e88d-4cb7-9318-d55a97d73eff","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.505349Z","iopub.execute_input":"2026-09-06T19:34:32.505619Z","iopub.status.idle":"2026-09-06T19:34:32.530587Z","shell.execute_reply.started":"2026-09-06T19:34:32.505591Z","shell.execute_reply":"2026-09-06T19:34:32.529766Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"ms_df.head()","metadata":{"_uuid":"e32cb805-7193-427b-9415-9e01862cab6e","_cell_guid":"4a65a9cd-e7c1-4114-8a80-1763a7f3835a","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.532605Z","iopub.execute_input":"2026-09-06T19:34:32.532897Z","iopub.status.idle":"2026-09-06T19:34:32.539814Z","shell.execute_reply.started":"2026-09-06T19:34:32.532875Z","shell.execute_reply":"2026-09-06T19:34:32.539102Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"idr_df.head()","metadata":{"_uuid":"8534dc53-f579-40e0-8ed7-09b78da63fe0","_cell_guid":"8e48ac0c-af87-4f41-9140-d3f0a9ab613d","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.540759Z","iopub.execute_input":"2026-09-06T19:34:32.541062Z","iopub.status.idle":"2026-09-06T19:34:32.555510Z","shell.execute_reply.started":"2026-09-06T19:34:32.541040Z","shell.execute_reply":"2026-09-06T19:34:32.554786Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Label & Shape Verification\n\nConfirms all three datasets use the same 0–4 ICDR-style severity scale, so no label harmonisation is required.","metadata":{"_uuid":"e65b01fd-7158-42b0-98b9-25f97b030aeb","_cell_guid":"f0705997-ce22-4e9b-a174-c1d8c539c74c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"print(\"APTOS-2019 target feature labels:\",np.sort(apt_df.diagnosis.unique()))\nprint(\"Messidor-2 target feature labels:\",np.sort(ms_df.diagnosis.unique()))\nprint(\"IDRid: DR-Grading target feature labels:\",np.sort(idr_df.diagnosis.unique()))","metadata":{"_uuid":"32f48c17-7236-44bb-9774-24af901f5e2b","_cell_guid":"357b88ae-a011-4cc8-8e0d-458f085aff84","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.556435Z","iopub.execute_input":"2026-09-06T19:34:32.556685Z","iopub.status.idle":"2026-09-06T19:34:32.573573Z","shell.execute_reply.started":"2026-09-06T19:34:32.556654Z","shell.execute_reply":"2026-09-06T19:34:32.572680Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"APTOS-2019 dataset shape:\", apt_df.shape)\nprint(\"Messidor-2 dataset shape:\", ms_df.shape)\nprint(\"IDRid: DR-Grading dataset shape:\", idr_df.shape)","metadata":{"_uuid":"e4c9cbda-c4c7-428e-9430-3edb2bc87a74","_cell_guid":"84096813-7362-4c37-8990-b604ff09d45e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.574604Z","iopub.execute_input":"2026-09-06T19:34:32.574898Z","iopub.status.idle":"2026-09-06T19:34:32.583121Z","shell.execute_reply.started":"2026-09-06T19:34:32.574866Z","shell.execute_reply":"2026-09-06T19:34:32.582277Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Data Cleaning — Deduplication\n\nAPTOS 2019 provides no patient/eye identifier, so patient-level splitting (as recommended in the\nliterature) is not verifiable from the available metadata. As a proxy safeguard, images are\ndeduplicated by MD5 content hash before splitting, to prevent the same photo appearing in both\ntrain and test partitions.","metadata":{"_uuid":"bff00983-5dee-499b-a7d0-d3b292be733f","_cell_guid":"ad289678-96a8-41fb-8ffe-ab1e96604ebc","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"apt_df.duplicated().sum()","metadata":{"_uuid":"cea628b6-9296-4b5e-9acc-9bc213072249","_cell_guid":"dce8215a-c3b7-4e52-a68c-b9e195bb378c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.584038Z","iopub.execute_input":"2026-09-06T19:34:32.584350Z","iopub.status.idle":"2026-09-06T19:34:32.600020Z","shell.execute_reply.started":"2026-09-06T19:34:32.584319Z","shell.execute_reply":"2026-09-06T19:34:32.599411Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img_dir = '/kaggle/input/competitions/aptos2019-blindness-detection/train_images'\n\nhashes = {}\nduplicate_ids = []\n\nfor filename in apt_df['id_code']:\n    filepath = os.path.join(img_dir, f\"{filename}.png\")\n    with open(filepath, 'rb') as f:\n        h = hashlib.md5(f.read()).hexdigest()\n    if h in hashes:\n        duplicate_ids.append(filename)\n    else:\n        hashes[h] = filename\n\nprint(f\"Duplicate images found in APTOS-2019 Dataset (by actual content): {len(duplicate_ids)}\")","metadata":{"_uuid":"1e3d41ab-de92-4459-8f2e-bb1b38e69ccd","_cell_guid":"2694e555-9b5b-452e-943e-d4d9f7c1ea37","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:34:32.600836Z","iopub.execute_input":"2026-09-06T19:34:32.601179Z","iopub.status.idle":"2026-09-06T19:36:02.074020Z","shell.execute_reply.started":"2026-09-06T19:34:32.601141Z","shell.execute_reply":"2026-09-06T19:36:02.073082Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Remove duplicates, keeping the first occurrence of each\nclean_df = apt_df[~apt_df['id_code'].isin(duplicate_ids)].reset_index(drop=True)\nprint(f\"Original: {len(apt_df)}  ->  After dedup: {len(clean_df)}\")","metadata":{"_uuid":"f369ea6f-befc-4dfa-9468-f22b0bde2bf7","_cell_guid":"3dda1871-a79b-4f35-8a28-8b90085f1b44","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:36:02.075101Z","iopub.execute_input":"2026-09-06T19:36:02.075423Z","iopub.status.idle":"2026-09-06T19:36:02.086980Z","shell.execute_reply.started":"2026-09-06T19:36:02.075386Z","shell.execute_reply":"2026-09-06T19:36:02.086277Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Exploratory Data Analysis","metadata":{"_uuid":"1a87de63-04ac-499a-ae23-cb3dd4e229b8","_cell_guid":"46f09169-a645-4ace-86d3-37e3d5b9eeea","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 4))\nfor ax, df, name in zip(axes, [clean_df, ms_df, idr_df], ['APTOS 2019 (deduped)', 'Messidor-2', 'IDRiD']):\n    sns.countplot(x='diagnosis', data=df, ax=ax, palette='viridis')\n    ax.set_title(f'{name} (n={len(df)})')\n    ax.set_xlabel('DR Severity Grade')\nplt.tight_layout()\n# plt.savefig('/kaggle/working/class_distributions.png', dpi=150)\nplt.show()","metadata":{"_uuid":"80c30389-bd16-4c71-9151-850861f239a0","_cell_guid":"d9b6c760-e3c9-4a31-ba3a-f5df8889ff51","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:36:02.088182Z","iopub.execute_input":"2026-09-06T19:36:02.088731Z","iopub.status.idle":"2026-09-06T19:36:02.583941Z","shell.execute_reply.started":"2026-09-06T19:36:02.088705Z","shell.execute_reply":"2026-09-06T19:36:02.583317Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 5, figsize=(15, 3))\nfor grade in range(5):\n    sample = clean_df[clean_df['diagnosis'] == grade].iloc[0]\n    img = Image.open(f\"{img_dir}/{sample['id_code']}.png\")\n    axes[grade].imshow(img)\n    axes[grade].set_title(f'Grade {grade}')\n    axes[grade].axis('off')\nplt.tight_layout()\n# plt.savefig('/kaggle/working/sample_images_by_grade.png', dpi=150)\nplt.show()","metadata":{"_uuid":"bcc311f9-2037-4a25-85c7-7e2f9faff318","_cell_guid":"425bda72-04cf-46a1-a5bd-9fd301823c1d","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:36:02.584948Z","iopub.execute_input":"2026-09-06T19:36:02.585318Z","iopub.status.idle":"2026-09-06T19:36:04.749216Z","shell.execute_reply.started":"2026-09-06T19:36:02.585286Z","shell.execute_reply":"2026-09-06T19:36:04.748317Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Train / Val / Test Split (APTOS)\n\nStratified 70/15/15 split on the deduplicated pool, preserving class proportions.","metadata":{"_uuid":"9c5e4126-cc96-4a4a-853d-905aa332e6c4","_cell_guid":"e4067258-e3e1-40c9-a258-23af9fdd8c37","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# First split off test (15%)\ntrain_val_df, test_df = train_test_split(\n    clean_df, test_size=0.15, stratify=clean_df['diagnosis'], random_state=42\n)\n\n# Then split remaining into train (70% of total) / val (15% of total)\ntrain_df_final, val_df = train_test_split(\n    train_val_df, test_size=0.1765,  # 0.15 / 0.85 ≈ 0.1765\n    stratify=train_val_df['diagnosis'], random_state=42\n)\n\nprint(f\"Train: {len(train_df_final)}, Val: {len(val_df)}, Test: {len(test_df)}\")\nfor name, df in [('Train', train_df_final), ('Val', val_df), ('Test', test_df)]:\n    print(f\"\\n{name} class distribution:\")\n    print(df['diagnosis'].value_counts(normalize=True).sort_index().round(3))\n\ntrain_df_final.to_csv('/kaggle/working/aptos_train.csv', index=False)\nval_df.to_csv('/kaggle/working/aptos_val.csv', index=False)\ntest_df.to_csv('/kaggle/working/aptos_test.csv', index=False)","metadata":{"_uuid":"d4c9b65e-35cd-473f-8909-11bf149bd94c","_cell_guid":"6c521c3f-6d4b-4b10-b311-a233a08172c7","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:36:04.752181Z","iopub.execute_input":"2026-09-06T19:36:04.752547Z","iopub.status.idle":"2026-09-06T19:36:04.787422Z","shell.execute_reply.started":"2026-09-06T19:36:04.752520Z","shell.execute_reply":"2026-09-06T19:36:04.786740Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Image Loading into Arrays","metadata":{"_uuid":"37eca1a5-d78b-40be-849b-f0a0464f5131","_cell_guid":"ae8557ad-5e10-4e23-b5ec-bd569f36cfc6","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"IMG_SIZE = 224\nIMG_DIR = '/kaggle/input/competitions/aptos2019-blindness-detection/train_images'\n\ntrain_df_final = pd.read_csv('/kaggle/working/aptos_train.csv')\nval_df = pd.read_csv('/kaggle/working/aptos_val.csv')\ntest_df = pd.read_csv('/kaggle/working/aptos_test.csv')\n\n# General image loading to array function\ndef load_images_to_array(df, img_dir, img_format, img_size):\n    images = np.zeros((len(df), img_size, img_size, 3), dtype=np.float32)\n    labels = df['diagnosis'].astype(int).values\n    for i, fname in enumerate(tqdm(df['id_code'])):\n        img = Image.open(f\"{img_dir}/{fname}.{img_format}\").convert('RGB').resize((img_size, img_size))\n        images[i] = np.array(img) / 255.0\n    return images, labels\n\n\n# Specific image loading to array function for Messidor-2 dataset\ndef load_images_to_array2(df, img_dir, img_size):\n    images = np.zeros((len(df), img_size, img_size, 3), dtype=np.float32)\n    labels = df['diagnosis'].astype(int).values\n    for i, fname in enumerate(tqdm(df['id_code'])):\n        img = Image.open(f\"{img_dir}/{fname}\").convert('RGB').resize((img_size, img_size))\n        images[i] = np.array(img) / 255.0\n    return images, labels\n\nstart = time.time()\nX_train, y_train = load_images_to_array(train_df_final, IMG_DIR, 'png',IMG_SIZE)\nprint(f\"Train images loaded in {time.time()-start:.1f}s, shape: {X_train.shape}\")\n\nstart = time.time()\nX_val, y_val = load_images_to_array(val_df, IMG_DIR, 'png', IMG_SIZE)\nX_test, y_test = load_images_to_array(test_df, IMG_DIR, 'png', IMG_SIZE)\nprint(f\"Val/test images loaded in {time.time()-start:.1f}s\")","metadata":{"_uuid":"95b984b6-d3bf-489e-ba7b-2d6b1f297002","_cell_guid":"570a284b-1b58-4992-838a-cb622f57e3c5","trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:36:04.788336Z","iopub.execute_input":"2026-09-06T19:36:04.788620Z","iopub.status.idle":"2026-09-06T19:43:09.066896Z","shell.execute_reply.started":"2026-09-06T19:36:04.788574Z","shell.execute_reply":"2026-09-06T19:43:09.066062Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"X_train:\", X_train.shape, \"y_train:\", y_train.shape)\nprint(\"X_val:\", X_val.shape, \"y_val:\", y_val.shape)\nprint(\"X_test:\", X_test.shape, \"y_test:\", y_test.shape)","metadata":{"_uuid":"1d1684c9-9f7f-41ea-b907-f9cd8cc3a785","_cell_guid":"f677b194-692a-4371-bfda-45e9bf086873","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:43:09.067869Z","iopub.execute_input":"2026-09-06T19:43:09.068116Z","iopub.status.idle":"2026-09-06T19:43:09.072817Z","shell.execute_reply.started":"2026-09-06T19:43:09.068094Z","shell.execute_reply":"2026-09-06T19:43:09.072106Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Class Weights\n\nCompensates for APTOS's class imbalance (~51% class 0 vs ~5% class 3) during training.","metadata":{"_uuid":"a97aabf0-7baa-4f44-b588-d5dc00e7c795","_cell_guid":"3e58202b-2587-451c-aa26-d99061a1f5c2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"class_weights = compute_class_weight('balanced', classes=np.unique(y_train), y=y_train)\nclass_weight_dict = dict(enumerate(class_weights))\nprint(class_weight_dict)","metadata":{"_uuid":"37e4dfd9-67aa-42a6-ae34-642abf4b34b4","_cell_guid":"97fda55b-9d16-4d9a-8299-76a7e573bbf1","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:43:09.073890Z","iopub.execute_input":"2026-09-06T19:43:09.074221Z","iopub.status.idle":"2026-09-06T19:43:09.088113Z","shell.execute_reply.started":"2026-09-06T19:43:09.074200Z","shell.execute_reply":"2026-09-06T19:43:09.087404Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9. Baseline Model — EfficientNetB0\n\nEfficientNetB0 chosen as the sole baseline architecture per the approved dissertation scope\n(single-model pipeline, protecting time for cross-dataset evaluation and the mitigation technique).\n\n**Note on preprocessing:** Keras's `EfficientNetB0` includes normalisation as part of the model\nitself and expects raw pixel values in `[0, 255]`. Since `X_train`/`X_val`/`X_test` were loaded as\n`[0, 1]` (divided by 255 above for consistency with the other model variants in the Appendix), they\nare multiplied back by 255 at `.fit()` / `.predict()` time below. This must be applied consistently\nat both training and inference time — forgetting it at inference is what caused the invalid,\nall-class-0 predictions in an earlier notebook version.","metadata":{"_uuid":"349b56e4-bfaf-48cf-9fbf-fb59fcaed969","_cell_guid":"f90754f7-9734-4006-b18a-b6e77aff13d9","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"NUM_CLASSES=5\n\nenb0_base_model=EfficientNetB0(input_shape=(IMG_SIZE, IMG_SIZE, 3), include_top=False, weights='imagenet')\nenb0_base_model.trainable=True\n\nfor layer in enb0_base_model.layers[:-10]:  # keep most of the base frozen, unfreeze last ~10 layers\n    layer.trainable = False\n\nx = GlobalAveragePooling2D()(enb0_base_model.output)\nx = Dropout(0.3)(x)\nx = Dense(128, activation='relu')(x)\nx = Dropout(0.2)(x)\noutput = Dense(NUM_CLASSES, activation='softmax')(x)\nenb0_model = Model(inputs=enb0_base_model.input, outputs=output)\n\nenb0_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n              loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n\n\n\nenb0_callbacks = [\n    EarlyStopping(monitor='val_accuracy', patience=6, verbose=1, restore_best_weights=True),\n    ModelCheckpoint('/kaggle/working/best_enb0_cnn.keras', monitor='val_accuracy', save_best_only=True),\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3)\n]\n\nenb0_history = enb0_model.fit(\n    X_train*255., y_train,\n    validation_data=(X_val*255., y_val),\n    batch_size=64,\n    epochs=40,                    # early stopping will cut this short automatically\n    class_weight=class_weight_dict,\n    callbacks=enb0_callbacks\n)","metadata":{"_uuid":"868b30a1-6c38-4434-b084-bbf300076a0d","_cell_guid":"8bae572d-cfda-458b-8547-1eea2b3abf04","trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:43:09.089221Z","iopub.execute_input":"2026-09-06T19:43:09.089621Z","iopub.status.idle":"2026-09-06T19:46:50.598100Z","shell.execute_reply.started":"2026-09-06T19:43:09.089600Z","shell.execute_reply":"2026-09-06T19:46:50.597379Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 10. Training Curves","metadata":{"_uuid":"dc5080e0-7351-4380-b6ef-1bd3c0965dc9","_cell_guid":"ae917fc1-8316-48b3-ab73-9cdac781b615","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"enb0_epoch=[epoch for epoch in range(len(enb0_history.history['accuracy']))]\n\nfig,ax=plt.subplots(1,2, figsize=(13,5))\n\n# EfficientNetB0 History\nsns.lineplot(x=enb0_epoch, y=enb0_history.history['accuracy'], ax=ax[0], label='Training Accuracy')\nsns.lineplot(x=enb0_epoch, y=enb0_history.history['val_accuracy'], ax=ax[0], label='Validation Accuracy')\nax[0].set_title('EfficientNetB0 Model Accuracy')\n\nsns.lineplot(x=enb0_epoch, y=enb0_history.history['loss'], ax=ax[1], label='Training Loss')\nsns.lineplot(x=enb0_epoch, y=enb0_history.history['val_loss'], ax=ax[1], label='Validation Loss')\nax[1].set_title('EfficientNetB0 Model Loss')\n\n\nplt.tight_layout()\nplt.legend()\nplt.show()","metadata":{"_uuid":"9ee3e38c-f783-4f72-b558-956b762adbb3","_cell_guid":"467f3c53-801e-4e98-aa5d-fa512af6774c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:46:50.599828Z","iopub.execute_input":"2026-09-06T19:46:50.600085Z","iopub.status.idle":"2026-09-06T19:46:50.957315Z","shell.execute_reply.started":"2026-09-06T19:46:50.600063Z","shell.execute_reply":"2026-09-06T19:46:50.956153Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 11. External Dataset Preparation (Messidor-2, IDRiD)\n\nBoth datasets are used purely as external, zero-shot test sets — no fine-tuning, no internal\ntrain/val/test split of their own. They are deduplicated for consistency with the APTOS\nmethodology, though this only matters if either is ever reused for training in future work.","metadata":{"_uuid":"82b2cf17-21e1-4b34-af26-c49a1f719d9a","_cell_guid":"5dceabc9-1aa3-4f4e-96c3-a35675e9f7e0","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# IDRiD: Diabetic Retinopathy - Grading\nidr_img_dir = '/kaggle/input/datasets/mariaherrerot/idrid-dataset/Imagenes/Imagenes'\n\nidr_hashes = {}\nidr_duplicate_ids = []\n\n\nfor filename in idr_df['id_code']:\n    idr_filepath = os.path.join(idr_img_dir, f\"{filename}.jpg\")\n    with open(idr_filepath, 'rb') as f:\n        h = hashlib.md5(f.read()).hexdigest()\n    if h in idr_hashes:\n        idr_duplicate_ids.append(filename)\n    else:\n        idr_hashes[h] = filename\n\nprint(f\"Duplicate images found (by actual content): {len(idr_duplicate_ids)}\")\n\n\n# Messidor-2\nms_img_dir = '/kaggle/input/datasets/mariaherrerot/messidor2preprocess/messidor-2/messidor-2/preprocess'\n\nms_hashes = {}\nms_duplicate_ids = []\n\n\nfor filename in ms_df['id_code']:\n    ms_filepath = os.path.join(ms_img_dir, f\"{filename}\")\n    with open(ms_filepath, 'rb') as f:\n        h = hashlib.md5(f.read()).hexdigest()\n    if h in ms_hashes:\n        ms_duplicate_ids.append(filename)\n    else:\n        ms_hashes[h] = filename\n\nprint(f\"Duplicate images found (by actual content): {len(ms_duplicate_ids)}\")","metadata":{"_uuid":"a69cd4f4-1e34-49b9-b23b-f01430bd2599","_cell_guid":"880d86ca-771d-4b7c-8057-6e0a0b5db863","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:46:50.958604Z","iopub.execute_input":"2026-09-06T19:46:50.958943Z","iopub.status.idle":"2026-09-06T19:47:09.428485Z","shell.execute_reply.started":"2026-09-06T19:46:50.958909Z","shell.execute_reply":"2026-09-06T19:47:09.427748Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# IDRiD dataset: Remove duplicates, keeping the first occurrence of each\nidr_clean_df = idr_df[~idr_df['id_code'].isin(idr_duplicate_ids)].reset_index(drop=True)\nprint(f\"Original: {len(idr_df)}  ->  After dedup: {len(idr_clean_df)}\")\n\n# Messidor dataset: Remove duplicates, keeping the first occurrence of each\nms_clean_df = ms_df[~ms_df['id_code'].isin(ms_duplicate_ids)].reset_index(drop=True)\nprint(f\"Original: {len(ms_df)}  ->  After dedup: {len(ms_clean_df)}\")\n\n\nidr_clean_df.to_csv('/kaggle/working/IDRiD_test.csv', index=False)\nms_clean_df.to_csv(\"/kaggle/working/Messidor_test.csv\", index=False)","metadata":{"_uuid":"bed7a532-6cfb-4166-9024-de91c95e760a","_cell_guid":"a7a823d5-8c6c-465a-bb37-11cc193c5cc1","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:47:09.429382Z","iopub.execute_input":"2026-09-06T19:47:09.429662Z","iopub.status.idle":"2026-09-06T19:47:09.441427Z","shell.execute_reply.started":"2026-09-06T19:47:09.429639Z","shell.execute_reply":"2026-09-06T19:47:09.440639Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Reload the best checkpointed model from disk (rather than relying on the in-memory object)\n# — more reproducible, since it confirms the saved artifact is what actually gets evaluated.\nenb0_model_saved=load_model('/kaggle/working/best_enb0_cnn.keras')","metadata":{"_uuid":"2324c74b-c972-4240-9730-e97b81bc2930","_cell_guid":"54f566f3-afbd-45a1-80ad-a114d1bdb0e5","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:47:09.442438Z","iopub.execute_input":"2026-09-06T19:47:09.442712Z","iopub.status.idle":"2026-09-06T19:47:10.745878Z","shell.execute_reply.started":"2026-09-06T19:47:09.442688Z","shell.execute_reply":"2026-09-06T19:47:10.744915Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"idr_clean_df=pd.read_csv('/kaggle/working/IDRiD_test.csv')\nms_clean_df=pd.read_csv('/kaggle/working/Messidor_test.csv')\n\nIDR_IMG_DIR='/kaggle/input/datasets/mariaherrerot/idrid-dataset/Imagenes/Imagenes'\nMS_IMG_DIR='/kaggle/input/datasets/mariaherrerot/messidor2preprocess/messidor-2/messidor-2/preprocess'\n\nidr_X_test, idr_y_test = load_images_to_array(idr_clean_df, IDR_IMG_DIR,'jpg', IMG_SIZE)\nms_X_test, ms_y_test = load_images_to_array2(ms_clean_df, MS_IMG_DIR, IMG_SIZE)","metadata":{"_uuid":"9df72b2d-0802-40bf-9907-9459afd9b5d8","_cell_guid":"a08facf1-347b-45ab-8d1d-bbf533a4f989","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:47:10.747171Z","iopub.execute_input":"2026-09-06T19:47:10.747720Z","iopub.status.idle":"2026-09-06T19:48:50.219046Z","shell.execute_reply.started":"2026-09-06T19:47:10.747685Z","shell.execute_reply":"2026-09-06T19:48:50.218442Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 12. Evaluation — RQ1 (APTOS, in-distribution)\n\n*RQ1: How does a baseline deep learning model perform on the APTOS 2019 dataset under standard\nevaluation?*\n\nEstablishes the baseline against which cross-dataset performance (Section 13) is compared.\nReports precision, recall, F1 per class, and quadratic weighted kappa — the same metric used\nby the original APTOS competition, for comparability with published results.","metadata":{"_uuid":"f7357be4-088a-46f2-96e7-ec5ad307abbb","_cell_guid":"44055078-a5ab-41f8-8432-2bd9e84244ab","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"aptos_pred_probs=enb0_model_saved.predict(X_test*255.)\naptos_pred=np.argmax(aptos_pred_probs, axis=1)\n\nprint(classification_report(y_test, aptos_pred, digits=3))\nprint(\"APTOS Quadratic Weighted kappa score:\", cohen_kappa_score(y_test, aptos_pred, weights='quadratic'))","metadata":{"_uuid":"6a938f19-fee4-444c-a17e-fc57a450c397","_cell_guid":"f161750b-758a-440c-8727-c2739534f149","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:48:50.220156Z","iopub.execute_input":"2026-09-06T19:48:50.220453Z","iopub.status.idle":"2026-09-06T19:48:56.458217Z","shell.execute_reply.started":"2026-09-06T19:48:50.220430Z","shell.execute_reply":"2026-09-06T19:48:56.457590Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 13. Evaluation — RQ2 & RQ4 (Cross-dataset generalization)\n\n*RQ2: How does performance change when evaluated on Messidor-2 and IDRiD without fine-tuning?*\n\n*RQ4: Which severity classes are most affected by the distribution shift?*\n\nThe baseline model from Section 12 is applied directly to Messidor-2 and IDRiD with no\nretraining or fine-tuning — a zero-shot test of generalization to unseen clinics, devices, and\ngrading protocols. Per-class metrics (not just accuracy) are used throughout to answer RQ4.","metadata":{"_uuid":"44f65581-a0f8-4cc6-aba8-ab135cc6aa38","_cell_guid":"863a70e3-9e04-458e-bc54-16d22be4f1da","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"idr_pred_probs = enb0_model_saved.predict(idr_X_test * 255.)\nidr_pred = np.argmax(idr_pred_probs, axis=1)\n\nms_pred_probs = enb0_model_saved.predict(ms_X_test * 255.)\nms_pred = np.argmax(ms_pred_probs, axis=1)\n\nprint(\"=== IDRiD ===\")\nprint(classification_report(idr_y_test, idr_pred, digits=3))\nprint(\"IDRiD Quadratic Weighted Kappa:\", cohen_kappa_score(idr_y_test, idr_pred, weights='quadratic'))\n\nprint(\"\\n=== Messidor-2 ===\")\nprint(classification_report(ms_y_test, ms_pred, digits=3))\nprint(\"Messidor-2 Quadratic Weighted Kappa:\", cohen_kappa_score(ms_y_test, ms_pred, weights='quadratic'))","metadata":{"_uuid":"e8168d44-6a56-4ff5-b721-1c60ba91591e","_cell_guid":"55f4c12e-434d-486a-96b5-bd4b1708175b","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:48:56.459309Z","iopub.execute_input":"2026-09-06T19:48:56.459627Z","iopub.status.idle":"2026-09-06T19:49:06.693293Z","shell.execute_reply.started":"2026-09-06T19:48:56.459587Z","shell.execute_reply":"2026-09-06T19:49:06.692321Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(18, 5))\ndatasets = [\n    ('APTOS (in-dist.)', y_test, aptos_pred),\n    ('IDRiD', idr_y_test, idr_pred),\n    ('Messidor-2', ms_y_test, ms_pred),\n]\nfor ax, (name, y_true, y_pred) in zip(axes, datasets):\n    cm = confusion_matrix(y_true, y_pred, normalize='true')\n    sns.heatmap(cm, annot=True, fmt='.2f', cmap='Blues', ax=ax,\n                xticklabels=range(5), yticklabels=range(5))\n    ax.set_title(name)\n    ax.set_xlabel('Predicted')\n    ax.set_ylabel('True')\nplt.tight_layout()\nplt.show()","metadata":{"_uuid":"11e450de-02a6-4098-8c87-8a7358702c0c","_cell_guid":"a3dfe702-bf0f-4f3f-9cd1-a883b57f43c2","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:49:06.694571Z","iopub.execute_input":"2026-09-06T19:49:06.694852Z","iopub.status.idle":"2026-09-06T19:49:07.370145Z","shell.execute_reply.started":"2026-09-06T19:49:06.694829Z","shell.execute_reply":"2026-09-06T19:49:07.369302Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"recall_data = pd.DataFrame({\n    'APTOS': recall_score(y_test, aptos_pred, average=None, labels=range(5)),\n    'IDRiD': recall_score(idr_y_test, idr_pred, average=None, labels=range(5)),\n    'Messidor-2': recall_score(ms_y_test, ms_pred, average=None, labels=range(5)),\n}, index=['0-None', '1-Mild', '2-Moderate', '3-Severe', '4-PDR'])\n\nrecall_data.plot(kind='bar', figsize=(10, 5), colormap='Set2')\nplt.title('Per-Class Recall Across Datasets (RQ4)')\nplt.ylabel('Recall')\nplt.xlabel('DR Severity Class')\nplt.legend(title='Dataset')\nplt.xticks(rotation=0)\nplt.tight_layout()\nplt.show()\n\nrecall_data.round(3)","metadata":{"_uuid":"d2014c59-5f1f-4e87-a2b2-8ab7c3b877ea","_cell_guid":"3d87f548-689a-4d61-a70d-cb287344a5fa","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:49:07.371175Z","iopub.execute_input":"2026-09-06T19:49:07.371565Z","iopub.status.idle":"2026-09-06T19:49:07.604857Z","shell.execute_reply.started":"2026-09-06T19:49:07.371540Z","shell.execute_reply":"2026-09-06T19:49:07.604155Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 14. Results Summary\n\nConsolidates RQ1/RQ2/RQ4 findings into a single comparison table (point estimates), followed\nby 95% bootstrapped confidence intervals (Section 14.1) to quantify the uncertainty in each\nestimate, particularly for the low-support Severe NPDR class.","metadata":{"_uuid":"f7f164c0-eaa9-4c45-800b-3f4f4cab26ba","_cell_guid":"4c8d4381-ebd7-473a-beec-6b10edec62ba","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"summary = pd.DataFrame({\n    'Dataset': ['APTOS (in-distribution)', 'IDRiD (external)', 'Messidor-2 (external)'],\n    'Accuracy': [\n        accuracy_score(y_test, aptos_pred),\n        accuracy_score(idr_y_test, idr_pred),\n        accuracy_score(ms_y_test, ms_pred),\n    ],\n    'Macro F1': [\n        f1_score(y_test, aptos_pred, average='macro'),\n        f1_score(idr_y_test, idr_pred, average='macro'),\n        f1_score(ms_y_test, ms_pred, average='macro'),\n    ],\n    'QW Kappa': [\n        cohen_kappa_score(y_test, aptos_pred, weights='quadratic'),\n        cohen_kappa_score(idr_y_test, idr_pred, weights='quadratic'),\n        cohen_kappa_score(ms_y_test, ms_pred, weights='quadratic'),\n    ],\n    'Class 3 (Severe) Recall': [\n        recall_score(y_test, aptos_pred, average=None, labels=range(5))[3],\n        recall_score(idr_y_test, idr_pred, average=None, labels=range(5))[3],\n        recall_score(ms_y_test, ms_pred, average=None, labels=range(5))[3],\n    ],\n})\nsummary = summary.round(3)\nsummary.to_csv('/kaggle/working/results_summary.csv', index=False)\nsummary","metadata":{"_uuid":"2ae72762-52fd-43a0-b890-5a1436899a65","_cell_guid":"937634bd-c13d-4e6d-82b6-e78a780bb2e3","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:49:07.605945Z","iopub.execute_input":"2026-09-06T19:49:07.606282Z","iopub.status.idle":"2026-09-06T19:49:07.635613Z","shell.execute_reply.started":"2026-09-06T19:49:07.606242Z","shell.execute_reply":"2026-09-06T19:49:07.634926Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 14.1. Bootstrapped Confidence Intervals\n\nPoint estimates alone can be misleading for classes with few test samples (e.g. n=28 for\nSevere NPDR). 1,000 bootstrap resamples give a 95% CI for kappa, macro F1, and Class 3\nrecall on each dataset.","metadata":{"_uuid":"ad3454d0-bc98-4cd8-8ba2-cfc7f30fdb43","_cell_guid":"66ad9842-066c-41ca-b6b5-595f00cb275c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def bootstrap_metric_ci(y_true, y_pred, metric_fn, n_bootstrap=1000, ci=95, **metric_kwargs):\n    \"\"\"\n    Computes a bootstrapped confidence interval for a given sklearn-style metric\n    by resampling (y_true, y_pred) pairs with replacement.\n    \"\"\"\n    y_true = np.array(y_true)\n    y_pred = np.array(y_pred)\n    n = len(y_true)\n    scores = []\n\n    for _ in range(n_bootstrap):\n        idx = np.random.choice(n, size=n, replace=True)\n        score = metric_fn(y_true[idx], y_pred[idx], **metric_kwargs)\n        scores.append(score)\n\n    lower = np.percentile(scores, (100 - ci) / 2)\n    upper = np.percentile(scores, 100 - (100 - ci) / 2)\n    return np.mean(scores), lower, upper","metadata":{"_uuid":"971e9693-bab8-4bc9-b1e5-4a6fa400a4e1","_cell_guid":"183709f6-a119-43f0-be19-576abd964847","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:49:07.636459Z","iopub.execute_input":"2026-09-06T19:49:07.636805Z","iopub.status.idle":"2026-09-06T19:49:07.643122Z","shell.execute_reply.started":"2026-09-06T19:49:07.636783Z","shell.execute_reply":"2026-09-06T19:49:07.642058Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import cohen_kappa_score, f1_score, recall_score\n\nresults = []\ndatasets = [\n    ('APTOS (in-dist.)', y_test, aptos_pred),\n    ('IDRiD', idr_y_test, idr_pred),\n    ('Messidor-2', ms_y_test, ms_pred),\n]\n\nfor name, y_true, y_pred in datasets:\n    kappa_mean, kappa_lo, kappa_hi = bootstrap_metric_ci(\n        y_true, y_pred, cohen_kappa_score, weights='quadratic')\n    f1_mean, f1_lo, f1_hi = bootstrap_metric_ci(\n        y_true, y_pred, f1_score, average='macro')\n    recall3_mean, recall3_lo, recall3_hi = bootstrap_metric_ci(\n        y_true, y_pred, lambda yt, yp: recall_score(yt, yp, average=None, labels=range(5))[3])\n\n    results.append({\n        'Dataset': name,\n        'QW Kappa': f\"{kappa_mean:.3f} [{kappa_lo:.3f}, {kappa_hi:.3f}]\",\n        'Macro F1': f\"{f1_mean:.3f} [{f1_lo:.3f}, {f1_hi:.3f}]\",\n        'Class 3 Recall': f\"{recall3_mean:.3f} [{recall3_lo:.3f}, {recall3_hi:.3f}]\",\n    })\n\nci_summary = pd.DataFrame(results)\nci_summary.to_csv('/kaggle/working/results_summary_with_ci.csv', index=False)\nci_summary","metadata":{"_uuid":"277d87c8-718e-4ecf-96a3-d91250860950","_cell_guid":"fb25d2b8-14d9-4674-b8fd-b52f91d3a6d5","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T19:49:07.644170Z","iopub.execute_input":"2026-09-06T19:49:07.644510Z","iopub.status.idle":"2026-09-06T19:49:18.029295Z","shell.execute_reply.started":"2026-09-06T19:49:07.644461Z","shell.execute_reply":"2026-09-06T19:49:18.028624Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 14.2. Ablation: Class Weighting vs. No Class Weighting\n\nTo strengthen the modelling contribution beyond a single baseline, this ablation\nretrains the identical EfficientNetB0 architecture (last 10 layers unfrozen) without\nclass-weighted loss, isolating the effect of class weighting on minority-class\n(Severe DR) performance against the class-weighted baseline reported in Section 9.","metadata":{}},{"cell_type":"code","source":"# Model B: WITHOUT class weighting (ablation)\ntf.random.set_seed(42)\nnp.random.seed(42)\n\nenb0_ablation_base = EfficientNetB0(input_shape=(IMG_SIZE, IMG_SIZE, 3), include_top=False, weights='imagenet')\nenb0_ablation_base.trainable = True\nfor layer in enb0_ablation_base.layers[:-10]:\n    layer.trainable = False\n\nx = GlobalAveragePooling2D()(enb0_ablation_base.output)\nx = Dropout(0.3)(x)\nx = Dense(128, activation='relu')(x)\nx = Dropout(0.2)(x)\noutput = Dense(NUM_CLASSES, activation='softmax')(x)\nenb0_ablation_model = Model(inputs=enb0_ablation_base.input, outputs=output)\n\nenb0_ablation_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n                             loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n\nablation_callbacks = [\n    EarlyStopping(monitor='val_accuracy', patience=6, verbose=1, restore_best_weights=True),\n    ModelCheckpoint('/kaggle/working/best_enb0_no_classweight.keras', monitor='val_accuracy', save_best_only=True),\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3)\n]\n\nablation_history = enb0_ablation_model.fit(\n    X_train * 255., y_train,\n    validation_data=(X_val * 255., y_val),\n    batch_size=64,\n    epochs=40,\n    callbacks=ablation_callbacks   # NO class_weight argument — this is the ablation\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:49:18.030163Z","iopub.execute_input":"2026-09-06T19:49:18.030510Z","iopub.status.idle":"2026-09-06T19:51:08.886166Z","shell.execute_reply.started":"2026-09-06T19:49:18.030485Z","shell.execute_reply":"2026-09-06T19:51:08.885541Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Evaluate Model B (no class weighting) — same three datasets as the baseline\nablation_model_saved = load_model('/kaggle/working/best_enb0_no_classweight.keras')\n\nablation_aptos_pred = np.argmax(ablation_model_saved.predict(X_test * 255.), axis=1)\nprint(\"=== No class weighting — APTOS ===\")\nprint(classification_report(y_test, ablation_aptos_pred, digits=3))\nprint(\"Kappa:\", cohen_kappa_score(y_test, ablation_aptos_pred, weights='quadratic'))\n\nablation_idr_pred = np.argmax(ablation_model_saved.predict(idr_X_test * 255.), axis=1)\nprint(\"\\n=== No class weighting — IDRiD ===\")\nprint(classification_report(idr_y_test, ablation_idr_pred, digits=3))\nprint(\"Kappa:\", cohen_kappa_score(idr_y_test, ablation_idr_pred, weights='quadratic'))\n\nablation_ms_pred = np.argmax(ablation_model_saved.predict(ms_X_test * 255.), axis=1)\nprint(\"\\n=== No class weighting — Messidor-2 ===\")\nprint(classification_report(ms_y_test, ablation_ms_pred, digits=3))\nprint(\"Kappa:\", cohen_kappa_score(ms_y_test, ablation_ms_pred, weights='quadratic'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:51:08.887841Z","iopub.execute_input":"2026-09-06T19:51:08.888060Z","iopub.status.idle":"2026-09-06T19:51:26.833094Z","shell.execute_reply.started":"2026-09-06T19:51:08.888040Z","shell.execute_reply":"2026-09-06T19:51:26.832473Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 14.3 Ablation: Frozen vs. Fine-Tuned Layers\n\nThis ablation retrains the identical EfficientNetB0 architecture (with class-weighted\nloss, matching the baseline) but with the entire backbone frozen — only the\nclassification head trains — compared against the baseline's partial fine-tuning\n(last 10 layers unfrozen). This tests whether allowing some backbone layers to adapt\nto retinal-specific features meaningfully improves in-distribution and cross-dataset\nperformance.","metadata":{}},{"cell_type":"code","source":"# Ablation: Fully frozen backbone (only the classification head trains)\ntf.random.set_seed(42)\nnp.random.seed(42)\n\nenb0_frozen_base = EfficientNetB0(input_shape=(IMG_SIZE, IMG_SIZE, 3), include_top=False, weights='imagenet')\nenb0_frozen_base.trainable = False   # entire backbone frozen — no layers unfrozen\n\nx = GlobalAveragePooling2D()(enb0_frozen_base.output)\nx = Dropout(0.3)(x)\nx = Dense(128, activation='relu')(x)\nx = Dropout(0.2)(x)\noutput = Dense(NUM_CLASSES, activation='softmax')(x)\nenb0_frozen_model = Model(inputs=enb0_frozen_base.input, outputs=output)\n\nenb0_frozen_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n                           loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n\nfrozen_callbacks = [\n    EarlyStopping(monitor='val_accuracy', patience=6, verbose=1, restore_best_weights=True),\n    ModelCheckpoint('/kaggle/working/best_enb0_frozen.keras', monitor='val_accuracy', save_best_only=True),\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3)\n]\n\nfrozen_history = enb0_frozen_model.fit(\n    X_train * 255., y_train,\n    validation_data=(X_val * 255., y_val),\n    batch_size=64,\n    epochs=40,\n    class_weight=class_weight_dict,   # kept identical to baseline\n    callbacks=frozen_callbacks\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:51:26.837926Z","iopub.execute_input":"2026-09-06T19:51:26.838183Z","iopub.status.idle":"2026-09-06T19:54:38.469969Z","shell.execute_reply.started":"2026-09-06T19:51:26.838161Z","shell.execute_reply":"2026-09-06T19:54:38.469333Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"frozen_model_saved = load_model('/kaggle/working/best_enb0_frozen.keras')\n\nfrozen_aptos_pred = np.argmax(frozen_model_saved.predict(X_test * 255.), axis=1)\nprint(\"=== Fully frozen backbone — APTOS ===\")\nprint(classification_report(y_test, frozen_aptos_pred, digits=3))\nprint(\"Kappa:\", cohen_kappa_score(y_test, frozen_aptos_pred, weights='quadratic'))\n\nfrozen_idr_pred = np.argmax(frozen_model_saved.predict(idr_X_test * 255.), axis=1)\nprint(\"\\n=== Fully frozen backbone — IDRiD ===\")\nprint(classification_report(idr_y_test, frozen_idr_pred, digits=3))\nprint(\"Kappa:\", cohen_kappa_score(idr_y_test, frozen_idr_pred, weights='quadratic'))\n\nfrozen_ms_pred = np.argmax(frozen_model_saved.predict(ms_X_test * 255.), axis=1)\nprint(\"\\n=== Fully frozen backbone — Messidor-2 ===\")\nprint(classification_report(ms_y_test, frozen_ms_pred, digits=3))\nprint(\"Kappa:\", cohen_kappa_score(ms_y_test, frozen_ms_pred, weights='quadratic'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:54:38.471850Z","iopub.execute_input":"2026-09-06T19:54:38.472156Z","iopub.status.idle":"2026-09-06T19:54:56.540996Z","shell.execute_reply.started":"2026-09-06T19:54:38.472131Z","shell.execute_reply":"2026-09-06T19:54:56.540386Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 14.4 Multi-Seed Stability Analysis\n\nTo address instability observed in Severe DR (Class 3) recall across training runs,\nthe baseline model (EfficientNetB0, 10 unfrozen layers, class-weighted loss) is\nretrained here across 3 fixed seeds (42, 43, 44), with deterministic GPU execution\nenabled, to report a genuinely reproducible mean/std/range rather than a single run.","metadata":{}},{"cell_type":"code","source":"seeds = [42, 43, 44]  # adjust count based on your time budget\n\nmultiseed_records = []\n\nfor seed in seeds:\n    tf.random.set_seed(seed)\n    np.random.seed(seed)\n\n    base = EfficientNetB0(input_shape=(IMG_SIZE, IMG_SIZE, 3), include_top=False, weights='imagenet')\n    base.trainable = True\n    for layer in base.layers[:-10]:\n        layer.trainable = False\n\n    x = GlobalAveragePooling2D()(base.output)\n    x = Dropout(0.3)(x)\n    x = Dense(128, activation='relu')(x)\n    x = Dropout(0.2)(x)\n    output = Dense(NUM_CLASSES, activation='softmax')(x)\n    model = Model(inputs=base.input, outputs=output)\n    model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n                  loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n\n    callbacks = [\n        EarlyStopping(monitor='val_accuracy', patience=6, restore_best_weights=True),\n        ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3)\n    ]\n\n    model.fit(X_train * 255., y_train, validation_data=(X_val * 255., y_val),\n              batch_size=64, epochs=40, class_weight=class_weight_dict,\n              callbacks=callbacks, verbose=0)\n\n    aptos_pred = np.argmax(model.predict(X_test * 255., verbose=0), axis=1)\n    idr_pred = np.argmax(model.predict(idr_X_test * 255., verbose=0), axis=1)\n    ms_pred = np.argmax(model.predict(ms_X_test * 255., verbose=0), axis=1)\n\n    multiseed_records.append({\n        'Seed': seed,\n        'APTOS_Accuracy': accuracy_score(y_test, aptos_pred),\n        'APTOS_MacroF1': f1_score(y_test, aptos_pred, average='macro'),\n        'APTOS_Kappa': cohen_kappa_score(y_test, aptos_pred, weights='quadratic'),\n        'APTOS_Class3_Recall': recall_score(y_test, aptos_pred, average=None, labels=range(5))[3],\n        'IDRiD_Class3_Recall': recall_score(idr_y_test, idr_pred, average=None, labels=range(5))[3],\n        'Messidor2_Class3_Recall': recall_score(ms_y_test, ms_pred, average=None, labels=range(5))[3],\n    })\n    print(f\"Seed {seed} done.\")\n\nmultiseed_results = pd.DataFrame(multiseed_records)\nprint(multiseed_results.to_string(index=False))\nprint(multiseed_results.drop(columns='Seed').agg(['mean', 'std', 'min', 'max']).round(3))\nmultiseed_results.to_csv('/kaggle/working/multiseed_results.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T19:54:56.542187Z","iopub.execute_input":"2026-09-06T19:54:56.542519Z","iopub.status.idle":"2026-09-06T20:04:10.000151Z","shell.execute_reply.started":"2026-09-06T19:54:56.542495Z","shell.execute_reply":"2026-09-06T20:04:09.999409Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 15. Mitigation Technique — Test-Time Augmentation (RQ3)\n\n*RQ3: Can a lightweight mitigation technique improve cross-dataset performance without\nretraining the model?*\n\nTest-time augmentation (TTA) was chosen over alternatives (e.g. feature normalisation) for its\nlow implementation risk and because it fits naturally with the inference-only framing of the\ncross-dataset evaluation above — no model weights are changed, only how predictions are\naggregated at inference time.","metadata":{"_uuid":"c1dafce5-bdc4-424e-bb12-2d1754ef01ca","_cell_guid":"994c38d0-80d7-44dd-8dab-7ea33aa97c17","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def predict_with_tta(model, images, n_augments=5):\n    \"\"\"\n    Averages predictions over n_augments random flip/rotation variants of each image,\n    plus the original (unaugmented) image. images must already be in [0, 255] scale.\n    \"\"\"\n    all_preds = []\n\n    # Original, unaugmented prediction\n    all_preds.append(model.predict(images, verbose=0))\n\n    for _ in range(n_augments):\n        augmented = images.copy()\n        # Random horizontal flip (applied to the whole batch for speed; fine since TTA\n        # just needs varied views per image, not a different flip choice per image)\n        if np.random.rand() > 0.5:\n            augmented = augmented[:, :, ::-1, :]\n        # Random vertical flip\n        if np.random.rand() > 0.5:\n            augmented = augmented[:, ::-1, :, :]\n        # Random 90-degree rotation\n        k = np.random.randint(0, 4)\n        augmented = np.rot90(augmented, k=k, axes=(1, 2))\n\n        all_preds.append(model.predict(augmented, verbose=0))\n\n    avg_preds = np.mean(all_preds, axis=0)\n    return avg_preds","metadata":{"_uuid":"32e03520-5458-49a1-974d-c7085d0dc4b4","_cell_guid":"be67f4f5-1543-4857-87a8-3c269f298d3f","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T20:04:10.001135Z","iopub.execute_input":"2026-09-06T20:04:10.001482Z","iopub.status.idle":"2026-09-06T20:04:10.008561Z","shell.execute_reply.started":"2026-09-06T20:04:10.001458Z","shell.execute_reply":"2026-09-06T20:04:10.007733Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"idr_pred_tta_probs = predict_with_tta(enb0_model_saved, idr_X_test * 255., n_augments=5)\nidr_pred_tta = np.argmax(idr_pred_tta_probs, axis=1)\n\nms_pred_tta_probs = predict_with_tta(enb0_model_saved, ms_X_test * 255., n_augments=5)\nms_pred_tta = np.argmax(ms_pred_tta_probs, axis=1)\n\nprint(\"=== IDRiD — Before vs After TTA ===\")\nprint(\"Without TTA — Macro F1:\", np.round(f1_score(idr_y_test, idr_pred, average='macro'),3),\n      \"| QW Kappa:\", np.round(cohen_kappa_score(idr_y_test, idr_pred, weights='quadratic'),3))\nprint(\"With TTA    — Macro F1:\", np.round(f1_score(idr_y_test, idr_pred_tta, average='macro'),3),\n      \"| QW Kappa:\", np.round(cohen_kappa_score(idr_y_test, idr_pred_tta, weights='quadratic'),3))\n\nprint(\"\\n=== Messidor-2 — Before vs After TTA ===\")\nprint(\"Without TTA — Macro F1:\", np.round(f1_score(ms_y_test, ms_pred, average='macro'),3),\n      \"| QW Kappa:\", np.round(cohen_kappa_score(ms_y_test, ms_pred, weights='quadratic'),3))\nprint(\"With TTA    — Macro F1:\", np.round(f1_score(ms_y_test, ms_pred_tta, average='macro'),3),\n      \"| QW Kappa:\", np.round(cohen_kappa_score(ms_y_test, ms_pred_tta, weights='quadratic'),3))","metadata":{"_uuid":"db042a99-6d12-427d-b2c7-95ee21fb3b9b","_cell_guid":"3378c97c-420b-49dc-a6e9-0af66fc466ba","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T20:04:10.009598Z","iopub.execute_input":"2026-09-06T20:04:10.009897Z","iopub.status.idle":"2026-09-06T20:05:14.272780Z","shell.execute_reply.started":"2026-09-06T20:04:10.009867Z","shell.execute_reply":"2026-09-06T20:05:14.271913Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 15.1. Repeated TTA Evaluation for Uncertainty","metadata":{}},{"cell_type":"code","source":"def repeated_tta_evaluation(model, images, y_true, n_repeats=5, n_augments=5):\n    results = []\n    for i in range(n_repeats):\n        preds = predict_with_tta(model, images, n_augments=n_augments)\n        y_pred = np.argmax(preds, axis=1)\n        results.append({\n            'repeat': i + 1,\n            'macro_f1': f1_score(y_true, y_pred, average='macro'),\n            'kappa': cohen_kappa_score(y_true, y_pred, weights='quadratic'),\n        })\n    return pd.DataFrame(results)\n\nidr_tta_repeated = repeated_tta_evaluation(enb0_model_saved, idr_X_test * 255., idr_y_test)\nms_tta_repeated = repeated_tta_evaluation(enb0_model_saved, ms_X_test * 255., ms_y_test)\n\nprint(\"IDRiD — 5 repeated TTA runs:\")\nprint(idr_tta_repeated)\nprint(idr_tta_repeated[['macro_f1', 'kappa']].agg(['mean', 'std']))\n\nprint(\"\\nMessidor-2 — 5 repeated TTA runs:\")\nprint(ms_tta_repeated)\nprint(ms_tta_repeated[['macro_f1', 'kappa']].agg(['mean', 'std']))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-06T20:05:14.274544Z","iopub.execute_input":"2026-09-06T20:05:14.274865Z","iopub.status.idle":"2026-09-06T20:10:26.426484Z","shell.execute_reply.started":"2026-09-06T20:05:14.274841Z","shell.execute_reply":"2026-09-06T20:10:26.425741Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 16. Appendix — Models Not Used\n\nTwo alternative architectures were explored before EfficientNetB0 was confirmed as the sole\nbaseline, per the dissertation's approved single-model scope. Included here for transparency\nonly — not part of the main results.\n\n- **Simple CNN (from scratch)** — trained faster but underperformed the pretrained baseline,\n  as expected without transfer learning.\n- **EfficientNetB7** — substantially higher compute cost for unclear accuracy benefit; dropped\n  to protect the 8-week timeline.","metadata":{"_uuid":"edb41379-43b0-4b62-9a02-0c5e9f5ab023","_cell_guid":"4a8a262e-a692-4014-9df6-c39a7804aba3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# Simple CNN — from scratch, no transfer learning\nsimple_cnn = Sequential([\n    Conv2D(64, (3,3), activation='relu', input_shape=(IMG_SIZE, IMG_SIZE, 3)),\n    BatchNormalization(),\n    MaxPooling2D(2,2),\n    Dropout(0.2),\n\n    Conv2D(128, (3,3), activation='relu'),\n    BatchNormalization(),\n    MaxPooling2D(2,2),\n    Dropout(0.2),\n\n    Conv2D(256, (3,3), activation='relu'),\n    BatchNormalization(),\n    MaxPooling2D(2,2),\n\n    GlobalAveragePooling2D(),\n    Dense(128, activation='relu'),\n    Dropout(0.3),\n    Dense(NUM_CLASSES, activation='softmax')\n])\n\nsimple_cnn.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n                    loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n\nsimple_cnn_callbacks = [\n    EarlyStopping(monitor='val_accuracy', patience=6, verbose=1, restore_best_weights=True),\n    ModelCheckpoint('/kaggle/working/best_simple_cnn.keras', monitor='val_accuracy', save_best_only=True),\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3)\n]\n\n# # Not run by default — uncomment to reproduce\n# simple_cnn_history = simple_cnn.fit(\n#     X_train, y_train, validation_data=(X_val, y_val),\n#     batch_size=64, epochs=40, class_weight=class_weight_dict, callbacks=simple_cnn_callbacks\n# )","metadata":{"_uuid":"2800702e-ee05-4b1d-ae72-c11f60a077ed","_cell_guid":"e8cc7f8e-0d82-4a55-b60e-55b0b86c511a","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T20:10:26.427572Z","iopub.execute_input":"2026-09-06T20:10:26.427898Z","iopub.status.idle":"2026-09-06T20:10:26.510259Z","shell.execute_reply.started":"2026-09-06T20:10:26.427874Z","shell.execute_reply":"2026-09-06T20:10:26.509725Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# EfficientNetB7 — larger backbone, explored and dropped for scope/timeline reasons\n\nenb7_base_model = EfficientNetB7(input_shape=(IMG_SIZE, IMG_SIZE, 3), include_top=False, weights='imagenet')\nenb7_base_model.trainable = True\nfor layer in enb7_base_model.layers[:-10]:\n    layer.trainable = False\n\nx = GlobalAveragePooling2D()(enb7_base_model.output)\nx = Dropout(0.3)(x)\nx = Dense(128, activation='relu')(x)\nx = Dropout(0.2)(x)\noutput = Dense(NUM_CLASSES, activation='softmax')(x)\nenb7_model = Model(inputs=enb7_base_model.input, outputs=output)\n\nenb7_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-4),\n                    loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n\nenb7_callbacks = [\n    EarlyStopping(monitor='val_accuracy', patience=6, verbose=1, restore_best_weights=True),\n    ModelCheckpoint('/kaggle/working/best_enb7_cnn.keras', monitor='val_accuracy', save_best_only=True),\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3)\n]\n\n# # Not run by default — uncomment to reproduce (slow: ~90s+ for the first epoch alone)\n# enb7_history = enb7_model.fit(\n#     X_train*255., y_train, validation_data=(X_val*255., y_val),\n#     batch_size=64, epochs=50, class_weight=class_weight_dict, callbacks=enb7_callbacks\n# )","metadata":{"_uuid":"6efabb0f-bdeb-458b-bbcf-fdd4de322079","_cell_guid":"e4d90a26-e785-487d-9494-6c1ab5e5ce87","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2026-09-06T20:10:26.511097Z","iopub.execute_input":"2026-09-06T20:10:26.511381Z","iopub.status.idle":"2026-09-06T20:10:34.878124Z","shell.execute_reply.started":"2026-09-06T20:10:26.511359Z","shell.execute_reply":"2026-09-06T20:10:34.877567Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"_uuid":"d370959c-1540-45b1-8ca2-d85d20af8717","_cell_guid":"9aa3381b-e5bb-473f-978a-4f8be6f56669","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null}]}