{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","isGpuEnabled":true,"isInternetEnabled":true,"language":"python","sourceType":"notebook"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"c03ae394","cell_type":"markdown","source":"# Diabetic Retinopathy Stage Detection using Transfer Learning\n**Computer Vision Coursework, BSc (Hons) Computer Science**\n**Author:** L.M. Navaratna\n**Dataset:** APTOS 2019 Blindness Detection (Kaggle)\n\n---\n\nThis notebook grades diabetic eye disease from photos of the back of the eye (fundus photos). Every design choice is **tested, not assumed**. Each stage compares a few options on the validation set, keeps the winner, and passes it to the next stage. Only when every choice is fixed are the final models trained and tested once on the untouched test set.\n\n1. Setup and the settings that are fixed up front\n2. Looking at the data\n3. Preprocessing methods (resize only, CLAHE, Ben Graham) and their tests\n4. Image size check\n5. Data splits and class weights\n6. Augmentation policies and their tests\n7. Feeding photos to the model\n8. Models, freezing and the training engine\n9. Experiments: preprocessing, augmentation, class balance, freezing, learning rate, regularisation\n10. Final configuration\n11. Final training of three models\n12. Testing the models\n13. Grad-CAM\n14. Error analysis\n15. Web app export\n\n**How to run on Kaggle:** add the APTOS 2019 competition data, turn on the **GPU T4**, then **Save Version** and **Save & Run All (Commit)** so everything runs once, top to bottom, in one clean session. Expect roughly **4 to 5 hours**. Every figure, table and model is saved under `/kaggle/working/`, in a folder named after its section.","metadata":{}},{"id":"abb9a079","cell_type":"markdown","source":"## 1. Setup","metadata":{}},{"id":"7dd2ec0f","cell_type":"code","source":"# Cell 1: load every library the notebook uses, all in one place\nimport os\nimport random\nimport time\nimport glob\nimport copy\nimport platform\nimport json as json_module\nfrom datetime import datetime\nfrom functools import partial\n\nimport numpy as np\nimport pandas as pd\nimport cv2\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nimport seaborn as seaborn_plotting_library\n\nimport torch\nimport torch.nn as neural_network_layers\nimport torch.optim as optimizers\nfrom torch.utils.data import Dataset, DataLoader\nimport torchvision\nfrom torchvision import transforms\nfrom torchvision.models import (\n    efficientnet_b3, EfficientNet_B3_Weights,\n    efficientnet_b0, EfficientNet_B0_Weights,\n    resnet50, ResNet50_Weights,\n)\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils.class_weight import compute_class_weight\nfrom sklearn.metrics import (\n    accuracy_score,\n    precision_recall_fscore_support,\n    cohen_kappa_score,\n    confusion_matrix,\n    classification_report,\n    roc_curve,\n    roc_auc_score,\n)\nfrom scipy.stats import chi2\n\nprint(\"Cell 1 finished: all libraries loaded.\")\nprint(\"Python version:\", platform.python_version())\nprint(\"Torch version:\", torch.__version__)\nprint(\"Torchvision version:\", torchvision.__version__)\nprint(\"OpenCV version:\", cv2.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:37.955147Z","iopub.execute_input":"2026-09-26T11:32:37.955541Z","iopub.status.idle":"2026-09-26T11:32:56.05743Z","shell.execute_reply.started":"2026-09-26T11:32:37.955509Z","shell.execute_reply":"2026-09-26T11:32:56.056695Z"}},"outputs":[],"execution_count":null},{"id":"9b3890a1","cell_type":"code","source":"# Cell 2: check which hardware this session is running on\nis_gpu_available = torch.cuda.is_available()\ncompute_device = torch.device(\"cuda\" if is_gpu_available else \"cpu\")\n\nprint(\"Cell 2 finished: hardware check complete.\")\nprint(\"GPU available:\", is_gpu_available)\nif is_gpu_available:\n    print(\"GPU name:\", torch.cuda.get_device_name(0))\n    print(\"GPU memory (GB):\", round(torch.cuda.get_device_properties(0).total_memory / 1e9, 2))\nelse:\n    print(\"No GPU found. On Kaggle, turn on the GPU accelerator in the settings panel before running.\")\nprint(\"Selected compute device:\", compute_device)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:56.058645Z","iopub.execute_input":"2026-09-26T11:32:56.058971Z","iopub.status.idle":"2026-09-26T11:32:56.391948Z","shell.execute_reply.started":"2026-09-26T11:32:56.05895Z","shell.execute_reply":"2026-09-26T11:32:56.391376Z"}},"outputs":[],"execution_count":null},{"id":"ff91bbbf","cell_type":"code","source":"# Cell 3: fix the random seed so every run splits, shuffles and starts the models the same way\nRANDOM_SEED = 42\n\n\ndef set_global_random_seed(seed_value):\n    random.seed(seed_value)\n    np.random.seed(seed_value)\n    torch.manual_seed(seed_value)\n    torch.cuda.manual_seed_all(seed_value)\n\n\nset_global_random_seed(RANDOM_SEED)\ntorch.backends.cudnn.deterministic = True\ntorch.backends.cudnn.benchmark = False\n\nprint(\"Cell 3 finished: random seed fixed to\", RANDOM_SEED, \"and the GPU told to use repeatable maths where it can.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:56.39276Z","iopub.execute_input":"2026-09-26T11:32:56.393087Z","iopub.status.idle":"2026-09-26T11:32:56.411604Z","shell.execute_reply.started":"2026-09-26T11:32:56.393025Z","shell.execute_reply":"2026-09-26T11:32:56.410851Z"}},"outputs":[],"execution_count":null},{"id":"b82f6e77","cell_type":"markdown","source":"### Settings fixed up front, and why\nMost choices in this project are **tested** later (section 9). A few are fixed at the start, either because something else forces them or because testing them would cost more GPU time than it is worth. Each one has a reason:\n\n| Setting | Value | Why it is fixed rather than tested |\n|---|---|---|\n| Data split | 70% train, 15% validation, 15% test, same grade mix in each | Leaves about 29 Severe photos in each of validation and test, the smallest amount that still gives a usable score for that grade |\n| Random seeds | 42, plus 7 for the repeat run of each experiment | Repeatable results. Two seeds show whether a difference is real or just luck |\n| Image size | 384 x 384 | Checked by the image size study in section 4 |\n| Batch size | 16 | The largest batch where all three models fit in the T4's memory at 384, measured in section 4 |\n| Normalisation | ImageNet mean and standard deviation | The pretrained weights expect it |\n| Black border threshold | 7 | Pixels darker than this are background. Checked by the preprocessing tests |\n| CLAHE settings | clip 2.0, 8 x 8 tiles | Common default for fundus photos. Tuning them is listed as future work |\n| Ben Graham blur | eye radius / 30 | Graham's original setting |\n| Experiment model | EfficientNet-B0, short schedule | Fastest model, so about 26 experiments fit in the GPU budget |\n| Short schedule (experiments) | 2 head epochs, up to 10 more, stop after 3 without gain | Enough to rank options fairly |\n| Full schedule (final models) | 5 head epochs, up to 25 more, stop after 6 without gain | Gives the final models time to settle |\n| Main metric | Validation QWK | The grades are ordered, and QWK punishes big grade mistakes more than small ones |\n\n**Everything else is tested:** the preprocessing method, whether and how to augment, class weights, how much of the model to freeze, the learning rate, and extra regularisation against overfitting.","metadata":{}},{"id":"bc8938af","cell_type":"code","source":"# Cell 4: store the fixed settings, the starting recipe that experiments will improve, and create the output folders\nproject_configuration = {\n    \"competition_data_directory\": \"/kaggle/input/competitions/aptos2019-blindness-detection\",\n    \"train_csv_filename\": \"train.csv\",\n    \"train_images_subfolder\": \"train_images\",\n    \"image_file_extension\": \".png\",\n    \"target_image_size\": 384,\n    \"number_of_dr_classes\": 5,\n    \"batch_size\": 16,\n    \"validation_split_fraction\": 0.15,\n    \"test_split_fraction\": 0.15,\n    \"dark_pixel_brightness_threshold\": 7,\n    \"clahe_clip_limit\": 2.0,\n    \"clahe_tile_grid_size\": 8,\n    \"ben_graham_blur_divisor\": 30,\n    \"minimum_images_required_in_rarest_class\": 50,\n    \"number_of_data_loader_workers\": 4,\n    \"screening_architecture\": \"efficientnet_b0\",\n    \"screening_phase_one_epochs\": 2,\n    \"screening_phase_two_max_epochs\": 10,\n    \"screening_early_stopping_patience\": 3,\n    \"final_phase_one_epochs\": 5,\n    \"final_phase_two_max_epochs\": 25,\n    \"final_early_stopping_patience\": 6,\n    \"experiment_seeds\": [42, 7],\n    \"tie_margin_qwk\": 0.005,\n    \"output_root_directory\": \"/kaggle/working\",\n    \"cache_root_directory\": \"/tmp/dr_preprocessing_cache\",\n    \"figure_dpi\": 200,\n}\n\nstarting_training_recipe = {\n    \"preprocessing_method\": \"clahe\",\n    \"augmentation_policy\": \"full\",\n    \"use_class_weights\": True,\n    \"transfer_strategy\": \"two_phase_full_unfreeze\",\n    \"phase_one_learning_rate\": 1e-3,\n    \"phase_two_learning_rate\": 1e-4,\n    \"label_smoothing\": 0.0,\n    \"weight_decay\": 0.0,\n    \"head_dropout\": None,\n    \"phase_two_schedule\": \"reduce_on_plateau\",\n}\n\nfor candidate_data_directory in [\"/kaggle/input/competitions/aptos2019-blindness-detection\", \"/kaggle/input/aptos2019-blindness-detection\"]:\n    if os.path.exists(candidate_data_directory):\n        project_configuration[\"competition_data_directory\"] = candidate_data_directory\n        break\n\noutput_root_directory = project_configuration[\"output_root_directory\"]\noutput_section_names = [\n    \"01_data_exploration\",\n    \"02_preprocessing_methods\",\n    \"03_image_size_study\",\n    \"04_splits_and_class_weights\",\n    \"05_augmentation_policies\",\n    \"06_models_and_freezing\",\n    \"07_experiment_preprocessing\",\n    \"08_experiment_augmentation\",\n    \"09_experiment_class_balance\",\n    \"10_experiment_transfer_learning\",\n    \"11_experiment_learning_rate\",\n    \"12_experiment_regularisation\",\n    \"13_final_configuration\",\n    \"14_final_training\",\n    \"15_evaluation\",\n    \"16_gradcam\",\n    \"17_error_analysis\",\n    \"18_app_export\",\n]\nfigure_folder_paths = {section_name: os.path.join(output_root_directory, \"figures\", section_name) for section_name in output_section_names}\nresults_folder_paths = {section_name: os.path.join(output_root_directory, \"results\", section_name) for section_name in output_section_names}\nmain_models_folder_path = os.path.join(output_root_directory, \"models\")\napp_export_folder_path = os.path.join(output_root_directory, \"app_export\")\ndemo_images_folder_path = os.path.join(output_root_directory, \"demo_images\")\npreprocessing_method_names = [\"resize_only\", \"clahe\", \"ben_graham\"]\ncache_folder_paths = {method_name: os.path.join(project_configuration[\"cache_root_directory\"], method_name) for method_name in preprocessing_method_names}\nexperiment_checkpoint_path = os.path.join(project_configuration[\"cache_root_directory\"], \"experiment_checkpoint.pth\")\n\nfor folder_path in [*figure_folder_paths.values(), *results_folder_paths.values(), main_models_folder_path,\n                    app_export_folder_path, demo_images_folder_path, *cache_folder_paths.values()]:\n    os.makedirs(folder_path, exist_ok=True)\n\n\ndef save_current_figure(section_name, file_name):\n    plt.tight_layout()\n    saved_figure_path = os.path.join(figure_folder_paths[section_name], file_name)\n    plt.savefig(saved_figure_path, dpi=project_configuration[\"figure_dpi\"], bbox_inches=\"tight\")\n    plt.show()\n    print(f\"  figure saved: figures/{section_name}/{file_name}\")\n    return saved_figure_path\n\n\ndef save_results_table(results_dataframe, section_name, file_name, keep_index=False):\n    saved_table_path = os.path.join(results_folder_paths[section_name], file_name)\n    results_dataframe.to_csv(saved_table_path, index=keep_index)\n    print(f\"  table saved: results/{section_name}/{file_name}\")\n    return saved_table_path\n\n\ndef convert_numpy_values_for_json(value):\n    return value.item() if hasattr(value, \"item\") else str(value)\n\n\ndef save_results_json(results_dictionary, section_name, file_name):\n    saved_json_path = os.path.join(results_folder_paths[section_name], file_name)\n    with open(saved_json_path, \"w\") as json_file:\n        json_module.dump(results_dictionary, json_file, indent=2, default=convert_numpy_values_for_json)\n    print(f\"  summary saved: results/{section_name}/{file_name}\")\n    return saved_json_path\n\n\nprint(\"Cell 4 finished: settings stored and output folders created.\")\nprint(\"Data folder in use:\", project_configuration[\"competition_data_directory\"])\nprint(\"Starting recipe (the experiments in section 9 will replace each part with the tested winner):\")\nprint(json_module.dumps(starting_training_recipe, indent=2))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:56.412564Z","iopub.execute_input":"2026-09-26T11:32:56.412887Z","iopub.status.idle":"2026-09-26T11:32:56.42798Z","shell.execute_reply.started":"2026-09-26T11:32:56.412865Z","shell.execute_reply":"2026-09-26T11:32:56.427279Z"}},"outputs":[],"execution_count":null},{"id":"b28b783e","cell_type":"markdown","source":"## 2. Looking at the Data\nQuick pass or fail checks (smoke tests) come first, so nothing is built on broken data. Then the data is explored: how many photos each grade has, what they look like, and how much their sizes vary.","metadata":{}},{"id":"f221eeb8","cell_type":"code","source":"# Cell 5: load the labels file and link each label to its photo\ntrain_labels_csv_path = os.path.join(project_configuration[\"competition_data_directory\"], project_configuration[\"train_csv_filename\"])\ntrain_images_directory_path = os.path.join(project_configuration[\"competition_data_directory\"], project_configuration[\"train_images_subfolder\"])\n\ndiabetic_retinopathy_labels_dataframe = pd.read_csv(train_labels_csv_path)\ndiagnosis_stage_names = {0: \"No DR\", 1: \"Mild\", 2: \"Moderate\", 3: \"Severe\", 4: \"Proliferative DR\"}\nlist_of_class_indices = list(range(project_configuration[\"number_of_dr_classes\"]))\nstage_names_in_order = [diagnosis_stage_names[class_index] for class_index in list_of_class_indices]\n\n\ndef build_image_path_from_id(image_id_code):\n    return os.path.join(train_images_directory_path, image_id_code + project_configuration[\"image_file_extension\"])\n\n\nprint(\"Cell 5 finished: labels file loaded.\")\nprint(\"Number of labelled photos:\", len(diabetic_retinopathy_labels_dataframe))\nprint(\"Duplicated photo IDs:\", int(diabetic_retinopathy_labels_dataframe[\"id_code\"].duplicated().sum()))\nprint(diabetic_retinopathy_labels_dataframe.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:56.429845Z","iopub.execute_input":"2026-09-26T11:32:56.430183Z","iopub.status.idle":"2026-09-26T11:32:56.495884Z","shell.execute_reply.started":"2026-09-26T11:32:56.43016Z","shell.execute_reply":"2026-09-26T11:32:56.495285Z"}},"outputs":[],"execution_count":null},{"id":"9923792b","cell_type":"code","source":"# Cell 6: SMOKE TESTS, make sure nothing is missing or broken in the data before building on it\nexpected_file_names = set(diabetic_retinopathy_labels_dataframe[\"id_code\"] + project_configuration[\"image_file_extension\"])\nactual_file_names_on_disk = set(os.path.basename(file_path) for file_path in glob.glob(os.path.join(train_images_directory_path, \"*\")))\nlabelled_but_missing_from_disk = expected_file_names - actual_file_names_on_disk\npresent_on_disk_but_not_labelled = actual_file_names_on_disk - expected_file_names\n\ndiagnosis_column = diabetic_retinopathy_labels_dataframe[\"diagnosis\"]\nlabels_are_valid_grades = bool(\n    diagnosis_column.notnull().all() and pd.api.types.is_integer_dtype(diagnosis_column) and diagnosis_column.between(0, 4).all()\n)\nimages_in_rarest_grade = int(diagnosis_column.value_counts().min())\nminimum_images_needed = project_configuration[\"minimum_images_required_in_rarest_class\"]\n\nsmoke_test_results_dataframe = pd.DataFrame([\n    {\"smoke_test\": \"1. Every label has a photo and every photo has a label\",\n     \"passed\": len(labelled_but_missing_from_disk) == 0 and len(present_on_disk_but_not_labelled) == 0,\n     \"details\": f\"{len(labelled_but_missing_from_disk)} labels without a photo, {len(present_on_disk_but_not_labelled)} photos without a label\"},\n    {\"smoke_test\": \"2. Every grade is a whole number from 0 to 4\",\n     \"passed\": labels_are_valid_grades,\n     \"details\": f\"grades found: {sorted(diagnosis_column.dropna().unique().tolist())}\"},\n    {\"smoke_test\": \"3. The rarest grade has enough photos to split safely\",\n     \"passed\": images_in_rarest_grade >= minimum_images_needed,\n     \"details\": f\"rarest grade has {images_in_rarest_grade} photos, at least {minimum_images_needed} needed\"},\n])\nsmoke_test_results_dataframe[\"result\"] = np.where(smoke_test_results_dataframe[\"passed\"], \"PASS\", \"FAIL\")\nall_data_smoke_tests_passed = bool(smoke_test_results_dataframe[\"passed\"].all())\nsave_results_table(smoke_test_results_dataframe[[\"smoke_test\", \"result\", \"details\"]], \"01_data_exploration\", \"data_smoke_tests.csv\")\n\nprint(\"Cell 6 finished: data smoke tests complete.\")\nprint(smoke_test_results_dataframe[[\"smoke_test\", \"result\", \"details\"]].to_string(index=False))\nprint(\"All data checks passed, safe to continue.\" if all_data_smoke_tests_passed\n      else \"At least one check failed. Stop here and fix the data before going any further.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:56.496763Z","iopub.execute_input":"2026-09-26T11:32:56.497113Z","iopub.status.idle":"2026-09-26T11:32:56.625481Z","shell.execute_reply.started":"2026-09-26T11:32:56.497077Z","shell.execute_reply":"2026-09-26T11:32:56.624626Z"}},"outputs":[],"execution_count":null},{"id":"bf50d77e","cell_type":"code","source":"# Cell 7: count the photos in each grade and chart how unbalanced the dataset is\nclass_counts_series = diabetic_retinopathy_labels_dataframe[\"diagnosis\"].value_counts().sort_index()\nclass_percentages_series = (class_counts_series / class_counts_series.sum() * 100).round(1)\nimbalance_ratio = class_counts_series.max() / class_counts_series.min()\n\nclass_distribution_dataframe = pd.DataFrame({\n    \"grade\": class_counts_series.index,\n    \"stage\": [diagnosis_stage_names[grade] for grade in class_counts_series.index],\n    \"photo_count\": class_counts_series.values,\n    \"percent_of_dataset\": class_percentages_series.values,\n})\nsave_results_table(class_distribution_dataframe, \"01_data_exploration\", \"class_distribution.csv\")\n\nfigure_object, axis_object = plt.subplots(figsize=(9, 5))\naxis_object.bar(class_distribution_dataframe[\"stage\"], class_distribution_dataframe[\"photo_count\"],\n                color=seaborn_plotting_library.color_palette(\"viridis\", len(class_counts_series)))\nfor bar_index, (photo_count, percent_value) in enumerate(zip(class_distribution_dataframe[\"photo_count\"], class_distribution_dataframe[\"percent_of_dataset\"])):\n    axis_object.text(bar_index, photo_count + 15, f\"{photo_count} ({percent_value}%)\", ha=\"center\", fontsize=9)\naxis_object.set_title(\"Number of Photos per DR Grade (APTOS 2019)\")\naxis_object.set_xlabel(\"Diabetic retinopathy grade\")\naxis_object.set_ylabel(\"Number of photos\")\nplt.xticks(rotation=15)\nsave_current_figure(\"01_data_exploration\", \"class_distribution.png\")\n\nprint(\"Cell 7 finished: class distribution counted and charted.\")\nprint(class_distribution_dataframe.to_string(index=False))\nprint(f\"The biggest grade has {imbalance_ratio:.1f} times more photos than the smallest one.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:56.62656Z","iopub.execute_input":"2026-09-26T11:32:56.627002Z","iopub.status.idle":"2026-09-26T11:32:57.100636Z","shell.execute_reply.started":"2026-09-26T11:32:56.626968Z","shell.execute_reply":"2026-09-26T11:32:57.099892Z"}},"outputs":[],"execution_count":null},{"id":"0535b4cb","cell_type":"code","source":"# Cell 8: show three raw photos from each grade to see how the grades differ\nnumber_of_examples_per_class = 3\nfigure_object, axis_grid = plt.subplots(len(list_of_class_indices), number_of_examples_per_class,\n                                        figsize=(number_of_examples_per_class * 3.2, len(list_of_class_indices) * 3.2))\nfor class_label in list_of_class_indices:\n    example_rows = diabetic_retinopathy_labels_dataframe[diabetic_retinopathy_labels_dataframe[\"diagnosis\"] == class_label].sample(\n        number_of_examples_per_class, random_state=RANDOM_SEED)\n    for column_index, (_, example_row) in enumerate(example_rows.iterrows()):\n        current_axis = axis_grid[class_label, column_index]\n        current_axis.imshow(cv2.cvtColor(cv2.imread(build_image_path_from_id(example_row[\"id_code\"])), cv2.COLOR_BGR2RGB))\n        current_axis.set_title(diagnosis_stage_names[class_label], fontsize=10)\n        current_axis.axis(\"off\")\nplt.suptitle(\"Raw Fundus Photos by DR Grade (before any preprocessing)\", y=1.0)\nsave_current_figure(\"01_data_exploration\", \"sample_photos_by_grade.png\")\n\nprint(\"Cell 8 finished: sample photos shown for every grade.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:32:57.101614Z","iopub.execute_input":"2026-09-26T11:32:57.101902Z","iopub.status.idle":"2026-09-26T11:33:11.45681Z","shell.execute_reply.started":"2026-09-26T11:32:57.10188Z","shell.execute_reply":"2026-09-26T11:33:11.456Z"}},"outputs":[],"execution_count":null},{"id":"8a003025","cell_type":"code","source":"# Cell 9: measure the width, height and shape of every photo to see how much they vary\nrecorded_widths = []\nrecorded_heights = []\nfor image_id_code in diabetic_retinopathy_labels_dataframe[\"id_code\"]:\n    with Image.open(build_image_path_from_id(image_id_code)) as opened_image:\n        image_width, image_height = opened_image.size\n    recorded_widths.append(image_width)\n    recorded_heights.append(image_height)\n\nimage_dimensions_dataframe = pd.DataFrame({\n    \"id_code\": diabetic_retinopathy_labels_dataframe[\"id_code\"], \"width\": recorded_widths, \"height\": recorded_heights,\n})\nimage_dimensions_dataframe[\"aspect_ratio\"] = image_dimensions_dataframe[\"width\"] / image_dimensions_dataframe[\"height\"]\nnumber_of_distinct_resolutions = image_dimensions_dataframe.groupby([\"width\", \"height\"]).ngroups\nresolution_summary_dataframe = image_dimensions_dataframe[[\"width\", \"height\", \"aspect_ratio\"]].describe().round(2)\nsave_results_table(resolution_summary_dataframe, \"01_data_exploration\", \"resolution_summary.csv\", keep_index=True)\n\nfigure_object, axis_array = plt.subplots(1, 3, figsize=(16, 4))\nfor current_axis, column_name, colour, chart_title in zip(\n    axis_array, [\"width\", \"height\", \"aspect_ratio\"], [\"steelblue\", \"darkorange\", \"seagreen\"],\n    [\"Width (pixels)\", \"Height (pixels)\", \"Aspect ratio (width / height)\"],\n):\n    current_axis.hist(image_dimensions_dataframe[column_name], bins=30, color=colour)\n    current_axis.set_title(chart_title)\n    current_axis.set_ylabel(\"Number of photos\")\nsave_current_figure(\"01_data_exploration\", \"resolution_and_aspect_ratio.png\")\n\nprint(\"Cell 9 finished: size and shape measured for every photo.\")\nprint(resolution_summary_dataframe)\nprint(\"Number of different resolutions in the dataset:\", number_of_distinct_resolutions)\nprint(\"Sizes and shapes vary a lot, so every photo must be brought to one square size without stretching the eye.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:33:11.457842Z","iopub.execute_input":"2026-09-26T11:33:11.458278Z","iopub.status.idle":"2026-09-26T11:34:42.175943Z","shell.execute_reply.started":"2026-09-26T11:33:11.458254Z","shell.execute_reply":"2026-09-26T11:34:42.175112Z"}},"outputs":[],"execution_count":null},{"id":"9f7e04ac","cell_type":"markdown","source":"## 3. Preprocessing Methods\nThree methods are built and compared. They all share the same first steps, so only the contrast step differs:\n\n1. **Crop the black border** so only the eye is left.\n2. **Pad to a square** with black. Many APTOS photos cut off the top and bottom of the eye, so the crop is wider than it is tall. Resizing that rectangle straight to a square would squash the eye into an oval, so it is padded to a square first and keeps its true shape.\n3. **Circular mask** so only the round eye is left. The circle now fits the square, so the sides of the eye are no longer clipped.\n4. **Contrast step**, the only part that differs:\n   - **Resize only:** no enhancement. This is the baseline that shows whether enhancement helps at all.\n   - **CLAHE:** boosts contrast in small patches of the brightness channel, which makes tiny spots easier to see without changing colours (Lee et al., 2023; Fatima et al., 2024).\n   - **Ben Graham:** subtracts a blurred copy of the photo, removing uneven lighting and camera colour differences and leaving sharp details like spots and new vessels.\n5. **Resize to 384 x 384.**\n\nWhich method wins is decided by experiment in section 9, not by opinion.","metadata":{}},{"id":"db653bf7","cell_type":"code","source":"# Cell 10: shared step 1, crop away the black border around the eye\ndef crop_away_dark_borders(image_array_rgb, brightness_threshold=None):\n    if brightness_threshold is None:\n        brightness_threshold = project_configuration[\"dark_pixel_brightness_threshold\"]\n    non_dark_pixel_mask = cv2.cvtColor(image_array_rgb, cv2.COLOR_RGB2GRAY) > brightness_threshold\n    if non_dark_pixel_mask.sum() == 0:\n        return image_array_rgb\n    non_dark_row_indices = np.where(non_dark_pixel_mask.any(axis=1))[0]\n    non_dark_column_indices = np.where(non_dark_pixel_mask.any(axis=0))[0]\n    cropped_image = image_array_rgb[non_dark_row_indices[0]:non_dark_row_indices[-1] + 1, non_dark_column_indices[0]:non_dark_column_indices[-1] + 1]\n    if cropped_image.shape[0] < 10 or cropped_image.shape[1] < 10:\n        return image_array_rgb\n    return cropped_image\n\n\nprint(\"Cell 10 finished: border cropping function ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.177073Z","iopub.execute_input":"2026-09-26T11:34:42.177355Z","iopub.status.idle":"2026-09-26T11:34:42.183995Z","shell.execute_reply.started":"2026-09-26T11:34:42.17733Z","shell.execute_reply":"2026-09-26T11:34:42.183163Z"}},"outputs":[],"execution_count":null},{"id":"a1111aa2","cell_type":"code","source":"# Cell 11: shared step 2, pad the crop to a centred square so the eye keeps its true round shape\ndef pad_to_centred_square(image_array_rgb, fill_value=0):\n    image_height, image_width = image_array_rgb.shape[:2]\n    square_side = max(image_height, image_width)\n    top_padding = (square_side - image_height) // 2\n    left_padding = (square_side - image_width) // 2\n    return cv2.copyMakeBorder(\n        image_array_rgb, top_padding, square_side - image_height - top_padding, left_padding, square_side - image_width - left_padding,\n        cv2.BORDER_CONSTANT, value=(fill_value, fill_value, fill_value),\n    )\n\n\nprint(\"Cell 11 finished: pad-to-square function ready. Padding keeps the eye round when it is resized.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.184919Z","iopub.execute_input":"2026-09-26T11:34:42.185318Z","iopub.status.idle":"2026-09-26T11:34:42.208231Z","shell.execute_reply.started":"2026-09-26T11:34:42.185278Z","shell.execute_reply":"2026-09-26T11:34:42.207667Z"}},"outputs":[],"execution_count":null},{"id":"aa2d2279","cell_type":"code","source":"# Cell 12: shared step 3, blank out everything outside the circle that fits the square\ndef apply_circular_field_of_view_mask(square_image_rgb):\n    image_height, image_width = square_image_rgb.shape[:2]\n    circular_mask = np.zeros((image_height, image_width), dtype=np.uint8)\n    cv2.circle(circular_mask, (image_width // 2, image_height // 2), min(image_height, image_width) // 2, 255, thickness=-1)\n    return cv2.bitwise_and(square_image_rgb, square_image_rgb, mask=circular_mask)\n\n\nprint(\"Cell 12 finished: circular mask function ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.209096Z","iopub.execute_input":"2026-09-26T11:34:42.20981Z","iopub.status.idle":"2026-09-26T11:34:42.224771Z","shell.execute_reply.started":"2026-09-26T11:34:42.20978Z","shell.execute_reply":"2026-09-26T11:34:42.223993Z"}},"outputs":[],"execution_count":null},{"id":"efb62731","cell_type":"code","source":"# Cell 13: contrast option A, CLAHE on the brightness channel only, then re-mask so the background stays pure black\ndef enhance_local_contrast_with_clahe(masked_image_rgb):\n    lightness_channel, a_channel, b_channel = cv2.split(cv2.cvtColor(masked_image_rgb, cv2.COLOR_RGB2LAB))\n    tile_count = project_configuration[\"clahe_tile_grid_size\"]\n    clahe_operator = cv2.createCLAHE(clipLimit=project_configuration[\"clahe_clip_limit\"], tileGridSize=(tile_count, tile_count))\n    enhanced_image = cv2.cvtColor(cv2.merge((clahe_operator.apply(lightness_channel), a_channel, b_channel)), cv2.COLOR_LAB2RGB)\n    return apply_circular_field_of_view_mask(enhanced_image)\n\n\nprint(\"Cell 13 finished: CLAHE function ready, clip limit\", project_configuration[\"clahe_clip_limit\"],\n      \"on a\", project_configuration[\"clahe_tile_grid_size\"], \"by\", project_configuration[\"clahe_tile_grid_size\"], \"grid.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.225719Z","iopub.execute_input":"2026-09-26T11:34:42.22618Z","iopub.status.idle":"2026-09-26T11:34:42.241202Z","shell.execute_reply.started":"2026-09-26T11:34:42.226155Z","shell.execute_reply":"2026-09-26T11:34:42.24063Z"}},"outputs":[],"execution_count":null},{"id":"7b27e3ed","cell_type":"code","source":"# Cell 14: contrast option B, the Ben Graham method (subtract the blurred background, keep the inner eye)\ndef subtract_local_average_colour(image_array_rgb, gaussian_sigma):\n    blurred_background_estimate = cv2.GaussianBlur(image_array_rgb, (0, 0), gaussian_sigma)\n    return cv2.addWeighted(image_array_rgb, 4, blurred_background_estimate, -4, 128)\n\n\ndef keep_inner_eye_with_neutral_fill(image_array_rgb, radius_fraction=0.9, neutral_fill_value=128):\n    image_height, image_width = image_array_rgb.shape[:2]\n    inner_circle_mask = np.zeros((image_height, image_width), dtype=np.uint8)\n    cv2.circle(inner_circle_mask, (image_width // 2, image_height // 2), int(min(image_width, image_height) // 2 * radius_fraction), 1, thickness=-1)\n    inner_circle_mask = inner_circle_mask[:, :, np.newaxis]\n    return (image_array_rgb * inner_circle_mask + neutral_fill_value * (1 - inner_circle_mask)).astype(np.uint8)\n\n\ndef apply_ben_graham_enhancement(resized_masked_image_rgb):\n    gaussian_sigma = (resized_masked_image_rgb.shape[0] / 2) / project_configuration[\"ben_graham_blur_divisor\"]\n    return keep_inner_eye_with_neutral_fill(subtract_local_average_colour(resized_masked_image_rgb, gaussian_sigma))\n\n\nprint(\"Cell 14 finished: Ben Graham function ready. The blur size is the eye radius divided by\",\n      project_configuration[\"ben_graham_blur_divisor\"], \"as in Graham's original method.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.244183Z","iopub.execute_input":"2026-09-26T11:34:42.244508Z","iopub.status.idle":"2026-09-26T11:34:42.262677Z","shell.execute_reply.started":"2026-09-26T11:34:42.244487Z","shell.execute_reply":"2026-09-26T11:34:42.262025Z"}},"outputs":[],"execution_count":null},{"id":"733b3b8c","cell_type":"code","source":"# Cell 15: join the steps into one pipeline with a method switch, and cache each cleaned photo to save time\ndef load_image_as_rgb(image_file_path):\n    return cv2.cvtColor(cv2.imread(image_file_path), cv2.COLOR_BGR2RGB)\n\n\ndef resize_to_target(image_rgb, target_size):\n    return cv2.resize(image_rgb, (target_size, target_size), interpolation=cv2.INTER_AREA)\n\n\ndef run_preprocessing_steps(image_rgb, target_size, preprocessing_method):\n    cropped_image = crop_away_dark_borders(image_rgb)\n    squared_image = pad_to_centred_square(cropped_image)\n    masked_image = apply_circular_field_of_view_mask(squared_image)\n    if preprocessing_method == \"clahe\":\n        enhanced_image = enhance_local_contrast_with_clahe(masked_image)\n        final_image = resize_to_target(enhanced_image, target_size)\n    elif preprocessing_method == \"ben_graham\":\n        final_image = apply_ben_graham_enhancement(resize_to_target(masked_image, target_size))\n        enhanced_image = final_image\n    elif preprocessing_method == \"resize_only\":\n        enhanced_image = masked_image\n        final_image = resize_to_target(masked_image, target_size)\n    else:\n        raise ValueError(f\"Unknown preprocessing method: {preprocessing_method}\")\n    return {\n        \"original\": image_rgb,\n        \"cropped\": cropped_image,\n        \"squared\": squared_image,\n        \"masked\": masked_image,\n        \"enhanced\": enhanced_image,\n        \"resized\": final_image,\n    }\n\n\ndef preprocess_fundus_image(image_file_path, target_size, preprocessing_method):\n    cache_file_name = os.path.splitext(os.path.basename(image_file_path))[0] + f\"_{target_size}.npy\"\n    cache_file_path = os.path.join(cache_folder_paths[preprocessing_method], cache_file_name)\n    if os.path.exists(cache_file_path):\n        return np.load(cache_file_path)\n    preprocessed_image = run_preprocessing_steps(load_image_as_rgb(image_file_path), target_size, preprocessing_method)[\"resized\"]\n    np.save(cache_file_path, preprocessed_image)\n    return preprocessed_image\n\n\nprint(\"Cell 15 finished: one pipeline, three methods:\", preprocessing_method_names)\nprint(\"Shared order: crop border, pad to square, circular mask, then the method's contrast step, then resize.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.263648Z","iopub.execute_input":"2026-09-26T11:34:42.263969Z","iopub.status.idle":"2026-09-26T11:34:42.281655Z","shell.execute_reply.started":"2026-09-26T11:34:42.263948Z","shell.execute_reply":"2026-09-26T11:34:42.281103Z"}},"outputs":[],"execution_count":null},{"id":"1f9a225d","cell_type":"code","source":"# Cell 16: PREPROCESSING TESTS, check every method really did what it should, including that the eye keeps its shape\ndef measure_eye_width_to_height_ratio(image_rgb):\n    eye_pixels = cv2.cvtColor(image_rgb, cv2.COLOR_RGB2GRAY) > project_configuration[\"dark_pixel_brightness_threshold\"]\n    eye_rows = np.where(eye_pixels.any(axis=1))[0]\n    eye_columns = np.where(eye_pixels.any(axis=0))[0]\n    return (eye_columns[-1] - eye_columns[0] + 1) / (eye_rows[-1] - eye_rows[0] + 1)\n\n\ndef measure_average_local_contrast_inside_eye(image_rgb, window_size=15):\n    image_height, image_width = image_rgb.shape[:2]\n    inner_eye_mask = np.zeros((image_height, image_width), dtype=np.uint8)\n    cv2.circle(inner_eye_mask, (image_width // 2, image_height // 2), int(min(image_height, image_width) // 2 * 0.85), 1, thickness=-1)\n    green_channel = image_rgb[:, :, 1].astype(np.float32)\n    local_mean = cv2.blur(green_channel, (window_size, window_size))\n    local_standard_deviation = np.sqrt(np.clip(cv2.blur(green_channel ** 2, (window_size, window_size)) - local_mean ** 2, 0, None))\n    return float(local_standard_deviation[inner_eye_mask.astype(bool)].mean())\n\n\ntarget_size_for_tests = project_configuration[\"target_image_size\"]\npreprocessing_test_sample = diabetic_retinopathy_labels_dataframe.sample(50, random_state=RANDOM_SEED)\nper_photo_test_rows = []\nfor _, sample_row in preprocessing_test_sample.iterrows():\n    raw_image_rgb = load_image_as_rgb(build_image_path_from_id(sample_row[\"id_code\"]))\n    steps_by_method = {method: run_preprocessing_steps(raw_image_rgb, target_size_for_tests, method) for method in preprocessing_method_names}\n    cropped_image = steps_by_method[\"resize_only\"][\"cropped\"]\n    ratio_before = cropped_image.shape[1] / cropped_image.shape[0]\n    ratio_after = measure_eye_width_to_height_ratio(steps_by_method[\"resize_only\"][\"resized\"])\n    clahe_final = steps_by_method[\"clahe\"][\"resized\"]\n    ben_graham_final = steps_by_method[\"ben_graham\"][\"resized\"]\n    inner_region = ben_graham_final[target_size_for_tests // 4: 3 * target_size_for_tests // 4, target_size_for_tests // 4: 3 * target_size_for_tests // 4]\n    per_photo_test_rows.append({\n        \"id_code\": sample_row[\"id_code\"],\n        \"eye_ratio_before\": ratio_before,\n        \"eye_ratio_after\": ratio_after,\n        \"shape_distortion_percent\": abs(ratio_after - ratio_before) / ratio_before * 100,\n        \"all_outputs_correct_size\": all(steps[\"resized\"].shape == (target_size_for_tests, target_size_for_tests, 3) and steps[\"resized\"].dtype == np.uint8 for steps in steps_by_method.values()),\n        \"black_corners_resize_only_and_clahe\": int(steps_by_method[\"resize_only\"][\"resized\"][0, 0].max()) == 0 and int(clahe_final[0, 0].max()) == 0,\n        \"neutral_corners_ben_graham\": bool(np.all(ben_graham_final[0, 0] == 128)),\n        \"clahe_raised_local_contrast\": measure_average_local_contrast_inside_eye(clahe_final) > measure_average_local_contrast_inside_eye(steps_by_method[\"resize_only\"][\"resized\"]),\n        \"ben_graham_background_removed\": abs(float(inner_region.mean()) - 128) < 30,\n    })\nper_photo_test_dataframe = pd.DataFrame(per_photo_test_rows)\nsave_results_table(per_photo_test_dataframe.round(4), \"02_preprocessing_methods\", \"preprocessing_tests_per_photo.csv\")\n\ncut_eye_rows = per_photo_test_dataframe[per_photo_test_dataframe[\"eye_ratio_before\"] > 1.05]\npreprocessing_test_summary_dataframe = pd.DataFrame([\n    {\"test\": \"Every method outputs a square colour image of the target size\", \"passed_percent\": per_photo_test_dataframe[\"all_outputs_correct_size\"].mean() * 100, \"required_percent\": 100},\n    {\"test\": \"The eye keeps its shape after resizing (distortion under 3%)\", \"passed_percent\": (per_photo_test_dataframe[\"shape_distortion_percent\"] < 3).mean() * 100, \"required_percent\": 100},\n    {\"test\": \"Background corners are pure black (resize only, CLAHE)\", \"passed_percent\": per_photo_test_dataframe[\"black_corners_resize_only_and_clahe\"].mean() * 100, \"required_percent\": 100},\n    {\"test\": \"Background corners are neutral grey (Ben Graham)\", \"passed_percent\": per_photo_test_dataframe[\"neutral_corners_ben_graham\"].mean() * 100, \"required_percent\": 100},\n    {\"test\": \"CLAHE raises local contrast inside the eye\", \"passed_percent\": per_photo_test_dataframe[\"clahe_raised_local_contrast\"].mean() * 100, \"required_percent\": 90},\n    {\"test\": \"Ben Graham removes the background colour (centre near grey 128)\", \"passed_percent\": per_photo_test_dataframe[\"ben_graham_background_removed\"].mean() * 100, \"required_percent\": 90},\n])\npreprocessing_test_summary_dataframe[\"result\"] = np.where(\n    preprocessing_test_summary_dataframe[\"passed_percent\"] >= preprocessing_test_summary_dataframe[\"required_percent\"], \"PASS\", \"FAIL\")\nsave_results_table(preprocessing_test_summary_dataframe.round(1), \"02_preprocessing_methods\", \"preprocessing_tests_summary.csv\")\n\nmost_cut_row = per_photo_test_dataframe.loc[per_photo_test_dataframe[\"eye_ratio_before\"].idxmax()]\nmost_cut_steps = run_preprocessing_steps(load_image_as_rgb(build_image_path_from_id(most_cut_row[\"id_code\"])), target_size_for_tests, \"resize_only\")\nfigure_object, axis_array = plt.subplots(1, 4, figsize=(18, 5))\nfor current_axis, step_name, panel_title in zip(axis_array, [\"original\", \"cropped\", \"squared\", \"resized\"], [\n    \"Raw photo\",\n    f\"Border cropped: eye is {most_cut_row['eye_ratio_before']:.2f} times wider than tall\",\n    \"Padded to a square with black\",\n    f\"Masked and resized: width to height still {most_cut_row['eye_ratio_after']:.2f}\",\n]):\n    current_axis.imshow(most_cut_steps[step_name])\n    current_axis.set_title(panel_title, fontsize=10)\n    current_axis.axis(\"off\")\nplt.suptitle(\"How Padding Keeps the Eye's True Shape on a Photo Cut Off at the Top and Bottom\", y=1.02)\nsave_current_figure(\"02_preprocessing_methods\", \"eye_shape_kept_by_padding.png\")\n\nprint(\"Cell 16 finished: preprocessing tests complete on\", len(per_photo_test_dataframe), \"photos.\")\nprint(preprocessing_test_summary_dataframe.round(1).to_string(index=False))\nprint()\nprint(f\"Photos in the test sample where the eye is cut at the top and bottom: {len(cut_eye_rows)} of {len(per_photo_test_dataframe)}\")\nif len(cut_eye_rows) > 0:\n    print(f\"Average shape distortion on those photos after preprocessing: {cut_eye_rows['shape_distortion_percent'].mean():.1f}%\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:34:42.282554Z","iopub.execute_input":"2026-09-26T11:34:42.282863Z","iopub.status.idle":"2026-09-26T11:35:06.130019Z","shell.execute_reply.started":"2026-09-26T11:34:42.282831Z","shell.execute_reply":"2026-09-26T11:35:06.129239Z"}},"outputs":[],"execution_count":null},{"id":"2e8dd74b","cell_type":"code","source":"# Cell 17: show every step for one photo per grade, and compare how much local contrast each method gives\nstep_preview_rows = [\n    diabetic_retinopathy_labels_dataframe[diabetic_retinopathy_labels_dataframe[\"diagnosis\"] == class_label].sample(1, random_state=RANDOM_SEED).iloc[0]\n    for class_label in list_of_class_indices\n]\npreview_column_titles = [\"Original\", \"1. Border cropped\", \"2. Padded to square\", \"3. Circular mask\", \"Final: resize only\", \"Final: CLAHE\", \"Final: Ben Graham\"]\nfigure_object, axis_grid = plt.subplots(len(step_preview_rows), len(preview_column_titles), figsize=(24, 3.6 * len(step_preview_rows)))\nfor row_index, preview_row in enumerate(step_preview_rows):\n    raw_image_rgb = load_image_as_rgb(build_image_path_from_id(preview_row[\"id_code\"]))\n    shared_steps = run_preprocessing_steps(raw_image_rgb, project_configuration[\"target_image_size\"], \"resize_only\")\n    panel_images = [\n        shared_steps[\"original\"], shared_steps[\"cropped\"], shared_steps[\"squared\"], shared_steps[\"masked\"], shared_steps[\"resized\"],\n        run_preprocessing_steps(raw_image_rgb, project_configuration[\"target_image_size\"], \"clahe\")[\"resized\"],\n        run_preprocessing_steps(raw_image_rgb, project_configuration[\"target_image_size\"], \"ben_graham\")[\"resized\"],\n    ]\n    for column_index, panel_image in enumerate(panel_images):\n        current_axis = axis_grid[row_index, column_index]\n        current_axis.imshow(panel_image)\n        grade_prefix = f\"{diagnosis_stage_names[int(preview_row['diagnosis'])]}: \" if column_index == 0 else \"\"\n        current_axis.set_title(grade_prefix + preview_column_titles[column_index], fontsize=10)\n        current_axis.axis(\"off\")\nplt.suptitle(\"Preprocessing Step by Step and the Three Methods, One Photo per Grade\", y=1.0)\nsave_current_figure(\"02_preprocessing_methods\", \"preprocessing_steps_and_methods_per_grade.png\")\n\ncontrast_sample = diabetic_retinopathy_labels_dataframe.sample(200, random_state=RANDOM_SEED)\ncontrast_rows = []\nfor _, sample_row in contrast_sample.iterrows():\n    raw_image_rgb = load_image_as_rgb(build_image_path_from_id(sample_row[\"id_code\"]))\n    contrast_row = {\"id_code\": sample_row[\"id_code\"], \"grade\": int(sample_row[\"diagnosis\"])}\n    for method_name in preprocessing_method_names:\n        contrast_row[f\"local_contrast_{method_name}\"] = measure_average_local_contrast_inside_eye(\n            run_preprocessing_steps(raw_image_rgb, project_configuration[\"target_image_size\"], method_name)[\"resized\"])\n    contrast_rows.append(contrast_row)\ncontrast_dataframe = pd.DataFrame(contrast_rows)\nsave_results_table(contrast_dataframe.round(3), \"02_preprocessing_methods\", \"local_contrast_per_photo_per_method.csv\")\ncontrast_summary_dataframe = pd.DataFrame({\n    \"method\": preprocessing_method_names,\n    \"average_local_contrast\": [contrast_dataframe[f\"local_contrast_{method_name}\"].mean() for method_name in preprocessing_method_names],\n})\ncontrast_summary_dataframe[\"times_resize_only\"] = contrast_summary_dataframe[\"average_local_contrast\"] / contrast_summary_dataframe.loc[0, \"average_local_contrast\"]\nsave_results_table(contrast_summary_dataframe.round(3), \"02_preprocessing_methods\", \"local_contrast_summary_per_method.csv\")\n\nprint(\"Cell 17 finished: step-by-step figure saved and local contrast compared on 200 photos.\")\nprint(contrast_summary_dataframe.round(3).to_string(index=False))\nprint(\"Higher local contrast makes small spots easier to see, but it does not prove better grading.\")\nprint(\"That is decided by the preprocessing experiment in section 9.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:35:06.131541Z","iopub.execute_input":"2026-09-26T11:35:06.131917Z","iopub.status.idle":"2026-09-26T11:36:45.060675Z","shell.execute_reply.started":"2026-09-26T11:35:06.131873Z","shell.execute_reply":"2026-09-26T11:36:45.059846Z"}},"outputs":[],"execution_count":null},{"id":"47d93dca","cell_type":"markdown","source":"## 4. Image Size Check\nA bigger photo keeps more of the tiny spots that show the disease, but it makes training slower. This check measures both sides **without training any model**:\n\n- **Detail kept:** 100 photos (20 per grade) are cleaned at a large size of 1000 x 1000, shrunk to each size from 224 to 512, then grown back. **SSIM** (structural similarity, 1.0 means identical) compares the result with the original, inside the eye only, on the green channel where lesions stand out most.\n- **Cost:** each model is timed on the GPU at each size with a batch of 16, and its memory use is checked. This also confirms that a batch size of 16 fits.\n\nSizes go up in steps of 32 because all three models shrink the photo 32 times before their last layer. The pipeline uses 384. These results show what 384 keeps and what a bigger size would cost.","metadata":{}},{"id":"42ce1fa8","cell_type":"code","source":"# Cell 18: image size check part 1, measure how much structure survives at each size using SSIM\ncandidate_image_sizes = list(range(224, 513, 32))\nreference_image_size = 1000\nimages_per_grade_for_size_study = 20\n\nsize_study_sample_dataframe = diabetic_retinopathy_labels_dataframe.groupby(\"diagnosis\").sample(n=images_per_grade_for_size_study, random_state=RANDOM_SEED)\ninner_region_mask = np.zeros((reference_image_size, reference_image_size), dtype=np.uint8)\ncv2.circle(inner_region_mask, (reference_image_size // 2, reference_image_size // 2), int(reference_image_size // 2 * 0.85), 1, thickness=-1)\ninner_region_mask = inner_region_mask.astype(bool)\n\n\ndef compute_structural_similarity_map(first_channel, second_channel):\n    first_channel = first_channel.astype(np.float64)\n    second_channel = second_channel.astype(np.float64)\n    constant_one, constant_two = (0.01 * 255) ** 2, (0.03 * 255) ** 2\n    blur = partial(cv2.GaussianBlur, ksize=(11, 11), sigmaX=1.5)\n    mean_first, mean_second = blur(first_channel), blur(second_channel)\n    variance_first = blur(first_channel ** 2) - mean_first ** 2\n    variance_second = blur(second_channel ** 2) - mean_second ** 2\n    covariance = blur(first_channel * second_channel) - mean_first * mean_second\n    return ((2 * mean_first * mean_second + constant_one) * (2 * covariance + constant_two)) / (\n        (mean_first ** 2 + mean_second ** 2 + constant_one) * (variance_first + variance_second + constant_two))\n\n\ndef build_reference_green_channel(image_file_path):\n    masked_image = run_preprocessing_steps(load_image_as_rgb(image_file_path), reference_image_size, \"resize_only\")[\"masked\"]\n    resize_method = cv2.INTER_AREA if masked_image.shape[0] >= reference_image_size else cv2.INTER_CUBIC\n    return cv2.resize(masked_image, (reference_image_size, reference_image_size), interpolation=resize_method)[:, :, 1]\n\n\ndef measure_structure_kept_at_size(reference_green_channel, candidate_size):\n    shrunk_channel = cv2.resize(reference_green_channel, (candidate_size, candidate_size), interpolation=cv2.INTER_AREA)\n    restored_channel = cv2.resize(shrunk_channel, (reference_image_size, reference_image_size), interpolation=cv2.INTER_CUBIC)\n    return float(compute_structural_similarity_map(reference_green_channel, restored_channel)[inner_region_mask].mean())\n\n\nsize_study_start_time = time.time()\ndetail_retention_rows = []\nfor _, sample_row in size_study_sample_dataframe.iterrows():\n    reference_green_channel = build_reference_green_channel(build_image_path_from_id(sample_row[\"id_code\"]))\n    for candidate_size in candidate_image_sizes:\n        detail_retention_rows.append({\n            \"id_code\": sample_row[\"id_code\"], \"diagnosis\": int(sample_row[\"diagnosis\"]), \"image_size\": candidate_size,\n            \"ssim_kept\": measure_structure_kept_at_size(reference_green_channel, candidate_size),\n        })\ndetail_retention_dataframe = pd.DataFrame(detail_retention_rows)\nsave_results_table(detail_retention_dataframe, \"03_image_size_study\", \"ssim_per_photo_per_size.csv\")\n\ndetail_retention_summary = detail_retention_dataframe.groupby(\"image_size\")[\"ssim_kept\"].agg([\"mean\", \"std\", \"count\"])\ndetail_retention_summary[\"ci95_half_width\"] = 1.96 * detail_retention_summary[\"std\"] / np.sqrt(detail_retention_summary[\"count\"])\n\nprint(f\"Cell 18 finished: structure measured on {len(size_study_sample_dataframe)} photos in {(time.time() - size_study_start_time) / 60:.1f} minutes.\")\nprint(\"Average SSIM at each size compared with the 1000 x 1000 original (1.0 means nothing was lost):\")\nprint(detail_retention_summary[[\"mean\", \"ci95_half_width\"]].round(4).to_string())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:36:45.061642Z","iopub.execute_input":"2026-09-26T11:36:45.061983Z","iopub.status.idle":"2026-09-26T11:38:05.261895Z","shell.execute_reply.started":"2026-09-26T11:36:45.061959Z","shell.execute_reply":"2026-09-26T11:38:05.261113Z"}},"outputs":[],"execution_count":null},{"id":"6bbdb6a0","cell_type":"code","source":"# Cell 19: image size check part 2, time one training step for each model at each size and check GPU memory with a batch of 16\narchitecture_builders_for_timing = {\"efficientnet_b3\": efficientnet_b3, \"efficientnet_b0\": efficientnet_b0, \"resnet50\": resnet50}\nnumber_of_warmup_steps = 2\nnumber_of_timed_steps = 5\nusable_gpu_memory_gigabytes = torch.cuda.get_device_properties(0).total_memory / 1e9 * 0.9 if is_gpu_available else float(\"nan\")\nexpected_training_images = round(len(diabetic_retinopathy_labels_dataframe) * (1 - project_configuration[\"validation_split_fraction\"] - project_configuration[\"test_split_fraction\"]))\nexpected_validation_images = round(len(diabetic_retinopathy_labels_dataframe) * project_configuration[\"validation_split_fraction\"])\ntraining_batches_per_epoch = int(np.ceil(expected_training_images / project_configuration[\"batch_size\"]))\nvalidation_batches_per_epoch = int(np.ceil(expected_validation_images / project_configuration[\"batch_size\"]))\nexpected_epochs_per_architecture = 20\n\ngpu_cost_rows = []\nif not is_gpu_available:\n    print(\"No GPU found, so the timing check is skipped. Turn on the GPU to get real timing numbers.\")\nelse:\n    for architecture_name, architecture_builder in architecture_builders_for_timing.items():\n        timing_model = architecture_builder(weights=None, num_classes=project_configuration[\"number_of_dr_classes\"]).to(compute_device)\n        timing_model.train()\n        timing_optimizer = optimizers.Adam(timing_model.parameters(), lr=1e-4)\n        timing_loss_function = neural_network_layers.CrossEntropyLoss()\n        for candidate_size in candidate_image_sizes:\n            torch.cuda.empty_cache()\n            torch.cuda.reset_peak_memory_stats()\n            try:\n                fake_images = torch.randn(project_configuration[\"batch_size\"], 3, candidate_size, candidate_size, device=compute_device)\n                fake_labels = torch.randint(0, project_configuration[\"number_of_dr_classes\"], (project_configuration[\"batch_size\"],), device=compute_device)\n                for step_index in range(number_of_warmup_steps + number_of_timed_steps):\n                    if step_index == number_of_warmup_steps:\n                        torch.cuda.synchronize()\n                        timing_start = time.time()\n                    timing_optimizer.zero_grad()\n                    timing_loss_function(timing_model(fake_images), fake_labels).backward()\n                    timing_optimizer.step()\n                torch.cuda.synchronize()\n                seconds_per_training_step = (time.time() - timing_start) / number_of_timed_steps\n                peak_memory_gigabytes = torch.cuda.max_memory_allocated() / 1e9\n                fits_in_memory = peak_memory_gigabytes <= usable_gpu_memory_gigabytes\n            except torch.cuda.OutOfMemoryError:\n                seconds_per_training_step, peak_memory_gigabytes, fits_in_memory = float(\"nan\"), float(\"nan\"), False\n                timing_optimizer.zero_grad(set_to_none=True)\n            fake_images = fake_labels = None\n            torch.cuda.empty_cache()\n            seconds_per_epoch = seconds_per_training_step * (training_batches_per_epoch + validation_batches_per_epoch / 3)\n            gpu_cost_rows.append({\n                \"architecture\": architecture_name, \"image_size\": candidate_size,\n                \"seconds_per_training_step\": round(seconds_per_training_step, 4), \"peak_memory_gigabytes\": round(peak_memory_gigabytes, 2),\n                \"fits_in_memory\": fits_in_memory, \"estimated_training_minutes\": round(seconds_per_epoch * expected_epochs_per_architecture / 60, 1),\n            })\n        timing_model = timing_optimizer = None\n        torch.cuda.empty_cache()\n\ngpu_cost_dataframe = pd.DataFrame(gpu_cost_rows)\nif len(gpu_cost_dataframe) > 0:\n    save_results_table(gpu_cost_dataframe, \"03_image_size_study\", \"gpu_cost_per_size.csv\")\n    print(\"Cell 19 finished: training speed and memory measured for every model and size.\")\n    print(gpu_cost_dataframe.pivot(index=\"image_size\", columns=\"architecture\", values=\"estimated_training_minutes\").to_string())\n    print(\"Peak GPU memory at 384 with a batch of 16 (GB):\")\n    print(gpu_cost_dataframe[gpu_cost_dataframe[\"image_size\"] == project_configuration[\"target_image_size\"]][[\"architecture\", \"peak_memory_gigabytes\", \"fits_in_memory\"]].to_string(index=False))\nelse:\n    print(\"Cell 19 finished: timing skipped because there is no GPU.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:38:05.263076Z","iopub.execute_input":"2026-09-26T11:38:05.263614Z","iopub.status.idle":"2026-09-26T11:39:32.913331Z","shell.execute_reply.started":"2026-09-26T11:38:05.263588Z","shell.execute_reply":"2026-09-26T11:39:32.912622Z"}},"outputs":[],"execution_count":null},{"id":"537cc3fb","cell_type":"code","source":"# Cell 20: image size check part 3, put structure kept and training cost side by side and show where 384 sits\nused_image_size = project_configuration[\"target_image_size\"]\nlargest_candidate_size = max(candidate_image_sizes)\nsmallest_candidate_size = min(candidate_image_sizes)\n\nsize_evidence_table = detail_retention_summary[[\"mean\", \"ci95_half_width\"]].rename(columns={\"mean\": \"mean_ssim_kept\"})\nper_image_retention_table = detail_retention_dataframe.pivot(index=\"id_code\", columns=\"image_size\", values=\"ssim_kept\")\nsize_evidence_table[\"gain_from_next_size_up\"] = [\n    (per_image_retention_table[size + 32] - per_image_retention_table[size]).mean() if size + 32 in per_image_retention_table.columns else np.nan\n    for size in size_evidence_table.index\n]\ntiming_is_available = len(gpu_cost_dataframe) > 0\nif timing_is_available:\n    cost_by_size = gpu_cost_dataframe.groupby(\"image_size\").agg(\n        all_three_models_fit_in_memory=(\"fits_in_memory\", \"all\"),\n        estimated_minutes_all_three_models=(\"estimated_training_minutes\", lambda minutes: minutes.sum(min_count=len(minutes))),\n    )\n    size_evidence_table = size_evidence_table.join(cost_by_size)\n    size_evidence_table[\"time_compared_with_size_used\"] = (\n        size_evidence_table[\"estimated_minutes_all_three_models\"] / size_evidence_table.loc[used_image_size, \"estimated_minutes_all_three_models\"])\nsave_results_table(size_evidence_table.round(4).reset_index(), \"03_image_size_study\", \"image_size_evidence_table.csv\")\n\nssim_by_size = size_evidence_table[\"mean_ssim_kept\"]\nimage_size_study_summary = {\n    \"size_used_in_pipeline\": used_image_size,\n    \"photos_measured\": int(len(size_study_sample_dataframe)),\n    f\"ssim_at_{smallest_candidate_size}\": round(float(ssim_by_size.loc[smallest_candidate_size]), 4),\n    f\"ssim_at_{used_image_size}\": round(float(ssim_by_size.loc[used_image_size]), 4),\n    f\"ssim_at_{largest_candidate_size}\": round(float(ssim_by_size.loc[largest_candidate_size]), 4),\n}\nif timing_is_available:\n    image_size_study_summary.update({\n        f\"estimated_minutes_three_models_at_{used_image_size}\": float(size_evidence_table.loc[used_image_size, \"estimated_minutes_all_three_models\"]),\n        f\"time_multiplier_{used_image_size}_to_{largest_candidate_size}\": round(float(size_evidence_table.loc[largest_candidate_size, \"time_compared_with_size_used\"]), 2),\n        f\"all_models_fit_in_memory_at_{largest_candidate_size}\": bool(size_evidence_table.loc[largest_candidate_size, \"all_three_models_fit_in_memory\"]),\n    })\nsave_results_json(image_size_study_summary, \"03_image_size_study\", \"image_size_study_summary.json\")\n\nfigure_object, axis_array = plt.subplots(1, 3 if timing_is_available else 2, figsize=(19 if timing_is_available else 13, 5))\naxis_array[0].plot(candidate_image_sizes, ssim_by_size, marker=\"o\", color=\"steelblue\", label=\"Average over 100 photos\")\naxis_array[0].fill_between(candidate_image_sizes, ssim_by_size - size_evidence_table[\"ci95_half_width\"], ssim_by_size + size_evidence_table[\"ci95_half_width\"],\n                           alpha=0.2, color=\"steelblue\", label=\"95% confidence interval\")\naxis_array[0].set_title(\"Structure Kept (SSIM) vs Image Size\")\naxis_array[0].set_ylabel(\"SSIM against the 1000 px original\")\nfor grade_index, grade_rows in detail_retention_dataframe.groupby(\"diagnosis\"):\n    grade_means = grade_rows.groupby(\"image_size\")[\"ssim_kept\"].mean()\n    axis_array[1].plot(grade_means.index, grade_means.values, marker=\"o\", markersize=3, label=diagnosis_stage_names[grade_index])\naxis_array[1].set_title(\"Structure Kept vs Image Size, per Grade\")\naxis_array[1].set_ylabel(\"SSIM\")\nif timing_is_available:\n    axis_array[2].plot(candidate_image_sizes, size_evidence_table[\"estimated_minutes_all_three_models\"], marker=\"o\", color=\"darkorange\", label=\"Estimated GPU minutes, all 3 models\")\n    for size_that_does_not_fit in [size for size in candidate_image_sizes if not size_evidence_table.loc[size, \"all_three_models_fit_in_memory\"]]:\n        axis_array[2].axvspan(size_that_does_not_fit - 16, size_that_does_not_fit + 16, color=\"red\", alpha=0.1)\n    axis_array[2].set_title(\"Training Cost vs Image Size (red = out of GPU memory)\")\n    axis_array[2].set_ylabel(\"Estimated minutes\")\nfor current_axis in axis_array:\n    current_axis.axvline(used_image_size, color=\"crimson\", linestyle=\":\", label=f\"Size used: {used_image_size}\")\n    current_axis.set_xlabel(\"Image size (pixels per side)\")\n    current_axis.legend(fontsize=8)\nsave_current_figure(\"03_image_size_study\", \"image_size_ssim_vs_cost.png\")\n\nprint(\"Cell 20 finished: image size evidence table and chart saved.\")\nprint(size_evidence_table.round(4).to_string())\nprint(f\"At {used_image_size}, SSIM is {ssim_by_size.loc[used_image_size]:.3f}. At {largest_candidate_size} it is {ssim_by_size.loc[largest_candidate_size]:.3f}.\")\nif timing_is_available:\n    print(f\"Going to {largest_candidate_size} would make training about {size_evidence_table.loc[largest_candidate_size, 'time_compared_with_size_used']:.2f} times slower.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:39:32.914305Z","iopub.execute_input":"2026-09-26T11:39:32.914624Z","iopub.status.idle":"2026-09-26T11:39:34.11547Z","shell.execute_reply.started":"2026-09-26T11:39:32.91459Z","shell.execute_reply":"2026-09-26T11:39:34.114573Z"}},"outputs":[],"execution_count":null},{"id":"e9668139","cell_type":"code","source":"# Cell 21: image size check part 4, show the same small patch of one eye at every size so the blur is visible\npatch_example_row = size_study_sample_dataframe[size_study_sample_dataframe[\"diagnosis\"] == 2].iloc[0]\npatch_reference_rgb = cv2.resize(\n    run_preprocessing_steps(load_image_as_rgb(build_image_path_from_id(patch_example_row[\"id_code\"])), reference_image_size, \"clahe\")[\"enhanced\"],\n    (reference_image_size, reference_image_size), interpolation=cv2.INTER_AREA,\n)\npatch_start, patch_end = int(reference_image_size * 0.375), int(reference_image_size * 0.625)\nnumber_of_patch_columns = 5\nnumber_of_patch_rows = int(np.ceil(len(candidate_image_sizes) / number_of_patch_columns))\nfigure_object, axis_grid = plt.subplots(number_of_patch_rows, number_of_patch_columns, figsize=(18, 3.8 * number_of_patch_rows), squeeze=False)\nfor panel_index, candidate_size in enumerate(candidate_image_sizes):\n    restored_image = cv2.resize(cv2.resize(patch_reference_rgb, (candidate_size, candidate_size), interpolation=cv2.INTER_AREA),\n                                (reference_image_size, reference_image_size), interpolation=cv2.INTER_NEAREST)\n    current_axis = axis_grid[panel_index // number_of_patch_columns, panel_index % number_of_patch_columns]\n    current_axis.imshow(restored_image[patch_start:patch_end, patch_start:patch_end])\n    is_used_size = candidate_size == used_image_size\n    current_axis.set_title(f\"{candidate_size} x {candidate_size}\" + (\"  (used)\" if is_used_size else \"\"), fontsize=10, color=\"crimson\" if is_used_size else \"black\")\n    current_axis.axis(\"off\")\nfor empty_panel_index in range(len(candidate_image_sizes), number_of_patch_rows * number_of_patch_columns):\n    axis_grid[empty_panel_index // number_of_patch_columns, empty_panel_index % number_of_patch_columns].axis(\"off\")\nplt.suptitle(\"Centre of One Moderate DR Photo at Each Candidate Size\", y=1.0)\nsave_current_figure(\"03_image_size_study\", \"image_size_patch_comparison.png\")\n\nprint(\"Cell 21 finished: patch comparison saved. Small red dots and vessel edges blur as the size drops.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:39:34.116598Z","iopub.execute_input":"2026-09-26T11:39:34.11692Z","iopub.status.idle":"2026-09-26T11:39:36.729441Z","shell.execute_reply.started":"2026-09-26T11:39:34.116896Z","shell.execute_reply":"2026-09-26T11:39:36.728685Z"}},"outputs":[],"execution_count":null},{"id":"22ca3c12","cell_type":"code","source":"# Cell 22: clean every photo once with each of the three methods and cache the results, using worker processes\nclass CacheBuilderDataset(Dataset):\n    def __init__(self, image_id_codes, target_image_size, preprocessing_method):\n        self.image_id_codes = list(image_id_codes)\n        self.target_image_size = target_image_size\n        self.preprocessing_method = preprocessing_method\n\n    def __len__(self):\n        return len(self.image_id_codes)\n\n    def __getitem__(self, item_index):\n        preprocess_fundus_image(build_image_path_from_id(self.image_id_codes[item_index]), self.target_image_size, self.preprocessing_method)\n        return 0\n\n\ncache_build_rows = []\nfor method_name in preprocessing_method_names:\n    cache_build_start_time = time.time()\n    cache_builder_loader = DataLoader(\n        CacheBuilderDataset(diabetic_retinopathy_labels_dataframe[\"id_code\"], project_configuration[\"target_image_size\"], method_name),\n        batch_size=32, shuffle=False, num_workers=project_configuration[\"number_of_data_loader_workers\"],\n    )\n    for _ in cache_builder_loader:\n        pass\n    cache_build_rows.append({\n        \"method\": method_name,\n        \"minutes\": round((time.time() - cache_build_start_time) / 60, 1),\n        \"cached_photos\": len(os.listdir(cache_folder_paths[method_name])),\n        \"expected_photos\": len(diabetic_retinopathy_labels_dataframe),\n    })\ncache_build_dataframe = pd.DataFrame(cache_build_rows)\nsave_results_table(cache_build_dataframe, \"02_preprocessing_methods\", \"cache_build_times.csv\")\n\nprint(\"Cell 22 finished: all three caches built.\")\nprint(cache_build_dataframe.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:39:36.730433Z","iopub.execute_input":"2026-09-26T11:39:36.730636Z","iopub.status.idle":"2026-09-26T11:54:48.725703Z","shell.execute_reply.started":"2026-09-26T11:39:36.730616Z","shell.execute_reply":"2026-09-26T11:54:48.724677Z"}},"outputs":[],"execution_count":null},{"id":"dbc8fde0","cell_type":"markdown","source":"## 5. Data Splits and Class Weights\n70% of photos train the models, 15% (validation) are used to compare options and pick the best version of each model, and 15% (test) are locked away until the very end. Every split keeps the same share of each grade.\n\nClass weights are worked out now from the training split only. **Whether to use them is tested** in the class balance experiment.","metadata":{}},{"id":"40e78875","cell_type":"code","source":"# Cell 23: split the photos into training, validation and test sets, keeping the same grade mix in each\nall_image_ids = diabetic_retinopathy_labels_dataframe[\"id_code\"].to_numpy(dtype=object)\nall_diagnosis_labels = diabetic_retinopathy_labels_dataframe[\"diagnosis\"].to_numpy()\n\ntrain_ids, remaining_ids, train_labels, remaining_labels = train_test_split(\n    all_image_ids, all_diagnosis_labels,\n    test_size=project_configuration[\"validation_split_fraction\"] + project_configuration[\"test_split_fraction\"],\n    stratify=all_diagnosis_labels, random_state=RANDOM_SEED,\n)\nrelative_test_fraction = project_configuration[\"test_split_fraction\"] / (project_configuration[\"validation_split_fraction\"] + project_configuration[\"test_split_fraction\"])\nvalidation_ids, test_ids, validation_labels, test_labels = train_test_split(\n    remaining_ids, remaining_labels, test_size=relative_test_fraction, stratify=remaining_labels, random_state=RANDOM_SEED,\n)\ntrain_split_dataframe = pd.DataFrame({\"id_code\": train_ids, \"diagnosis\": train_labels})\nvalidation_split_dataframe = pd.DataFrame({\"id_code\": validation_ids, \"diagnosis\": validation_labels})\ntest_split_dataframe = pd.DataFrame({\"id_code\": test_ids, \"diagnosis\": test_labels})\nsplit_dataframes_by_name = {\"train\": train_split_dataframe, \"validation\": validation_split_dataframe, \"test\": test_split_dataframe}\n\nsave_results_table(pd.concat([split_dataframe.assign(split=split_name) for split_name, split_dataframe in split_dataframes_by_name.items()]),\n                   \"04_splits_and_class_weights\", \"split_membership.csv\")\nsplit_counts_dataframe = pd.DataFrame({split_name: split_dataframe[\"diagnosis\"].value_counts().sort_index() for split_name, split_dataframe in split_dataframes_by_name.items()})\nsplit_counts_dataframe.index = [diagnosis_stage_names[grade] for grade in split_counts_dataframe.index]\nsplit_counts_dataframe.loc[\"Total\"] = split_counts_dataframe.sum()\nsave_results_table(split_counts_dataframe, \"04_splits_and_class_weights\", \"split_counts_per_grade.csv\", keep_index=True)\nsplit_proportions_dataframe = pd.DataFrame({\n    \"full_dataset\": diabetic_retinopathy_labels_dataframe[\"diagnosis\"].value_counts(normalize=True).sort_index(),\n    **{split_name: split_dataframe[\"diagnosis\"].value_counts(normalize=True).sort_index() for split_name, split_dataframe in split_dataframes_by_name.items()},\n}).round(3)\nsplit_proportions_dataframe.index = [diagnosis_stage_names[grade] for grade in split_proportions_dataframe.index]\nsave_results_table(split_proportions_dataframe, \"04_splits_and_class_weights\", \"split_proportions_per_grade.csv\", keep_index=True)\n\nprint(\"Cell 23 finished: data split into training, validation and test sets.\")\nprint(split_counts_dataframe.to_string())\nprint()\nprint(\"Share of each grade in every split (should be almost identical):\")\nprint(split_proportions_dataframe.to_string())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:48.727018Z","iopub.execute_input":"2026-09-26T11:54:48.727435Z","iopub.status.idle":"2026-09-26T11:54:48.763598Z","shell.execute_reply.started":"2026-09-26T11:54:48.727406Z","shell.execute_reply":"2026-09-26T11:54:48.762836Z"}},"outputs":[],"execution_count":null},{"id":"454530ec","cell_type":"code","source":"# Cell 24: work out the class weights from the training split only, so rare grades can count for more if chosen\ncomputed_class_weights = compute_class_weight(\n    class_weight=\"balanced\", classes=np.arange(project_configuration[\"number_of_dr_classes\"]), y=train_split_dataframe[\"diagnosis\"].values,\n)\nclass_weights_tensor = torch.tensor(computed_class_weights, dtype=torch.float32).to(compute_device)\nclass_weights_dataframe = pd.DataFrame({\n    \"grade\": list_of_class_indices, \"stage\": stage_names_in_order,\n    \"training_photos\": train_split_dataframe[\"diagnosis\"].value_counts().sort_index().values,\n    \"loss_weight\": np.round(computed_class_weights, 3),\n})\nsave_results_table(class_weights_dataframe, \"04_splits_and_class_weights\", \"class_weights.csv\")\n\nprint(\"Cell 24 finished: class weights worked out from the training split only.\")\nprint(class_weights_dataframe.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:48.764663Z","iopub.execute_input":"2026-09-26T11:54:48.765064Z","iopub.status.idle":"2026-09-26T11:54:48.795783Z","shell.execute_reply.started":"2026-09-26T11:54:48.765006Z","shell.execute_reply":"2026-09-26T11:54:48.795246Z"}},"outputs":[],"execution_count":null},{"id":"da4db62e","cell_type":"markdown","source":"## 6. Augmentation Policies\nAugmentation makes small random changes to a training photo each time the model sees it, so the model learns the disease and not the exact photo. The changed photos are synthetic, made from real ones. That is acceptable because they are **only used for training**. Validation and test photos are never changed, so every score comes from real photos.\n\nFour policies are compared in the augmentation experiment, from none at all to the strongest:\n\n| Policy | What it adds | Why |\n|---|---|---|\n| none | Nothing | Baseline: is augmentation needed at all? |\n| flips_and_rotation | Left/right and up/down flips (50% each), rotation up to 25 degrees | An eye photo has no fixed orientation, and patients tilt their head |\n| full | The above, plus zoom in 1.0 to 1.1, brightness and contrast up to 15%, saturation up to 5% | Cameras sit at different distances and expose photos differently. Saturation is kept small because colour helps tell lesions apart |\n| full_with_zoom_out | Same as full, but zoom from 0.9 to 1.1 | Tests whether zooming out also helps, or hurts because it shrinks the eye and hides small lesions |","metadata":{}},{"id":"4354c851","cell_type":"code","source":"# Cell 25: define the four augmentation policies and the plain transform for validation and test photos\nimagenet_normalisation_mean = [0.485, 0.456, 0.406]\nimagenet_normalisation_std = [0.229, 0.224, 0.225]\naugmentation_policy_names = [\"none\", \"flips_and_rotation\", \"full\", \"full_with_zoom_out\"]\n\n\ndef build_training_transform(augmentation_policy):\n    transform_steps = [transforms.ToPILImage()]\n    if augmentation_policy in (\"flips_and_rotation\", \"full\", \"full_with_zoom_out\"):\n        transform_steps += [transforms.RandomHorizontalFlip(p=0.5), transforms.RandomVerticalFlip(p=0.5), transforms.RandomRotation(degrees=25)]\n    if augmentation_policy == \"full\":\n        transform_steps.append(transforms.RandomAffine(degrees=0, scale=(1.0, 1.1)))\n    if augmentation_policy == \"full_with_zoom_out\":\n        transform_steps.append(transforms.RandomAffine(degrees=0, scale=(0.9, 1.1)))\n    if augmentation_policy in (\"full\", \"full_with_zoom_out\"):\n        transform_steps.append(transforms.ColorJitter(brightness=0.15, contrast=0.15, saturation=0.05))\n    transform_steps += [transforms.ToTensor(), transforms.Normalize(mean=imagenet_normalisation_mean, std=imagenet_normalisation_std)]\n    return transforms.Compose(transform_steps)\n\n\nevaluation_time_image_transform = transforms.Compose([\n    transforms.ToPILImage(), transforms.ToTensor(), transforms.Normalize(mean=imagenet_normalisation_mean, std=imagenet_normalisation_std),\n])\n\nfor policy_name in augmentation_policy_names:\n    print(f\"  {policy_name}: {[type(step).__name__ for step in build_training_transform(policy_name).transforms]}\")\nprint(\"Cell 25 finished: four augmentation policies and the evaluation transform ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:48.796751Z","iopub.execute_input":"2026-09-26T11:54:48.797141Z","iopub.status.idle":"2026-09-26T11:54:48.805707Z","shell.execute_reply.started":"2026-09-26T11:54:48.797108Z","shell.execute_reply":"2026-09-26T11:54:48.804942Z"}},"outputs":[],"execution_count":null},{"id":"226898fa","cell_type":"code","source":"# Cell 26: AUGMENTATION TESTS, check each policy changes photos the way it should and never changes their size\ntest_photo_for_augmentation = preprocess_fundus_image(\n    build_image_path_from_id(train_split_dataframe[\"id_code\"].iloc[0]), project_configuration[\"target_image_size\"], starting_training_recipe[\"preprocessing_method\"])\nexpected_tensor_shape = (3, project_configuration[\"target_image_size\"], project_configuration[\"target_image_size\"])\naugmentation_test_rows = []\nfor policy_name in augmentation_policy_names:\n    policy_transform = build_training_transform(policy_name)\n    torch.manual_seed(RANDOM_SEED)\n    repeated_outputs = [policy_transform(test_photo_for_augmentation) for _ in range(5)]\n    outputs_differ = any(not torch.equal(repeated_outputs[0], other_output) for other_output in repeated_outputs[1:])\n    augmentation_test_rows.append({\n        \"policy\": policy_name,\n        \"size_unchanged\": all(tuple(output.shape) == expected_tensor_shape for output in repeated_outputs),\n        \"values_are_finite\": all(torch.isfinite(output).all().item() for output in repeated_outputs),\n        \"random_changes_seen\": outputs_differ,\n        \"random_changes_expected\": policy_name != \"none\",\n    })\naugmentation_test_dataframe = pd.DataFrame(augmentation_test_rows)\naugmentation_test_dataframe[\"result\"] = np.where(\n    augmentation_test_dataframe[\"size_unchanged\"] & augmentation_test_dataframe[\"values_are_finite\"]\n    & (augmentation_test_dataframe[\"random_changes_seen\"] == augmentation_test_dataframe[\"random_changes_expected\"]), \"PASS\", \"FAIL\")\n\nalways_flip_transform = transforms.Compose([transforms.ToPILImage(), transforms.RandomHorizontalFlip(p=1.0), transforms.PILToTensor()])\nflip_is_exact_mirror = torch.equal(always_flip_transform(test_photo_for_augmentation), torch.from_numpy(np.ascontiguousarray(test_photo_for_augmentation[:, ::-1])).permute(2, 0, 1))\nevaluation_is_repeatable = torch.equal(evaluation_time_image_transform(test_photo_for_augmentation), evaluation_time_image_transform(test_photo_for_augmentation))\nextra_augmentation_checks = pd.DataFrame([\n    {\"check\": \"A horizontal flip gives an exact mirror image\", \"result\": \"PASS\" if flip_is_exact_mirror else \"FAIL\"},\n    {\"check\": \"Validation and test transform gives the same output every time\", \"result\": \"PASS\" if evaluation_is_repeatable else \"FAIL\"},\n])\nsave_results_table(augmentation_test_dataframe, \"05_augmentation_policies\", \"augmentation_tests_per_policy.csv\")\nsave_results_table(extra_augmentation_checks, \"05_augmentation_policies\", \"augmentation_extra_checks.csv\")\n\nprint(\"Cell 26 finished: augmentation tests complete.\")\nprint(augmentation_test_dataframe.to_string(index=False))\nprint(extra_augmentation_checks.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:48.806591Z","iopub.execute_input":"2026-09-26T11:54:48.806889Z","iopub.status.idle":"2026-09-26T11:54:49.018Z","shell.execute_reply.started":"2026-09-26T11:54:48.806868Z","shell.execute_reply":"2026-09-26T11:54:49.016842Z"}},"outputs":[],"execution_count":null},{"id":"3ddfc58f","cell_type":"code","source":"# Cell 27: show five random training versions of the same photo under each policy\nnumber_of_augmented_previews = 5\nnormalisation_mean_tensor = torch.tensor(imagenet_normalisation_mean).view(3, 1, 1)\nnormalisation_std_tensor = torch.tensor(imagenet_normalisation_std).view(3, 1, 1)\nfigure_object, axis_grid = plt.subplots(len(augmentation_policy_names), number_of_augmented_previews + 1, figsize=(18, 3.2 * len(augmentation_policy_names)))\ntorch.manual_seed(RANDOM_SEED)\nfor row_index, policy_name in enumerate(augmentation_policy_names):\n    policy_transform = build_training_transform(policy_name)\n    axis_grid[row_index, 0].imshow(test_photo_for_augmentation)\n    axis_grid[row_index, 0].set_title(f\"{policy_name}: cleaned photo\", fontsize=9)\n    for preview_index in range(number_of_augmented_previews):\n        displayable_image = (policy_transform(test_photo_for_augmentation) * normalisation_std_tensor + normalisation_mean_tensor).permute(1, 2, 0).clamp(0, 1).numpy()\n        axis_grid[row_index, preview_index + 1].imshow(displayable_image)\n        axis_grid[row_index, preview_index + 1].set_title(f\"Random version {preview_index + 1}\", fontsize=9)\n    for current_axis in axis_grid[row_index]:\n        current_axis.axis(\"off\")\nplt.suptitle(\"What Each Augmentation Policy Does to One Training Photo\", y=1.0)\nsave_current_figure(\"05_augmentation_policies\", \"augmentation_policies_preview.png\")\n\nprint(\"Cell 27 finished: augmentation preview saved for all four policies.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:49.019322Z","iopub.execute_input":"2026-09-26T11:54:49.019551Z","iopub.status.idle":"2026-09-26T11:54:55.250009Z","shell.execute_reply.started":"2026-09-26T11:54:49.019529Z","shell.execute_reply":"2026-09-26T11:54:55.249238Z"}},"outputs":[],"execution_count":null},{"id":"db1b73a2","cell_type":"markdown","source":"## 7. Feeding Photos to the Model\nA **Dataset** reads one cleaned photo and its grade. A **DataLoader** groups them into batches of 16. Both the preprocessing method and the augmentation policy are settings, so every experiment reuses the same code with only those settings changed.","metadata":{}},{"id":"e75a238c","cell_type":"code","source":"# Cell 28: define how one photo and its grade are read, cleaned and turned into numbers\nclass DiabeticRetinopathyFundusImageDataset(Dataset):\n    def __init__(self, labels_dataframe, target_image_size, image_transform, preprocessing_method):\n        self.labels_dataframe = labels_dataframe.reset_index(drop=True)\n        self.target_image_size = target_image_size\n        self.image_transform = image_transform\n        self.preprocessing_method = preprocessing_method\n\n    def __len__(self):\n        return len(self.labels_dataframe)\n\n    def __getitem__(self, item_index):\n        row = self.labels_dataframe.iloc[item_index]\n        preprocessed_image_array = preprocess_fundus_image(build_image_path_from_id(row[\"id_code\"]), self.target_image_size, self.preprocessing_method)\n        return self.image_transform(preprocessed_image_array), torch.tensor(row[\"diagnosis\"], dtype=torch.long)\n\n\nprint(\"Cell 28 finished: dataset class ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:55.251294Z","iopub.execute_input":"2026-09-26T11:54:55.251593Z","iopub.status.idle":"2026-09-26T11:54:55.260447Z","shell.execute_reply.started":"2026-09-26T11:54:55.251563Z","shell.execute_reply":"2026-09-26T11:54:55.259706Z"}},"outputs":[],"execution_count":null},{"id":"3ca26f81","cell_type":"code","source":"# Cell 29: build training, validation and test loaders for any method and policy, then pull one real batch through as a check\ndef build_data_loaders(preprocessing_method, augmentation_policy):\n    split_settings = {\n        \"train\": (train_split_dataframe, build_training_transform(augmentation_policy), True),\n        \"validation\": (validation_split_dataframe, evaluation_time_image_transform, False),\n        \"test\": (test_split_dataframe, evaluation_time_image_transform, False),\n    }\n    return {\n        split_name: DataLoader(\n            DiabeticRetinopathyFundusImageDataset(split_dataframe, project_configuration[\"target_image_size\"], image_transform, preprocessing_method),\n            batch_size=project_configuration[\"batch_size\"], shuffle=shuffle_each_epoch,\n            num_workers=project_configuration[\"number_of_data_loader_workers\"], pin_memory=is_gpu_available,\n        )\n        for split_name, (split_dataframe, image_transform, shuffle_each_epoch) in split_settings.items()\n    }\n\n\nsanity_check_loaders = build_data_loaders(starting_training_recipe[\"preprocessing_method\"], starting_training_recipe[\"augmentation_policy\"])\nsanity_check_images, sanity_check_labels = next(iter(sanity_check_loaders[\"train\"]))\nexpected_batch_shape = (project_configuration[\"batch_size\"], 3, project_configuration[\"target_image_size\"], project_configuration[\"target_image_size\"])\n\nprint(\"Cell 29 finished: loader builder ready and one batch checked.\")\nprint(\"Batches per epoch, training:\", len(sanity_check_loaders[\"train\"]), \"| validation:\", len(sanity_check_loaders[\"validation\"]), \"| test:\", len(sanity_check_loaders[\"test\"]))\nprint(\"Batch image shape:\", tuple(sanity_check_images.shape), \"| shape check:\", \"PASS\" if tuple(sanity_check_images.shape) == expected_batch_shape else \"FAIL\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:55.261502Z","iopub.execute_input":"2026-09-26T11:54:55.262176Z","iopub.status.idle":"2026-09-26T11:54:55.992563Z","shell.execute_reply.started":"2026-09-26T11:54:55.262143Z","shell.execute_reply":"2026-09-26T11:54:55.991577Z"}},"outputs":[],"execution_count":null},{"id":"ddc6693b","cell_type":"markdown","source":"## 8. Models, Freezing and the Training Engine\n**Transfer learning** starts from a model that already learned edges, shapes and textures from over a million everyday photos (ImageNet), then teaches it to grade eyes. Only the last layer is swapped so it gives 5 grades.\n\nThree models are used in the final stage:\n- **EfficientNet-B3:** accurate for its size and widely used for eye photos (Bhimavarapu & Battineni, 2022; Acharya et al., 2023).\n- **EfficientNet-B0:** a smaller, faster version. It also runs every experiment, because it trains fastest.\n- **ResNet50:** built differently, with shortcut connections, so it tends to make different mistakes, which is what makes combining the three worthwhile.\n\n**Freezing** means locking layers so they do not change during training. The freezing experiment tests four strategies:\n\n| Strategy | What learns | Idea |\n|---|---|---|\n| head_only | Only the new last layer | Safest against overfitting, but the model cannot adapt its features to eyes |\n| two_phase_partial_unfreeze | Last layer first, then the last half of the model too | Early layers keep general edge detectors, later layers adapt to eyes |\n| two_phase_full_unfreeze | Last layer first, then everything | The usual transfer learning recipe, and the starting point for the experiments |\n| full_fine_tune | Everything from the start | Fastest to adapt, but the random new last layer can damage the pretrained weights early on |","metadata":{}},{"id":"c9261d7e","cell_type":"code","source":"# Cell 30: compare how big the three pretrained models are before anything is changed\narchitecture_names_to_compare = [\"efficientnet_b3\", \"efficientnet_b0\", \"resnet50\"]\npretrained_model_builders = {\n    \"efficientnet_b3\": lambda: efficientnet_b3(weights=EfficientNet_B3_Weights.IMAGENET1K_V1),\n    \"efficientnet_b0\": lambda: efficientnet_b0(weights=EfficientNet_B0_Weights.IMAGENET1K_V1),\n    \"resnet50\": lambda: resnet50(weights=ResNet50_Weights.IMAGENET1K_V2),\n}\nparameter_count_dataframe = pd.DataFrame([\n    {\"architecture\": architecture_name, \"total_parameters_with_imagenet_head\": sum(parameter.numel() for parameter in pretrained_model_builders[architecture_name]().parameters())}\n    for architecture_name in architecture_names_to_compare\n])\nsave_results_table(parameter_count_dataframe, \"06_models_and_freezing\", \"parameter_counts_before_changes.csv\")\n\nprint(\"Cell 30 finished: model sizes compared (original ImageNet versions, 1000 outputs).\")\nprint(parameter_count_dataframe.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:55.993968Z","iopub.execute_input":"2026-09-26T11:54:55.994331Z","iopub.status.idle":"2026-09-26T11:54:57.987325Z","shell.execute_reply.started":"2026-09-26T11:54:55.994301Z","shell.execute_reply":"2026-09-26T11:54:57.986545Z"}},"outputs":[],"execution_count":null},{"id":"86dd4524","cell_type":"code","source":"# Cell 31: the model factory, plus the helpers that freeze or unfreeze parts of any model\ndef build_transfer_learning_model(architecture_name, number_of_output_classes, head_dropout=None):\n    model = pretrained_model_builders[architecture_name]()\n    if architecture_name in (\"efficientnet_b3\", \"efficientnet_b0\"):\n        if head_dropout is not None:\n            model.classifier[0].p = head_dropout\n        model.classifier[1] = neural_network_layers.Linear(model.classifier[1].in_features, number_of_output_classes)\n    elif architecture_name == \"resnet50\":\n        model.fc = neural_network_layers.Sequential(\n            neural_network_layers.Dropout(head_dropout if head_dropout is not None else 0.0),\n            neural_network_layers.Linear(model.fc.in_features, number_of_output_classes),\n        )\n    else:\n        raise ValueError(f\"Unknown architecture name: {architecture_name}\")\n    return model\n\n\ndef get_backbone_blocks(model, architecture_name):\n    if architecture_name == \"resnet50\":\n        return [model.conv1, model.bn1, model.layer1, model.layer2, model.layer3, model.layer4]\n    return list(model.features.children())\n\n\ndef get_classifier_head(model, architecture_name):\n    return model.fc if architecture_name == \"resnet50\" else model.classifier\n\n\ndef set_trainable_part(model, architecture_name, trainable_part):\n    for parameter in model.parameters():\n        parameter.requires_grad = trainable_part == \"everything\"\n    if trainable_part != \"everything\":\n        for parameter in get_classifier_head(model, architecture_name).parameters():\n            parameter.requires_grad = True\n    if trainable_part == \"last_half\":\n        backbone_blocks = get_backbone_blocks(model, architecture_name)\n        for backbone_block in backbone_blocks[len(backbone_blocks) // 2:]:\n            for parameter in backbone_block.parameters():\n                parameter.requires_grad = True\n    return sum(parameter.numel() for parameter in model.parameters() if parameter.requires_grad)\n\n\nprint(\"Cell 31 finished: model factory and freezing helpers ready.\")\nprint(\"Trainable parts: head_only, last_half (the head plus the last half of the backbone), everything.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:57.988348Z","iopub.execute_input":"2026-09-26T11:54:57.98879Z","iopub.status.idle":"2026-09-26T11:54:57.997803Z","shell.execute_reply.started":"2026-09-26T11:54:57.988756Z","shell.execute_reply":"2026-09-26T11:54:57.997217Z"}},"outputs":[],"execution_count":null},{"id":"b6338732","cell_type":"code","source":"# Cell 32: SMOKE TESTS, check every model's output, pretrained weights, ability to learn, and that freezing really freezes\nprobe_batch_images, probe_batch_labels = sanity_check_images[:8], sanity_check_labels[:8]\nmodel_smoke_test_rows = []\nfreezing_rows = []\nfor architecture_name in architecture_names_to_compare:\n    probe_model = build_transfer_learning_model(architecture_name, project_configuration[\"number_of_dr_classes\"]).to(compute_device)\n    probe_model.eval()\n    with torch.no_grad():\n        output_shape_passed = tuple(probe_model(probe_batch_images.to(compute_device)).shape) == (len(probe_batch_images), project_configuration[\"number_of_dr_classes\"])\n    first_backbone_parameter = next(parameter for name, parameter in probe_model.named_parameters() if not name.startswith((\"fc\", \"classifier\")))\n    pretrained_weights_passed = bool(first_backbone_parameter.std().item() > 1e-6)\n\n    total_parameters = sum(parameter.numel() for parameter in probe_model.parameters())\n    head_parameters = sum(parameter.numel() for parameter in get_classifier_head(probe_model, architecture_name).parameters())\n    trainable_counts = {part: set_trainable_part(probe_model, architecture_name, part) for part in [\"head_only\", \"last_half\", \"everything\"]}\n    freezing_passed = trainable_counts[\"head_only\"] == head_parameters and head_parameters < trainable_counts[\"last_half\"] < total_parameters and trainable_counts[\"everything\"] == total_parameters\n    for part, trainable_count in trainable_counts.items():\n        freezing_rows.append({\"architecture\": architecture_name, \"trainable_part\": part, \"trainable_parameters\": trainable_count,\n                              \"percent_of_model\": round(trainable_count / total_parameters * 100, 2)})\n\n    probe_model.train()\n    learning_optimizer = optimizers.Adam(probe_model.parameters(), lr=1e-3)\n    learning_loss_function = neural_network_layers.CrossEntropyLoss()\n    loss_values = []\n    for _ in range(25):\n        learning_optimizer.zero_grad()\n        step_loss = learning_loss_function(probe_model(probe_batch_images.to(compute_device)), probe_batch_labels.to(compute_device))\n        step_loss.backward()\n        learning_optimizer.step()\n        loss_values.append(step_loss.item())\n    loss_reduction_fraction = (loss_values[0] - loss_values[-1]) / loss_values[0]\n    model_smoke_test_rows.append({\n        \"architecture\": architecture_name,\n        \"output_shape_test\": \"PASS\" if output_shape_passed else \"FAIL\",\n        \"pretrained_weight_test\": \"PASS\" if pretrained_weights_passed else \"FAIL\",\n        \"freezing_test\": \"PASS\" if freezing_passed else \"FAIL\",\n        \"single_batch_learning_test\": \"PASS\" if loss_reduction_fraction > 0.5 else \"FAIL\",\n        \"loss_reduction_percent\": round(loss_reduction_fraction * 100, 1),\n    })\n    del probe_model, learning_optimizer\n    if is_gpu_available:\n        torch.cuda.empty_cache()\n\nmodel_smoke_test_dataframe = pd.DataFrame(model_smoke_test_rows)\nfreezing_dataframe = pd.DataFrame(freezing_rows)\nsave_results_table(model_smoke_test_dataframe, \"06_models_and_freezing\", \"model_smoke_tests.csv\")\nsave_results_table(freezing_dataframe, \"06_models_and_freezing\", \"trainable_parameters_per_freezing_option.csv\")\ntest_columns = [\"output_shape_test\", \"pretrained_weight_test\", \"freezing_test\", \"single_batch_learning_test\"]\nall_model_smoke_tests_passed = bool((model_smoke_test_dataframe[test_columns] == \"PASS\").all().all())\n\nprint(\"Cell 32 finished: model smoke tests complete.\")\nprint(model_smoke_test_dataframe.to_string(index=False))\nprint()\nprint(\"How much of each model learns under each freezing option:\")\nprint(freezing_dataframe.to_string(index=False))\nprint(\"All model checks passed, safe to start training.\" if all_model_smoke_tests_passed else \"At least one model check failed. Stop and fix it before training.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:54:57.998901Z","iopub.execute_input":"2026-09-26T11:54:57.999233Z","iopub.status.idle":"2026-09-26T11:55:17.118616Z","shell.execute_reply.started":"2026-09-26T11:54:57.9992Z","shell.execute_reply":"2026-09-26T11:55:17.117715Z"}},"outputs":[],"execution_count":null},{"id":"6d2fa489","cell_type":"code","source":"# Cell 33: the loss function builder and the helper that trains for one epoch and tracks training QWK\ndef build_loss_function(training_recipe):\n    return neural_network_layers.CrossEntropyLoss(\n        weight=class_weights_tensor if training_recipe[\"use_class_weights\"] else None, label_smoothing=training_recipe[\"label_smoothing\"],\n    )\n\n\ndef run_one_training_epoch(model, data_loader, optimizer, loss_function, device):\n    model.train()\n    running_loss_total = 0.0\n    all_true_labels = []\n    all_predicted_labels = []\n    for batch_images, batch_labels in data_loader:\n        batch_images = batch_images.to(device, non_blocking=True)\n        batch_labels = batch_labels.to(device, non_blocking=True)\n        optimizer.zero_grad()\n        predicted_logits = model(batch_images)\n        batch_loss = loss_function(predicted_logits, batch_labels)\n        batch_loss.backward()\n        optimizer.step()\n        running_loss_total += batch_loss.item() * batch_images.size(0)\n        all_true_labels.extend(batch_labels.cpu().numpy().tolist())\n        all_predicted_labels.extend(predicted_logits.argmax(dim=1).cpu().numpy().tolist())\n    return (\n        running_loss_total / len(all_true_labels),\n        accuracy_score(all_true_labels, all_predicted_labels),\n        cohen_kappa_score(all_true_labels, all_predicted_labels, weights=\"quadratic\"),\n    )\n\n\nprint(\"Cell 33 finished: loss builder and one-epoch helper ready.\")\nprint(\"Training QWK is measured on the augmented batches as the model trains, so it costs no extra time.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.119767Z","iopub.execute_input":"2026-09-26T11:55:17.120148Z","iopub.status.idle":"2026-09-26T11:55:17.127801Z","shell.execute_reply.started":"2026-09-26T11:55:17.120124Z","shell.execute_reply":"2026-09-26T11:55:17.127153Z"}},"outputs":[],"execution_count":null},{"id":"77f5b94c","cell_type":"code","source":"# Cell 34: the helper that scores a model on validation or test photos without changing it\ndef run_evaluation_pass(model, data_loader, loss_function, device):\n    model.eval()\n    running_loss_total = 0.0\n    all_true_labels = []\n    all_predicted_labels = []\n    all_predicted_probabilities = []\n    with torch.no_grad():\n        for batch_images, batch_labels in data_loader:\n            batch_images = batch_images.to(device, non_blocking=True)\n            batch_labels = batch_labels.to(device, non_blocking=True)\n            predicted_logits = model(batch_images)\n            running_loss_total += loss_function(predicted_logits, batch_labels).item() * batch_images.size(0)\n            all_true_labels.extend(batch_labels.cpu().numpy().tolist())\n            all_predicted_labels.extend(predicted_logits.argmax(dim=1).cpu().numpy().tolist())\n            all_predicted_probabilities.extend(torch.softmax(predicted_logits, dim=1).cpu().numpy().tolist())\n    return (\n        running_loss_total / len(all_true_labels),\n        accuracy_score(all_true_labels, all_predicted_labels),\n        cohen_kappa_score(all_true_labels, all_predicted_labels, weights=\"quadratic\"),\n        np.array(all_true_labels), np.array(all_predicted_labels), np.array(all_predicted_probabilities),\n    )\n\n\nprint(\"Cell 34 finished: evaluation helper ready. It also returns the probability of every grade, which the ensemble needs.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.128749Z","iopub.execute_input":"2026-09-26T11:55:17.129073Z","iopub.status.idle":"2026-09-26T11:55:17.219739Z","shell.execute_reply.started":"2026-09-26T11:55:17.129015Z","shell.execute_reply":"2026-09-26T11:55:17.219099Z"}},"outputs":[],"execution_count":null},{"id":"35326417","cell_type":"code","source":"# Cell 35: early stopping on validation QWK, best-checkpoint saving, and a helper that records each epoch\nclass EarlyStoppingOnValidationQWK:\n    def __init__(self, patience, checkpoint_file_path):\n        self.patience = patience\n        self.checkpoint_file_path = checkpoint_file_path\n        self.best_validation_qwk_so_far = -1.0\n        self.epochs_since_last_improvement = 0\n        self.should_stop_training = False\n\n    def reset_patience(self):\n        self.epochs_since_last_improvement = 0\n        self.should_stop_training = False\n\n    def check_after_epoch(self, current_validation_qwk, model):\n        if current_validation_qwk > self.best_validation_qwk_so_far:\n            self.best_validation_qwk_so_far = current_validation_qwk\n            self.epochs_since_last_improvement = 0\n            torch.save(model.state_dict(), self.checkpoint_file_path)\n            return True\n        self.epochs_since_last_improvement += 1\n        if self.epochs_since_last_improvement >= self.patience:\n            self.should_stop_training = True\n        return False\n\n\nhistory_column_names = [\"epoch\", \"phase\", \"learning_rate\", \"train_loss\", \"train_accuracy\", \"train_qwk\", \"val_loss\", \"val_accuracy\", \"val_qwk\"]\n\n\ndef create_empty_history():\n    return {column_name: [] for column_name in history_column_names}\n\n\nprint(\"Cell 35 finished: early stopping, checkpoint saving and history recording ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.220917Z","iopub.execute_input":"2026-09-26T11:55:17.221185Z","iopub.status.idle":"2026-09-26T11:55:17.242582Z","shell.execute_reply.started":"2026-09-26T11:55:17.221162Z","shell.execute_reply":"2026-09-26T11:55:17.241974Z"}},"outputs":[],"execution_count":null},{"id":"ee7c8445","cell_type":"code","source":"# Cell 36: run one training phase for a set number of epochs, with an optional learning rate schedule and early stopping\ndef run_training_phase(model, data_loaders, loss_function, optimizer, learning_rate_scheduler, number_of_epochs, first_epoch_number,\n                       phase_name, early_stopping_tracker, history_records, run_label, allow_early_stop, print_every_epoch):\n    for epoch_offset in range(number_of_epochs):\n        epoch_number = first_epoch_number + epoch_offset\n        current_learning_rate = optimizer.param_groups[0][\"lr\"]\n        train_loss, train_accuracy, train_qwk = run_one_training_epoch(model, data_loaders[\"train\"], optimizer, loss_function, compute_device)\n        val_loss, val_accuracy, val_qwk, _, _, _ = run_evaluation_pass(model, data_loaders[\"validation\"], loss_function, compute_device)\n        if isinstance(learning_rate_scheduler, optimizers.lr_scheduler.ReduceLROnPlateau):\n            learning_rate_scheduler.step(val_qwk)\n        elif learning_rate_scheduler is not None:\n            learning_rate_scheduler.step()\n        improved = early_stopping_tracker.check_after_epoch(val_qwk, model)\n        for column_name, column_value in zip(history_column_names, [epoch_number, phase_name, current_learning_rate, train_loss, train_accuracy, train_qwk, val_loss, val_accuracy, val_qwk]):\n            history_records[column_name].append(column_value)\n        if print_every_epoch:\n            print(f\"[{run_label}][{phase_name}][epoch {epoch_number}] lr={current_learning_rate:.1e} \"\n                  f\"train loss={train_loss:.4f} acc={train_accuracy:.4f} qwk={train_qwk:.4f} | \"\n                  f\"val loss={val_loss:.4f} acc={val_accuracy:.4f} qwk={val_qwk:.4f} {'(new best, saved)' if improved else ''}\")\n        if allow_early_stop and early_stopping_tracker.should_stop_training:\n            print(f\"[{run_label}] early stopping: validation QWK has not improved for {early_stopping_tracker.patience} epochs.\")\n            break\n\n\nprint(\"Cell 36 finished: training phase helper ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.243433Z","iopub.execute_input":"2026-09-26T11:55:17.243732Z","iopub.status.idle":"2026-09-26T11:55:17.271195Z","shell.execute_reply.started":"2026-09-26T11:55:17.243711Z","shell.execute_reply":"2026-09-26T11:55:17.270603Z"}},"outputs":[],"execution_count":null},{"id":"726d45f9","cell_type":"code","source":"# Cell 37: one function that trains any model with any recipe, used for every experiment and every final model\nscreening_schedule = {\n    \"phase_one_epochs\": project_configuration[\"screening_phase_one_epochs\"],\n    \"phase_two_max_epochs\": project_configuration[\"screening_phase_two_max_epochs\"],\n    \"patience\": project_configuration[\"screening_early_stopping_patience\"],\n}\nfinal_schedule = {\n    \"phase_one_epochs\": project_configuration[\"final_phase_one_epochs\"],\n    \"phase_two_max_epochs\": project_configuration[\"final_phase_two_max_epochs\"],\n    \"patience\": project_configuration[\"final_early_stopping_patience\"],\n}\n\n\ndef train_with_recipe(architecture_name, training_recipe, training_schedule, run_label, checkpoint_file_path, seed_value, print_every_epoch):\n    set_global_random_seed(seed_value)\n    run_start_time = time.time()\n    data_loaders = build_data_loaders(training_recipe[\"preprocessing_method\"], training_recipe[\"augmentation_policy\"])\n    model = build_transfer_learning_model(architecture_name, project_configuration[\"number_of_dr_classes\"], training_recipe[\"head_dropout\"]).to(compute_device)\n    loss_function = build_loss_function(training_recipe)\n    early_stopping_tracker = EarlyStoppingOnValidationQWK(training_schedule[\"patience\"], checkpoint_file_path)\n    history_records = create_empty_history()\n    total_epochs = training_schedule[\"phase_one_epochs\"] + training_schedule[\"phase_two_max_epochs\"]\n    strategy = training_recipe[\"transfer_strategy\"]\n\n    def make_optimizer(learning_rate):\n        return optimizers.AdamW([parameter for parameter in model.parameters() if parameter.requires_grad], lr=learning_rate, weight_decay=training_recipe[\"weight_decay\"])\n\n    def make_scheduler(optimizer, number_of_epochs):\n        if training_recipe[\"phase_two_schedule\"] == \"cosine\":\n            return optimizers.lr_scheduler.CosineAnnealingLR(optimizer, T_max=max(number_of_epochs, 1))\n        return optimizers.lr_scheduler.ReduceLROnPlateau(optimizer, mode=\"max\", factor=0.5, patience=2)\n\n    common_arguments = dict(model=model, data_loaders=data_loaders, loss_function=loss_function, early_stopping_tracker=early_stopping_tracker,\n                            history_records=history_records, run_label=run_label, print_every_epoch=print_every_epoch)\n    if strategy in (\"head_only\", \"full_fine_tune\"):\n        trainable_parameter_count = set_trainable_part(model, architecture_name, \"head_only\" if strategy == \"head_only\" else \"everything\")\n        learning_rate = training_recipe[\"phase_one_learning_rate\"] if strategy == \"head_only\" else training_recipe[\"phase_two_learning_rate\"]\n        optimizer = make_optimizer(learning_rate)\n        run_training_phase(optimizer=optimizer, learning_rate_scheduler=make_scheduler(optimizer, total_epochs), number_of_epochs=total_epochs,\n                           first_epoch_number=1, phase_name=strategy, allow_early_stop=True, **common_arguments)\n    elif strategy in (\"two_phase_full_unfreeze\", \"two_phase_partial_unfreeze\"):\n        set_trainable_part(model, architecture_name, \"head_only\")\n        run_training_phase(optimizer=make_optimizer(training_recipe[\"phase_one_learning_rate\"]), learning_rate_scheduler=None,\n                           number_of_epochs=training_schedule[\"phase_one_epochs\"], first_epoch_number=1, phase_name=\"phase_1_head_only\",\n                           allow_early_stop=False, **common_arguments)\n        early_stopping_tracker.reset_patience()\n        trainable_parameter_count = set_trainable_part(model, architecture_name, \"everything\" if strategy == \"two_phase_full_unfreeze\" else \"last_half\")\n        phase_two_optimizer = make_optimizer(training_recipe[\"phase_two_learning_rate\"])\n        run_training_phase(optimizer=phase_two_optimizer, learning_rate_scheduler=make_scheduler(phase_two_optimizer, training_schedule[\"phase_two_max_epochs\"]),\n                           number_of_epochs=training_schedule[\"phase_two_max_epochs\"], first_epoch_number=training_schedule[\"phase_one_epochs\"] + 1,\n                           phase_name=\"phase_2_fine_tuning\", allow_early_stop=True, **common_arguments)\n    else:\n        raise ValueError(f\"Unknown transfer strategy: {strategy}\")\n\n    history_dataframe = pd.DataFrame(history_records)\n    best_row = history_dataframe.loc[history_dataframe[\"val_qwk\"].idxmax()]\n    training_result = {\n        \"architecture_name\": architecture_name,\n        \"run_label\": run_label,\n        \"seed\": seed_value,\n        \"recipe\": copy.deepcopy(training_recipe),\n        \"checkpoint_path\": checkpoint_file_path,\n        \"history_dataframe\": history_dataframe,\n        \"best_validation_qwk\": float(best_row[\"val_qwk\"]),\n        \"best_epoch\": int(best_row[\"epoch\"]),\n        \"train_qwk_at_best_epoch\": float(best_row[\"train_qwk\"]),\n        \"train_accuracy_at_best_epoch\": float(best_row[\"train_accuracy\"]),\n        \"val_accuracy_at_best_epoch\": float(best_row[\"val_accuracy\"]),\n        \"val_loss_at_best_epoch\": float(best_row[\"val_loss\"]),\n        \"epoch_with_lowest_val_loss\": int(history_dataframe.loc[history_dataframe[\"val_loss\"].idxmin(), \"epoch\"]),\n        \"epochs_trained\": int(len(history_dataframe)),\n        \"trainable_parameters\": int(trainable_parameter_count),\n        \"total_parameters\": int(sum(parameter.numel() for parameter in model.parameters())),\n        \"training_duration_minutes\": (time.time() - run_start_time) / 60,\n    }\n    del model, data_loaders\n    if is_gpu_available:\n        torch.cuda.empty_cache()\n    return training_result\n\n\nprint(\"Cell 37 finished: recipe-driven training function ready.\")\nprint(\"It supports four freezing strategies, two learning rate schedules, class weights, label smoothing, weight decay and head dropout.\")\nprint(\"The seed is reset at the start of every run, so runs with the same seed start from the same point.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.271977Z","iopub.execute_input":"2026-09-26T11:55:17.272247Z","iopub.status.idle":"2026-09-26T11:55:17.300569Z","shell.execute_reply.started":"2026-09-26T11:55:17.272208Z","shell.execute_reply":"2026-09-26T11:55:17.299812Z"}},"outputs":[],"execution_count":null},{"id":"3356eeb2","cell_type":"markdown","source":"## 9. Experiments\nEach experiment changes **one** part of the recipe, keeps everything else the same, and trains EfficientNet-B0 with the short schedule. Every option runs **twice, with two different seeds**, because a difference smaller than the seed-to-seed wobble is just luck.\n\n**How a winner is picked (fixed before any results):**\n1. The option with the highest **average validation QWK** across both seeds leads.\n2. If a simpler or safer option is within **0.005 QWK** of the leader, the simpler or safer option wins, because a change that does not clearly help is not worth the extra complexity.\n3. The winner replaces that part of the recipe, and the next experiment starts from there.\n\n**The test set is never used in this section.** Each run also records validation macro F1, recall per grade, and the gap between training QWK and validation QWK, which shows overfitting.\n\n**Honest limit:** the experiments use B0 only, to fit the GPU budget. The final recipe is assumed to carry over to B3 and ResNet50. Choosing on the validation set many times can also slightly flatter validation scores, which is exactly why the test set stays locked until section 12.","metadata":{}},{"id":"b98546d4","cell_type":"code","source":"# Cell 38: the experiment runner, which trains every option with every seed, summarises, charts and picks a winner\nexperiment_run_cache = {}\nexperiment_decisions = []\ncurrent_training_recipe = copy.deepcopy(starting_training_recipe)\n\n\ndef run_one_screening_experiment(training_recipe, seed_value, run_label):\n    cache_key = json_module.dumps(training_recipe, sort_keys=True) + f\"|seed={seed_value}\"\n    if cache_key in experiment_run_cache:\n        print(f\"  {run_label}: identical run already done in a previous experiment, reusing its result\")\n        return experiment_run_cache[cache_key]\n    training_result = train_with_recipe(project_configuration[\"screening_architecture\"], training_recipe, screening_schedule, run_label,\n                                        experiment_checkpoint_path, seed_value, print_every_epoch=False)\n    evaluation_model = build_transfer_learning_model(project_configuration[\"screening_architecture\"], project_configuration[\"number_of_dr_classes\"],\n                                                     training_recipe[\"head_dropout\"]).to(compute_device)\n    evaluation_model.load_state_dict(torch.load(experiment_checkpoint_path, map_location=compute_device))\n    validation_loader = build_data_loaders(training_recipe[\"preprocessing_method\"], \"none\")[\"validation\"]\n    _, _, _, true_labels, predicted_labels, _ = run_evaluation_pass(evaluation_model, validation_loader, neural_network_layers.CrossEntropyLoss(), compute_device)\n    del evaluation_model\n    _, recall_per_grade, _, _ = precision_recall_fscore_support(true_labels, predicted_labels, labels=list_of_class_indices, zero_division=0)\n    experiment_row = {\n        \"best_val_qwk\": training_result[\"best_validation_qwk\"],\n        \"best_epoch\": training_result[\"best_epoch\"],\n        \"train_qwk_at_best_epoch\": training_result[\"train_qwk_at_best_epoch\"],\n        \"train_minus_val_qwk_gap\": training_result[\"train_qwk_at_best_epoch\"] - training_result[\"best_validation_qwk\"],\n        \"val_macro_f1\": precision_recall_fscore_support(true_labels, predicted_labels, average=\"macro\", zero_division=0)[2],\n        \"val_accuracy\": training_result[\"val_accuracy_at_best_epoch\"],\n        \"trainable_parameters\": training_result[\"trainable_parameters\"],\n        \"minutes\": training_result[\"training_duration_minutes\"],\n        **{f\"val_recall_{stage_name.lower().replace(' ', '_')}\": recall_per_grade[grade] for grade, stage_name in enumerate(stage_names_in_order)},\n    }\n    print(f\"  {run_label}: best val QWK {experiment_row['best_val_qwk']:.4f} at epoch {experiment_row['best_epoch']}, \"\n          f\"train-val QWK gap {experiment_row['train_minus_val_qwk_gap']:+.4f}, {experiment_row['minutes']:.1f} min\")\n    experiment_run_cache[cache_key] = (experiment_row, training_result[\"history_dataframe\"])\n    return experiment_run_cache[cache_key]\n\n\ndef run_experiment_stage(stage_title, section_name, option_recipes, preference_order, preference_reason):\n    run_rows = []\n    first_seed_histories = {}\n    for option_name, option_recipe in option_recipes.items():\n        for seed_value in project_configuration[\"experiment_seeds\"]:\n            experiment_row, history_dataframe = run_one_screening_experiment(option_recipe, seed_value, f\"{stage_title} | {option_name} | seed {seed_value}\")\n            run_rows.append({\"option\": option_name, \"seed\": seed_value, **experiment_row})\n            if seed_value == project_configuration[\"experiment_seeds\"][0]:\n                first_seed_histories[option_name] = history_dataframe\n    runs_dataframe = pd.DataFrame(run_rows)\n    save_results_table(runs_dataframe.round(4), section_name, \"all_runs.csv\")\n\n    summary_dataframe = runs_dataframe.groupby(\"option\", sort=False).agg(\n        mean_val_qwk=(\"best_val_qwk\", \"mean\"), lowest_val_qwk=(\"best_val_qwk\", \"min\"), highest_val_qwk=(\"best_val_qwk\", \"max\"),\n        mean_val_macro_f1=(\"val_macro_f1\", \"mean\"), mean_val_accuracy=(\"val_accuracy\", \"mean\"),\n        mean_train_minus_val_qwk_gap=(\"train_minus_val_qwk_gap\", \"mean\"), mean_val_recall_severe=(\"val_recall_severe\", \"mean\"),\n        mean_best_epoch=(\"best_epoch\", \"mean\"), trainable_parameters=(\"trainable_parameters\", \"first\"), mean_minutes=(\"minutes\", \"mean\"),\n    )\n    save_results_table(summary_dataframe.round(4).reset_index(), section_name, \"summary_per_option.csv\")\n\n    leader_name = summary_dataframe[\"mean_val_qwk\"].idxmax()\n    leader_qwk = summary_dataframe.loc[leader_name, \"mean_val_qwk\"]\n    winner_name = leader_name\n    for preferred_option in preference_order:\n        if summary_dataframe.loc[preferred_option, \"mean_val_qwk\"] >= leader_qwk - project_configuration[\"tie_margin_qwk\"]:\n            winner_name = preferred_option\n            break\n    if winner_name == leader_name:\n        decision_reason = f\"{winner_name} had the highest average validation QWK ({leader_qwk:.4f}).\"\n    else:\n        decision_reason = (f\"{leader_name} led narrowly ({leader_qwk:.4f}), but {winner_name} was within {project_configuration['tie_margin_qwk']} \"\n                           f\"({summary_dataframe.loc[winner_name, 'mean_val_qwk']:.4f}), so it wins because {preference_reason}.\")\n\n    figure_object, axis_array = plt.subplots(1, 3, figsize=(20, 5))\n    option_positions = np.arange(len(summary_dataframe))\n    bar_colours = [\"crimson\" if option_name == winner_name else \"steelblue\" for option_name in summary_dataframe.index]\n    axis_array[0].bar(option_positions, summary_dataframe[\"mean_val_qwk\"], color=bar_colours, alpha=0.8)\n    for option_index, option_name in enumerate(summary_dataframe.index):\n        seed_values = runs_dataframe.loc[runs_dataframe[\"option\"] == option_name, \"best_val_qwk\"]\n        axis_array[0].scatter([option_index] * len(seed_values), seed_values, color=\"black\", zorder=5, s=18)\n        axis_array[0].text(option_index, summary_dataframe.loc[option_name, \"mean_val_qwk\"] + 0.003, f\"{summary_dataframe.loc[option_name, 'mean_val_qwk']:.3f}\", ha=\"center\", fontsize=8)\n    lowest_value = runs_dataframe[\"best_val_qwk\"].min()\n    axis_array[0].set_ylim(max(0, lowest_value - 0.05), min(1, runs_dataframe[\"best_val_qwk\"].max() + 0.03))\n    axis_array[0].set_xticks(option_positions)\n    axis_array[0].set_xticklabels(summary_dataframe.index, rotation=15, fontsize=8)\n    axis_array[0].set_title(\"Best Validation QWK (bar = average, dots = each seed, red = winner)\")\n    axis_array[0].set_ylabel(\"Validation QWK\")\n    for option_name, history_dataframe in first_seed_histories.items():\n        axis_array[1].plot(history_dataframe[\"epoch\"], history_dataframe[\"val_qwk\"], marker=\"o\", markersize=3, label=option_name)\n    axis_array[1].set_title(f\"Validation QWK per Epoch (seed {project_configuration['experiment_seeds'][0]})\")\n    axis_array[1].set_xlabel(\"Epoch\")\n    axis_array[1].set_ylabel(\"Validation QWK\")\n    axis_array[1].legend(fontsize=8)\n    axis_array[2].bar(option_positions, summary_dataframe[\"mean_train_minus_val_qwk_gap\"], color=bar_colours, alpha=0.8)\n    axis_array[2].set_xticks(option_positions)\n    axis_array[2].set_xticklabels(summary_dataframe.index, rotation=15, fontsize=8)\n    axis_array[2].set_title(\"Overfitting Gap: Training QWK minus Validation QWK (lower is better)\")\n    axis_array[2].set_ylabel(\"QWK gap\")\n    plt.suptitle(stage_title, y=1.02)\n    save_current_figure(section_name, \"experiment_comparison.png\")\n\n    stage_decision = {\n        \"stage\": stage_title,\n        \"options_tested\": list(option_recipes.keys()),\n        \"seeds\": project_configuration[\"experiment_seeds\"],\n        \"leader_by_mean_val_qwk\": leader_name,\n        \"winner\": winner_name,\n        \"winner_mean_val_qwk\": round(float(summary_dataframe.loc[winner_name, \"mean_val_qwk\"]), 4),\n        \"reason\": decision_reason,\n        \"summary\": summary_dataframe.round(4).reset_index().to_dict(orient=\"records\"),\n    }\n    save_results_json(stage_decision, section_name, \"decision.json\")\n    experiment_decisions.append(stage_decision)\n    print()\n    print(summary_dataframe[[\"mean_val_qwk\", \"lowest_val_qwk\", \"highest_val_qwk\", \"mean_val_macro_f1\", \"mean_train_minus_val_qwk_gap\", \"mean_val_recall_severe\"]].round(4).to_string())\n    print(\"Decision:\", decision_reason)\n    return winner_name, summary_dataframe\n\n\nprint(\"Cell 38 finished: experiment runner ready.\")\nprint(f\"Each option runs with seeds {project_configuration['experiment_seeds']} on {project_configuration['screening_architecture']}, \"\n      f\"with ties decided by a margin of {project_configuration['tie_margin_qwk']} QWK.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.301557Z","iopub.execute_input":"2026-09-26T11:55:17.302154Z","iopub.status.idle":"2026-09-26T11:55:17.324786Z","shell.execute_reply.started":"2026-09-26T11:55:17.302124Z","shell.execute_reply":"2026-09-26T11:55:17.324223Z"}},"outputs":[],"execution_count":null},{"id":"2466f698","cell_type":"markdown","source":"### Experiment 1: Preprocessing method\nResize only (baseline), CLAHE and Ben Graham, with everything else unchanged. On a tie, the simpler method wins: resize only first, then CLAHE (one step, keeps natural colours, which makes Grad-CAM easier for clinicians to read), then Ben Graham.","metadata":{}},{"id":"f8366280","cell_type":"code","source":"# Cell 39: EXPERIMENT 1, which preprocessing method gives the best validation QWK?\npreprocessing_options = {method_name: {**current_training_recipe, \"preprocessing_method\": method_name} for method_name in preprocessing_method_names}\npreprocessing_winner, preprocessing_summary = run_experiment_stage(\n    \"Experiment 1: Preprocessing Method\", \"07_experiment_preprocessing\", preprocessing_options,\n    preference_order=[\"resize_only\", \"clahe\", \"ben_graham\"], preference_reason=\"it is the simpler preprocessing\",\n)\ncurrent_training_recipe[\"preprocessing_method\"] = preprocessing_winner\nprint(f\"Cell 39 finished: preprocessing method set to {preprocessing_winner}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T11:55:17.325637Z","iopub.execute_input":"2026-09-26T11:55:17.325877Z","iopub.status.idle":"2026-09-26T12:29:22.840591Z","shell.execute_reply.started":"2026-09-26T11:55:17.32585Z","shell.execute_reply":"2026-09-26T12:29:22.839751Z"}},"outputs":[],"execution_count":null},{"id":"3272134f","cell_type":"markdown","source":"### Experiment 2: Augmentation\nIs augmentation needed, and how much? On a tie the lighter policy wins, because it changes the photos less.","metadata":{}},{"id":"34389fb0","cell_type":"code","source":"# Cell 40: EXPERIMENT 2, is augmentation needed, and which policy gives the best validation QWK?\naugmentation_options = {policy_name: {**current_training_recipe, \"augmentation_policy\": policy_name} for policy_name in augmentation_policy_names}\naugmentation_winner, augmentation_summary = run_experiment_stage(\n    \"Experiment 2: Augmentation Policy\", \"08_experiment_augmentation\", augmentation_options,\n    preference_order=augmentation_policy_names, preference_reason=\"it changes the photos less\",\n)\ncurrent_training_recipe[\"augmentation_policy\"] = augmentation_winner\nprint(f\"Cell 40 finished: augmentation policy set to {augmentation_winner}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T12:29:22.841905Z","iopub.execute_input":"2026-09-26T12:29:22.842293Z","iopub.status.idle":"2026-09-26T13:01:53.252758Z","shell.execute_reply.started":"2026-09-26T12:29:22.84226Z","shell.execute_reply":"2026-09-26T13:01:53.252078Z"}},"outputs":[],"execution_count":null},{"id":"1ec01e46","cell_type":"markdown","source":"### Experiment 3: Class balance\nPlain loss vs class-weighted loss, where mistakes on rare grades count for more. Look at Severe recall in the table as well as QWK. On a tie the plain loss wins, because it is simpler.","metadata":{}},{"id":"b3f4cb15","cell_type":"code","source":"# Cell 41: EXPERIMENT 3, does weighting the loss towards rare grades help?\nclass_balance_options = {\n    \"no_class_weights\": {**current_training_recipe, \"use_class_weights\": False},\n    \"class_weighted_loss\": {**current_training_recipe, \"use_class_weights\": True},\n}\nclass_balance_winner, class_balance_summary = run_experiment_stage(\n    \"Experiment 3: Class Balance\", \"09_experiment_class_balance\", class_balance_options,\n    preference_order=[\"no_class_weights\", \"class_weighted_loss\"], preference_reason=\"the plain loss is simpler\",\n)\ncurrent_training_recipe[\"use_class_weights\"] = class_balance_winner == \"class_weighted_loss\"\nprint(f\"Cell 41 finished: class weights {'on' if current_training_recipe['use_class_weights'] else 'off'}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T13:01:53.258951Z","iopub.execute_input":"2026-09-26T13:01:53.259331Z","iopub.status.idle":"2026-09-26T13:15:40.683307Z","shell.execute_reply.started":"2026-09-26T13:01:53.259299Z","shell.execute_reply":"2026-09-26T13:15:40.682546Z"}},"outputs":[],"execution_count":null},{"id":"d11d7fce","cell_type":"markdown","source":"### Experiment 4: Transfer learning and freezing\nWhich parts of the model should learn? On a tie, the option that trains fewer weights wins, because fewer changing weights means less room to overfit.","metadata":{}},{"id":"5b20f37d","cell_type":"code","source":"# Cell 42: EXPERIMENT 4, how much of the pretrained model should be frozen?\ntransfer_strategy_names = [\"head_only\", \"two_phase_partial_unfreeze\", \"two_phase_full_unfreeze\", \"full_fine_tune\"]\ntransfer_strategy_options = {strategy_name: {**current_training_recipe, \"transfer_strategy\": strategy_name} for strategy_name in transfer_strategy_names}\ntransfer_strategy_winner, transfer_strategy_summary = run_experiment_stage(\n    \"Experiment 4: Transfer Learning and Freezing\", \"10_experiment_transfer_learning\", transfer_strategy_options,\n    preference_order=transfer_strategy_names, preference_reason=\"it trains fewer weights, leaving less room to overfit\",\n)\ncurrent_training_recipe[\"transfer_strategy\"] = transfer_strategy_winner\nprint(f\"Cell 42 finished: transfer strategy set to {transfer_strategy_winner}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T13:15:40.684554Z","iopub.execute_input":"2026-09-26T13:15:40.684799Z","iopub.status.idle":"2026-09-26T13:43:19.66414Z","shell.execute_reply.started":"2026-09-26T13:15:40.684771Z","shell.execute_reply":"2026-09-26T13:43:19.663384Z"}},"outputs":[],"execution_count":null},{"id":"13826546","cell_type":"markdown","source":"### Experiment 5: Learning rate\nThe learning rate of the phase where the model adapts is tested at the default and at roughly three times lower and higher. If only the last layer learns, that is the phase 1 rate. Otherwise it is the fine-tuning rate. On a tie the default wins, because it was the starting point.","metadata":{}},{"id":"54cef6ae","cell_type":"code","source":"# Cell 43: EXPERIMENT 5, which learning rate works best for the chosen strategy?\nlearning_rate_setting_name = \"phase_one_learning_rate\" if current_training_recipe[\"transfer_strategy\"] == \"head_only\" else \"phase_two_learning_rate\"\ndefault_learning_rate = current_training_recipe[learning_rate_setting_name]\nlearning_rate_values = [default_learning_rate, default_learning_rate / 3, default_learning_rate * 3]\nlearning_rate_options = {f\"lr_{learning_rate_value:.0e}\": {**current_training_recipe, learning_rate_setting_name: learning_rate_value} for learning_rate_value in learning_rate_values}\nlearning_rate_winner, learning_rate_summary = run_experiment_stage(\n    f\"Experiment 5: Learning Rate ({learning_rate_setting_name})\", \"11_experiment_learning_rate\", learning_rate_options,\n    preference_order=list(learning_rate_options.keys()), preference_reason=\"it is the default starting value\",\n)\ncurrent_training_recipe = copy.deepcopy(learning_rate_options[learning_rate_winner])\nprint(f\"Cell 43 finished: {learning_rate_setting_name} set to {current_training_recipe[learning_rate_setting_name]:.0e}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T13:43:19.66554Z","iopub.execute_input":"2026-09-26T13:43:19.665928Z","iopub.status.idle":"2026-09-26T14:08:57.987668Z","shell.execute_reply.started":"2026-09-26T13:43:19.665894Z","shell.execute_reply":"2026-09-26T14:08:57.986961Z"}},"outputs":[],"execution_count":null},{"id":"31f69888","cell_type":"markdown","source":"### Experiment 6: Regularisation against overfitting\nDeep models trained on a few thousand photos tend to overfit: training accuracy climbs well above validation accuracy, and validation loss starts rising while QWK is still creeping up, a sign of over-confident wrong answers. This experiment adds four standard fixes together:\n\n- **Label smoothing (0.1):** the model is trained to be 90% sure instead of 100%, which directly targets over-confidence.\n- **Weight decay (0.01, with AdamW):** keeps weights small, so the model cannot memorise individual photos.\n- **Higher dropout (0.4) before the last layer:** randomly switches off features during training, so the model cannot rely on any single one.\n- **Cosine learning rate schedule:** lowers the rate smoothly instead of in sudden halvings.\n\nQWK stays the metric that picks the best version. **On a tie the regularised recipe wins**, because generalising to new photos is the goal. The chart's right panel shows the overfitting gap directly.","metadata":{}},{"id":"547497fe","cell_type":"code","source":"# Cell 44: EXPERIMENT 6, does regularisation reduce overfitting without losing validation QWK?\nregularisation_options = {\n    \"regularised\": {**current_training_recipe, \"label_smoothing\": 0.1, \"weight_decay\": 0.01, \"head_dropout\": 0.4, \"phase_two_schedule\": \"cosine\"},\n    \"no_extra_regularisation\": {**current_training_recipe},\n}\nregularisation_winner, regularisation_summary = run_experiment_stage(\n    \"Experiment 6: Regularisation\", \"12_experiment_regularisation\", regularisation_options,\n    preference_order=[\"regularised\", \"no_extra_regularisation\"], preference_reason=\"it generalises better, with a smaller training vs validation gap\",\n)\ncurrent_training_recipe = copy.deepcopy(regularisation_options[regularisation_winner])\n\nregularisation_histories = {option_name: experiment_run_cache[json_module.dumps(option_recipe, sort_keys=True) + f\"|seed={project_configuration['experiment_seeds'][0]}\"][1]\n                            for option_name, option_recipe in regularisation_options.items()}\nfigure_object, axis_array = plt.subplots(1, 2, figsize=(15, 5))\nfor axis_index, (option_name, history_dataframe) in enumerate(regularisation_histories.items()):\n    axis_array[axis_index].plot(history_dataframe[\"epoch\"], history_dataframe[\"train_qwk\"], marker=\"o\", markersize=3, label=\"Training QWK\")\n    axis_array[axis_index].plot(history_dataframe[\"epoch\"], history_dataframe[\"val_qwk\"], marker=\"o\", markersize=3, label=\"Validation QWK\")\n    axis_array[axis_index].plot(history_dataframe[\"epoch\"], history_dataframe[\"val_loss\"], linestyle=\"--\", color=\"grey\", label=\"Validation loss\")\n    axis_array[axis_index].set_title(f\"{option_name}: training vs validation\")\n    axis_array[axis_index].set_xlabel(\"Epoch\")\n    axis_array[axis_index].legend(fontsize=8)\nsave_current_figure(\"12_experiment_regularisation\", \"train_vs_validation_qwk_with_and_without_regularisation.png\")\nprint(f\"Cell 44 finished: regularisation set to {regularisation_winner}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T14:08:57.988832Z","iopub.execute_input":"2026-09-26T14:08:57.989546Z","iopub.status.idle":"2026-09-26T14:21:30.115635Z","shell.execute_reply.started":"2026-09-26T14:08:57.989518Z","shell.execute_reply":"2026-09-26T14:21:30.114789Z"}},"outputs":[],"execution_count":null},{"id":"cac9ac89","cell_type":"markdown","source":"## 10. Final Configuration\nEvery tested choice is now fixed. The table below is the justification trail for the report: what was tested, what won and why.","metadata":{}},{"id":"2b379f20","cell_type":"code","source":"# Cell 45: gather every experiment decision into the final configuration and chart how validation QWK moved stage by stage\nfinal_training_recipe = copy.deepcopy(current_training_recipe)\ndecision_table_dataframe = pd.DataFrame([\n    {\"experiment\": decision[\"stage\"], \"options_tested\": \", \".join(decision[\"options_tested\"]), \"winner\": decision[\"winner\"],\n     \"winner_mean_val_qwk\": decision[\"winner_mean_val_qwk\"], \"reason\": decision[\"reason\"]}\n    for decision in experiment_decisions\n])\nsave_results_table(decision_table_dataframe, \"13_final_configuration\", \"experiment_decisions_table.csv\")\nfinal_project_configuration = {\n    \"fixed_settings\": project_configuration,\n    \"tested_and_chosen_recipe\": final_training_recipe,\n    \"starting_recipe_before_experiments\": starting_training_recipe,\n    \"final_training_schedule\": final_schedule,\n    \"experiment_decisions\": experiment_decisions,\n}\nsave_results_json(final_project_configuration, \"13_final_configuration\", \"final_project_configuration.json\")\n\nfigure_object, axis_object = plt.subplots(figsize=(11, 5))\naxis_object.plot(range(1, len(decision_table_dataframe) + 1), decision_table_dataframe[\"winner_mean_val_qwk\"], marker=\"o\", color=\"crimson\")\nfor stage_index, row in decision_table_dataframe.iterrows():\n    axis_object.annotate(f\"{row['winner']}\\n{row['winner_mean_val_qwk']:.3f}\", (stage_index + 1, row[\"winner_mean_val_qwk\"]),\n                         textcoords=\"offset points\", xytext=(0, 10), ha=\"center\", fontsize=8)\naxis_object.set_xticks(range(1, len(decision_table_dataframe) + 1))\naxis_object.set_xticklabels([f\"Exp {number}\" for number in range(1, len(decision_table_dataframe) + 1)])\naxis_object.set_ylabel(\"Winner's average validation QWK (B0, short schedule)\")\naxis_object.set_title(\"How Validation QWK Moved as Each Choice Was Fixed\")\nsave_current_figure(\"13_final_configuration\", \"validation_qwk_by_experiment_stage.png\")\n\nprint(\"Cell 45 finished: final configuration fixed and saved.\")\nprint(decision_table_dataframe[[\"experiment\", \"winner\", \"winner_mean_val_qwk\"]].to_string(index=False))\nprint()\nprint(\"Final recipe used for all three models:\")\nprint(json_module.dumps(final_training_recipe, indent=2))\nprint(\"Total experiment runs actually trained (reused runs not counted):\", len(experiment_run_cache))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T14:21:30.116878Z","iopub.execute_input":"2026-09-26T14:21:30.117246Z","iopub.status.idle":"2026-09-26T14:21:30.566537Z","shell.execute_reply.started":"2026-09-26T14:21:30.117216Z","shell.execute_reply":"2026-09-26T14:21:30.565907Z"}},"outputs":[],"execution_count":null},{"id":"2ae84bab","cell_type":"markdown","source":"## 11. Final Training\nAll three models are trained with the final recipe and the full schedule, using seed 42. The best version of each is kept by **validation QWK**. Training QWK is plotted next to validation QWK for every model, so the overfitting gap is shown in the metric that actually matters.","metadata":{}},{"id":"8929dc96","cell_type":"code","source":"# Cell 46: train EfficientNet-B3 with the final recipe\ndef build_main_checkpoint_path(architecture_name):\n    return os.path.join(main_models_folder_path, f\"best_{architecture_name}_by_validation_qwk.pth\")\n\n\nefficientnet_b3_training_result = train_with_recipe(\"efficientnet_b3\", final_training_recipe, final_schedule, \"efficientnet_b3\",\n                                                    build_main_checkpoint_path(\"efficientnet_b3\"), RANDOM_SEED, print_every_epoch=True)\nprint(f\"Cell 46 finished: EfficientNet-B3 trained in {efficientnet_b3_training_result['training_duration_minutes']:.1f} minutes, \"\n      f\"best validation QWK {efficientnet_b3_training_result['best_validation_qwk']:.4f} at epoch {efficientnet_b3_training_result['best_epoch']}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T14:21:30.567473Z","iopub.execute_input":"2026-09-26T14:21:30.567668Z","iopub.status.idle":"2026-09-26T14:40:40.570527Z","shell.execute_reply.started":"2026-09-26T14:21:30.567648Z","shell.execute_reply":"2026-09-26T14:40:40.569712Z"}},"outputs":[],"execution_count":null},{"id":"f216d5f6","cell_type":"code","source":"# Cell 47: train EfficientNet-B0 with the final recipe\nefficientnet_b0_training_result = train_with_recipe(\"efficientnet_b0\", final_training_recipe, final_schedule, \"efficientnet_b0\",\n                                                    build_main_checkpoint_path(\"efficientnet_b0\"), RANDOM_SEED, print_every_epoch=True)\nprint(f\"Cell 47 finished: EfficientNet-B0 trained in {efficientnet_b0_training_result['training_duration_minutes']:.1f} minutes, \"\n      f\"best validation QWK {efficientnet_b0_training_result['best_validation_qwk']:.4f} at epoch {efficientnet_b0_training_result['best_epoch']}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T14:40:40.571832Z","iopub.execute_input":"2026-09-26T14:40:40.572209Z","iopub.status.idle":"2026-09-26T14:52:39.870957Z","shell.execute_reply.started":"2026-09-26T14:40:40.572178Z","shell.execute_reply":"2026-09-26T14:52:39.870192Z"}},"outputs":[],"execution_count":null},{"id":"e7a779b7","cell_type":"code","source":"# Cell 48: train ResNet50 with the final recipe\nresnet50_training_result = train_with_recipe(\"resnet50\", final_training_recipe, final_schedule, \"resnet50\",\n                                             build_main_checkpoint_path(\"resnet50\"), RANDOM_SEED, print_every_epoch=True)\nprint(f\"Cell 48 finished: ResNet50 trained in {resnet50_training_result['training_duration_minutes']:.1f} minutes, \"\n      f\"best validation QWK {resnet50_training_result['best_validation_qwk']:.4f} at epoch {resnet50_training_result['best_epoch']}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T14:52:39.8724Z","iopub.execute_input":"2026-09-26T14:52:39.872763Z","iopub.status.idle":"2026-09-26T15:16:01.464154Z","shell.execute_reply.started":"2026-09-26T14:52:39.872733Z","shell.execute_reply":"2026-09-26T15:16:01.463249Z"}},"outputs":[],"execution_count":null},{"id":"cbe7a99b","cell_type":"code","source":"# Cell 49: training QWK vs validation QWK per epoch for every model, the main overfitting evidence\nall_training_results = [efficientnet_b3_training_result, efficientnet_b0_training_result, resnet50_training_result]\narchitecture_display_colours = {\"efficientnet_b3\": \"steelblue\", \"efficientnet_b0\": \"darkorange\", \"resnet50\": \"seagreen\"}\nphase_boundary_epoch_position = final_schedule[\"phase_one_epochs\"] + 0.5\nuses_two_phases = final_training_recipe[\"transfer_strategy\"].startswith(\"two_phase\")\n\nfigure_object, axis_array = plt.subplots(1, len(all_training_results) + 1, figsize=(24, 5))\nfor axis_index, training_result in enumerate(all_training_results):\n    history_dataframe = training_result[\"history_dataframe\"]\n    current_axis = axis_array[axis_index]\n    current_axis.plot(history_dataframe[\"epoch\"], history_dataframe[\"train_qwk\"], marker=\"o\", markersize=3, label=\"Training QWK\")\n    current_axis.plot(history_dataframe[\"epoch\"], history_dataframe[\"val_qwk\"], marker=\"o\", markersize=3, label=\"Validation QWK\")\n    current_axis.axvline(training_result[\"best_epoch\"], color=\"crimson\", linestyle=\":\", label=f\"Kept version (epoch {training_result['best_epoch']})\")\n    if uses_two_phases:\n        current_axis.axvline(phase_boundary_epoch_position, color=\"grey\", linestyle=\"--\", label=\"Fine-tuning starts\")\n    current_axis.set_title(f\"{training_result['architecture_name']}: training vs validation QWK\")\n    current_axis.set_xlabel(\"Epoch\")\n    current_axis.set_ylabel(\"QWK\")\n    current_axis.legend(fontsize=8)\nfor training_result in all_training_results:\n    axis_array[-1].plot(training_result[\"history_dataframe\"][\"epoch\"], training_result[\"history_dataframe\"][\"val_qwk\"], marker=\"o\", markersize=3,\n                        color=architecture_display_colours[training_result[\"architecture_name\"]], label=training_result[\"architecture_name\"])\naxis_array[-1].set_title(\"Validation QWK, all three models\")\naxis_array[-1].set_xlabel(\"Epoch\")\naxis_array[-1].legend(fontsize=8)\nsave_current_figure(\"14_final_training\", \"train_vs_validation_qwk_per_model.png\")\n\nprint(\"Cell 49 finished: QWK curves saved.\")\nprint(\"Training QWK is measured on augmented photos, which are harder than clean ones, so a small gap is expected even without overfitting.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:01.465744Z","iopub.execute_input":"2026-09-26T15:16:01.466151Z","iopub.status.idle":"2026-09-26T15:16:02.8961Z","shell.execute_reply.started":"2026-09-26T15:16:01.466115Z","shell.execute_reply":"2026-09-26T15:16:02.895388Z"}},"outputs":[],"execution_count":null},{"id":"2bfee13b","cell_type":"code","source":"# Cell 50: training vs validation loss and accuracy per model, plus a table of the gaps at the kept epoch\nfigure_object, axis_grid = plt.subplots(len(all_training_results), 2, figsize=(14, 4.2 * len(all_training_results)), squeeze=False)\noverfitting_gap_rows = []\nfor row_index, training_result in enumerate(all_training_results):\n    history_dataframe = training_result[\"history_dataframe\"]\n    for column_index, (train_column, val_column, chart_label) in enumerate([(\"train_loss\", \"val_loss\", \"loss\"), (\"train_accuracy\", \"val_accuracy\", \"accuracy\")]):\n        current_axis = axis_grid[row_index, column_index]\n        current_axis.plot(history_dataframe[\"epoch\"], history_dataframe[train_column], marker=\"o\", markersize=3, label=f\"Training {chart_label}\")\n        current_axis.plot(history_dataframe[\"epoch\"], history_dataframe[val_column], marker=\"o\", markersize=3, label=f\"Validation {chart_label}\")\n        current_axis.axvline(training_result[\"best_epoch\"], color=\"crimson\", linestyle=\":\", label=\"Kept version\")\n        current_axis.set_title(f\"{training_result['architecture_name']}: {chart_label}\")\n        current_axis.set_xlabel(\"Epoch\")\n        current_axis.legend(fontsize=8)\n    kept_row = history_dataframe[history_dataframe[\"epoch\"] == training_result[\"best_epoch\"]].iloc[0]\n    overfitting_gap_rows.append({\n        \"architecture\": training_result[\"architecture_name\"],\n        \"kept_epoch\": training_result[\"best_epoch\"],\n        \"epochs_trained\": training_result[\"epochs_trained\"],\n        \"train_qwk\": round(float(kept_row[\"train_qwk\"]), 4),\n        \"val_qwk\": round(float(kept_row[\"val_qwk\"]), 4),\n        \"qwk_gap\": round(float(kept_row[\"train_qwk\"] - kept_row[\"val_qwk\"]), 4),\n        \"train_accuracy\": round(float(kept_row[\"train_accuracy\"]), 4),\n        \"val_accuracy\": round(float(kept_row[\"val_accuracy\"]), 4),\n        \"accuracy_gap\": round(float(kept_row[\"train_accuracy\"] - kept_row[\"val_accuracy\"]), 4),\n        \"train_loss\": round(float(kept_row[\"train_loss\"]), 4),\n        \"val_loss\": round(float(kept_row[\"val_loss\"]), 4),\n        \"epoch_with_lowest_val_loss\": training_result[\"epoch_with_lowest_val_loss\"],\n    })\nsave_current_figure(\"14_final_training\", \"train_vs_validation_loss_and_accuracy_per_model.png\")\noverfitting_gap_dataframe = pd.DataFrame(overfitting_gap_rows)\nsave_results_table(overfitting_gap_dataframe, \"14_final_training\", \"overfitting_gaps_at_kept_epoch.csv\")\n\nprint(\"Cell 50 finished: loss and accuracy curves and the overfitting gap table saved.\")\nprint(overfitting_gap_dataframe.to_string(index=False))\nprint(\"Compare these gaps with the no_extra_regularisation option in experiment 6 to see how much the regularisation helped.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:02.897268Z","iopub.execute_input":"2026-09-26T15:16:02.897572Z","iopub.status.idle":"2026-09-26T15:16:04.979031Z","shell.execute_reply.started":"2026-09-26T15:16:02.89755Z","shell.execute_reply":"2026-09-26T15:16:04.978246Z"}},"outputs":[],"execution_count":null},{"id":"1c937c4a","cell_type":"code","source":"# Cell 51: save every model's epoch-by-epoch history and a size, speed and score table, and pick the best single model by validation QWK\nfor training_result in all_training_results:\n    save_results_table(training_result[\"history_dataframe\"], \"14_final_training\", f\"training_history_{training_result['architecture_name']}.csv\")\n\narchitecture_comparison_dataframe = pd.DataFrame([{\n    \"architecture\": result[\"architecture_name\"], \"total_parameters\": result[\"total_parameters\"], \"trainable_parameters\": result[\"trainable_parameters\"],\n    \"epochs_trained\": result[\"epochs_trained\"], \"kept_epoch\": result[\"best_epoch\"], \"best_validation_qwk\": round(result[\"best_validation_qwk\"], 4),\n    \"training_minutes\": round(result[\"training_duration_minutes\"], 1),\n} for result in all_training_results])\nsave_results_table(architecture_comparison_dataframe, \"14_final_training\", \"architecture_comparison.csv\")\n\nbest_single_architecture_name = architecture_comparison_dataframe.loc[architecture_comparison_dataframe[\"best_validation_qwk\"].idxmax(), \"architecture\"]\n\nprint(\"Cell 51 finished: histories and comparison table saved.\")\nprint(architecture_comparison_dataframe.to_string(index=False))\nprint(f\"Best single model, chosen on validation QWK only: {best_single_architecture_name}. The test set has still not been used.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:04.980819Z","iopub.execute_input":"2026-09-26T15:16:04.98127Z","iopub.status.idle":"2026-09-26T15:16:04.994584Z","shell.execute_reply.started":"2026-09-26T15:16:04.981245Z","shell.execute_reply":"2026-09-26T15:16:04.993692Z"}},"outputs":[],"execution_count":null},{"id":"767fb592","cell_type":"code","source":"# Cell 52: time how fast each model predicts, on the GPU in batches and on the CPU one photo at a time like the web app\nfinal_data_loaders = build_data_loaders(final_training_recipe[\"preprocessing_method\"], final_training_recipe[\"augmentation_policy\"])\nlatency_probe_images, _ = next(iter(final_data_loaders[\"test\"]))\n\n\ndef time_model_forward_passes(model, input_batch, number_of_repeats, device):\n    model = model.to(device).eval()\n    input_batch = input_batch.to(device)\n    with torch.no_grad():\n        for _ in range(2):\n            model(input_batch)\n        if device.type == \"cuda\":\n            torch.cuda.synchronize()\n        timing_start = time.time()\n        for _ in range(number_of_repeats):\n            model(input_batch)\n        if device.type == \"cuda\":\n            torch.cuda.synchronize()\n    return (time.time() - timing_start) / number_of_repeats * 1000 / input_batch.shape[0]\n\n\ndef load_trained_model(architecture_name, device):\n    model = build_transfer_learning_model(architecture_name, project_configuration[\"number_of_dr_classes\"], final_training_recipe[\"head_dropout\"])\n    model.load_state_dict(torch.load(build_main_checkpoint_path(architecture_name), map_location=\"cpu\"))\n    return model.to(device).eval()\n\n\ninference_latency_rows = []\nfor architecture_name in architecture_names_to_compare:\n    timing_model = load_trained_model(architecture_name, torch.device(\"cpu\"))\n    inference_latency_rows.append({\n        \"model\": architecture_name,\n        \"gpu_milliseconds_per_photo_in_batches_of_16\": round(time_model_forward_passes(timing_model, latency_probe_images, 30, compute_device), 2) if is_gpu_available else float(\"nan\"),\n        \"cpu_milliseconds_for_one_photo\": round(time_model_forward_passes(timing_model, latency_probe_images[:1].clone(), 5, torch.device(\"cpu\")), 1),\n    })\n    del timing_model\ninference_latency_dataframe = pd.DataFrame(inference_latency_rows)\ninference_latency_dataframe.loc[len(inference_latency_dataframe)] = {\n    \"model\": \"ensemble (all three)\",\n    \"gpu_milliseconds_per_photo_in_batches_of_16\": round(inference_latency_dataframe[\"gpu_milliseconds_per_photo_in_batches_of_16\"].sum(), 2),\n    \"cpu_milliseconds_for_one_photo\": round(inference_latency_dataframe[\"cpu_milliseconds_for_one_photo\"].sum(), 1),\n}\nsave_results_table(inference_latency_dataframe, \"14_final_training\", \"inference_speed.csv\")\n\nprint(\"Cell 52 finished: prediction speed measured.\")\nprint(inference_latency_dataframe.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:04.995761Z","iopub.execute_input":"2026-09-26T15:16:04.996189Z","iopub.status.idle":"2026-09-26T15:16:19.142912Z","shell.execute_reply.started":"2026-09-26T15:16:04.996156Z","shell.execute_reply":"2026-09-26T15:16:19.141773Z"}},"outputs":[],"execution_count":null},{"id":"0db876f1","cell_type":"markdown","source":"## 12. Testing the Models\nThe test set is used **now, and only once**. Each model and the ensemble (the average of the three models' grade probabilities, called soft voting) are scored on:\n\n- **QWK** with a 95% confidence interval from bootstrapping (re-sampling the test photos 1000 times), so the report shows how certain each number is\n- **Accuracy, precision, recall and F1**, overall and per grade\n- **Confusion matrices**, which show which grades get mixed up\n- A **screening view**: are patients who need a specialist caught?\n- **Model agreement** and **McNemar's test**, which checks whether the ensemble is really better than the best single model","metadata":{}},{"id":"33dda33e","cell_type":"code","source":"# Cell 53: test all three models once on the held-out test photos and build the soft-voting ensemble\nevaluation_loss_function = build_loss_function(final_training_recipe)\nper_architecture_test_results = {}\nfor architecture_name in architecture_names_to_compare:\n    evaluation_model = load_trained_model(architecture_name, compute_device)\n    test_loss, test_accuracy, test_qwk, true_labels, predicted_labels, predicted_probabilities = run_evaluation_pass(\n        evaluation_model, final_data_loaders[\"test\"], evaluation_loss_function, compute_device)\n    per_architecture_test_results[architecture_name] = {\n        \"test_loss\": test_loss, \"test_accuracy\": test_accuracy, \"test_qwk\": test_qwk,\n        \"test_true_labels\": true_labels, \"test_predicted_labels\": predicted_labels, \"test_predicted_probabilities\": predicted_probabilities,\n    }\n    del evaluation_model\n\nensemble_true_labels = per_architecture_test_results[\"efficientnet_b3\"][\"test_true_labels\"]\nassert all(np.array_equal(per_architecture_test_results[name][\"test_true_labels\"], ensemble_true_labels) for name in architecture_names_to_compare), \\\n    \"The test photos came out in a different order for different models, so averaging would be wrong.\"\nensemble_averaged_probabilities = np.stack([per_architecture_test_results[name][\"test_predicted_probabilities\"] for name in architecture_names_to_compare]).mean(axis=0)\nensemble_predicted_labels = ensemble_averaged_probabilities.argmax(axis=1)\nensemble_test_accuracy = accuracy_score(ensemble_true_labels, ensemble_predicted_labels)\nensemble_test_qwk = cohen_kappa_score(ensemble_true_labels, ensemble_predicted_labels, weights=\"quadratic\")\n\nall_model_predictions_by_name = {\n    name: (per_architecture_test_results[name][\"test_true_labels\"], per_architecture_test_results[name][\"test_predicted_labels\"], per_architecture_test_results[name][\"test_predicted_probabilities\"])\n    for name in architecture_names_to_compare\n}\nall_model_predictions_by_name[\"ensemble (soft vote)\"] = (ensemble_true_labels, ensemble_predicted_labels, ensemble_averaged_probabilities)\ntest_split_in_loader_order = test_split_dataframe.reset_index(drop=True)\n\nprint(\"Cell 53 finished: all models tested on the same\", len(ensemble_true_labels), \"test photos, and the ensemble built.\")\nfor model_name, (true_labels_array, predicted_labels_array, _) in all_model_predictions_by_name.items():\n    print(f\"  {model_name}: accuracy {accuracy_score(true_labels_array, predicted_labels_array):.4f}, QWK {cohen_kappa_score(true_labels_array, predicted_labels_array, weights='quadratic'):.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:19.144731Z","iopub.execute_input":"2026-09-26T15:16:19.145181Z","iopub.status.idle":"2026-09-26T15:16:32.27051Z","shell.execute_reply.started":"2026-09-26T15:16:19.145146Z","shell.execute_reply":"2026-09-26T15:16:32.269772Z"}},"outputs":[],"execution_count":null},{"id":"29bb944f","cell_type":"code","source":"# Cell 54: headline metrics for all four models, with 95% bootstrap confidence intervals for QWK, accuracy and macro F1\nnumber_of_bootstrap_resamples = 1000\nbootstrap_random_generator = np.random.default_rng(RANDOM_SEED)\nbootstrap_index_sets = [bootstrap_random_generator.integers(0, len(ensemble_true_labels), len(ensemble_true_labels)) for _ in range(number_of_bootstrap_resamples)]\n\n\ndef bootstrap_confidence_interval(true_labels_array, predicted_labels_array, metric_function):\n    resampled_scores = [metric_function(true_labels_array[index_set], predicted_labels_array[index_set]) for index_set in bootstrap_index_sets]\n    return np.percentile(resampled_scores, 2.5), np.percentile(resampled_scores, 97.5)\n\n\nmetric_functions = {\n    \"qwk\": lambda true_values, predicted_values: cohen_kappa_score(true_values, predicted_values, weights=\"quadratic\"),\n    \"accuracy\": accuracy_score,\n    \"macro_f1\": lambda true_values, predicted_values: precision_recall_fscore_support(true_values, predicted_values, average=\"macro\", zero_division=0)[2],\n}\nheadline_metric_rows = []\nfor model_name, (true_labels_array, predicted_labels_array, _) in all_model_predictions_by_name.items():\n    macro_precision, macro_recall, macro_f1, _ = precision_recall_fscore_support(true_labels_array, predicted_labels_array, average=\"macro\", zero_division=0)\n    weighted_precision, weighted_recall, weighted_f1, _ = precision_recall_fscore_support(true_labels_array, predicted_labels_array, average=\"weighted\", zero_division=0)\n    headline_row = {\n        \"model\": model_name,\n        \"qwk\": cohen_kappa_score(true_labels_array, predicted_labels_array, weights=\"quadratic\"),\n        \"accuracy\": accuracy_score(true_labels_array, predicted_labels_array),\n        \"macro_precision\": macro_precision, \"macro_recall\": macro_recall, \"macro_f1\": macro_f1,\n        \"weighted_precision\": weighted_precision, \"weighted_recall\": weighted_recall, \"weighted_f1\": weighted_f1,\n    }\n    for metric_name, metric_function in metric_functions.items():\n        headline_row[f\"{metric_name}_ci_low\"], headline_row[f\"{metric_name}_ci_high\"] = bootstrap_confidence_interval(true_labels_array, predicted_labels_array, metric_function)\n    headline_metric_rows.append(headline_row)\nheadline_metrics_dataframe = pd.DataFrame(headline_metric_rows).round(4)\nsave_results_table(headline_metrics_dataframe, \"15_evaluation\", \"headline_metrics_with_confidence_intervals.csv\")\n\nfigure_object, axis_object = plt.subplots(figsize=(12, 5.5))\nbar_x_positions = np.arange(len(headline_metrics_dataframe))\nfor metric_index, (metric_name, metric_label) in enumerate([(\"qwk\", \"QWK\"), (\"accuracy\", \"Accuracy\"), (\"macro_f1\", \"Macro F1\")]):\n    metric_values = headline_metrics_dataframe[metric_name]\n    error_bars = [metric_values - headline_metrics_dataframe[f\"{metric_name}_ci_low\"], headline_metrics_dataframe[f\"{metric_name}_ci_high\"] - metric_values]\n    metric_bars = axis_object.bar(bar_x_positions + (metric_index - 1) * 0.27, metric_values, 0.27, yerr=error_bars, capsize=3, label=metric_label)\n    axis_object.bar_label(metric_bars, fmt=\"%.3f\", fontsize=7, padding=6)\naxis_object.set_xticks(bar_x_positions)\naxis_object.set_xticklabels(headline_metrics_dataframe[\"model\"], rotation=10)\naxis_object.set_ylim(0, 1.1)\naxis_object.set_title(\"Test Set Results with 95% Bootstrap Confidence Intervals\")\naxis_object.legend(loc=\"lower right\")\nsave_current_figure(\"15_evaluation\", \"headline_metrics_with_confidence_intervals.png\")\n\nprint(\"Cell 54 finished: headline metrics with confidence intervals saved.\")\nprint(headline_metrics_dataframe[[\"model\", \"qwk\", \"qwk_ci_low\", \"qwk_ci_high\", \"accuracy\", \"macro_f1\", \"weighted_f1\"]].to_string(index=False))\nprint(\"Overlapping confidence intervals mean the difference between two models may not be real on this test set.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:32.271856Z","iopub.execute_input":"2026-09-26T15:16:32.272563Z","iopub.status.idle":"2026-09-26T15:16:42.878443Z","shell.execute_reply.started":"2026-09-26T15:16:32.272532Z","shell.execute_reply":"2026-09-26T15:16:42.877784Z"}},"outputs":[],"execution_count":null},{"id":"0dc986c7","cell_type":"code","source":"# Cell 55: precision, recall and F1 for every grade and every model, plus the full ensemble report\nper_class_report_rows = []\nfor model_name, (true_labels_array, predicted_labels_array, _) in all_model_predictions_by_name.items():\n    precision_per_class, recall_per_class, f1_per_class, support_per_class = precision_recall_fscore_support(\n        true_labels_array, predicted_labels_array, labels=list_of_class_indices, zero_division=0)\n    for class_index in list_of_class_indices:\n        per_class_report_rows.append({\"model\": model_name, \"stage\": diagnosis_stage_names[class_index], \"precision\": precision_per_class[class_index],\n                                      \"recall\": recall_per_class[class_index], \"f1\": f1_per_class[class_index], \"test_photos\": int(support_per_class[class_index])})\nper_class_report_dataframe = pd.DataFrame(per_class_report_rows).round(3)\nsave_results_table(per_class_report_dataframe, \"15_evaluation\", \"per_grade_precision_recall_f1_all_models.csv\")\nensemble_classification_report_text = classification_report(ensemble_true_labels, ensemble_predicted_labels, labels=list_of_class_indices,\n                                                            target_names=stage_names_in_order, digits=3, zero_division=0)\nwith open(os.path.join(results_folder_paths[\"15_evaluation\"], \"classification_report_ensemble.txt\"), \"w\") as report_file:\n    report_file.write(ensemble_classification_report_text)\n\nfigure_object, axis_array = plt.subplots(1, 2, figsize=(20, 5))\nper_class_report_dataframe.pivot(index=\"stage\", columns=\"model\", values=\"f1\").reindex(stage_names_in_order).plot(kind=\"bar\", ax=axis_array[0])\naxis_array[0].set_title(\"F1 per DR Grade, All Models (test set)\")\naxis_array[0].set_ylim(0, 1.05)\naxis_array[0].legend(title=\"Model\", fontsize=8)\nensemble_rows = per_class_report_dataframe[per_class_report_dataframe[\"model\"] == \"ensemble (soft vote)\"].set_index(\"stage\").reindex(stage_names_in_order)\nensemble_rows[[\"precision\", \"recall\", \"f1\"]].plot(kind=\"bar\", ax=axis_array[1])\nfor bar_container in axis_array[1].containers:\n    axis_array[1].bar_label(bar_container, fmt=\"%.2f\", fontsize=7)\naxis_array[1].set_title(\"Ensemble: Precision, Recall and F1 per DR Grade (test set)\")\naxis_array[1].set_ylim(0, 1.1)\nfor current_axis in axis_array:\n    current_axis.tick_params(axis=\"x\", rotation=15)\n    current_axis.set_xlabel(\"DR grade\")\nsave_current_figure(\"15_evaluation\", \"per_grade_scores.png\")\n\nprint(\"Cell 55 finished: per-grade scores saved for all models.\")\nprint(per_class_report_dataframe.pivot(index=\"stage\", columns=\"model\", values=\"recall\").reindex(stage_names_in_order).to_string())\nprint()\nprint(ensemble_classification_report_text)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:42.87955Z","iopub.execute_input":"2026-09-26T15:16:42.879855Z","iopub.status.idle":"2026-09-26T15:16:43.961844Z","shell.execute_reply.started":"2026-09-26T15:16:42.879829Z","shell.execute_reply":"2026-09-26T15:16:43.961105Z"}},"outputs":[],"execution_count":null},{"id":"d78c4ce3","cell_type":"code","source":"# Cell 56: confusion matrices for all four models, as raw counts and as percentages of each true grade\nfor matrix_style, value_format, file_name in [(\"counts\", \"d\", \"confusion_matrices_counts.png\"), (\"row_percent\", \".2f\", \"confusion_matrices_row_normalised.png\")]:\n    figure_object, axis_grid = plt.subplots(2, 2, figsize=(14, 12))\n    for plot_index, (model_name, (true_labels_array, predicted_labels_array, _)) in enumerate(all_model_predictions_by_name.items()):\n        current_axis = axis_grid[plot_index // 2, plot_index % 2]\n        matrix_values = confusion_matrix(true_labels_array, predicted_labels_array, labels=list_of_class_indices, normalize=\"true\" if matrix_style == \"row_percent\" else None)\n        seaborn_plotting_library.heatmap(matrix_values, annot=True, fmt=value_format, cmap=\"Blues\", vmin=0, vmax=1 if matrix_style == \"row_percent\" else None,\n                                         xticklabels=stage_names_in_order, yticklabels=stage_names_in_order, ax=current_axis)\n        current_axis.set_title(f\"{model_name} (QWK = {cohen_kappa_score(true_labels_array, predicted_labels_array, weights='quadratic'):.3f})\")\n        current_axis.set_xlabel(\"Predicted grade\")\n        current_axis.set_ylabel(\"True grade\")\n    save_current_figure(\"15_evaluation\", file_name)\nensemble_confusion_matrix_dataframe = pd.DataFrame(confusion_matrix(ensemble_true_labels, ensemble_predicted_labels, labels=list_of_class_indices),\n                                                   index=[f\"true {stage}\" for stage in stage_names_in_order], columns=[f\"predicted {stage}\" for stage in stage_names_in_order])\nsave_results_table(ensemble_confusion_matrix_dataframe, \"15_evaluation\", \"confusion_matrix_ensemble_counts.csv\", keep_index=True)\n\nprint(\"Cell 56 finished: confusion matrices saved as counts and as row percentages.\")\nprint(ensemble_confusion_matrix_dataframe.to_string())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:43.962806Z","iopub.execute_input":"2026-09-26T15:16:43.963582Z","iopub.status.idle":"2026-09-26T15:16:48.295744Z","shell.execute_reply.started":"2026-09-26T15:16:43.963556Z","shell.execute_reply":"2026-09-26T15:16:48.294667Z"}},"outputs":[],"execution_count":null},{"id":"87ca476f","cell_type":"code","source":"# Cell 57: screening view, turning the 5 grades into two yes or no questions a clinic actually asks\ndef compute_binary_screening_metrics(true_binary_labels, predicted_binary_labels, positive_class_scores):\n    true_positive_count = int(np.sum((true_binary_labels == 1) & (predicted_binary_labels == 1)))\n    false_negative_count = int(np.sum((true_binary_labels == 1) & (predicted_binary_labels == 0)))\n    true_negative_count = int(np.sum((true_binary_labels == 0) & (predicted_binary_labels == 0)))\n    false_positive_count = int(np.sum((true_binary_labels == 0) & (predicted_binary_labels == 1)))\n    safe_ratio = lambda numerator, denominator: numerator / denominator if denominator > 0 else float(\"nan\")\n    return {\n        \"true_positives\": true_positive_count, \"false_negatives\": false_negative_count, \"true_negatives\": true_negative_count, \"false_positives\": false_positive_count,\n        \"sensitivity\": safe_ratio(true_positive_count, true_positive_count + false_negative_count),\n        \"specificity\": safe_ratio(true_negative_count, true_negative_count + false_positive_count),\n        \"positive_predictive_value\": safe_ratio(true_positive_count, true_positive_count + false_positive_count),\n        \"negative_predictive_value\": safe_ratio(true_negative_count, true_negative_count + false_negative_count),\n        \"roc_auc\": roc_auc_score(true_binary_labels, positive_class_scores),\n    }\n\n\nscreening_metric_rows = []\nfor model_name, (true_labels_array, predicted_labels_array, predicted_probability_array) in all_model_predictions_by_name.items():\n    for screening_question_name, grade_threshold in {\"referable_dr_grade_2_or_above\": 2, \"any_dr_grade_1_or_above\": 1}.items():\n        screening_metric_rows.append({\"model\": model_name, \"screening_question\": screening_question_name, **compute_binary_screening_metrics(\n            (true_labels_array >= grade_threshold).astype(int), (predicted_labels_array >= grade_threshold).astype(int), predicted_probability_array[:, grade_threshold:].sum(axis=1))})\nscreening_metrics_dataframe = pd.DataFrame(screening_metric_rows).round(4)\nsave_results_table(screening_metrics_dataframe, \"15_evaluation\", \"screening_metrics_referable_and_any_dr.csv\")\n\nfigure_object, axis_object = plt.subplots(figsize=(7, 6))\nfor model_name, (true_labels_array, _, predicted_probability_array) in all_model_predictions_by_name.items():\n    true_binary_labels = (true_labels_array >= 2).astype(int)\n    referable_scores = predicted_probability_array[:, 2:].sum(axis=1)\n    false_positive_rates, true_positive_rates, _ = roc_curve(true_binary_labels, referable_scores)\n    axis_object.plot(false_positive_rates, true_positive_rates, label=f\"{model_name} (AUC = {roc_auc_score(true_binary_labels, referable_scores):.3f})\")\naxis_object.plot([0, 1], [0, 1], linestyle=\"--\", color=\"grey\", label=\"Chance\")\naxis_object.set_xlabel(\"False positive rate (1 minus specificity)\")\naxis_object.set_ylabel(\"True positive rate (sensitivity)\")\naxis_object.set_title(\"ROC: Referable DR (Moderate or worse) vs Not Referable\")\naxis_object.legend(fontsize=8, loc=\"lower right\")\nsave_current_figure(\"15_evaluation\", \"roc_referable_dr_all_models.png\")\n\nprint(\"Cell 57 finished: screening metrics and referable DR ROC curves saved.\")\nprint(screening_metrics_dataframe[[\"model\", \"screening_question\", \"sensitivity\", \"specificity\", \"positive_predictive_value\", \"negative_predictive_value\", \"roc_auc\"]].to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:48.297099Z","iopub.execute_input":"2026-09-26T15:16:48.297473Z","iopub.status.idle":"2026-09-26T15:16:48.821229Z","shell.execute_reply.started":"2026-09-26T15:16:48.297445Z","shell.execute_reply":"2026-09-26T15:16:48.820607Z"}},"outputs":[],"execution_count":null},{"id":"ac89df80","cell_type":"code","source":"# Cell 58: how often do the three models agree, and how accurate is the ensemble when they do and do not?\npredictions_by_architecture = {name: per_architecture_test_results[name][\"test_predicted_labels\"] for name in architecture_names_to_compare}\npairwise_agreement_rates = {\n    f\"{first_name} vs {second_name}\": float((predictions_by_architecture[first_name] == predictions_by_architecture[second_name]).mean())\n    for first_index, first_name in enumerate(architecture_names_to_compare) for second_name in architecture_names_to_compare[first_index + 1:]\n}\nall_three_agree_mask = (predictions_by_architecture[\"efficientnet_b3\"] == predictions_by_architecture[\"efficientnet_b0\"]) & (predictions_by_architecture[\"efficientnet_b0\"] == predictions_by_architecture[\"resnet50\"])\nensemble_answer_is_correct = ensemble_predicted_labels == ensemble_true_labels\naccuracy_when_models_agree = float(ensemble_answer_is_correct[all_three_agree_mask].mean())\naccuracy_when_models_disagree = float(ensemble_answer_is_correct[~all_three_agree_mask].mean()) if (~all_three_agree_mask).any() else float(\"nan\")\nmodel_agreement_dataframe = pd.DataFrame(\n    [{\"comparison\": pair_name, \"agreement_percent\": round(rate * 100, 1)} for pair_name, rate in pairwise_agreement_rates.items()]\n    + [{\"comparison\": \"all three models\", \"agreement_percent\": round(float(all_three_agree_mask.mean()) * 100, 1)}]\n)\nsave_results_table(model_agreement_dataframe, \"15_evaluation\", \"model_agreement.csv\")\n\nprint(\"Cell 58 finished: model agreement measured.\")\nprint(model_agreement_dataframe.to_string(index=False))\nprint(f\"Ensemble accuracy when all three agree: {accuracy_when_models_agree:.3f} ({int(all_three_agree_mask.sum())} photos)\")\nprint(f\"Ensemble accuracy when they disagree:   {accuracy_when_models_disagree:.3f} ({int((~all_three_agree_mask).sum())} photos)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:48.822027Z","iopub.execute_input":"2026-09-26T15:16:48.822445Z","iopub.status.idle":"2026-09-26T15:16:48.834759Z","shell.execute_reply.started":"2026-09-26T15:16:48.82242Z","shell.execute_reply":"2026-09-26T15:16:48.833905Z"}},"outputs":[],"execution_count":null},{"id":"3981fe88","cell_type":"code","source":"# Cell 59: McNemar's test, is the ensemble really better than the best single model (chosen on validation) or just lucky?\ndef run_mcnemar_test_on_paired_predictions(first_model_correct, second_model_correct):\n    first_right_second_wrong_count = int(np.sum(first_model_correct & ~second_model_correct))\n    second_right_first_wrong_count = int(np.sum(~first_model_correct & second_model_correct))\n    total_disagreements = first_right_second_wrong_count + second_right_first_wrong_count\n    if total_disagreements == 0:\n        return first_right_second_wrong_count, second_right_first_wrong_count, 0.0, 1.0\n    chi_squared_statistic = (abs(first_right_second_wrong_count - second_right_first_wrong_count) - 1) ** 2 / total_disagreements\n    return first_right_second_wrong_count, second_right_first_wrong_count, chi_squared_statistic, float(chi2.sf(chi_squared_statistic, df=1))\n\n\nbest_single_model_correct = per_architecture_test_results[best_single_architecture_name][\"test_predicted_labels\"] == ensemble_true_labels\nensemble_right_single_wrong_count, single_right_ensemble_wrong_count, mcnemar_statistic, mcnemar_p_value = run_mcnemar_test_on_paired_predictions(\n    ensemble_answer_is_correct, best_single_model_correct)\nensemble_mcnemar_summary = {\n    \"best_single_model_by_validation_qwk\": best_single_architecture_name,\n    \"ensemble_right_single_model_wrong\": ensemble_right_single_wrong_count,\n    \"single_model_right_ensemble_wrong\": single_right_ensemble_wrong_count,\n    \"chi_squared_statistic\": round(float(mcnemar_statistic), 4),\n    \"p_value\": round(float(mcnemar_p_value), 4),\n    \"significant_at_0_05\": bool(mcnemar_p_value < 0.05),\n}\nsave_results_json(ensemble_mcnemar_summary, \"15_evaluation\", \"mcnemar_ensemble_vs_best_single.json\")\n\nprint(\"Cell 59 finished: McNemar's test done.\")\nfor summary_key, summary_value in ensemble_mcnemar_summary.items():\n    print(f\"  {summary_key}: {summary_value}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:48.836302Z","iopub.execute_input":"2026-09-26T15:16:48.836657Z","iopub.status.idle":"2026-09-26T15:16:48.860182Z","shell.execute_reply.started":"2026-09-26T15:16:48.83662Z","shell.execute_reply":"2026-09-26T15:16:48.85957Z"}},"outputs":[],"execution_count":null},{"id":"af4a775b","cell_type":"code","source":"# Cell 60: everything about QWK in one place, from training through validation to test, plus how QWK scores mistakes\nqwk_detail_rows = []\nfor training_result in all_training_results:\n    architecture_name = training_result[\"architecture_name\"]\n    headline_row = headline_metrics_dataframe.set_index(\"model\").loc[architecture_name]\n    qwk_detail_rows.append({\n        \"model\": architecture_name, \"kept_epoch\": training_result[\"best_epoch\"],\n        \"train_qwk_at_kept_epoch\": round(training_result[\"train_qwk_at_best_epoch\"], 4),\n        \"best_validation_qwk\": round(training_result[\"best_validation_qwk\"], 4),\n        \"train_minus_validation_gap\": round(training_result[\"train_qwk_at_best_epoch\"] - training_result[\"best_validation_qwk\"], 4),\n        \"test_qwk\": headline_row[\"qwk\"], \"test_qwk_ci_low\": headline_row[\"qwk_ci_low\"], \"test_qwk_ci_high\": headline_row[\"qwk_ci_high\"],\n        \"validation_to_test_change\": round(headline_row[\"qwk\"] - training_result[\"best_validation_qwk\"], 4),\n    })\nensemble_headline_row = headline_metrics_dataframe.set_index(\"model\").loc[\"ensemble (soft vote)\"]\nqwk_detail_rows.append({\"model\": \"ensemble (soft vote)\", \"test_qwk\": ensemble_headline_row[\"qwk\"],\n                        \"test_qwk_ci_low\": ensemble_headline_row[\"qwk_ci_low\"], \"test_qwk_ci_high\": ensemble_headline_row[\"qwk_ci_high\"]})\nqwk_detail_dataframe = pd.DataFrame(qwk_detail_rows)\nsave_results_table(qwk_detail_dataframe, \"15_evaluation\", \"qwk_detail_table.csv\")\n\nnumber_of_grades = len(list_of_class_indices)\nquadratic_penalty_weights = np.array([[(row - column) ** 2 / (number_of_grades - 1) ** 2 for column in list_of_class_indices] for row in list_of_class_indices])\nfigure_object, axis_array = plt.subplots(1, 2, figsize=(16, 5.5))\nseaborn_plotting_library.heatmap(quadratic_penalty_weights, annot=True, fmt=\".2f\", cmap=\"Reds\", xticklabels=stage_names_in_order, yticklabels=stage_names_in_order, ax=axis_array[0])\naxis_array[0].set_title(\"How QWK Penalises Mistakes (0 = no penalty, 1 = worst)\")\naxis_array[0].set_xlabel(\"Predicted grade\")\naxis_array[0].set_ylabel(\"True grade\")\nmodel_positions = np.arange(len(qwk_detail_dataframe))\naxis_array[1].bar(model_positions - 0.27, qwk_detail_dataframe[\"train_qwk_at_kept_epoch\"].fillna(0), 0.27, label=\"Training QWK at kept epoch\")\naxis_array[1].bar(model_positions, qwk_detail_dataframe[\"best_validation_qwk\"].fillna(0), 0.27, label=\"Best validation QWK\")\naxis_array[1].bar(model_positions + 0.27, qwk_detail_dataframe[\"test_qwk\"], 0.27,\n                  yerr=[qwk_detail_dataframe[\"test_qwk\"] - qwk_detail_dataframe[\"test_qwk_ci_low\"], qwk_detail_dataframe[\"test_qwk_ci_high\"] - qwk_detail_dataframe[\"test_qwk\"]],\n                  capsize=3, label=\"Test QWK (95% CI)\")\naxis_array[1].set_xticks(model_positions)\naxis_array[1].set_xticklabels(qwk_detail_dataframe[\"model\"], rotation=10)\naxis_array[1].set_ylim(0.6, 1.02)\naxis_array[1].set_title(\"QWK from Training to Validation to Test\")\naxis_array[1].legend(fontsize=8, loc=\"lower right\")\nsave_current_figure(\"15_evaluation\", \"qwk_explained_and_train_val_test.png\")\n\nprint(\"Cell 60 finished: QWK detail table and chart saved.\")\nprint(qwk_detail_dataframe.to_string(index=False))\nprint(\"A one-grade mistake costs 0.06 of the worst-case penalty, while mixing up No DR and Proliferative costs 1.0.\")\nprint(\"That is why QWK suits ordered grades better than plain accuracy.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:48.861119Z","iopub.execute_input":"2026-09-26T15:16:48.861352Z","iopub.status.idle":"2026-09-26T15:16:49.812325Z","shell.execute_reply.started":"2026-09-26T15:16:48.861332Z","shell.execute_reply":"2026-09-26T15:16:49.811465Z"}},"outputs":[],"execution_count":null},{"id":"57020874","cell_type":"code","source":"# Cell 61: should the app use the ensemble or the best single model? Put all the evidence for that decision in one place\nheadline_by_model = headline_metrics_dataframe.set_index(\"model\")\nlatency_by_model = inference_latency_dataframe.set_index(\"model\")\nensemble_qwk_gain = float(headline_by_model.loc[\"ensemble (soft vote)\", \"qwk\"] - headline_by_model.loc[best_single_architecture_name, \"qwk\"])\nconfidence_intervals_overlap = bool(\n    headline_by_model.loc[\"ensemble (soft vote)\", \"qwk_ci_low\"] <= headline_by_model.loc[best_single_architecture_name, \"qwk_ci_high\"])\ndisagreement_signal_strength = accuracy_when_models_agree - accuracy_when_models_disagree\nif ensemble_mcnemar_summary[\"significant_at_0_05\"] and ensemble_qwk_gain > 0:\n    ensemble_decision = \"use_ensemble_because_it_is_significantly_better\"\nelif disagreement_signal_strength > 0.1:\n    ensemble_decision = \"use_ensemble_for_its_disagreement_warning_not_for_accuracy\"\nelse:\n    ensemble_decision = \"use_best_single_model\"\nensemble_decision_summary = {\n    \"best_single_model_by_validation\": best_single_architecture_name,\n    \"test_qwk_gain_of_ensemble\": round(ensemble_qwk_gain, 4),\n    \"qwk_confidence_intervals_overlap\": confidence_intervals_overlap,\n    \"mcnemar_p_value\": ensemble_mcnemar_summary[\"p_value\"],\n    \"ensemble_accuracy_when_models_agree\": round(accuracy_when_models_agree, 3),\n    \"ensemble_accuracy_when_models_disagree\": round(accuracy_when_models_disagree, 3),\n    \"cpu_milliseconds_best_single\": float(latency_by_model.loc[best_single_architecture_name, \"cpu_milliseconds_for_one_photo\"]),\n    \"cpu_milliseconds_ensemble\": float(latency_by_model.loc[\"ensemble (all three)\", \"cpu_milliseconds_for_one_photo\"]),\n    \"decision\": ensemble_decision,\n}\nsave_results_json(ensemble_decision_summary, \"15_evaluation\", \"ensemble_vs_single_model_decision.json\")\n\nprint(\"Cell 61 finished: ensemble decision evidence saved.\")\nfor summary_key, summary_value in ensemble_decision_summary.items():\n    print(f\"  {summary_key}: {summary_value}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:49.813541Z","iopub.execute_input":"2026-09-26T15:16:49.813821Z","iopub.status.idle":"2026-09-26T15:16:49.823837Z","shell.execute_reply.started":"2026-09-26T15:16:49.813796Z","shell.execute_reply":"2026-09-26T15:16:49.823274Z"}},"outputs":[],"execution_count":null},{"id":"08ef3054","cell_type":"markdown","source":"## 13. Where Is the Model Looking? (Grad-CAM)\nGrad-CAM draws a heat map over the photo. Red areas pushed the model towards its answer the most, so a doctor can check whether it looked at real lesions or at something irrelevant. It is shown for the best single model (picked on validation QWK), because the ensemble is an average of three models and has no single set of layers to look inside.","metadata":{}},{"id":"8dc6901b","cell_type":"code","source":"# Cell 62: build Grad-CAM with hooks on the last convolution block of the best single model\nclass GradientWeightedClassActivationMapping:\n    def __init__(self, model, target_convolutional_layer):\n        self.model = model\n        self.captured_activations = None\n        self.captured_gradients = None\n        target_convolutional_layer.register_forward_hook(self._store_activations)\n        target_convolutional_layer.register_full_backward_hook(self._store_gradients)\n\n    def _store_activations(self, module, layer_input, layer_output):\n        self.captured_activations = layer_output.detach()\n\n    def _store_gradients(self, module, grad_input, grad_output):\n        self.captured_gradients = grad_output[0].detach()\n\n    def generate_heatmap_for_image(self, input_image_tensor, target_class_index=None):\n        self.model.eval()\n        predicted_logits = self.model(input_image_tensor.unsqueeze(0).to(compute_device))\n        if target_class_index is None:\n            target_class_index = predicted_logits.argmax(dim=1).item()\n        self.model.zero_grad()\n        predicted_logits[0, target_class_index].backward()\n        channel_importance_weights = self.captured_gradients.mean(dim=(2, 3), keepdim=True)\n        heatmap_as_numpy = torch.relu((channel_importance_weights * self.captured_activations).sum(dim=1).squeeze()).cpu().numpy()\n        if heatmap_as_numpy.max() > 0:\n            heatmap_as_numpy = heatmap_as_numpy / heatmap_as_numpy.max()\n        return heatmap_as_numpy, target_class_index\n\n\ndef overlay_heatmap_on_image(original_image_rgb, heatmap_array):\n    resized_heatmap = cv2.resize(heatmap_array, (original_image_rgb.shape[1], original_image_rgb.shape[0]))\n    heatmap_as_colour_rgb = cv2.cvtColor(cv2.applyColorMap(np.uint8(255 * resized_heatmap), cv2.COLORMAP_JET), cv2.COLOR_BGR2RGB)\n    return cv2.addWeighted(original_image_rgb, 0.6, heatmap_as_colour_rgb, 0.4, 0)\n\n\ndef load_final_preprocessed_photo(image_id_code):\n    return preprocess_fundus_image(build_image_path_from_id(image_id_code), project_configuration[\"target_image_size\"], final_training_recipe[\"preprocessing_method\"])\n\n\nbest_single_model_for_gradcam = load_trained_model(best_single_architecture_name, compute_device)\ngrad_cam_target_layer = best_single_model_for_gradcam.layer4[-1] if best_single_architecture_name == \"resnet50\" else best_single_model_for_gradcam.features[-1]\ngrad_cam_generator = GradientWeightedClassActivationMapping(best_single_model_for_gradcam, grad_cam_target_layer)\n\nprint(\"Cell 62 finished: Grad-CAM ready on\", best_single_architecture_name, \"using its last convolution block.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:49.824862Z","iopub.execute_input":"2026-09-26T15:16:49.825837Z","iopub.status.idle":"2026-09-26T15:16:50.326655Z","shell.execute_reply.started":"2026-09-26T15:16:49.825803Z","shell.execute_reply":"2026-09-26T15:16:50.325971Z"}},"outputs":[],"execution_count":null},{"id":"f7e9aa7d","cell_type":"code","source":"# Cell 63: Grad-CAM heat maps for one test photo of each grade\nfigure_object, axis_grid = plt.subplots(len(list_of_class_indices), 2, figsize=(7.5, len(list_of_class_indices) * 3.4), squeeze=False)\nfor class_label in list_of_class_indices:\n    chosen_row = test_split_in_loader_order[test_split_in_loader_order[\"diagnosis\"] == class_label].iloc[0]\n    preprocessed_image = load_final_preprocessed_photo(chosen_row[\"id_code\"])\n    generated_heatmap, predicted_class_index = grad_cam_generator.generate_heatmap_for_image(evaluation_time_image_transform(preprocessed_image))\n    axis_grid[class_label, 0].imshow(preprocessed_image)\n    axis_grid[class_label, 0].set_title(f\"True grade: {diagnosis_stage_names[class_label]}\", fontsize=10)\n    axis_grid[class_label, 1].imshow(overlay_heatmap_on_image(preprocessed_image, generated_heatmap))\n    axis_grid[class_label, 1].set_title(f\"Grad-CAM, predicted: {diagnosis_stage_names[predicted_class_index]}\", fontsize=10)\n    axis_grid[class_label, 0].axis(\"off\")\n    axis_grid[class_label, 1].axis(\"off\")\nplt.suptitle(f\"Grad-CAM ({best_single_architecture_name}): Where the Model Looks\", y=1.0)\nsave_current_figure(\"16_gradcam\", \"gradcam_one_photo_per_grade.png\")\n\nprint(\"Cell 63 finished: Grad-CAM heat maps saved for one test photo of each grade.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:50.327512Z","iopub.execute_input":"2026-09-26T15:16:50.327829Z","iopub.status.idle":"2026-09-26T15:16:54.06153Z","shell.execute_reply.started":"2026-09-26T15:16:50.327788Z","shell.execute_reply":"2026-09-26T15:16:54.059127Z"}},"outputs":[],"execution_count":null},{"id":"39c4a673","cell_type":"markdown","source":"## 14. Error Analysis\nHow far off are the wrong answers, which grades get mixed up, does the ensemble know when it is unsure, and what was the model looking at when it got things badly wrong?","metadata":{}},{"id":"74d04433","cell_type":"code","source":"# Cell 64: how far off are the mistakes, and which grade pairs get mixed up most often\nerror_distance_rows = []\nfor model_name, (true_labels_array, predicted_labels_array, _) in all_model_predictions_by_name.items():\n    absolute_grade_distance = np.abs(true_labels_array - predicted_labels_array)\n    number_of_errors = int((absolute_grade_distance > 0).sum())\n    error_distance_rows.append({\n        \"model\": model_name,\n        \"exact_grade_percent\": round(float((absolute_grade_distance == 0).mean() * 100), 1),\n        \"one_grade_off_percent\": round(float((absolute_grade_distance == 1).mean() * 100), 1),\n        \"two_or_more_grades_off_percent\": round(float((absolute_grade_distance >= 2).mean() * 100), 1),\n        \"share_of_mistakes_one_grade_off_percent\": round(float((absolute_grade_distance == 1).sum() / number_of_errors * 100), 1) if number_of_errors > 0 else 0.0,\n    })\nerror_distance_dataframe = pd.DataFrame(error_distance_rows)\nsave_results_table(error_distance_dataframe, \"17_error_analysis\", \"error_distance_all_models.csv\")\n\nensemble_mistake_pairs = pd.DataFrame({\"true_grade\": ensemble_true_labels, \"predicted_grade\": ensemble_predicted_labels})\nensemble_mistake_pairs = ensemble_mistake_pairs[ensemble_mistake_pairs[\"true_grade\"] != ensemble_mistake_pairs[\"predicted_grade\"]]\nmost_common_confusions_dataframe = ensemble_mistake_pairs.value_counts().rename(\"number_of_photos\").reset_index()\nmost_common_confusions_dataframe[\"true_stage\"] = most_common_confusions_dataframe[\"true_grade\"].map(diagnosis_stage_names)\nmost_common_confusions_dataframe[\"predicted_stage\"] = most_common_confusions_dataframe[\"predicted_grade\"].map(diagnosis_stage_names)\nmost_common_confusions_dataframe[\"share_of_all_mistakes_percent\"] = (most_common_confusions_dataframe[\"number_of_photos\"] / max(len(ensemble_mistake_pairs), 1) * 100).round(1)\nmost_common_confusions_dataframe[\"direction\"] = np.where(most_common_confusions_dataframe[\"predicted_grade\"] > most_common_confusions_dataframe[\"true_grade\"], \"over-graded\", \"under-graded\")\nsave_results_table(most_common_confusions_dataframe, \"17_error_analysis\", \"most_common_confusions_ensemble.csv\")\n\nfigure_object, axis_object = plt.subplots(figsize=(10, 4.5))\nerror_distance_dataframe.set_index(\"model\")[[\"exact_grade_percent\", \"one_grade_off_percent\", \"two_or_more_grades_off_percent\"]].plot(\n    kind=\"barh\", stacked=True, ax=axis_object, color=[\"seagreen\", \"orange\", \"firebrick\"])\naxis_object.set_xlabel(\"Percentage of test photos\")\naxis_object.set_ylabel(\"\")\naxis_object.set_title(\"How Far Off Are the Predictions? (test set)\")\naxis_object.legend([\"Exact grade\", \"One grade off\", \"Two or more grades off\"], loc=\"lower right\", fontsize=8)\nsave_current_figure(\"17_error_analysis\", \"error_distance_all_models.png\")\n\nprint(\"Cell 64 finished: mistake distance and most common confusions saved.\")\nprint(error_distance_dataframe.to_string(index=False))\nprint()\nprint(most_common_confusions_dataframe[[\"true_stage\", \"predicted_stage\", \"number_of_photos\", \"share_of_all_mistakes_percent\", \"direction\"]].head(8).to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:54.063218Z","iopub.execute_input":"2026-09-26T15:16:54.063816Z","iopub.status.idle":"2026-09-26T15:16:54.451359Z","shell.execute_reply.started":"2026-09-26T15:16:54.063783Z","shell.execute_reply":"2026-09-26T15:16:54.450732Z"}},"outputs":[],"execution_count":null},{"id":"b5eceef5","cell_type":"code","source":"# Cell 65: does the ensemble know when it is unsure? Compare its confidence on right and wrong answers\nlow_confidence_threshold = 0.6\nensemble_top_probability = ensemble_averaged_probabilities.max(axis=1)\nphoto_is_flagged_for_review = ensemble_top_probability < low_confidence_threshold\nconfidence_band_labels = [\"under 50%\", \"50 to 60%\", \"60 to 70%\", \"70 to 80%\", \"80 to 90%\", \"90% or more\"]\nconfidence_band_series = pd.cut(ensemble_top_probability, bins=[0.0, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0001], labels=confidence_band_labels, right=False)\nconfidence_band_dataframe = pd.DataFrame({\"band\": confidence_band_series, \"correct\": ensemble_answer_is_correct}).groupby(\"band\", observed=False).agg(\n    photos=(\"correct\", \"size\"), accuracy=(\"correct\", \"mean\")).reset_index()\nconfidence_band_dataframe[\"share_of_test_photos_percent\"] = (confidence_band_dataframe[\"photos\"] / len(ensemble_top_probability) * 100).round(1)\nconfidence_band_dataframe[\"accuracy\"] = confidence_band_dataframe[\"accuracy\"].round(3)\nsave_results_table(confidence_band_dataframe, \"17_error_analysis\", \"ensemble_accuracy_by_confidence.csv\")\n\nconfidence_summary = {\n    \"low_confidence_threshold_used_by_app\": low_confidence_threshold,\n    \"average_confidence_when_right\": round(float(ensemble_top_probability[ensemble_answer_is_correct].mean()), 3),\n    \"average_confidence_when_wrong\": round(float(ensemble_top_probability[~ensemble_answer_is_correct].mean()), 3) if (~ensemble_answer_is_correct).any() else None,\n    \"share_of_photos_flagged_percent\": round(float(photo_is_flagged_for_review.mean() * 100), 1),\n    \"accuracy_on_flagged_photos\": round(float(ensemble_answer_is_correct[photo_is_flagged_for_review].mean()), 3) if photo_is_flagged_for_review.any() else None,\n    \"accuracy_on_unflagged_photos\": round(float(ensemble_answer_is_correct[~photo_is_flagged_for_review].mean()), 3) if (~photo_is_flagged_for_review).any() else None,\n    \"share_of_mistakes_caught_by_confidence_flag_percent\": round(float(photo_is_flagged_for_review[~ensemble_answer_is_correct].mean() * 100), 1) if (~ensemble_answer_is_correct).any() else None,\n    \"share_of_mistakes_caught_by_confidence_or_disagreement_percent\": round(float((photo_is_flagged_for_review | ~all_three_agree_mask)[~ensemble_answer_is_correct].mean() * 100), 1) if (~ensemble_answer_is_correct).any() else None,\n}\nsave_results_json(confidence_summary, \"17_error_analysis\", \"ensemble_confidence_summary.json\")\n\nfigure_object, axis_array = plt.subplots(1, 2, figsize=(15, 4.8))\nhistogram_bins = np.linspace(0.2, 1.0, 17)\naxis_array[0].hist(ensemble_top_probability[ensemble_answer_is_correct], bins=histogram_bins, alpha=0.7, color=\"seagreen\", label=\"Right answers\")\naxis_array[0].hist(ensemble_top_probability[~ensemble_answer_is_correct], bins=histogram_bins, alpha=0.7, color=\"firebrick\", label=\"Wrong answers\")\naxis_array[0].axvline(low_confidence_threshold, color=\"black\", linestyle=\"--\", label=f\"Review flag below {low_confidence_threshold:.0%}\")\naxis_array[0].set_title(\"Ensemble Confidence on Right vs Wrong Answers\")\naxis_array[0].set_xlabel(\"Probability of the chosen grade\")\naxis_array[0].set_ylabel(\"Number of test photos\")\naxis_array[0].legend(fontsize=8)\nconfidence_bars = axis_array[1].bar(confidence_band_dataframe[\"band\"].astype(str), confidence_band_dataframe[\"accuracy\"].fillna(0), color=\"steelblue\")\naxis_array[1].bar_label(confidence_bars, labels=[f\"{accuracy:.0%}\\n(n={photos})\" for accuracy, photos in zip(confidence_band_dataframe[\"accuracy\"].fillna(0), confidence_band_dataframe[\"photos\"])], fontsize=8)\naxis_array[1].set_ylim(0, 1.15)\naxis_array[1].set_title(\"Ensemble Accuracy by Confidence Band\")\naxis_array[1].set_xlabel(\"Confidence band\")\nsave_current_figure(\"17_error_analysis\", \"ensemble_confidence_right_vs_wrong.png\")\n\nprint(\"Cell 65 finished: confidence analysis saved.\")\nprint(confidence_band_dataframe.to_string(index=False))\nfor summary_key, summary_value in confidence_summary.items():\n    print(f\"  {summary_key}: {summary_value}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:54.45237Z","iopub.execute_input":"2026-09-26T15:16:54.452748Z","iopub.status.idle":"2026-09-26T15:16:55.273454Z","shell.execute_reply.started":"2026-09-26T15:16:54.452724Z","shell.execute_reply":"2026-09-26T15:16:55.272789Z"}},"outputs":[],"execution_count":null},{"id":"bfe2c184","cell_type":"code","source":"# Cell 66: save every test photo's prediction from every model into one table for the report\nper_image_predictions_dataframe = test_split_in_loader_order.rename(columns={\"diagnosis\": \"true_grade\"})\nassert np.array_equal(per_image_predictions_dataframe[\"true_grade\"].values, ensemble_true_labels), \"Test loader order does not match the test table.\"\nfor architecture_name in architecture_names_to_compare:\n    per_image_predictions_dataframe[f\"{architecture_name}_predicted_grade\"] = per_architecture_test_results[architecture_name][\"test_predicted_labels\"]\nper_image_predictions_dataframe[\"ensemble_predicted_grade\"] = ensemble_predicted_labels\nfor class_index in list_of_class_indices:\n    per_image_predictions_dataframe[\"ensemble_probability_\" + diagnosis_stage_names[class_index].lower().replace(\" \", \"_\")] = ensemble_averaged_probabilities[:, class_index].round(4)\nper_image_predictions_dataframe[\"ensemble_confidence\"] = ensemble_top_probability.round(4)\nper_image_predictions_dataframe[\"flagged_low_confidence\"] = photo_is_flagged_for_review\nper_image_predictions_dataframe[\"all_three_models_agree\"] = all_three_agree_mask\nper_image_predictions_dataframe[\"ensemble_is_correct\"] = ensemble_answer_is_correct\nsave_results_table(per_image_predictions_dataframe, \"17_error_analysis\", \"test_set_predictions_per_photo.csv\")\n\nprint(\"Cell 66 finished: per-photo predictions saved for\", len(per_image_predictions_dataframe), \"test photos.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:55.274474Z","iopub.execute_input":"2026-09-26T15:16:55.274821Z","iopub.status.idle":"2026-09-26T15:16:55.293628Z","shell.execute_reply.started":"2026-09-26T15:16:55.274796Z","shell.execute_reply":"2026-09-26T15:16:55.292724Z"}},"outputs":[],"execution_count":null},{"id":"cd5c9cd9","cell_type":"code","source":"# Cell 67: Grad-CAM on the test photos the best single model got most wrong\nbest_single_model_predicted_labels = per_architecture_test_results[best_single_architecture_name][\"test_predicted_labels\"]\ngrade_distance_of_each_prediction = np.abs(ensemble_true_labels - best_single_model_predicted_labels)\nmisclassified_test_positions = np.where(grade_distance_of_each_prediction > 0)[0]\nmisclassified_positions_largest_first = misclassified_test_positions[np.argsort(-grade_distance_of_each_prediction[misclassified_test_positions], kind=\"stable\")]\nnumber_of_mistakes_to_show = min(4, len(misclassified_positions_largest_first))\nif number_of_mistakes_to_show == 0:\n    print(\"Cell 67 finished: the best single model made no mistakes on the test set, nothing to show.\")\nelse:\n    figure_object, axis_grid = plt.subplots(number_of_mistakes_to_show, 2, figsize=(7.5, number_of_mistakes_to_show * 3.4), squeeze=False)\n    for row_index, test_position in enumerate(misclassified_positions_largest_first[:number_of_mistakes_to_show]):\n        chosen_row = test_split_in_loader_order.iloc[test_position]\n        preprocessed_image = load_final_preprocessed_photo(chosen_row[\"id_code\"])\n        generated_heatmap, predicted_class_index = grad_cam_generator.generate_heatmap_for_image(evaluation_time_image_transform(preprocessed_image))\n        axis_grid[row_index, 0].imshow(preprocessed_image)\n        axis_grid[row_index, 0].set_title(f\"True grade: {diagnosis_stage_names[int(chosen_row['diagnosis'])]}\", fontsize=10)\n        axis_grid[row_index, 1].imshow(overlay_heatmap_on_image(preprocessed_image, generated_heatmap))\n        axis_grid[row_index, 1].set_title(f\"Grad-CAM, predicted: {diagnosis_stage_names[predicted_class_index]}\", fontsize=10)\n        axis_grid[row_index, 0].axis(\"off\")\n        axis_grid[row_index, 1].axis(\"off\")\n    plt.suptitle(f\"Grad-CAM on the Biggest Mistakes ({best_single_architecture_name})\", y=1.0)\n    save_current_figure(\"17_error_analysis\", \"gradcam_biggest_mistakes.png\")\n    print(\"Cell 67 finished: Grad-CAM saved for the\", number_of_mistakes_to_show, \"biggest mistakes.\")\n    print(\"Mistakes by the best single model:\", len(misclassified_test_positions), \"| two or more grades off:\", int((grade_distance_of_each_prediction >= 2).sum()))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:55.294742Z","iopub.execute_input":"2026-09-26T15:16:55.295149Z","iopub.status.idle":"2026-09-26T15:16:58.052508Z","shell.execute_reply.started":"2026-09-26T15:16:55.295118Z","shell.execute_reply":"2026-09-26T15:16:58.051651Z"}},"outputs":[],"execution_count":null},{"id":"bc02f383","cell_type":"markdown","source":"## 15. Final Summary and Web App Export\nAll key numbers go into one summary file for the report. The three models and every setting the app needs, including the chosen preprocessing method, are packed for the Streamlit app, and a final check makes sure the app's copy gives the same answers as the notebook.\n\nDownload from the Kaggle output: `figures/` and `results/` for the report, `app_export/` into the app folder, and `demo_images/` for the video (keep those outside the app folder).","metadata":{}},{"id":"5c94a7b3","cell_type":"code","source":"# Cell 68: gather every key number from every section into one summary file for the report\nreferable_screening_row = screening_metrics_dataframe[(screening_metrics_dataframe[\"model\"] == \"ensemble (soft vote)\")\n                                                      & (screening_metrics_dataframe[\"screening_question\"] == \"referable_dr_grade_2_or_above\")].iloc[0]\nconsolidated_results_summary = {\n    \"run_metadata\": {\"run_finished_at\": datetime.now().isoformat(timespec=\"minutes\"), \"torch_version\": torch.__version__,\n                     \"torchvision_version\": torchvision.__version__, \"gpu_name\": torch.cuda.get_device_name(0) if is_gpu_available else \"cpu\", \"random_seed\": RANDOM_SEED},\n    \"dataset\": {\"total_photos\": int(len(diabetic_retinopathy_labels_dataframe)), \"imbalance_ratio\": round(float(imbalance_ratio), 1),\n                \"split_sizes\": {name: int(len(frame)) for name, frame in split_dataframes_by_name.items()}},\n    \"preprocessing_tests\": preprocessing_test_summary_dataframe.round(1).to_dict(orient=\"records\"),\n    \"local_contrast_by_method\": contrast_summary_dataframe.round(3).to_dict(orient=\"records\"),\n    \"image_size_study\": image_size_study_summary,\n    \"experiment_decisions\": decision_table_dataframe.to_dict(orient=\"records\"),\n    \"final_training_recipe\": final_training_recipe,\n    \"overfitting_gaps\": overfitting_gap_dataframe.to_dict(orient=\"records\"),\n    \"qwk_detail\": qwk_detail_dataframe.to_dict(orient=\"records\"),\n    \"headline_metrics\": headline_metrics_dataframe.set_index(\"model\").to_dict(orient=\"index\"),\n    \"screening_metrics\": screening_metrics_dataframe.to_dict(orient=\"records\"),\n    \"mcnemar_ensemble_vs_best_single\": ensemble_mcnemar_summary,\n    \"ensemble_decision\": ensemble_decision_summary,\n    \"error_distance\": error_distance_dataframe.set_index(\"model\").to_dict(orient=\"index\"),\n    \"confidence_analysis\": confidence_summary,\n}\nsave_results_json(consolidated_results_summary, \"18_app_export\", \"consolidated_results_summary.json\")\n\nprint(\"Cell 68 finished: consolidated summary saved.\")\nprint(\"Sections in the file:\", list(consolidated_results_summary.keys()))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:58.053625Z","iopub.execute_input":"2026-09-26T15:16:58.054365Z","iopub.status.idle":"2026-09-26T15:16:58.07109Z","shell.execute_reply.started":"2026-09-26T15:16:58.054339Z","shell.execute_reply":"2026-09-26T15:16:58.070295Z"}},"outputs":[],"execution_count":null},{"id":"0f40c084","cell_type":"code","source":"# Cell 69: pack the three models and every setting the web app needs into the app_export folder\napp_export_models_folder_path = os.path.join(app_export_folder_path, \"models\")\nos.makedirs(app_export_models_folder_path, exist_ok=True)\nexported_model_rows = []\nfor architecture_name in architecture_names_to_compare:\n    full_precision_state = torch.load(build_main_checkpoint_path(architecture_name), map_location=\"cpu\")\n    export_file_name = f\"{architecture_name}_fp16.pth\"\n    export_file_path = os.path.join(app_export_models_folder_path, export_file_name)\n    torch.save({key: (value.half() if value.is_floating_point() else value) for key, value in full_precision_state.items()}, export_file_path)\n    exported_model_rows.append({\"architecture\": architecture_name, \"exported_file\": export_file_name,\n                                \"original_size_mb\": round(os.path.getsize(build_main_checkpoint_path(architecture_name)) / 1e6, 1),\n                                \"exported_size_mb\": round(os.path.getsize(export_file_path) / 1e6, 1)})\n\napp_configuration = {\n    \"model_version\": f\"aptos2019-ensemble-{datetime.now():%Y%m%d}\",\n    \"created_on\": datetime.now().isoformat(timespec=\"minutes\"),\n    \"dataset\": \"APTOS 2019 Blindness Detection (Kaggle)\",\n    \"target_image_size\": project_configuration[\"target_image_size\"],\n    \"preprocessing_method\": final_training_recipe[\"preprocessing_method\"],\n    \"dark_pixel_brightness_threshold\": project_configuration[\"dark_pixel_brightness_threshold\"],\n    \"clahe_clip_limit\": project_configuration[\"clahe_clip_limit\"],\n    \"clahe_tile_grid_size\": project_configuration[\"clahe_tile_grid_size\"],\n    \"ben_graham_blur_divisor\": project_configuration[\"ben_graham_blur_divisor\"],\n    \"normalisation_mean\": imagenet_normalisation_mean,\n    \"normalisation_std\": imagenet_normalisation_std,\n    \"stage_names\": stage_names_in_order,\n    \"referable_grade_threshold\": 2,\n    \"low_confidence_threshold\": low_confidence_threshold,\n    \"architectures\": [{\"name\": row[\"architecture\"], \"weights_file\": row[\"exported_file\"]} for row in exported_model_rows],\n    \"grad_cam_architecture\": best_single_architecture_name,\n    \"ensemble_decision\": ensemble_decision,\n    \"test_set_performance\": {\n        \"test_photos\": int(len(test_split_dataframe)),\n        \"ensemble_accuracy\": round(float(ensemble_test_accuracy), 4),\n        \"ensemble_qwk\": round(float(ensemble_test_qwk), 4),\n        \"ensemble_macro_f1\": round(float(headline_by_model.loc[\"ensemble (soft vote)\", \"macro_f1\"]), 4),\n        \"referable_dr_sensitivity\": round(float(referable_screening_row[\"sensitivity\"]), 4),\n        \"referable_dr_specificity\": round(float(referable_screening_row[\"specificity\"]), 4),\n        \"referable_dr_auc\": round(float(referable_screening_row[\"roc_auc\"]), 4),\n    },\n    \"training_environment\": {\"torch\": torch.__version__, \"torchvision\": torchvision.__version__},\n}\nwith open(os.path.join(app_export_folder_path, \"app_config.json\"), \"w\") as app_config_file:\n    json_module.dump(app_configuration, app_config_file, indent=2, default=convert_numpy_values_for_json)\nexported_model_dataframe = pd.DataFrame(exported_model_rows)\nsave_results_table(exported_model_dataframe, \"18_app_export\", \"exported_app_models.csv\")\n\nprint(\"Cell 69 finished: app bundle written to\", app_export_folder_path)\nprint(exported_model_dataframe.to_string(index=False))\nprint(\"Preprocessing method written to the app settings:\", app_configuration[\"preprocessing_method\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:58.072391Z","iopub.execute_input":"2026-09-26T15:16:58.072725Z","iopub.status.idle":"2026-09-26T15:16:58.488347Z","shell.execute_reply.started":"2026-09-26T15:16:58.072688Z","shell.execute_reply":"2026-09-26T15:16:58.487492Z"}},"outputs":[],"execution_count":null},{"id":"ea2f35c7","cell_type":"code","source":"# Cell 70: rebuild the web app's prediction steps using only the exported files, on the CPU like the app\ndef load_exported_models_for_cpu(exported_configuration, models_folder_path):\n    loaded_models = {}\n    for architecture_entry in exported_configuration[\"architectures\"]:\n        model = build_transfer_learning_model(architecture_entry[\"name\"], len(exported_configuration[\"stage_names\"]), final_training_recipe[\"head_dropout\"])\n        half_precision_state = torch.load(os.path.join(models_folder_path, architecture_entry[\"weights_file\"]), map_location=\"cpu\")\n        model.load_state_dict({key: (value.float() if value.is_floating_point() else value) for key, value in half_precision_state.items()})\n        loaded_models[architecture_entry[\"name\"]] = model.eval()\n    return loaded_models\n\n\ndef predict_stage_with_ensemble_from_raw_image(image_rgb, loaded_models, exported_configuration):\n    preprocessed_image = run_preprocessing_steps(image_rgb, exported_configuration[\"target_image_size\"], exported_configuration[\"preprocessing_method\"])[\"resized\"]\n    input_tensor = evaluation_time_image_transform(preprocessed_image).unsqueeze(0)\n    with torch.no_grad():\n        per_model_probabilities = {name: torch.softmax(model(input_tensor), dim=1)[0].numpy() for name, model in loaded_models.items()}\n    averaged_probabilities = np.mean(list(per_model_probabilities.values()), axis=0)\n    return {\"predicted_stage_index\": int(averaged_probabilities.argmax()), \"averaged_probabilities\": averaged_probabilities}\n\n\nexported_cpu_models = load_exported_models_for_cpu(app_configuration, app_export_models_folder_path)\nprint(\"Cell 70 finished: exported models loaded on the CPU and the app's prediction function is ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:58.489522Z","iopub.execute_input":"2026-09-26T15:16:58.490208Z","iopub.status.idle":"2026-09-26T15:16:59.465484Z","shell.execute_reply.started":"2026-09-26T15:16:58.490181Z","shell.execute_reply":"2026-09-26T15:16:59.464595Z"}},"outputs":[],"execution_count":null},{"id":"598e0850","cell_type":"code","source":"# Cell 71: check the app copy agrees with the notebook on 50 test photos, and save one demo photo per grade for the video\nparity_check_positions = []\nfor class_label in list_of_class_indices:\n    parity_check_positions.extend(np.where(ensemble_true_labels == class_label)[0][:10].tolist())\nparity_rows = []\nfor test_position in parity_check_positions:\n    test_row = test_split_in_loader_order.iloc[test_position]\n    app_prediction = predict_stage_with_ensemble_from_raw_image(load_image_as_rgb(build_image_path_from_id(test_row[\"id_code\"])), exported_cpu_models, app_configuration)\n    parity_rows.append({\"id_code\": test_row[\"id_code\"], \"true_grade\": int(test_row[\"diagnosis\"]),\n                        \"notebook_prediction\": int(ensemble_predicted_labels[test_position]), \"app_prediction\": app_prediction[\"predicted_stage_index\"],\n                        \"largest_probability_difference\": float(np.abs(app_prediction[\"averaged_probabilities\"] - ensemble_averaged_probabilities[test_position]).max())})\nparity_dataframe = pd.DataFrame(parity_rows)\nparity_dataframe[\"same_grade\"] = parity_dataframe[\"notebook_prediction\"] == parity_dataframe[\"app_prediction\"]\nsave_results_table(parity_dataframe.round(5), \"18_app_export\", \"app_vs_notebook_parity_check.csv\")\n\ndemo_image_rows = []\nfor class_label in list_of_class_indices:\n    grade_rows = per_image_predictions_dataframe[(per_image_predictions_dataframe[\"true_grade\"] == class_label) & per_image_predictions_dataframe[\"ensemble_is_correct\"]]\n    if len(grade_rows) == 0:\n        grade_rows = per_image_predictions_dataframe[per_image_predictions_dataframe[\"true_grade\"] == class_label]\n    chosen_demo_row = grade_rows.iloc[0]\n    demo_file_name = f\"demo_grade{class_label}_{diagnosis_stage_names[class_label].lower().replace(' ', '_')}_{chosen_demo_row['id_code']}.png\"\n    Image.fromarray(load_image_as_rgb(build_image_path_from_id(chosen_demo_row[\"id_code\"]))).save(os.path.join(demo_images_folder_path, demo_file_name))\n    demo_image_rows.append({\"file_name\": demo_file_name, \"true_grade\": diagnosis_stage_names[class_label],\n                            \"ensemble_prediction\": diagnosis_stage_names[int(chosen_demo_row[\"ensemble_predicted_grade\"])],\n                            \"ensemble_confidence\": float(chosen_demo_row[\"ensemble_confidence\"])})\nsave_results_table(pd.DataFrame(demo_image_rows), \"18_app_export\", \"demo_images_for_video.csv\")\n\nprint(\"Cell 71 finished: app parity check and demo photos saved.\")\nprint(f\"Same grade from the app copy and the notebook: {parity_dataframe['same_grade'].mean():.1%} of {len(parity_dataframe)} photos\")\nprint(f\"Largest probability difference: {parity_dataframe['largest_probability_difference'].max():.4f}\")\nprint(\"Tiny differences come from half-precision weights and CPU vs GPU maths. A grade can only flip when two grades were almost tied.\")\nprint(pd.DataFrame(demo_image_rows).to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:16:59.466609Z","iopub.execute_input":"2026-09-26T15:16:59.466977Z","iopub.status.idle":"2026-09-26T15:17:42.102229Z","shell.execute_reply.started":"2026-09-26T15:16:59.466953Z","shell.execute_reply":"2026-09-26T15:17:42.101492Z"}},"outputs":[],"execution_count":null},{"id":"747185a2","cell_type":"code","source":"# Cell 72: list every saved figure, table, model and app file so they are easy to find for the report\nfor output_folder_name in [\"figures\", \"results\", \"models\", \"app_export\", \"demo_images\"]:\n    output_folder_path = os.path.join(output_root_directory, output_folder_name)\n    print(f\"\\n{output_folder_path}\")\n    for current_directory, _, file_names in sorted(os.walk(output_folder_path)):\n        for file_name in sorted(file_names):\n            file_path = os.path.join(current_directory, file_name)\n            print(f\"  {os.path.relpath(file_path, output_folder_path)}  ({os.path.getsize(file_path) / 1e6:.2f} MB)\")\nprint(\"\\nCell 72 finished: every saved output listed above. Photo caches stayed in /tmp, so they are not in the saved output.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-26T15:17:42.103267Z","iopub.execute_input":"2026-09-26T15:17:42.103562Z","iopub.status.idle":"2026-09-26T15:17:42.11427Z","shell.execute_reply.started":"2026-09-26T15:17:42.103537Z","shell.execute_reply":"2026-09-26T15:17:42.113499Z"}},"outputs":[],"execution_count":null}]}