{"cells":[{"cell_type":"markdown","id":"f545629c","metadata":{},"source":"# APTOS: Three Predictions I Wrote Before the Data\n\nA diabetic retinopathy grader on APTOS-2019 reaches a five-fold ensemble QWK of **0.909**. Public solutions reach 0.93, so the score is not why this notebook exists.\n\nWhat is worth showing is what happened when every claim about that model was tested, and each prediction was committed to GitHub **before** the data it concerned was touched:\n\n1. A shortcut exists: file metadata alone predicts the grade. **Recomputed live below on this dataset.**\n2. On **IDRiD**, where that shortcut cannot work, I predicted the model would collapse if it depended on it. It did not - the prediction failed in the model's favour.\n3. On **Messidor-2**, graded by a panel of retina specialists, I predicted referable AUC of at least 0.93. It came out **0.819** - the prediction failed against the model.\n4. A fine-tuning experiment with a **control arm** then traced that failure to the training labels, not the camera.\n\nFull project: [GitHub](https://github.com/senanurcetin/APTOS-2019-diabetic-retinopathy) | [live demo](https://aptos-2019-diabetic-retinopathy.onrender.com) | [model card](https://huggingface.co/senanurcetin/aptos-retinopathy-grader) | [write-up](https://github.com/senanurcetin/APTOS-2019-diabetic-retinopathy/blob/main/docs/WRITEUP.md)\n\nThe images are the training set of the [APTOS 2019 Blindness Detection](https://www.kaggle.com/competitions/aptos2019-blindness-detection) competition, used here through a copy pre-split into train, validation and test ([mariaherrerot/aptos2019](https://www.kaggle.com/datasets/mariaherrerot/aptos2019)).\n\n*A research project, not a medical device.*"},{"cell_type":"code","execution_count":null,"id":"c70ddd3d","metadata":{},"outputs":[],"source":"import glob\nimport io\nimport json\nimport os\nimport urllib.request\n\nimport cv2\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nfrom joblib import Parallel, delayed\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.model_selection import StratifiedKFold, cross_val_predict\n\n# Results of the GPU experiments are read from the repository at the v1.0.0\n# release tag, so every number below matches the published reports exactly.\nRELEASE = \"https://raw.githubusercontent.com/senanurcetin/APTOS-2019-diabetic-retinopathy/v1.0.0/reports/\"\n\ndef report(name):\n    with urllib.request.urlopen(RELEASE + name, timeout=30) as response:\n        data = response.read().decode(\"utf-8\")\n    return json.loads(data) if name.endswith(\".json\") else pd.read_csv(io.StringIO(data))\n\ncalibration = report(\"calibration.json\")\nfinetune = report(\"finetune_messidor2.json\")\nmessidor = report(\"external_validation_messidor2.json\")\nruns = report(\"runs.csv\")\nprint(\"reports loaded from the v1.0.0 release\")"},{"cell_type":"markdown","id":"fe7a85be","metadata":{},"source":"## The whole story in one chart\n\nReferable DR (grade 2 or above) is the screening decision, and ROC AUC measures how well the model ranks eyes for it, independent of any threshold."},{"cell_type":"code","execution_count":null,"id":"349b1d93","metadata":{},"outputs":[],"source":"applied = {row[\"set\"]: row for row in calibration[\"applied\"]}\nbars = [\n    (\"APTOS test\\n(in-distribution)\", applied[\"APTOS test\"][\"ROC AUC\"], \"#4c72b0\"),\n    (\"IDRiD\\n(external)\", applied[\"IDRiD\"][\"ROC AUC\"], \"#4c72b0\"),\n    (\"Messidor-2\\n(external)\", applied[\"Messidor-2\"][\"ROC AUC\"], \"#c44e52\"),\n    (\"Messidor-2 held-out half\\nafter fine-tuning\\non specialist labels\", finetune[\"adjudicated\"][\"Messidor-2 test half\"][\"referable_auc\"], \"#55a868\"),\n    (\"Messidor-2 held-out half\\ncontrol: same images,\\nmodel's own labels\", finetune[\"control\"][\"Messidor-2 test half\"][\"referable_auc\"], \"#8c8c8c\"),\n]\nfig, ax = plt.subplots(figsize=(11, 5.2))\nlabels, values, colours = zip(*bars)\npositions = range(len(bars))\nax.bar(positions, values, color=colours, width=0.62)\nfor x, v in zip(positions, values):\n    ax.text(x, v + 0.005, f\"{v:.3f}\", ha=\"center\", va=\"bottom\", fontsize=12, fontweight=\"bold\")\nax.axhline(0.93, ls=\"--\", color=\"#333333\", lw=1)\nax.text(1.62, 0.947, \"pre-registered prediction for Messidor-2: AUC >= 0.93\", ha=\"left\", va=\"bottom\", fontsize=9, color=\"#333333\")\nax.set_xticks(list(positions))\nax.set_xticklabels(labels, fontsize=9.5)\nax.set_ylim(0.75, 1.0)\nax.set_ylabel(\"referable ROC AUC\")\nax.set_title(\"A retinopathy model that reads the retina - through its training labels\", fontsize=13, loc=\"left\")\nfor side in (\"top\", \"right\"):\n    ax.spines[side].set_visible(False)\nplt.tight_layout()\nplt.show()"},{"cell_type":"markdown","id":"85335798","metadata":{},"source":"## 1. The shortcut, recomputed on this dataset\n\nAPTOS was collected at several sites with different cameras, and the camera correlates with disease. If that is true, a model that never sees a retinal pixel should still predict the grade.\n\nThe cell below reads every image in this Kaggle dataset, extracts **only file properties** - width, height, aspect ratio, megapixels, file size, mean brightness and contrast - and cross-validates a random forest on them. Nothing here is loaded from the repository."},{"cell_type":"code","execution_count":null,"id":"7373e4b8","metadata":{},"outputs":[],"source":"# The mirror below splits this competition's 3,662 training images into\n# train/valid/test. Search two levels deep only: the competition data is also\n# mounted, and a recursive walk would visit every one of its images.\nroots = (glob.glob(\"/kaggle/input/*/train_1.csv\") + glob.glob(\"/kaggle/input/*/*/train_1.csv\")\n         or glob.glob(\"/kaggle/input/**/train_1.csv\", recursive=True))\nROOT = os.path.dirname(roots[0]) if roots else os.environ.get(\"APTOS_DATA\", \"data/images\")\n\nfolders = {\"train_1.csv\": \"train_images/train_images\",\n           \"valid.csv\": \"val_images/val_images\",\n           \"test.csv\": \"test_images/test_images\"}\nlabels = pd.concat([\n    pd.read_csv(os.path.join(ROOT, name)).dropna(subset=[\"id_code\"]).assign(folder=folder)\n    for name, folder in folders.items()\n], ignore_index=True)\nlabels[\"diagnosis\"] = labels[\"diagnosis\"].astype(int)\nprint(f\"{len(labels)} labelled images\")\n\ndef metadata(row):\n    path = os.path.join(ROOT, row.folder, f\"{row.id_code}.png\")\n    gray = cv2.imread(path, cv2.IMREAD_GRAYSCALE)\n    h, w = gray.shape\n    return {\"id_code\": row.id_code, \"width\": w, \"height\": h,\n            \"aspect_ratio\": w / h, \"megapixels\": w * h / 1e6,\n            \"file_kb\": os.path.getsize(path) / 1024,\n            \"brightness\": float(gray.mean()), \"contrast_std\": float(gray.std())}\n\nmeta = pd.DataFrame(Parallel(n_jobs=-1)(delayed(metadata)(row) for row in labels.itertuples()))\ndata = labels.merge(meta, on=\"id_code\")\nprint(f\"{data[['width', 'height']].drop_duplicates().shape[0]} distinct resolutions\")"},{"cell_type":"code","execution_count":null,"id":"d90178ce","metadata":{},"outputs":[],"source":"features = [\"width\", \"height\", \"aspect_ratio\", \"megapixels\", \"file_kb\", \"brightness\", \"contrast_std\"]\nfolds = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)\npred = cross_val_predict(RandomForestClassifier(n_estimators=200, random_state=0, n_jobs=-1),\n                         data[features], data[\"diagnosis\"], cv=folds)\nshortcut_qwk = cohen_kappa_score(data[\"diagnosis\"], pred, weights=\"quadratic\")\nprint(f\"QWK from file metadata alone, 5-fold CV: {shortcut_qwk:.3f}\")\nprint(f\"QWK from always predicting 'No DR':      0.000\")\nprint(\"(the repository reports 0.652 with its own fold split)\")"},{"cell_type":"code","execution_count":null,"id":"2d3d3a18","metadata":{},"outputs":[],"source":"square = (data[\"width\"] == 1050) & (data[\"height\"] == 1050)\ntable = pd.DataFrame({\n    \"images\": [square.sum(), (~square).sum()],\n    \"share graded No DR\": [(data.loc[square, \"diagnosis\"] == 0).mean(),\n                           (data.loc[~square, \"diagnosis\"] == 0).mean()],\n}, index=[\"1050x1050 (most common resolution)\", \"every other resolution\"])\ntable.style.format({\"share graded No DR\": \"{:.1%}\"})"},{"cell_type":"markdown","id":"f5a3a13f","metadata":{},"source":"Most images at the most common resolution are healthy; elsewhere the mix is very different. Resizing every image to 512px does not remove this, because aspect ratio and sharpness survive it. **A QWK of 0.9 therefore has a floor of roughly 0.65 that requires no medicine at all** - and every later question becomes: how much of the score is the retina?\n\n## 2. Cross-validation, and two ideas that did nothing\n\nThe model is an EfficientNet-B0 ordinal regressor with validation-fitted thresholds, trained in five folds. Two preprocessing ideas with good arguments behind them - CLAHE contrast enhancement, and squashing images to square instead of padding - were compared on paired folds:"},{"cell_type":"code","execution_count":null,"id":"a528681c","metadata":{},"outputs":[],"source":"cv = runs[(runs[\"split\"] == \"cv-fold\") & (runs[\"stratify\"] == \"diagnosis\")]\nper_fold = cv.pivot_table(index=\"fold\", columns=\"variant\", values=\"test_qwk\")[[\"baseline\", \"squash\", \"clahe\"]]\nsummary = runs[(runs[\"split\"] == \"cv-summary\") & (runs[\"stratify\"] == \"diagnosis\")].set_index(\"variant\")\npd.DataFrame({\n    \"mean test QWK per fold\": per_fold.mean(),\n    \"five-fold ensemble QWK\": summary[\"ensemble_qwk\"],\n}).round(4)"},{"cell_type":"markdown","id":"8a218474","metadata":{},"source":"Both differences are within fold-to-fold noise (paired t-test p = 0.47 for squash, p = 0.87 for CLAHE). The largest effect in the whole project is ensembling the five folds: about +0.019 QWK.\n\n## 3. External test 1 - IDRiD: a prediction that failed in the model's favour\n\nIDRiD (455 images from another Indian clinic) was captured at **one resolution**, so the metadata shortcut carries no information there. Better still, it inverts it: in APTOS every image at IDRiD's resolution is diseased; in IDRiD 28% are healthy.\n\n**Written before the run:** if the model leans on acquisition cues, it will call IDRiD's healthy eyes diseased and specificity will collapse."},{"cell_type":"code","execution_count":null,"id":"5e98f01c","metadata":{},"outputs":[],"source":"rows = [\"APTOS test\", \"IDRiD\"]\npd.DataFrame({name: {\"referable ROC AUC\": applied[name][\"ROC AUC\"],\n                     \"sensitivity at the APTOS cut\": applied[name][\"sensitivity\"],\n                     \"specificity at the APTOS cut\": applied[name][\"specificity\"]}\n              for name in rows}).round(3)"},{"cell_type":"markdown","id":"11552357","metadata":{},"source":"Specificity went *up*. On the dataset built to expose the shortcut, the model ranked patients as well as at home. **It reads the retina.** What drifted was calibration: a cut chosen for 90% sensitivity on APTOS gave 82% on IDRiD.\n\n## 4. External test 2 - Messidor-2: a prediction that failed against the model\n\nOne external set can be luck. Messidor-2 (France, 1744 images) adds something new: every image was graded by a **panel of three retina specialists** who adjudicated disagreements (Krause et al., 2018). APTOS labels come from single graders, and duplicate images in APTOS disagree with themselves 29% of the time.\n\n**Written before the run:** referable AUC >= 0.93."},{"cell_type":"code","execution_count":null,"id":"c9676b47","metadata":{},"outputs":[],"source":"print(f\"Messidor-2 referable ROC AUC: {applied['Messidor-2']['ROC AUC']:.3f}\")\nprint(f\"Sensitivity at the APTOS cut: {applied['Messidor-2']['sensitivity']:.3f}\")\nrecall = pd.Series(messidor[\"per_class_recall\"]).astype(float)\nrecall.index = [\"0 No DR\", \"1 Mild\", \"2 Moderate\", \"3 Severe\", \"4 Proliferative\"]\nrecall.round(3).to_frame(\"recall on Messidor-2\")"},{"cell_type":"markdown","id":"8b7019a2","metadata":{},"source":"The prediction failed. The loss sits at **grade 2, Moderate** - the referral boundary: most Moderate eyes are graded below 2. In the repository's analysis, Moderate eyes with hard exudates were caught, and those defined by subtler signs were scored like Mild.\n\nThe reading after the fact: the model learned where APTOS's single graders draw the Mild/Moderate line, and the specialist panel draws it lower. That is a hypothesis, so it got its own experiment.\n\n## 5. Testing the cause, with a control arm\n\nMessidor-2 was split in half **by patient**, so no one's two eyes landed on both sides. The five fold models were fine-tuned on one half and evaluated on the other.\n\nThe obvious objection: fine-tuning also shows the model Messidor-2's cameras, and that alone might help. So a **control** arm was fine-tuned on the same images with the original model's own grades as labels - new cameras, old boundary."},{"cell_type":"code","execution_count":null,"id":"e7f9d6d5","metadata":{},"outputs":[],"source":"arms = {\"unchanged\": \"none\", \"specialist labels\": \"adjudicated\", \"control (model's own labels)\": \"control\"}\npd.DataFrame({\n    label: {\"Messidor-2 held-out AUC\": finetune[key][\"Messidor-2 test half\"][\"referable_auc\"],\n            \"Moderate eyes graded below 2\": finetune[key][\"Messidor-2 test half\"][\"moderate_below_2\"],\n            \"IDRiD AUC\": finetune[key][\"IDRiD\"][\"referable_auc\"],\n            \"APTOS test AUC\": finetune[key][\"APTOS test\"][\"referable_auc\"],\n            \"APTOS test accuracy\": finetune[key][\"APTOS test\"][\"accuracy\"]}\n    for label, key in arms.items()\n}).round(3)"},{"cell_type":"markdown","id":"ee621f6d","metadata":{},"source":"The control gains nothing; the specialist labels move the boundary. **The failure was the training labels, not the camera.**\n\nTwo things did not go the way the hypothesis said, and both are reported in the repository:\n- The whole Moderate grade shifted up, rather than the subtle cases specifically - the model redrew a line; there is no sign it learned to see anything new.\n- The fine-tuned model over-grades APTOS (accuracy falls), so it stays an experiment and the deployed model is unchanged.\n\n## What I take from it\n\n- **A benchmark score is a claim, not a result.** A floor of 0.65 was sitting under the 0.9.\n- **External data answers questions internal splits cannot.** IDRiD settled the shortcut question in one run.\n- **Models inherit their labels.** The ceiling here is label quality, not architecture.\n- **Writing the prediction down first** is what made the failures informative instead of embarrassing.\n\nEverything - code, the pre-registered predictions with their commit timestamps, 177 tests, and the serving stack - is in the [repository](https://github.com/senanurcetin/APTOS-2019-diabetic-retinopathy). You can try the deployed model in the [live demo](https://aptos-2019-diabetic-retinopathy.onrender.com).\n\n*Data: APTOS 2019 Blindness Detection; IDRiD; Messidor-2 with adjudicated grades published by Google (Krause et al., 2018). Not a medical device.*"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}