{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"3823b17e","cell_type":"markdown","source":"# EDL + Graham (final_model_v2, Plain variant) -- Kaggle run against the original raw 2015 Diabetic Retinopathy Detection dataset\n\nThis notebook runs `final_model_v2`'s **Plain** variant (`train.py` / `config.yaml`) against the *original* raw 2015 Kaggle DR competition data -- heavily class-imbalanced by nature (grade 0 ~73%, grade 4 ~2%), unlike the artificially-balanced dr-88gb-augmented dataset this same model family was also run against. `final_model_v2` itself is never modified -- everything dataset-specific lives in this fork (`final_model_v2_2015`) and this notebook. See `REVISIONAL_EDITS.md` in the fork for the full diagnosis: archive-mode dataset support (this competition's data is served on Kaggle as split zip archives, not a pre-extracted directory), patient/eye leakage prevention, and softened class-balance sampling.\n\n**Before running**: attach the \"Diabetic Retinopathy Detection\" competition data to this notebook (Add Input -> Competitions), alongside this fork's code dataset.","metadata":{}},{"id":"0df641ee","cell_type":"code","source":"import os\nfrom pathlib import Path\n\n# Folders named '0'..'4' would be per-class image folders on a class-per-folder\n# dataset -- not this one (flat <id>.jpeg files inside a zip), but pruning them\n# out generically costs nothing and matches every other notebook this session.\nPRUNE_DIR_NAMES = {\"0\", \"1\", \"2\", \"3\", \"4\"}\n\n\ndef bounded_walk(root):\n    for dirpath, dirnames, filenames in os.walk(root):\n        yield dirpath, dirnames, filenames\n        dirnames[:] = [d for d in dirnames if d not in PRUNE_DIR_NAMES]\n\n\nprint(\"Top-level of /kaggle/input:\")\nfor p in sorted(Path(\"/kaggle/input\").iterdir()):\n    print(\" \", p)","metadata":{},"outputs":[],"execution_count":null},{"id":"16acdf67","cell_type":"markdown","source":"## Locate and copy the `final_model_v2` code package into `/kaggle/working/`","metadata":{}},{"id":"fa4a4e96","cell_type":"code","source":"import shutil\nfrom pathlib import Path\n\n# Content-based discovery (not by dataset/folder name) -- robust to whatever\n# name you gave the uploaded code dataset.\ndef find_package_root():\n    for dirpath, dirnames, filenames in bounded_walk(\"/kaggle/input\"):\n        if \"train.py\" in filenames and \"config.yaml\" in filenames:\n            return Path(dirpath)\n    return None\n\npkg_root = find_package_root()\nassert pkg_root is not None, (\n    \"Could not find a folder containing both train.py and config.yaml under \"\n    \"/kaggle/input -- check that the final_model_v2_2015 code dataset is attached.\")\nprint(f\"Found package: {pkg_root}\")\n\nwork_dir = Path(\"/kaggle/working/final_model_v2_2015\")\nif not work_dir.exists():\n    shutil.copytree(pkg_root, work_dir)\n    print(f\"Copied to {work_dir}\")\nelse:\n    print(f\"{work_dir} already exists -- leaving it as-is (delete it first if you \"\n          \"want a completely fresh copy of the code).\")\n\nprint(\"train.py exists:\", (work_dir / \"train.py\").exists())\nprint(\"config.yaml exists:\", (work_dir / \"config.yaml\").exists())","metadata":{},"outputs":[],"execution_count":null},{"id":"ffcdd05c","cell_type":"code","source":"%cd /kaggle/working/final_model_v2_2015\n!ls -la","metadata":{},"outputs":[],"execution_count":null},{"id":"69e0ff8c","cell_type":"markdown","source":"## Resume a previous session (skip this on your very first run)\n\nIf a prior Kaggle session was cut off mid-training, copy its `outputs/` folder back in before re-running -- `training.resume: true` (the default) then continues from the last fully-completed epoch instead of starting over.","metadata":{}},{"id":"6d8b6c61","cell_type":"code","source":"def find_outputs_checkpoints_dirs():\n    found = []\n    for dirpath, dirnames, filenames in bounded_walk(\"/kaggle/input\"):\n        if os.path.basename(dirpath) == \"checkpoints\" and os.path.basename(os.path.dirname(dirpath)) == \"outputs\":\n            found.append(Path(dirpath))\n    return found\n\nprior = find_outputs_checkpoints_dirs()\nif prior:\n    print(\"Found previous-session checkpoints under /kaggle/input:\")\n    for p in prior:\n        print(\" \", p)\n    print(\"\\nIf one of these is your prior run, copy its parent outputs/ folder into \")\n    print(\"/kaggle/working/final_model_v2_2015/outputs/ before running the training cell below.\")\nelse:\n    print(\"No previous-session outputs/ found under /kaggle/input -- starting fresh \"\n          \"(this is normal on a first run).\")","metadata":{},"outputs":[],"execution_count":null},{"id":"c72c9f9e","cell_type":"code","source":"!pip install -q timm albumentations","metadata":{},"outputs":[],"execution_count":null},{"id":"3ff495e4","cell_type":"markdown","source":"## Confirm the GPU actually works with this session's PyTorch build\n\nA GPU/PyTorch compute-capability mismatch fails LATE and confusingly otherwise -- this fails fast, clearly, before any real work starts.","metadata":{}},{"id":"67841511","cell_type":"code","source":"import torch\n\nprint(\"CUDA available:\", torch.cuda.is_available())\nif torch.cuda.is_available():\n    print(f\"GPU: {torch.cuda.get_device_name(0)} (compute capability \"\n          f\"{'.'.join(map(str, torch.cuda.get_device_capability(0)))})\")\n    try:\n        (torch.randn(4, 4, device=\"cuda\") @ torch.randn(4, 4, device=\"cuda\")).sum().item()\n        print(\"PASS: a trivial CUDA kernel actually ran on this GPU with the installed PyTorch build.\")\n    except RuntimeError as e:\n        raise RuntimeError(\n            \"A trivial CUDA operation failed on this GPU/PyTorch combination -- likely a \"\n            \"compute-capability mismatch. Switch this notebook's Accelerator to GPU T4 x2 \"\n            f\"(Settings -> Accelerator) and re-run. Original error: {e}\") from e\nelse:\n    print(\"WARNING: no GPU detected -- check Settings -> Accelerator. Training will be \"\n          \"extremely slow on CPU for a dataset this size.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"b21b9cd2","cell_type":"markdown","source":"## Check `config.yaml`'s dataset paths and anti-leakage/anti-overfitting settings\n\nUnlike the dr-88gb-augmented notebook, this dataset's path is a well-known, fixed Kaggle competition mount -- no dynamic discovery needed, just verified below. If this assertion fails, attach the \"Diabetic Retinopathy Detection\" competition data (Add Input -> Competitions) rather than editing the path -- the path itself is Kaggle's own standard mount location for it.","metadata":{}},{"id":"f538f17e","cell_type":"code","source":"import sys\nsys.path.insert(0, \"src\")\nfrom utils.config import load_config\n\ncfg = load_config()\nprint(\"raw_images_dir:     \", cfg.paths.raw_images_dir)\nprint(\"raw_images_archive: \", cfg.paths.raw_images_archive)\nprint(\"labels_csv:         \", cfg.paths.labels_csv)\nprint(\"id_col:             \", cfg.data.id_col)\nprint(\"label_col:          \", cfg.data.label_col)\nprint(\"subject_group_regex:\", getattr(cfg.data, \"subject_group_regex\", None),\n      \"(patient/eye leakage guard -- see REVISIONAL_EDITS.md)\")\nprint(\"sampler_balance_power:\", getattr(cfg.data, \"sampler_balance_power\", 1.0),\n      \"(1.0 = full balancing, which overfit on this exact dataset before -- should be < 1.0)\")\nprint(\"run_name:           \", cfg.experiment.run_name)\n\nimport os\nfor name, path in [(\"raw_images_dir\", cfg.paths.raw_images_dir),\n                    (\"raw_images_archive\", cfg.paths.raw_images_archive),\n                    (\"labels_csv\", cfg.paths.labels_csv)]:\n    if path is not None:\n        print(f\"{name} exists: {os.path.exists(path)}\")\n\nassert (cfg.paths.raw_images_dir and os.path.exists(cfg.paths.raw_images_dir)) or \\\n       (cfg.paths.raw_images_archive and os.path.exists(cfg.paths.raw_images_archive)), (\n    \"Neither raw_images_dir nor raw_images_archive exists on disk -- attach the \"\n    \"'Diabetic Retinopathy Detection' competition data to this notebook.\")\nassert os.path.exists(cfg.paths.labels_csv), (\n    f\"labels_csv not found at {cfg.paths.labels_csv} -- check the competition data is attached.\")\nassert getattr(cfg.data, \"subject_group_regex\", None), (\n    \"data.subject_group_regex is not set -- this dataset's ids encode shared patient/eye \"\n    \"identity ('<patient>_left'/'<patient>_right'), so training without this risks patient-\"\n    \"level leakage across splits. Do not proceed without it unless that's genuinely intended.\")\nassert getattr(cfg.data, \"sampler_balance_power\", 1.0) < 1.0, (\n    \"data.sampler_balance_power is 1.0 (full balancing) -- this exact setting overfit almost \"\n    \"immediately on this exact dataset before (see REVISIONAL_EDITS.md). Confirm this is \"\n    \"intentional before proceeding, or lower it (e.g. 0.5).\")\nprint(\"\\nPASS: dataset paths exist, and both the patient-leakage guard and the softened \"\n      \"class-balance sampler are active.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"419c30ef","cell_type":"markdown","source":"## Step 1 -- dedup + patient/eye-grouped split","metadata":{}},{"id":"cba68d19","cell_type":"code","source":"!python src/preprocessing/dedup_and_split.py --config config.yaml","metadata":{},"outputs":[],"execution_count":null},{"id":"9e7e066b","cell_type":"code","source":"# Verify the split is actually patient/eye-disjoint before spending any GPU time --\n# don't just trust the script's own print above, confirm it on the actual output.\nimport pandas as pd\n\nsplits = pd.read_csv(\"outputs/splits/splits.csv\")\nprint(\"splits.csv columns:\", list(splits.columns))\nprint(splits[\"split\"].value_counts())\n\ngroup_col = \"subject_id\" if \"subject_id\" in splits.columns else \"group_id\"\nby_split = {s: set(g[group_col]) for s, g in splits.groupby(\"split\")}\nnames = list(by_split.keys())\noverlaps = {f\"{a} vs {b}\": len(by_split[a] & by_split[b])\n            for i, a in enumerate(names) for b in names[i + 1:]}\nprint(f\"\\n{group_col} overlap across splits (must all be 0):\", overlaps)\nassert all(v == 0 for v in overlaps.values()), (\n    \"Patient/eye leakage detected across splits -- do not proceed to training. \"\n    \"Check config.yaml's data.subject_group_regex and re-run Step 1.\")\nprint(f\"PASS: no {group_col} overlap across train/val/test.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"c023630d","cell_type":"markdown","source":"## Step 2 -- Graham preprocessing (writes preprocessed images to `outputs/processed_images/`)\n\nUnlike the sibling `final_model_kaggle_v2` package (which does Graham preprocessing in-memory per-sample, since it was built around never writing a second copy of this large dataset to disk), `final_model_v2`'s architecture processes every split image ONCE here into local 512x512 PNGs, then reads those flat files at train time. This step reads through the same archive-aware `ImageSource` as Step 1, so it works directly against the competition's zip archives too -- expect this step to take a while over the full ~35k-image dataset.","metadata":{}},{"id":"08b19c1b","cell_type":"code","source":"!python src/preprocessing/graham_preprocess.py --config config.yaml","metadata":{},"outputs":[],"execution_count":null},{"id":"24f4eaaa","cell_type":"markdown","source":"## Step 3 -- train","metadata":{}},{"id":"71a5d735","cell_type":"code","source":"!python train.py --config config.yaml","metadata":{},"outputs":[],"execution_count":null},{"id":"48ded951","cell_type":"markdown","source":"## Check the generated report\n\n`final_model_v2` writes a single JSON report (no figure-generation step) -- check it exists and print the headline test metrics.","metadata":{}},{"id":"6ea07c70","cell_type":"code","source":"import json\nimport os\n\nreport_path = f\"outputs/reports/{cfg.experiment.run_name}_test_report.json\"\nprint(f\"{report_path}: {'FOUND' if os.path.exists(report_path) else 'MISSING'}\")\n\nassert os.path.exists(report_path), (\n    \"Expected test report is missing -- check train.py's log output above for \"\n    \"an earlier failure.\")\n\nwith open(report_path) as f:\n    report = json.load(f)\nprint(\"\\nbest_epoch:   \", report.get(\"best_epoch\"))\nprint(\"best_val_qwk: \", report.get(\"best_val_qwk\"))\nprint(\"TEST metrics: \", report.get(\"prediction_metrics\"))","metadata":{},"outputs":[],"execution_count":null}]}