{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# train_final_model.ipynb — Final YOLOv8 Training (Kaggle)\n\nGreat Barrier Reef COTS Detection Project\n\nWired to Person A's real data pipeline (`dataset.py`). Run cells top to\nbottom, in order — each one prepares something the next one needs.\n\nRunning directly in the competition notebook — no Kaggle API auth or\ndownload needed, data is already mounted at `/kaggle/input/`.\n\nTurn on **Internet** in the Settings panel (right sidebar) so the\n`git clone` and `pip install` cells below can reach GitHub/PyPI.\nAlso set **Accelerator → GPU T4 x2**.\\n\\n---\\n**Updated:** now uses `yolo11s.pt` (was `yolov8n.pt`), trains for up to 100 epochs with early stopping (was a fixed 8), uses both GPUs, fixes an imgsz/optimizer mismatch from the previous run, and adds augmentation tuned for small underwater objects. See `model.py` for details.","metadata":{}},{"cell_type":"markdown","source":"## Step 1 — Get the repo code (`dataset.py`, `model.py`)\n\nClones fresh every run so you always get the latest pushed code.\nChecks out your branch, then merges in `main` locally (inside this\nsession only — nothing gets pushed anywhere) so you have both your\nfiles and anything Person A has pushed to `main`.","metadata":{}},{"cell_type":"code","source":"import sys\n\n!rm -rf /kaggle/working/repo\n!git clone https://github.com/RuzannaMkhitaryan/great-barrier-reef.git /kaggle/working/repo\n!git -C /kaggle/working/repo checkout anahit\n!git -C /kaggle/working/repo merge origin/main --no-edit\n\nsys.path.append('/kaggle/working/repo/src')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:52:23.506829Z","iopub.execute_input":"2026-08-22T07:52:23.507332Z","iopub.status.idle":"2026-08-22T07:52:25.424224Z","shell.execute_reply.started":"2026-08-22T07:52:23.507281Z","shell.execute_reply":"2026-08-22T07:52:25.423340Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2 — Copy `splits.csv` into the working data folder\n\nAdjust the source path below if it moves in the repo — check with\n`!find /kaggle/working/repo -name splits.csv` if this cell errors.","metadata":{}},{"cell_type":"code","source":"!mkdir -p /kaggle/working/data\n!cp /kaggle/working/repo/data/splits/splits.csv /kaggle/working/data/splits.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:52:25.425769Z","iopub.execute_input":"2026-08-22T07:52:25.426121Z","iopub.status.idle":"2026-08-22T07:52:25.667489Z","shell.execute_reply.started":"2026-08-22T07:52:25.426085Z","shell.execute_reply":"2026-08-22T07:52:25.666773Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3 — Install dependencies","metadata":{}},{"cell_type":"code","source":"!pip install -q ultralytics","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:52:25.668648Z","iopub.execute_input":"2026-08-22T07:52:25.668951Z","iopub.status.idle":"2026-08-22T07:52:33.567750Z","shell.execute_reply.started":"2026-08-22T07:52:25.668910Z","shell.execute_reply":"2026-08-22T07:52:33.567013Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 4 — Imports and config","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:52:33.569569Z","iopub.execute_input":"2026-08-22T07:52:33.569940Z","iopub.status.idle":"2026-08-22T07:52:33.691718Z","shell.execute_reply.started":"2026-08-22T07:52:33.569911Z","shell.execute_reply":"2026-08-22T07:52:33.691028Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport glob\nimport shutil\n\nimport dataset\nfrom model import load_model, get_train_config\n\n# ---- CONFIG ----\n# Competition data can be nested (e.g. /kaggle/input/competitions/<slug>/),\n# so search recursively for train.csv instead of assuming a fixed depth.\n_candidates = glob.glob(\"/kaggle/input/**/train.csv\", recursive=True)\n_reef_candidates = [c for c in _candidates if \"barrier\" in c.lower() or \"reef\" in c.lower()]\nif _reef_candidates:\n    _candidates = _reef_candidates\nif not _candidates:\n    raise FileNotFoundError(\"No train.csv under /kaggle/input/**. Run `!ls -R /kaggle/input` to check the Input panel, then set COMP_DIR manually.\")\nCOMP_DIR = os.path.dirname(_candidates[0])\nprint(f\"Using competition data at: {COMP_DIR}\")\n\nTRAIN_CSV = f\"{COMP_DIR}/train.csv\"\nSPLITS_CSV = \"/kaggle/working/data/splits.csv\"\nRAW_IMAGES_ROOT = f\"{COMP_DIR}/train_images\"\n\nLABELS_TRAIN_DIR = \"/kaggle/working/data/labels/train\"\nLABELS_VAL_DIR = \"/kaggle/working/data/labels/val\"\nIMAGES_TRAIN_DIR = \"/kaggle/working/data/images/train\"\nIMAGES_VAL_DIR = \"/kaggle/working/data/images/val\"\nDATA_YAML_PATH = \"/kaggle/working/data/data.yaml\"\n\nOUTPUT_DIR = \"/kaggle/working/great-barrier-reef-checkpoints\"\nRUN_NAME = \"final_model_v3\"  # new name - v2's 8-epoch run and this 60-epoch run aren't comparable/resumable from each other\n\n# Model checkpoint to start from. To try a different YOLO generation, just\n# change this string - load_model() / Ultralytics handle the rest the same\n# way for YOLOv8/v9/v10/YOLO11/YOLO26. Avoid yolo12*.pt (Ultralytics flags\n# it as unstable to train). Bumped from yolov8n.pt (nano, 3M params) up to\n# yolo11s.pt (small, ~9M params) for more capacity on this small-object task.\nMODEL_WEIGHTS = \"yolo11s.pt\"\n\n# Kaggle gave us GPU T4 x2 - actually use both instead of defaulting to one.\nDEVICE = [0, 1]\n\n# Person A's experiment found ratio=all (None) gave the best mAP50, but that\n# was presumably measured with a long, converged training run. At only 8\n# epochs we never converged, so ratio=None (67% background frames in train)\n# mostly just slowed early learning. Start with ratio=2.0 for a faster real\n# convergence check, then try None again once epochs/imgsz are fixed and you\n# have time budget for a longer run.\nNEGATIVE_RATIO = 2.0\n\n# Path Ultralytics writes rolling checkpoints to when save_period is set,\n# and whether we should resume from one. NOTE: /kaggle/working does NOT\n# persist between separate interactive Kaggle sessions - if your session\n# gets cut, you need to \"Save Version\" (or manually download weights/) so\n# last.pt survives, then re-upload it as an input to resume in a new\n# session. Within a single session, this lets you resume after e.g. an\n# accidental interrupt/kernel restart.\nLAST_CHECKPOINT = f\"{OUTPUT_DIR}/{RUN_NAME}/weights/last.pt\"\nRESUME = os.path.exists(LAST_CHECKPOINT)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:52:33.693033Z","iopub.execute_input":"2026-08-22T07:52:33.693437Z","iopub.status.idle":"2026-08-22T07:53:34.220432Z","shell.execute_reply.started":"2026-08-22T07:52:33.693408Z","shell.execute_reply":"2026-08-22T07:53:34.219528Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 5 — Dataset preparation functions","metadata":{}},{"cell_type":"code","source":"def clean_yolo_dataset():\n    \"\"\"\n    Remove any previously generated YOLO dataset folders so every run\n    starts clean - avoids stale labels/images from a prior ratio experiment\n    leaking into the current run.\n    \"\"\"\n    dirs_to_clean = [IMAGES_TRAIN_DIR, IMAGES_VAL_DIR, LABELS_TRAIN_DIR, LABELS_VAL_DIR]\n    for directory in dirs_to_clean:\n        if os.path.exists(directory):\n            shutil.rmtree(directory)\n        os.makedirs(directory, exist_ok=True)\n\n\ndef build_yolo_dataset(ratio=NEGATIVE_RATIO):\n    \"\"\"\n    Runs Person A's full pipeline: load -> filter negatives (train only)\n    -> write YOLO labels -> symlink images -> write data.yaml.\n    Prepares files on disk for YOLO to read; returns nothing.\n    \"\"\"\n    clean_yolo_dataset()\n\n    full_df = dataset.load_data(TRAIN_CSV, SPLITS_CSV)\n    train_df = full_df[full_df[\"split\"] == \"train\"]\n    val_df = full_df[full_df[\"split\"] == \"val\"]  # never filtered - keep val untouched\n\n    train_df_filtered = dataset.filter_negatives(train_df, ratio=ratio)\n\n    print(f\"Train frames after filtering: {len(train_df_filtered)} (ratio={ratio})\")\n    print(f\"Val frames (untouched):       {len(val_df)}\")\n\n    dataset.write_yolo_labels(train_df_filtered, labels_dir=LABELS_TRAIN_DIR)\n    dataset.write_yolo_labels(val_df, labels_dir=LABELS_VAL_DIR)\n\n    dataset.link_images(train_df_filtered, RAW_IMAGES_ROOT, images_dir=IMAGES_TRAIN_DIR)\n    dataset.link_images(val_df, RAW_IMAGES_ROOT, images_dir=IMAGES_VAL_DIR)\n\n    dataset.write_data_yaml(DATA_YAML_PATH, IMAGES_TRAIN_DIR, IMAGES_VAL_DIR)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:53:34.221599Z","iopub.execute_input":"2026-08-22T07:53:34.222104Z","iopub.status.idle":"2026-08-22T07:53:34.229469Z","shell.execute_reply.started":"2026-08-22T07:53:34.222030Z","shell.execute_reply":"2026-08-22T07:53:34.228803Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 6 — Training function","metadata":{}},{"cell_type":"code","source":"def train():\n    build_yolo_dataset(ratio=NEGATIVE_RATIO)\n\n    if RESUME:\n        # Ultralytics restores imgsz/epochs/optimizer/augmentation/etc from\n        # that run's saved args.yaml automatically - don't pass them again.\n        print(f\"Found existing checkpoint at {LAST_CHECKPOINT} - resuming from there.\")\n        model = load_model(pretrained_weights=LAST_CHECKPOINT)\n        results = model.train(resume=True)\n    else:\n        model = load_model(pretrained_weights=MODEL_WEIGHTS)\n\n        # epochs=60 (was 8): the 8-epoch run\\'s loss was still dropping hard\n        # and mAP was still noisy/non-monotonic (close_mosaic=10 never even\n        # triggered) - 8 epochs was cut way too short.\n        # batch_size=24 (was 16): the 8-epoch run only used ~8.6GB of the\n        # 14.9GB on each T4, so there\\'s room to raise it. Drop back to 16\n        # if you hit an out-of-memory error.\n        # save_period=5: writes epoch5.pt, epoch10.pt, ... to weights/ in\n        # addition to last.pt/best.pt, so a long run can survive a Kaggle\n        # session interruption - see RESUME logic above.\n        config = get_train_config(\n            image_size=1280,\n            epochs=60,\n            batch_size=24,\n            learning_rate=0.01,\n            patience=10,\n            optimizer=\"SGD\",\n            device=DEVICE,\n            save_period=5,\n        )\n        print(\"Requested train config:\", config)\n\n\n        results = model.train(\n            data=DATA_YAML_PATH,\n            imgsz=config[\"imgsz\"],\n            epochs=config[\"epochs\"],\n            batch=config[\"batch\"],\n            optimizer=config[\"optimizer\"],\n            lr0=config[\"lr0\"],\n            patience=config[\"patience\"],\n            device=config[\"device\"],\n            save_period=config[\"save_period\"],\n            mosaic=config[\"mosaic\"],\n            mixup=config[\"mixup\"],\n            copy_paste=config[\"copy_paste\"],\n            hsv_h=config[\"hsv_h\"],\n            hsv_s=config[\"hsv_s\"],\n            hsv_v=config[\"hsv_v\"],\n            project=OUTPUT_DIR,\n            name=RUN_NAME,\n            workers=4,\n            cache=\"disk\",\n            plots=True,\n        )\n\n    # Sanity check: confirm what Ultralytics actually trained with matches\n    # what we asked for.\n    actual = model.trainer.args\n    print(f\"Actual imgsz used:     {actual.imgsz}\")\n    print(f\"Actual optimizer used: {actual.optimizer}\")\n    print(f\"Actual lr0 used:       {actual.lr0}\")\n    print(f\"Actual epochs used:    {actual.epochs}\")\n    print(f\"Actual save_period:    {actual.save_period}\")\n\n    print(\"Training complete.\")\n    print(f\"Results and checkpoints saved to: {OUTPUT_DIR}/{RUN_NAME}\")\n    return results","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:53:34.230421Z","iopub.execute_input":"2026-08-22T07:53:34.230721Z","iopub.status.idle":"2026-08-22T07:53:34.249282Z","shell.execute_reply.started":"2026-08-22T07:53:34.230697Z","shell.execute_reply":"2026-08-22T07:53:34.248636Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 7 — Run training\n\nCheckpoints save to `/kaggle/working/`, so they persist as notebook\noutput when you click **Save Version**.","metadata":{}},{"cell_type":"code","source":"results = train()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-22T07:53:34.250178Z","iopub.execute_input":"2026-08-22T07:53:34.250682Z","iopub.status.idle":"2026-08-22T14:21:05.210391Z","shell.execute_reply.started":"2026-08-22T07:53:34.250656Z","shell.execute_reply":"2026-08-22T14:21:05.204888Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import shutil\nshutil.copy(\n    \"/kaggle/working/great-barrier-reef-checkpoints/final_model_v3/weights/best.pt\",\n    \"/kaggle/working/yolo11s_v3_best_mAP50-0.619.pt\"\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}