{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":123966,"databundleVersionId":14902028,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🛰️ A Practical Engineering Pipeline for Geospatial Building Detection (YOLOv8)\n\n### Lessons from an Ongoing Kaggle Competition\n\n**Keywords:** Object Detection · Remote Sensing · YOLOv8 · Dataset Engineering · Negative Mining · Submission Hygiene · Error Analysis\n\n---\n\nThis notebook documents a **complete, end-to-end engineering workflow** for a geospatial building detection task in an **ongoing Kaggle competition**.\n\nRather than focusing on hyperparameter tuning alone, this work emphasizes:\n\n- 📦 Dataset engineering and validation  \n- 🔁 Iterative training–inference–diagnosis loops  \n- 🧪 Careful evaluation of *what helps* and *what hurts* model performance  \n- 🧼 Strict submission format hygiene  \n- ⚠️ Explicit discussion of **failed or misleading approaches**, to help others avoid common pitfalls\n\n> ⚠️ **Competition Status Notice**  \n> This competition is still ongoing.  \n> Therefore, this notebook intentionally avoids sharing any rule-violating content (e.g. private test data details or unreleased final solutions).  \n> The focus is on **methodology, reproducible scaffolding, and engineering insights**, not leaderboard exploitation.\n\n---\n\n🧭 **Target Audience**\n\nThis notebook is written for:\n- Practitioners working on remote sensing object detection\n- Kaggle competitors seeking a robust engineering workflow\n- Readers who want to understand *why* certain approaches fail, not just *what* works\n\n---\n\n📌 **Notebook Philosophy**\n\n> *Improve systems, not just scores.*\n\nThe same pipeline applies whether you run on:\n- Kaggle GPU\n- Local GPU\n- CPU-only environments\n\nHardware differences affect **trade-offs**, not the core methodology.\n","metadata":{}},{"cell_type":"markdown","source":"## 🎯 Problem Definition (Engineering View)\n\nThe task is **geospatial building detection** from overhead imagery.\n\n- **Input**: High-resolution remote sensing images\n- **Output**: Bounding boxes for buildings\n- **Formulation**: Single-class object detection (building only)\n\nThis is *not* a standard natural-image detection problem.  \nSeveral properties make it fundamentally different:\n\n### 🌍 Characteristics of Remote Sensing Detection\n\n- 🧱 **Extreme scale variation**  \n  Buildings range from large urban blocks to tiny rural houses.\n- 🧩 **High background complexity**  \n  Terraces, grids, roads, water reflections, boats, poles, and textures often resemble buildings.\n- 🧠 **Low semantic cues**  \n  Color and texture are unreliable; geometry dominates.\n- 📐 **Small-object dominance**  \n  Most valid objects occupy a very small fraction of the image.\n\n---\n\n## 📊 Evaluation Metric: Practical Implications\n\nWhile the exact metric formulation is not discussed here, several **empirical behaviors** strongly influence engineering decisions:\n\n- 📈 **Recall is heavily rewarded**  \n  Missing true buildings is often more harmful than introducing some false positives.\n- 🎚️ **Confidence ranking matters**  \n  The *ordering* of predictions (not just their existence) impacts final score.\n- 🧼 **Submission format is unforgiving**  \n  Formatting errors (IDs, coordinate normalization, empty predictions) can invalidate otherwise good models.\n\n> **Key takeaway:**  \n> This is a competition where *engineering discipline* can matter as much as model architecture.\n\n---\n\n## 🔁 Consequence for System Design\n\nThese observations directly shape the workflow adopted in this notebook:\n\n- Prefer **robust inference pipelines** over aggressive pruning\n- Treat `confidence threshold` as a **model behavior probe**, not a fixed hyperparameter\n- Invest early in **dataset validation and submission hygiene**\n- Always analyze both **false negatives and false positives** visually\n\nThis mindset guided every decision described in the following sections.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔁 Overall Workflow (System-Level View)\n\nThis notebook follows a **closed engineering loop**, not a linear training script.\n\n**Pipeline overview:**\n\n1. 📦 Dataset preparation & validation  \n2. 🩺 Dataset health check (before training anything)  \n3. 🏋️ Model training (baseline → variants)  \n4. 🔮 Inference with controlled confidence thresholds  \n5. 🧼 Submission generation & strict format validation  \n6. 👀 Visual error analysis (FP vs FN patterns)  \n7. 🔄 Data strategy update (what to add / what to avoid)\n\nThis loop repeats until:\n- performance saturates, or\n- further changes become unstable or non-transferable.\n\n> **Design principle:**  \n> Every iteration must answer *why* a change helped or hurt — not just *whether* it changed the score.\n","metadata":{}},{"cell_type":"markdown","source":"## 🖥️ Hardware-Agnostic Design Philosophy\n\nAll code in this notebook is written to be **device-independent**.\n\n- Works on:\n  - Kaggle GPU notebooks\n  - Local GPU machines\n  - CPU-only environments\n- No assumption about CUDA availability\n- No hardcoded device IDs\n\nHardware differences only affect:\n- training speed\n- feasible number of epochs\n- experiment breadth\n\nThey **do not** change:\n- dataset validation logic\n- inference methodology\n- submission hygiene\n- diagnostic principles\n","metadata":{}},{"cell_type":"code","source":"import torch\n\ndef get_device():\n    if torch.cuda.is_available():\n        return \"cuda\"\n    if hasattr(torch.backends, \"mps\") and torch.backends.mps.is_available():\n        return \"mps\"\n    return \"cpu\"\n\nDEVICE = get_device()\nprint(\"Using device:\", DEVICE)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🩺 Why Dataset Validation Comes *Before* Training\n\nMany failures in object detection are **silent failures**:\n\n- Wrong directory depth → zero test images\n- Missing label files → implicit negatives\n- Illegal class IDs → undefined behavior\n- Broken images → sporadic crashes\n- Over-cleaning → recall collapse\n\nTraining a model **before** validating data often wastes:\n- time\n- compute\n- leaderboard submissions\n\nTherefore, every experiment in this workflow starts with:\n> **“Can I trust my dataset?”**\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nfrom collections import Counter\n\ndef scan_dataset(root, name=\"dataset\", max_list=20):\n    root = Path(root)\n    img_exts = {\".png\", \".jpg\", \".jpeg\", \".tif\", \".tiff\", \".webp\"}\n\n    files = [p for p in root.rglob(\"*\") if p.is_file()]\n    imgs = [p for p in files if p.suffix.lower() in img_exts]\n    labels = [p for p in files if p.suffix.lower() == \".txt\"]\n\n    print(\"=\" * 70)\n    print(f\"📁 {name}\")\n    print(\"Path:\", root)\n    print(\"Total files:\", len(files))\n    print(\"Images:\", len(imgs))\n    print(\"Labels:\", len(labels))\n\n    ext_cnt = Counter(p.suffix.lower() for p in imgs)\n    print(\"Image extensions:\", dict(ext_cnt))\n\n    print(\"Top-level contents:\")\n    if root.exists():\n        for p in sorted(root.iterdir())[:max_list]:\n            print(\"  📂\" if p.is_dir() else \"  📄\", p.name)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🩺 Dataset Health Check: What Matters (and What Doesn’t)\n\nBefore any training, the dataset is audited with a **clear priority order**.\n\n### ✅ Must-check (critical)\n- Image–label one-to-one matching\n- Broken or unreadable images\n- Label format correctness (YOLO format)\n- Illegal class IDs\n\n### ⚠️ Observe but usually keep\n- Empty label files (negative samples)\n- Very small bounding boxes\n\n### ❌ Do NOT over-clean\n- Removing all small boxes\n- Removing all empty-label images\n- Aggressively filtering dense scenes\n\n> **Principle:**  \n> Only remove what is *provably invalid*.  \n> Everything else is part of the real data distribution.\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nfrom collections import Counter\nfrom PIL import Image\n\ndef health_check_yolo_dataset(root):\n    root = Path(root)\n    img_dirs = [root/\"images\"/\"train\", root/\"images\"/\"val\"]\n    lbl_dirs = [root/\"labels\"/\"train\", root/\"labels\"/\"val\"]\n\n    imgs, lbls = [], []\n    for d in img_dirs:\n        if d.exists():\n            imgs += list(d.glob(\"*.png\")) + list(d.glob(\"*.jpg\"))\n    for d in lbl_dirs:\n        if d.exists():\n            lbls += list(d.glob(\"*.txt\"))\n\n    print(\"=\" * 70)\n    print(\"🔍 DATASET HEALTH CHECK\")\n    print(\"Dataset:\", root)\n    print(\"=\" * 70)\n\n    print(\"\\n[1] Basic counts\")\n    print(\"Images:\", len(imgs))\n    print(\"Labels:\", len(lbls))\n\n    img_stems = {p.stem for p in imgs}\n    lbl_stems = {p.stem for p in lbls}\n\n    print(\"\\n[2] Image–Label matching\")\n    print(\"Images without labels:\", len(img_stems - lbl_stems))\n    print(\"Labels without images:\", len(lbl_stems - img_stems))\n\n    print(\"\\n[3] Image integrity (sampled)\")\n    broken = 0\n    for p in imgs[:2000]:\n        try:\n            with Image.open(p) as im:\n                im.verify()\n        except:\n            broken += 1\n    print(\"Broken images (sample):\", broken)\n\n    print(\"\\n[4] Label format & class check\")\n    empty = 0\n    bad_format = 0\n    class_cnt = Counter()\n\n    for p in lbls:\n        txt = p.read_text().strip()\n        if txt == \"\":\n            empty += 1\n            continue\n        for line in txt.splitlines():\n            parts = line.split()\n            if len(parts) != 5:\n                bad_format += 1\n                continue\n            cls = parts[0]\n            class_cnt[cls] += 1\n\n    print(\"Empty labels:\", empty)\n    print(\"Bad-format lines:\", bad_format)\n    print(\"Class distribution:\", dict(class_cnt))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## 🏷️ Single-Class Decision: Why Only “Building”\n\nAlthough some public label schemas include multiple classes\n(e.g. building, vehicle), this workflow intentionally uses:\n\n> **Single-class detection: building only**\n\n### Reasons:\n- Evaluation focuses on building localization quality\n- Multi-class training without explicit metric support can:\n  - dilute objectness learning\n  - introduce noisy gradients\n- Illegal class IDs silently break training/inference\n\n### Cleaning rule adopted:\n- ❌ Remove label lines with class ≠ 0\n- ✅ Keep empty labels\n- ✅ Keep small boxes\n\nThis keeps the task **well-defined and stable**.\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚠️ Common Data Cleaning Pitfalls (Real Failures)\n\nThese are **tempting but harmful** operations:\n\n### ❌ Deleting all empty-label images\n- Removes negative context\n- Increases false positives\n- Makes background “too easy”\n\n### ❌ Removing all small bounding boxes\n- Small buildings are *real* in remote sensing\n- Causes systematic recall loss\n\n### ❌ Re-splitting train/val randomly after merging\n- Breaks original distribution assumptions\n- Can leak correlated samples\n\n> **Lesson:**  \n> Dataset cleaning should be *surgical*, not aggressive.\n","metadata":{}},{"cell_type":"markdown","source":"## ✅ Dataset Validation Checklist (Reusable)\n\nBefore training any model, confirm:\n\n- [ ] Image and label counts match\n- [ ] No broken images\n- [ ] Label format is valid YOLO\n- [ ] All classes are expected (single-class = 0)\n- [ ] Empty labels are intentional\n- [ ] No accidental data leakage between splits\n\nIf any box above is unchecked:\n> **Fix data first. Do not train.**\n","metadata":{}},{"cell_type":"markdown","source":"## 🏋️ Baseline Model: Training Philosophy (Not Hyperparameters)\n\nThis notebook uses **YOLOv8** as the detection framework.\n\nThe focus is **not** on finding a perfect set of hyperparameters,\nbut on establishing a **stable baseline** that can be reasoned about.\n\n### Training philosophy:\n- Prefer **reasonable defaults** over aggressive tuning\n- Avoid frequent changes to multiple variables at once\n- Treat each training run as a *diagnostic tool*, not a one-shot solution\n\n> A stable, interpretable baseline is more valuable than an unstable high-score spike.\n","metadata":{}},{"cell_type":"code","source":"from ultralytics import YOLO\n\ndef train_yolo(\n    model_ckpt,\n    data_yaml,\n    epochs=50,\n    imgsz=1024,\n    batch=8,\n    project=\"runs/detect\",\n    name=\"baseline\",\n):\n    model = YOLO(model_ckpt)\n    model.train(\n        data=data_yaml,\n        epochs=epochs,\n        imgsz=imgsz,\n        batch=batch,\n        project=project,\n        name=name,\n        verbose=True,\n    )\n\n# Example (adjust paths):\n# train_yolo(\n#     model_ckpt=\"yolov8n.pt\",\n#     data_yaml=\"super_dataset.yaml\",\n#     epochs=50,\n#     name=\"train_baseline\"\n# )\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔮 Inference Is Not Training: Understanding `conf`\n\nA critical realization:\n\n> **`conf` is an inference-time gate, not a training parameter.**\n\nWhat `conf` does:\n- Filters predictions *before* they appear in results\n- Affects recall–precision trade-off\n- Changes the distribution of boxes fed into submission\n\nWhat `conf` does NOT do:\n- It cannot recover boxes that were never produced\n- It cannot fix poor ranking quality\n\nLowering `conf` is **probing model behavior**, not “cheating the metric”.\n","metadata":{}},{"cell_type":"code","source":"from ultralytics import YOLO\nfrom pathlib import Path\n\ndef run_inference(\n    model_path,\n    image_dir,\n    conf=0.01,\n    iou=0.6,\n    imgsz=1024,\n):\n    model = YOLO(model_path)\n    results = model.predict(\n        source=str(image_dir),\n        conf=conf,\n        iou=iou,\n        imgsz=imgsz,\n        verbose=False,\n    )\n    return results\n\n# Example:\n# results = run_inference(\n#     model_path=\"best.pt\",\n#     image_dir=\"test_images\",\n#     conf=0.01\n# )\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧼 Submission Hygiene: Why Good Models Get Score = 0\n\nMany low scores are caused by **formatting errors**, not model quality.\n\n### Common failure modes:\n- `image_id` includes file extension or leading zeros\n- Bounding boxes are in **pixel** coordinates instead of normalized\n- Wrong field order in prediction string\n- Missing class ID\n- Empty predictions not replaced by `\"no box\"`\n\nThese errors **invalidate predictions silently**.\n\n> Submission formatting must be treated as production code.\n","metadata":{}},{"cell_type":"code","source":"import csv\nfrom pathlib import Path\n\ndef generate_submission(results, out_csv):\n    rows = []\n\n    for r in results:\n        image_id = int(Path(r.path).stem)  # critical: numeric ID only\n\n        if r.boxes is None or len(r.boxes) == 0:\n            rows.append((image_id, \"no box\"))\n            continue\n\n        parts = []\n        for b in r.boxes:\n            cx, cy, w, h = b.xywhn[0].tolist()\n            conf = float(b.conf.item())\n            cls = int(b.cls.item())\n            parts.append(\n                f\"{cls} {conf:.6f} {cx:.6f} {cy:.6f} {w:.6f} {h:.6f}\"\n            )\n\n        rows.append((image_id, \" \".join(parts)))\n\n    rows.sort(key=lambda x: x[0])\n\n    with open(out_csv, \"w\", newline=\"\") as f:\n        writer = csv.writer(f)\n        writer.writerow([\"image_id\", \"prediction_string\"])\n        writer.writerows(rows)\n\n    print(\"Saved submission:\", out_csv)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎚️ Low `conf`: When It Works\n\nLowering `conf` is useful **only if** the model's score ranking is still meaningful.\n\nEngineering symptom of a “healthy low-conf regime”:\n- many added boxes are real buildings\n- FP increases but does not explode\n- per-image box counts stay within a reasonable range\n- top-confidence boxes remain mostly correct\n\nNext cell: a **conf sweep diagnostic** to quantify this.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nfrom pathlib import Path\nfrom ultralytics import YOLO\nfrom tqdm import tqdm\n\ndef infer_box_counts(model_path, image_dir, conf_list, iou=0.6, imgsz=1024, max_images=None):\n    model = YOLO(model_path)\n    image_dir = Path(image_dir)\n    imgs = sorted([p for p in image_dir.glob(\"*\") if p.suffix.lower() in [\".jpg\",\".png\",\".jpeg\"]])\n    if max_images:\n        imgs = imgs[:max_images]\n\n    report = {}  # conf -> dict(stats)\n    for conf in conf_list:\n        counts = []\n        paths  = []\n        results = model.predict(\n            source=[str(p) for p in imgs],\n            conf=conf,\n            iou=iou,\n            imgsz=imgsz,\n            verbose=False,\n        )\n        for r in results:\n            n = 0 if (r.boxes is None) else len(r.boxes)\n            counts.append(n)\n            paths.append(r.path)\n\n        counts = np.array(counts)\n        report[conf] = {\n            \"n_images\": len(counts),\n            \"mean\": float(counts.mean()),\n            \"p50\": float(np.percentile(counts, 50)),\n            \"p90\": float(np.percentile(counts, 90)),\n            \"p99\": float(np.percentile(counts, 99)),\n            \"max\": int(counts.max()),\n            \"top_fp_candidates\": [paths[i] for i in np.argsort(-counts)[:10]],\n            \"top_counts\": [int(counts[i]) for i in np.argsort(-counts)[:10]],\n        }\n    return report\n\n# Example (fill your paths):\n# MODEL_PATH = \"best.pt\"\n# TEST_DIR   = \"/kaggle/input/.../testImages/images\"\n# conf_list  = [0.001, 0.003, 0.005, 0.01, 0.02]\n# rep = infer_box_counts(MODEL_PATH, TEST_DIR, conf_list, max_images=None)\n# for c, s in rep.items():\n#     print(c, s[\"mean\"], s[\"p90\"], s[\"p99\"], s[\"max\"])\n#     print(\"Top-10 (count):\", list(zip(s[\"top_counts\"], s[\"top_fp_candidates\"][:3])))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚠️ Low `conf`: When It Fails\n\nLow `conf` fails when ranking quality degrades.\n\nEngineering symptoms:\n- per-image box counts explode (p99 / max extremely high)\n- top-confidence boxes include many “structured textures”\n- lowering conf adds mostly junk, not true buildings\n\nNext cell: visualize **Top FP candidates** quickly.\n","metadata":{}},{"cell_type":"code","source":"import cv2\nimport random\nfrom pathlib import Path\n\ndef save_visualizations(results, out_dir, max_save=50, thickness=2):\n    out_dir = Path(out_dir)\n    out_dir.mkdir(parents=True, exist_ok=True)\n\n    # pick subset\n    idxs = list(range(len(results)))\n    random.shuffle(idxs)\n    idxs = idxs[:max_save]\n\n    saved = 0\n    for i in idxs:\n        r = results[i]\n        img = cv2.imread(r.path)\n        if img is None:\n            continue\n\n        if r.boxes is not None and len(r.boxes) > 0:\n            for b in r.boxes:\n                x1,y1,x2,y2 = b.xyxy[0].cpu().numpy().astype(int).tolist()\n                conf = float(b.conf.item())\n                cv2.rectangle(img, (x1,y1), (x2,y2), (0,255,0), thickness)\n                cv2.putText(img, f\"{conf:.3f}\", (x1, max(0,y1-5)),\n                            cv2.FONT_HERSHEY_SIMPLEX, 0.6, (0,255,0), 2)\n\n        out_path = out_dir / f\"{Path(r.path).stem}.jpg\"\n        cv2.imwrite(str(out_path), img)\n        saved += 1\n\n    print(f\"Saved {saved} visualizations to {out_dir}\")\n\n# Usage:\n# model = YOLO(MODEL_PATH)\n# results = model.predict(source=str(TEST_DIR), conf=0.005, iou=0.6, imgsz=1024, verbose=False)\n# save_visualizations(results, out_dir=WORKDIR/\"viz_conf_0005\", max_save=60)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧩 FP Taxonomy → Structured Logging\n\nA common mistake: “I saw many false positives” (too vague).\n\nBetter:\n- define FP categories (terrace, grid, stripe, pole, boat, truck…)\n- record which images exhibit which pattern\n- use this log to drive data strategy (synthetic or negative mining)\n\nNext cell: a lightweight annotation log tool (manual but fast).\n","metadata":{}},{"cell_type":"code","source":"import json\nfrom pathlib import Path\n\nFP_TAGS = [\"terrace\", \"stripe\", \"grid_water\", \"pole\", \"boat\", \"truck\", \"other\"]\n\ndef init_fp_log(log_path):\n    log_path = Path(log_path)\n    if not log_path.exists():\n        log_path.write_text(json.dumps({\"items\": []}, indent=2))\n    print(\"Log:\", log_path)\n\ndef add_fp_entry(log_path, image_id, tags, note=\"\"):\n    log_path = Path(log_path)\n    data = json.loads(log_path.read_text())\n    data[\"items\"].append({\n        \"image_id\": str(image_id),\n        \"tags\": tags,\n        \"note\": note\n    })\n    log_path.write_text(json.dumps(data, indent=2))\n\n# Example:\n# LOG = WORKDIR/\"fp_taxonomy_log.json\"\n# init_fp_log(LOG)\n# add_fp_entry(LOG, \"000123\", [\"grid_water\",\"boat\"], note=\"dock-like bright grid, many boxes\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚖️ Great Attempt: Negative Mining (Controlled Version)\n\nNegative mining is not “add as much as possible”.\n\nA safe engineering version:\n- add negatives gradually\n- monitor:\n  - TP score distribution shift\n  - per-image box count changes at low conf\n  - whether recall collapses visually\n\nNext cell: a *controlled* negative merge tool (ratio-aware + collision-safe).\n","metadata":{}},{"cell_type":"code","source":"import shutil\nfrom pathlib import Path\nimport uuid\n\ndef merge_negatives_ratio_aware(\n    super_root,\n    neg_dirs,\n    max_add=None,\n    rename_on_collision=True,\n    dry_run=False\n):\n    super_root = Path(super_root)\n    train_img = super_root/\"images\"/\"train\"\n    train_lbl = super_root/\"labels\"/\"train\"\n    train_img.mkdir(parents=True, exist_ok=True)\n    train_lbl.mkdir(parents=True, exist_ok=True)\n\n    existing = set(p.name for p in train_img.glob(\"*\"))\n    added, skipped = 0, 0\n\n    for neg_dir in map(Path, neg_dirs):\n        imgs = sorted([p for p in neg_dir.glob(\"*\") if p.suffix.lower() in [\".png\",\".jpg\",\".jpeg\"]])\n        for img in imgs:\n            if max_add is not None and added >= max_add:\n                break\n\n            dst_name = img.name\n            if dst_name in existing and rename_on_collision:\n                dst_name = f\"{img.stem}_{uuid.uuid4().hex[:8]}{img.suffix.lower()}\"\n\n            dst_img = train_img/dst_name\n            dst_lbl = train_lbl/(Path(dst_name).stem + \".txt\")\n\n            if dst_img.exists() or dst_lbl.exists():\n                skipped += 1\n                continue\n\n            if not dry_run:\n                shutil.copy2(img, dst_img)\n                dst_lbl.write_text(\"\")  # empty label = negative sample\n\n            existing.add(dst_name)\n            added += 1\n\n    print(f\"Added negatives: {added}, skipped: {skipped}\")\n\n# Example:\n# merge_negatives_ratio_aware(\n#     super_root=SUPER_DATASET,\n#     neg_dirs=[\"/kaggle/input/road_and_shore\", \"/kaggle/input/oceanData\"],\n#     max_add=500,  # incremental\n#     dry_run=False\n# )\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📉 Monitoring Ranking Quality via Score Distributions\n\nIf negative mining hurts ranking quality, a common symptom is:\n\n- confidence scores for true positives shift downward\n- score distribution compresses\n- low-conf regime becomes dominated by junk\n\nWe can monitor this without labels by tracking:\n- confidence histogram\n- top-K confidence per image\n- box count vs confidence\n\nNext cell: confidence statistics.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef confidence_stats(results, topk=5):\n    all_conf = []\n    topk_conf = []\n\n    for r in results:\n        if r.boxes is None or len(r.boxes) == 0:\n            topk_conf.append([])\n            continue\n        confs = [float(b.conf.item()) for b in r.boxes]\n        all_conf.extend(confs)\n        topk_conf.append(sorted(confs, reverse=True)[:topk])\n\n    all_conf = np.array(all_conf) if len(all_conf) else np.array([0.0])\n    print(\"Total boxes:\", len(all_conf))\n    print(\"Conf mean:\", float(all_conf.mean()))\n    print(\"Conf p50:\", float(np.percentile(all_conf, 50)))\n    print(\"Conf p90:\", float(np.percentile(all_conf, 90)))\n    print(\"Conf p99:\", float(np.percentile(all_conf, 99)))\n    return all_conf, topk_conf\n\n# Usage:\n# results = run_inference(MODEL_PATH, TEST_DIR, conf=0.005)\n# all_conf, topk_conf = confidence_stats(results, topk=5)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧪 Great Attempt: Segmentation Filters (Attractive but Risky)\n\nSegmentation seems ideal for suppressing texture-based false positives.\n\nBut in remote sensing:\n- small isolated buildings are fragile\n- background dominates pixels\n- a slightly underfit segmentation mask can delete real buildings\n\nSo we keep segmentation **as an optional analysis tool**:\n- use it to study where the model fires\n- avoid hard filtering until proven safe\n\nNext cell: a safe skeleton (no hard-coded “final filter rules”).\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef apply_soft_mask_to_boxes(boxes_xyxy, mask, min_overlap=0.05):\n    \"\"\"\n    boxes_xyxy: (N,4) in pixel coords\n    mask: (H,W) bool or 0/1\n    min_overlap: keep box if overlap ratio >= threshold\n    NOTE: This is a *skeleton* for research, not a final competition trick.\n    \"\"\"\n    H, W = mask.shape[:2]\n    keep = []\n    for (x1,y1,x2,y2) in boxes_xyxy:\n        x1 = max(0, min(W-1, int(x1)))\n        x2 = max(0, min(W,   int(x2)))\n        y1 = max(0, min(H-1, int(y1)))\n        y2 = max(0, min(H,   int(y2)))\n        if x2 <= x1 or y2 <= y1:\n            keep.append(False); continue\n\n        region = mask[y1:y2, x1:x2]\n        overlap = float(region.mean())  # ratio\n        keep.append(overlap >= min_overlap)\n    return np.array(keep, dtype=bool)\n\n# NOTE:\n# - You must validate visually before using any filtering.\n# - Over-filtering can destroy recall.\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔁 Duplicate Detections: A Silent Score Killer\n\nA common failure mode in geospatial detection:\n\n> **The same building is detected multiple times**, often with:\n- slightly shifted boxes\n- different confidence scores\n- overlapping but non-identical extents\n\nThis happens frequently when:\n- using low `conf`\n- buildings are small\n- textures are repetitive\n- object boundaries are ambiguous\n\nImportantly:\n- These duplicates are **not obvious false positives**\n- But they **harm ranking-based metrics**\n","metadata":{}},{"cell_type":"markdown","source":"## ❓ Why Built-in NMS Is Often Not Enough\n\nYOLO already applies NMS internally.\n\nHowever, duplicate detections still occur because:\n- IoU between boxes may be just below NMS threshold\n- Boxes differ in scale (tight vs loose)\n- Low `conf` allows many borderline boxes to survive\n- Dense scenes amplify overlaps\n\nTherefore:\n> **NMS is necessary, but not always sufficient.**\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef iou_xyxy(a, b):\n    x1 = max(a[0], b[0])\n    y1 = max(a[1], b[1])\n    x2 = min(a[2], b[2])\n    y2 = min(a[3], b[3])\n    inter = max(0, x2 - x1) * max(0, y2 - y1)\n    area_a = (a[2] - a[0]) * (a[3] - a[1])\n    area_b = (b[2] - b[0]) * (b[3] - b[1])\n    union = area_a + area_b - inter + 1e-6\n    return inter / union\n\n\ndef simple_post_nms(boxes_xyxy, scores, iou_thresh=0.5):\n    \"\"\"\n    boxes_xyxy: (N,4) numpy array\n    scores:     (N,) confidence scores\n    returns indices to keep\n    \"\"\"\n    order = np.argsort(-scores)\n    keep = []\n\n    while order.size > 0:\n        i = order[0]\n        keep.append(i)\n        rest = order[1:]\n\n        ious = np.array([iou_xyxy(boxes_xyxy[i], boxes_xyxy[j]) for j in rest])\n        order = rest[ious < iou_thresh]\n\n    return keep\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧠 A More Important Principle Than NMS\n\nIn practice, the most robust way to reduce duplicates is:\n\n> **Maintain ranking quality first.**\n\nBecause:\n- If ranking is good → duplicates usually rank lower\n- If ranking collapses → no amount of NMS will save the submission\n\nThis connects directly to earlier lessons:\n- overly aggressive negative mining hurts ranking\n- segmentation over-filtering hurts ranking\n- low `conf` only works when ranking is preserved\n\n> Post-processing should *support* ranking,\n> not attempt to replace it.\n","metadata":{}},{"cell_type":"code","source":"def boxes_to_submission_strings(boxes_xyxy, scores, classes, img_w, img_h, iou_thresh=0.5):\n    \"\"\"\n    Apply post-NMS and convert to normalized cxcywh strings.\n    \"\"\"\n    keep = simple_post_nms(boxes_xyxy, scores, iou_thresh=iou_thresh)\n\n    parts = []\n    for i in keep:\n        x1,y1,x2,y2 = boxes_xyxy[i]\n        cx = ((x1 + x2) / 2) / img_w\n        cy = ((y1 + y2) / 2) / img_h\n        w  = (x2 - x1) / img_w\n        h  = (y2 - y1) / img_h\n        parts.append(\n            f\"{int(classes[i])} {scores[i]:.6f} {cx:.6f} {cy:.6f} {w:.6f} {h:.6f}\"\n        )\n    return parts\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #1: Never Inplace-Modify Tensors from `r.boxes`\n\nIf you write postprocess like:\n- `b.conf *= 0.8`\n- `scores[i] *= penalty`\n- `tensor[mask] *= ...`\n\nyou may hit runtime errors under inference mode, or silently produce wrong results.\n\n✅ Safe pattern:\n- convert values to Python floats\n- or clone tensors before modification\n- or store into new lists/arrays (recommended)\n","metadata":{}},{"cell_type":"code","source":"def extract_boxes_py(r):\n    \"\"\"\n    Convert Ultralytics result `r` into pure-Python lists for safe postprocess.\n    Returns:\n      boxes_xyxy: list of [x1,y1,x2,y2] (pixel)\n      confs:      list of float\n      clss:       list of int\n      orig_shape: (h,w)\n    \"\"\"\n    if r.boxes is None or len(r.boxes) == 0:\n        return [], [], [], r.orig_shape  # (h,w)\n\n    boxes_xyxy = r.boxes.xyxy.cpu().numpy().tolist()\n    confs      = r.boxes.conf.cpu().numpy().astype(float).tolist()\n    clss       = r.boxes.cls.cpu().numpy().astype(int).tolist()\n    return boxes_xyxy, confs, clss, r.orig_shape\n\n# After this, ALL postprocess operations are safe pure python / numpy ops.\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #2: Do NOT Assume Images Are 1024×1024\n\nEven if you set `imgsz=1024`, the original images can be:\n- 864×1024, 1024×992, 960×1024, ...\n\nIf you normalize cx/cy/w/h using a hardcoded `1024`,\nyour submission boxes can become wrong.\n\n✅ Rule:\n- Always normalize using the *actual image width/height*.\n- Best source: `r.orig_shape` (h, w).\n","metadata":{}},{"cell_type":"code","source":"def xyxy_to_xywhn(xyxy, orig_shape):\n    \"\"\"\n    xyxy: [x1,y1,x2,y2] in pixels\n    orig_shape: (h,w)\n    returns normalized (cx,cy,w,h)\n    \"\"\"\n    h, w = orig_shape\n    x1, y1, x2, y2 = xyxy\n    cx = ((x1 + x2) / 2.0) / w\n    cy = ((y1 + y2) / 2.0) / h\n    bw = (x2 - x1) / w\n    bh = (y2 - y1) / h\n    return cx, cy, bw, bh\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #3: “Found 0 test images” Usually Means Extension Mismatch\n\nA silent failure mode:\n- your code globs `*.png`\n- but test set is `*.jpg`\n- result: 0 images → submission becomes all `\"no box\"`\n\n✅ Rule:\n- Always scan with multiple extensions\n- Always print first 5 filenames to confirm\n- Always assert count > 0 before inference\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\n\ndef list_images(image_dir):\n    image_dir = Path(image_dir)\n    exts = [\".jpg\", \".jpeg\", \".png\", \".tif\", \".tiff\", \".webp\"]\n    imgs = []\n    for e in exts:\n        imgs.extend(image_dir.glob(f\"*{e}\"))\n        imgs.extend(image_dir.glob(f\"*{e.upper()}\"))\n    imgs = sorted(set(imgs))\n    return imgs\n\ndef assert_test_images_ok(image_dir, min_expected=1):\n    imgs = list_images(image_dir)\n    print(\"✅ Found images:\", len(imgs))\n    print(\"Sample files:\", [p.name for p in imgs[:5]])\n    assert len(imgs) >= min_expected, \"No test images found — check path and extensions!\"\n    return imgs\n\n# Usage:\n# TEST_DIR = Path(\"/kaggle/input/.../testImages/images\")\n# test_imgs = assert_test_images_ok(TEST_DIR, min_expected=10)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #4: `rect=True` Silently Disables Shuffle\n\nIn YOLO training, enabling `rect=True` triggers this behavior:\n\n- DataLoader shuffle is disabled\n- Images are grouped by aspect ratio\n- Training order becomes structured and repeatable\n\nThis is not always bad, but it is **not free**.\n","metadata":{}},{"cell_type":"markdown","source":"## ⚖️ When `rect=True` Helps — and When It Hurts\n\n### Can help:\n- GPU memory efficiency\n- Faster convergence on uniform datasets\n- Stable training for large images\n\n### Can hurt:\n- Reduced stochasticity\n- Worse generalization on diverse scenes\n- Overfitting to spatial patterns\n\n**Recommendation**:\n- Default to `rect=False`\n- Enable only if you *measure* a benefit\n","metadata":{}},{"cell_type":"code","source":"def train_yolo_explicit_rect(\n    model_ckpt,\n    data_yaml,\n    epochs=50,\n    imgsz=1024,\n    batch=8,\n    rect=False,\n    name=\"exp\",\n):\n    from ultralytics import YOLO\n    model = YOLO(model_ckpt)\n    model.train(\n        data=data_yaml,\n        epochs=epochs,\n        imgsz=imgsz,\n        batch=batch,\n        rect=rect,\n        name=name,\n    )\n\n# Always log rect=True/False in experiment name or notes.\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #5: Label Cache Can Lie After Manual Cleaning\n\nUltralytics creates `.cache` files for labels to speed up scanning.\n\nIf you:\n- manually edit labels\n- delete illegal classes\n- merge datasets\n\nbut keep old `.cache` files,\n\nthe loader may reuse outdated statistics.\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\n\ndef clear_label_cache(dataset_root):\n    root = Path(dataset_root)\n    removed = 0\n    for p in root.rglob(\"*.cache\"):\n        p.unlink()\n        removed += 1\n    print(f\"🧹 Removed {removed} cache files\")\n\n# Usage:\n# clear_label_cache(\"/kaggle/working/super_dataset\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #6: Not Every `.txt` in `labels/` Is a Label\n\nCommon non-label files:\n- `classes.txt`\n- `README.txt`\n- export metadata\n\nIf your parser treats all `.txt` as labels:\n- format checks fail\n- illegal class counts appear\n- statistics become misleading\n\nAlways explicitly skip known non-label files.\n","metadata":{}},{"cell_type":"code","source":"def is_valid_label_file(path):\n    name = path.name.lower()\n    if name in {\"classes.txt\", \"readme.txt\"}:\n        return False\n    return True\n\ndef scan_label_files(label_dir):\n    label_dir = Path(label_dir)\n    labels = []\n    for p in label_dir.glob(\"*.txt\"):\n        if is_valid_label_file(p):\n            labels.append(p)\n    return labels\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #7: String Sorting Breaks `image_id` Order\n\nIf image filenames are numeric but stored as strings:\n\n- `\"1.jpg\", \"10.jpg\", \"100.jpg\", \"2.jpg\"`\n\nLexicographic sorting is **wrong**.\n\nWrong order can:\n- misalign predictions\n- break leaderboard evaluation\n\nRule:\n- Always sort by `int(stem)`\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\n\ndef sort_images_by_id(img_paths):\n    return sorted(img_paths, key=lambda p: int(Path(p).stem))\n\n# Example:\n# imgs = sort_images_by_id(list_images(TEST_DIR))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #8: Synthetic Data Is Not Free Generalization\n\nSynthetic data (e.g. Falcon / Duality-generated imagery) is powerful,\nbut **it encodes assumptions**.\n\nIf those assumptions are too narrow, the model may:\n- overfit synthetic regularities\n- underperform on real-world edge cases\n- show high confidence on wrong structures\n\nSynthetic data helps **only if its diversity matches failure modes**.\n","metadata":{}},{"cell_type":"code","source":"## 🛰️ Observed Limitations of Falcon/Duality Synthetic Data\n\nThrough repeated error analysis, several patterns emerged:\n\n### Synthetic scenes are often:\n- visually clean\n- geometrically regular\n- lighting-consistent\n- texture-simplified\n\n### Missing or underrepresented:\n- complex water reflections\n- color ambiguity (blue roofs vs water)\n- rocky/mountain backgrounds similar to roofs\n- cluttered coastal artifacts\n- real-world noise and partial occlusions\n\nThese gaps directly map to failure cases.\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ❌ False Negatives Linked to Synthetic Bias\n\nTypical missed detections after heavy synthetic training:\n\n- 🟦 Small houses near water  \n  (roof color blends with water surface)\n- 🪨 Houses embedded in rocky or mountainous terrain  \n  (roof vs rock color confusion)\n- 🌲 Sparse rural houses partially occluded by terrain\n\nThese cases are **rare or simplified** in synthetic datasets,\nbut common in real imagery.\n","metadata":{}},{"cell_type":"markdown","source":"## ❌ False Positives Amplified by Synthetic Regularity\n\nSynthetic data often reinforces *geometric priors*:\n\nThe model learns:\n> “Rectilinear + high contrast = building”\n\nResulting false positives:\n- 🚢 Boats\n- 🟦 Floating grid-like water structures\n- 🚚 Vehicles and containers\n- 🌾 Square-like agricultural terraces\n- 🪵 Poles and thin vertical structures\n\nThese objects are **geometrically similar but semantically different**.\n","metadata":{}},{"cell_type":"markdown","source":"## ⚠️ Why Adding More Synthetic Data Can Hurt Recall\n\nCounterintuitive but observed behavior:\n\nAfter adding large volumes of synthetic or curated water/shore negatives:\n- false positives decreased\n- but recall also dropped\n- leaderboard score degraded\n\nReason:\n- objectness calibration shifts\n- the model becomes conservative\n- borderline true buildings are suppressed\n\n> Cleaner predictions ≠ better ranking.\n","metadata":{}},{"cell_type":"markdown","source":"## 🌊 Official Ocean / Shore Datasets: A Cautionary Note\n\nAdding official ocean or shoreline datasets seems logical.\n\nHowever, in practice:\n- background dominates gradients\n- building probability mass shrinks\n- small coastal houses suffer first\n\nObserved effect:\n- fewer water-related false positives\n- more missed coastal buildings\n- overall score may drop\n\nThis does NOT mean such data is useless,\nbut it must be **balanced and incremental**.\n","metadata":{}},{"cell_type":"markdown","source":"## ✅ A Safer Strategy for Synthetic & Shore Data\n\nInstead of bulk addition:\n\n- Add synthetic data **in small batches**\n- Monitor:\n  - box count distribution at low `conf`\n  - recall on known hard cases\n- Mix with:\n  - real cluttered scenes\n  - ambiguous color cases\n  - partial occlusions\n\nSynthetic data should **target failure modes**, not replace reality.\n","metadata":{}},{"cell_type":"markdown","source":"## 🧠 Designing Synthetic Data Around Failure Modes\n\nBased on observed errors, useful synthetic scenes should emphasize:\n\n### For false negatives:\n- blue/gray roofs near water\n- low-contrast roofs against rocks\n- small isolated houses\n\n### For false positives:\n- boats near docks\n- floating grids on water\n- vehicles aligned with roads\n- square agricultural patterns\n- thin poles and vertical clutter\n\nSynthetic data should be *adversarial*, not idealized.\n","metadata":{}},{"cell_type":"markdown","source":"## 🧭 Key Takeaway: Data Is a Hypothesis\n\nEvery dataset encodes a belief about the world.\n\n- Synthetic data encodes **what we think matters**\n- The model learns those beliefs literally\n\nIf failures persist, the problem is often not:\n- learning rate\n- architecture\n- number of epochs\n\nbut **what the data tells the model is important**.\n\n> Data design is hypothesis testing.\n","metadata":{}},{"cell_type":"markdown","source":"## 🚧 Limits of Digital-Twin Synthetic Data (Falcon/Duality)\n\nWhile digital-twin generators are powerful, they have **structural limits**.\n\nIn this competition, Falcon-style synthetic data shows difficulty in generating:\n- fine-grained color ambiguity (e.g. blue roofs vs water)\n- realistic water reflections and clutter\n- complex rocky/mountain backgrounds\n- rare man-made non-building objects (boats, floating grids, poles)\n\nThese are not tuning issues — they are **expressiveness limits**.\n","metadata":{}},{"cell_type":"markdown","source":"## ❓ Why Falcon Cannot Fully Target Observed Failure Modes\n\nThe main issue is **control granularity**.\n\nFalcon can control:\n- object presence\n- rough geometry\n- coarse scene layout\n\nBut it has limited control over:\n- spectral similarity between objects\n- subtle texture confusion\n- rare real-world clutter patterns\n- context-specific ambiguity (e.g. water + roof + shadow)\n\nAs a result, many real failure modes remain underrepresented.\n","metadata":{}},{"cell_type":"markdown","source":"## ⚠️ Common Mistake: Blind Negative Sample Injection\n\nA tempting strategy is:\n> \"The model has false positives → add lots of negative samples.\"\n\nIn practice, this often leads to:\n- reduced false positives\n- **but also reduced recall**\n- suppressed low-confidence true positives\n- worse leaderboard score\n\nMore negatives ≠ better model.\n","metadata":{}},{"cell_type":"markdown","source":"## 📉 Why Too Many Negatives Hurt Recall\n\nAdding large volumes of negative samples shifts:\n- objectness calibration\n- confidence distribution\n- ranking behavior\n\nEffects observed:\n- borderline buildings are filtered out earlier\n- small or ambiguous buildings disappear first\n- low-conf recall collapses\n\nThis is especially harmful in recall-weighted metrics.\n","metadata":{}},{"cell_type":"markdown","source":"## ⚖️ Observed Trade-off in This Competition\n\nEmpirical observation:\n\n- Cleaner predictions (fewer FP)\n- ❌ do NOT guarantee higher leaderboard score\n- Slightly noisy predictions\n- ✅ often rank better if recall is preserved\n\nThis indicates the metric favors:\n- coverage over purity\n- ranking quality over visual cleanliness\n","metadata":{}},{"cell_type":"markdown","source":"## ✅ Recommended Data Strategy\n\nInstead of aggressively adding negatives:\n\n- Prioritize **scene diversity**\n- Preserve ambiguous positives\n- Add negatives **incrementally**\n- Continuously monitor recall impact\n\nGoal:\n> Let the model see *more kinds of scenes*, not just *fewer buildings*.\n","metadata":{}},{"cell_type":"markdown","source":"## 🌍 What “Learning More Scenes” Means in Practice\n\nUseful training diversity includes:\n- different lighting conditions\n- mixed natural + man-made backgrounds\n- partial occlusions\n- ambiguous colors and textures\n- small, isolated structures\n\nThe objective is not balance by count,\nbut **coverage by scenario**.\n","metadata":{}},{"cell_type":"markdown","source":"## 🧭 Practical Guideline for Using Synthetic Data\n\nSynthetic data is best used to:\n- complement real-world variability\n- fill obvious gaps\n- stress-test specific hypotheses\n\nSynthetic data should NOT be used to:\n- replace real complexity\n- over-regularize the model\n- enforce overly clean decision boundaries\n","metadata":{}},{"cell_type":"markdown","source":"## 🧠 Key Takeaway\n\nIn this competition:\n\n- Falcon synthetic data is useful but limited\n- Blind negative mining can backfire\n- Recall loss is often irreversible\n- Scene diversity matters more than label balance\n\nA better model is not one that is stricter,\nbut one that has *seen more of the world*.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #9: Falcon Synthetic Data Has an Expressiveness Ceiling\n\nBased on repeated experiments, Falcon/Duality synthetic data shows\na clear *expressiveness ceiling*.\n\nIt cannot reliably generate:\n- fine-grained spectral ambiguity (e.g. roof color ≈ water color)\n- realistic water reflections and clutter\n- complex rocky or mountainous terrain visually similar to rooftops\n- rare but common real-world non-building objects at scale\n\nThis is not a parameter issue — it is a generator capability limit.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #10: Failure Modes Can Be Clearly Observed but Not Synthesized\n\nA key realization from this project:\n\nSome failure modes are easy to identify through visualization,\nbut extremely difficult to recreate with digital-twin data.\n\nExamples include:\n- small blue or gray roofs near water being missed\n- buildings embedded in rocky or mountainous backgrounds\n- coastal houses partially blended into shoreline textures\n\nKnowing *where* the model fails does not guarantee\nthat synthetic data can be generated to fix it.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #11: Synthetic Data Often Encodes an Over-Clean Prior\n\nFalcon-generated scenes tend to be:\n- visually clean\n- geometrically regular\n- low in background clutter\n\nAs a result, the model may learn an implicit prior:\n\"clean, rectilinear structures are buildings\".\n\nThis prior performs well on synthetic validation data,\nbut generalizes poorly to cluttered real-world imagery.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #12: Many False Positives Are Geometric, Not Semantic\n\nMost false positives observed shared strong geometric cues,\nrather than true semantic similarity to buildings.\n\nCommon examples:\n- boats near docks\n- floating grid-like structures on water\n- vehicles aligned with roads\n- square agricultural terraces\n- thin vertical poles or linear clutter\n\nThese structures are underrepresented or simplified in synthetic data.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #13: Blind Negative Mining Backfires on Recall\n\nA natural reaction to false positives is to add more negative samples.\n\nIn practice, this repeatedly led to:\n- reduced false positives\n- but a noticeable drop in recall\n- especially for small or ambiguous buildings\n\nCleaner predictions did not translate to better leaderboard scores.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #14: Recall Loss Is Not Uniform\n\nWhen recall degradation occurred, it was highly asymmetric.\n\nThe first structures to disappear were:\n- small isolated houses\n- water-adjacent buildings\n- low-contrast roofs\n\nThese cases are both hard and disproportionately important\nfor leaderboard performance.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #15: Cleaner Predictions Can Score Worse\n\nAn unintuitive but consistent observation:\n\nModels producing visually cleaner predictions\n(fewer false positives, stricter filtering)\noften scored worse on the leaderboard.\n\nThis suggests the metric favors:\n- coverage over visual purity\n- ranking quality over strict objectness\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #16: Synthetic Data Should Increase Scene Diversity, Not Strictness\n\nThe primary value of synthetic data in this competition was:\n- exposing the model to more scene types\n\nIt was less effective when used to:\n- aggressively suppress false positives\n- enforce overly strict decision boundaries\n\nScene diversity mattered more than label balance.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #17: Over-Regularization Is Hard to Undo\n\nOnce recall was lost due to aggressive negative mining\nor over-regularization:\n\n- lowering confidence thresholds no longer helped\n- borderline true positives vanished early\n- leaderboard scores stagnated or dropped\n\nPreventing recall loss proved easier than recovering from it.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #18: The Goal Is Exposure, Not Suppression\n\nThe most effective models in this project were not the strictest ones.\n\nThey were models that had:\n- seen more kinds of scenes\n- learned to rank ambiguous cases\n- preserved low-confidence true positives\n\nThe objective was exposure to reality,\nnot suppression of uncertainty.\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #19: Why Segmentation Did Not Improve the Leaderboard\n\nSegmentation was explored to suppress texture-driven false positives\n(e.g. farmland patterns, water grids, rocky textures).\n\nDespite reasonable qualitative masks, segmentation-based filtering\ndid not improve leaderboard score and often reduced recall.\n\nThis was a consistent outcome across multiple trials.\n","metadata":{}},{"cell_type":"markdown","source":"def filter_boxes_by_mask(boxes, mask, threshold=0.3):\n    \"\"\"\n    boxes: list of [x1,y1,x2,y2] in pixel coords\n    mask: binary segmentation mask (H,W)\n    threshold: minimum mask coverage inside box\n    \"\"\"\n    kept = []\n    H, W = mask.shape\n    for b in boxes:\n        x1,y1,x2,y2 = map(int, b)\n        x1,y1 = max(0,x1), max(0,y1)\n        x2,y2 = min(W-1,x2), min(H-1,y2)\n\n        region = mask[y1:y2, x1:x2]\n        if region.size == 0:\n            continue\n\n        coverage = region.mean()\n        if coverage >= threshold:\n            kept.append(b)\n    return kept\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #20: Background Dominance Breaks Segmentation Utility\n\nIn remote sensing imagery:\n- background pixels dominate\n- buildings occupy a tiny fraction of area\n\nSegmentation models quickly learn to:\n\"predict background well\"\n\nThis leads to visually clean masks,\nbut poor discrimination for rare or ambiguous buildings.\n","metadata":{}},{"cell_type":"code","source":"def foreground_ratio(mask):\n    \"\"\"\n    mask: binary segmentation mask\n    returns fraction of foreground pixels\n    \"\"\"\n    return mask.mean()\n\n# Typical observed values:\n# foreground_ratio(mask) < 0.02  # <2% building pixels\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #21: Segmentation Is a One-Way Gate\n\nWhen segmentation is used as a pre- or post-filter:\n\n- false positives can be removed\n- but false negatives cannot be recovered\n\nOnce a building is masked out,\nlowering detection confidence cannot bring it back.\n","metadata":{}},{"cell_type":"code","source":"# Detection (soft gate)\nif conf >= 0.001:\n    keep_box = True  # ranking still preserved\n\n# Segmentation (hard gate)\nif mask_coverage < 0.3:\n    keep_box = False  # irreversible\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #22: YOLO-seg High mAP ≠ Leaderboard Gain\n\nYOLO-seg showed:\n- high mAP on validation\n- stable training curves\n\nBut leaderboard score did not improve.\n\nReason:\nvalidation segmentation metrics do not reflect\nranking quality of detection boxes.\n","metadata":{}},{"cell_type":"code","source":"# Compare number of predicted boxes at low confidence\ndef count_boxes(results, conf=0.001):\n    return sum(len(r.boxes) for r in results if r.boxes is not None)\n\n# Often observed:\n# YOLO-seg -> fewer boxes at low conf than YOLO-det\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #23: Post-processing Can Only Trim, Not Repair\n\nPost-processing (NMS variants, heuristics) helped slightly,\nbut only when applied conservatively.\n\nAggressive post-processing:\n- removed false positives\n- but damaged ranking and recall\n","metadata":{}},{"cell_type":"code","source":"def light_postprocess(boxes, confs, iou_thresh=0.6):\n    \"\"\"\n    Apply only standard NMS.\n    Avoid score decay or aggressive filtering.\n    \"\"\"\n    import torchvision.ops as ops\n    import torch\n\n    boxes_t = torch.tensor(boxes)\n    confs_t = torch.tensor(confs)\n    keep = ops.nms(boxes_t, confs_t, iou_thresh)\n    return keep.tolist()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #24: Confidence Sweeps Must Be Interpreted Carefully\n\nLowering confidence threshold:\n- can recover recall\n- but only if ranking quality is intact\n\nIf ranking collapses,\nlower confidence only amplifies noise.\n","metadata":{}},{"cell_type":"code","source":"def box_count_distribution(results, confs=[0.001, 0.01, 0.05]):\n    stats = {}\n    for c in confs:\n        stats[c] = [len(r.boxes) if r.boxes else 0 for r in results]\n    return stats\n\n# You looked at:\n# - mean box count\n# - extreme outliers\n# - stability across conf levels","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #25: A Pipeline Is Not Complete Without Submission Closure\n\nUp to this point, the pipeline covered:\n- data preparation\n- training\n- inference\n- error analysis\n\nHowever, in practice, a detection pipeline is **not complete**\nuntil predictions are transformed into a valid submission artifact.\n\nIn this competition, submission handling itself caused:\n- multiple zero-score submissions\n- misleading evaluations\n- false conclusions about model quality\n","metadata":{}},{"cell_type":"markdown","source":"## 🔎 Tiny Detail #26: Submission Is an Engineering Interface\n\nThe submission file is the interface between:\n- your model\n- and the competition evaluator\n\nAny mismatch in:\n- format\n- ordering\n- normalization\n- empty cases\n\ncan invalidate otherwise good predictions.\n\nSubmission must be treated as a first-class engineering component.\n","metadata":{}},{"cell_type":"code","source":"def build_submission(results, image_paths):\n    \"\"\"\n    results: model inference results (ordered)\n    image_paths: list of image paths used for inference\n    returns list of (image_id, prediction_string)\n    \"\"\"\n    submission = []\n\n    for r, p in zip(results, image_paths):\n        image_id = int(p.stem)  # critical: numeric ID\n\n        if r.boxes is None or len(r.boxes) == 0:\n            submission.append((image_id, \"no box\"))\n            continue\n\n        preds = []\n        h, w = r.orig_shape\n\n        for b in r.boxes:\n            cx, cy, bw, bh = b.xywhn[0].tolist()\n            conf = float(b.conf.item())\n            cls  = int(b.cls.item())\n\n            preds.append(\n                f\"{cls} {conf:.6f} {cx:.6f} {cy:.6f} {bw:.6f} {bh:.6f}\"\n            )\n\n        submission.append((image_id, \" \".join(preds)))\n\n    return submission\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def sanity_check_submission(rows, expected_images=None):\n    \"\"\"\n    rows: list of (image_id, prediction_string)\n    \"\"\"\n    ids = [r[0] for r in rows]\n\n    assert all(isinstance(i, int) for i in ids)\n    assert len(ids) == len(set(ids)), \"Duplicate image_id detected\"\n\n    for _, s in rows:\n        assert s == \"no box\" or len(s.split()) % 6 == 0\n\n    if expected_images is not None:\n        assert len(rows) == expected_images\n\n    print(\"✅ Submission sanity check passed\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## 🔎 Tiny Detail #27: Debugging Must Happen Before Submission\n\nOnce a submission is uploaded:\n- feedback is delayed\n- errors are opaque\n- iteration cost is high\n\nEffective debugging happens **before submission**:\n- visualize low-confidence predictions\n- inspect box count distribution\n- validate formatting and ordering\n\nTreat submission as the final test of the pipeline,\nnot the first.\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Beyond the Leaderboard: Transferable Lessons from a Remote Sensing Detection Project\n\nThis section summarizes lessons that extend beyond this specific competition.\n\nThe goal is not to optimize a leaderboard score,\nbut to extract principles that remain valid in:\n- other remote sensing competitions\n- applied geospatial analytics\n- real-world industrial pipelines\n\nAll conclusions here are grounded in experiments conducted in this project.\n","metadata":{}},{"cell_type":"markdown","source":"With the competition pipeline fully closed and evaluated,\nthe remaining question is how these findings translate beyond a leaderboard setting.\n","metadata":{}},{"cell_type":"markdown","source":"## 1. Detection Is a Ranking Problem, Not a Classification Problem\n\nA key realization from this project:\n\nObject detection performance is driven less by\n\"correct vs incorrect classification\"\nand more by the *ranking quality* of predictions.\n\nThis is especially true in recall-weighted metrics,\ncommon in remote sensing tasks.\n","metadata":{}},{"cell_type":"code","source":"# High-ranking but noisy predictions\n# often outperform clean but conservative ones\n\n# Example intuition (pseudo):\n# Model A: ranks true buildings early, includes noise later\n# Model B: suppresses noise, but ranks true buildings late\n\n# In recall-weighted evaluation:\n# Model A > Model B\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Recall Loss Is Often Irrecoverable\n\nOnce recall is lost due to:\n- aggressive negative mining\n- hard filtering\n- over-regularization\n\nit is difficult to recover by:\n- lowering confidence thresholds\n- post-processing\n- ensembling\n\nThis was repeatedly observed in this project.\n","metadata":{}},{"cell_type":"code","source":"# Bad pattern (irreversible):\nif confidence < threshold:\n    discard_detection()\n\n# Better pattern:\n# keep detections, rank them later, filter downstream\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Synthetic Data Is a Tool, Not a Substitute for Reality\n\nSynthetic data was valuable for:\n- expanding scene coverage\n- accelerating early training\n- filling obvious data gaps\n\nHowever, it consistently failed to reproduce:\n- fine-grained ambiguity\n- complex clutter\n- rare but impactful edge cases\n\nBlind reliance on synthetic data reduced generalization.\n","metadata":{}},{"cell_type":"code","source":"# Pseudo-strategy:\n# Use synthetic data to:\n#   - initialize\n#   - diversify\n# But rely on real data to:\n#   - calibrate\n#   - rank\n#   - generalize\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Segmentation Is a Support Signal, Not a Gatekeeper\n\nSegmentation models produced visually convincing masks,\nbut failed when used as hard filters.\n\nIn remote sensing:\n- segmentation is background-dominated\n- small objects are fragile\n- hard masks destroy recall\n\nSegmentation works best as a *supporting signal*,\nnot a decision gate.\n","metadata":{}},{"cell_type":"code","source":"# Recommended usage:\n# segmentation -> context / confidence adjustment\n# NOT -> binary acceptance / rejection\n\nadjusted_conf = conf * context_factor\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Submission Hygiene Mirrors Production Interfaces\n\nThe submission file acted as:\n- a strict interface\n- an unforgiving contract\n- a source of silent failures\n\nThis mirrors real-world production systems:\n- APIs\n- data contracts\n- downstream consumers\n\nEngineering rigor at interfaces is non-negotiable.\n","metadata":{}},{"cell_type":"code","source":"def validate_output_schema(output):\n    \"\"\"\n    Validate format, range, completeness.\n    \"\"\"\n    pass  # explicit validation is always cheaper than debugging downstream\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Visual Cleanliness Is Not Operational Quality\n\nA recurring trap:\n\nModels that *look* good in visualizations\noften performed worse quantitatively.\n\nIn operational settings:\n- missing critical objects is more costly\n- false positives can often be filtered later\n\nMetrics should reflect operational risk, not aesthetics.\n","metadata":{}},{"cell_type":"markdown","source":"## 7. Scene Diversity Matters More Than Dataset Balance\n\nTraditional dataset advice emphasizes class balance.\n\nIn remote sensing detection, this project showed:\n- scene diversity matters more\n- rare contexts dominate failure modes\n- balancing counts can hurt ranking\n\nCoverage of environments beats label symmetry.\n","metadata":{}},{"cell_type":"markdown","source":"## 8. A Practical Industrial Mindset for Remote Sensing AI\n\nEffective remote sensing systems are built by:\n- preserving uncertainty\n- ranking candidates\n- deferring hard decisions downstream\n\nThis mindset aligns better with:\n- large-scale monitoring\n- noisy environments\n- human-in-the-loop workflows\n","metadata":{}},{"cell_type":"markdown","source":"## Final Reflection\n\nThis project reinforced a core principle:\n\nSuccess in remote sensing AI is not about\nthe most complex model,\nbut about respecting uncertainty,\nunderstanding data limitations,\nand designing pipelines that fail gracefully.\n\nThese lessons extend far beyond a single competition.\n","metadata":{}},{"cell_type":"markdown","source":"This notebook focuses on methodological insights rather than leaderboard optimization,\nand avoids sharing competition-sensitive details.\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}