{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":123966,"databundleVersionId":14902028,"sourceType":"competition"},{"sourceId":14243819,"sourceType":"datasetVersion","datasetId":9087514},{"sourceId":14282123,"sourceType":"datasetVersion","datasetId":9115730},{"sourceId":14477767,"sourceType":"datasetVersion","datasetId":9087518},{"sourceId":14482902,"sourceType":"datasetVersion","datasetId":9097071}],"dockerImageVersionId":31236,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from IPython.display import Image\nImage(\"/kaggle/input/geospatialdetection-dualityai-images/Title_GOD.png\")","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-01-13T07:57:17.139558Z","iopub.execute_input":"2026-01-13T07:57:17.139806Z","iopub.status.idle":"2026-01-13T07:57:17.281184Z","shell.execute_reply.started":"2026-01-13T07:57:17.139779Z","shell.execute_reply":"2026-01-13T07:57:17.280477Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Here, **outside of the competition**, I was interested in testing how the model would behave outside of DualityAI\n- Through **HailuoAI**, using only one image, it turned out to generate a good video for 5 seconds.\n- Within the framework of this notebook, a comparison of the results will not be provided here.","metadata":{}},{"cell_type":"code","source":"from IPython.display import HTML\nfrom base64 import b64encode\n\ndef play(filename):\n    html = ''\n    video = open(filename,'rb').read()\n    src = 'data:video/mp4;base64,' + b64encode(video).decode()\n    html += '<video width=1000 controls autoplay loop><source src=\"%s\" type=\"video/mp4\"></video>' % src\n    return HTML(html)\n\nplay('/kaggle/input/geospatialdetection-dualityai-images/video_DualityAI.mp4')","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-01-13T07:57:31.259729Z","iopub.execute_input":"2026-01-13T07:57:31.260343Z","iopub.status.idle":"2026-01-13T07:57:31.326348Z","shell.execute_reply.started":"2026-01-13T07:57:31.260315Z","shell.execute_reply":"2026-01-13T07:57:31.325174Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Geospatial Object Detection with YOLO\n\n## 📌 Introduction\n\nThis notebook addresses the task of **geospatial object detection** as part of the  \n**Duality AI & Lunate AI Geospatial Object Detection** competition on Kaggle.\n\nThe goal of the competition is to develop a computer vision model capable of **detecting and localizing objects in aerial and satellite imagery** using bounding box annotations in YOLO format.  \nSuch tasks are highly relevant for real-world applications, including:\n\n- urban development and land-use analysis,\n- infrastructure monitoring,\n- remote sensing and Earth observation,\n- automated interpretation of large-scale geospatial data.\n\n## 📂 Dataset\n\nThe competition dataset consists of:\n- **geospatial images** (aerial / satellite imagery),\n- **object annotations in YOLO format** (bounding boxes),\n- multiple object classes of interest representing elements within the observed scenes.\n\nThe dataset follows the standard YOLO directory structure:\n\n```\n\ntrain_ALL/\n├── images/\n└── labels/\nval_ALL/\n├── images/\n└── labels/\n\n```\n\n## 🧠 Model Architecture\n\nFor this task, the **YOLOv8x, YOLOv10x, YOLO12x,** models from the **Ultralytics YOLO** framework was selected. Why x, because when testing nano and small models, I got a very low LB, no more than 0.32.\n\n## 🔄 Data Augmentation Strategy\n\nTo improve generalization and reduce overfitting, **online data augmentation** provided by YOLOv8 was extensively used during training.\n\nAn **adaptive augmentation strategy** was applied, where augmentation parameters are randomly sampled from predefined ranges (conservative / balanced / aggressive) at the start of training.\n\nThe augmentation pipeline includes:\n\n- **Color augmentations (HSV)**  \n  Adjustments of hue, saturation, and value simulate variations in illumination, atmospheric conditions, sensor properties, and roofing materials.  \n  These augmentations reduce sensitivity to lighting and color differences while encouraging the model to focus on structural features.\n\n- **Geometric transformations**  \n  - **Rotations (`degrees`)**: handle arbitrary building orientations in aerial views.  \n  - **Scaling (`scale`)**: account for different ground sampling distances and flight altitudes.  \n  - **Translations (`translate`)**: reduce positional bias of objects within the image.  \n  - **Perspective distortion (`perspective`)**: simulate slight camera tilt and off-nadir views.  \n  - **Shear (`shear`)**: increase robustness to affine distortions and minor geometric deformations.\n\n\n- **Flipping operations**  \n  - **Horizontal flip (`fliplr`)** and **vertical flip (`flipud`)** remove directional bias, as buildings have no inherent orientation in geospatial imagery.\n\n\n- **Composite augmentations**  \n  - **Mosaic**: combines multiple images into one, increasing object density and improving multi-scale detection.  \n  - **MixUp**: blends images to improve regularization and reduce overfitting.  \n  - **CutMix**: inserts regions from one image into another to increase spatial diversity (used cautiously).  \n  - **Close Mosaic**: disables Mosaic in the final training epochs to refine bounding box localization.  \n  - **Copy-Paste**: replicates building instances across scenes to enhance object-level diversity, especially beneficial for limited datasets.\n\nAll augmentations are applied **only during training** and are disabled for validation and testing.\n\n## ⚙️ Training Setup\n\nModel training was performed using:\n- the **AdamW and SGD** optimizer,\n- a fixed random seed to ensure reproducibility,\n- built-in YOLOv8x, YOLOv10l, YOLO12x logging and visualization tools.\n\nModel performance is evaluated using standard object detection metrics:\n- **Precision**\n- **Recall**\n- **mAP@0.5**\n- **mAP@0.5:0.95**","metadata":{}},{"cell_type":"markdown","source":"# Checking the GPU device","metadata":{}},{"cell_type":"code","source":"import torch\n# Check environment\nprint(\"Environment Check:\")\nprint(\"=\" * 50)\nprint(f\"PyTorch version: {torch.__version__}\")\nprint(f\"CUDA available: {torch.cuda.is_available()}\")\n\nif torch.cuda.is_available():\n    print(f\"GPU: {torch.cuda.get_device_name(0)}\")\n    print(\n        f\"GPU Memory: {torch.cuda.get_device_properties(0).total_memory / 1024**3:.1f} GB\"\n    )\nelse:\n    print(\"!!! No GPU detected - training will be slower on CPU\")\n\nprint(\"\\n All systems ready! Let's begin.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:06:36.638763Z","iopub.execute_input":"2026-01-13T08:06:36.639584Z","iopub.status.idle":"2026-01-13T08:06:40.291573Z","shell.execute_reply.started":"2026-01-13T08:06:36.639553Z","shell.execute_reply":"2026-01-13T08:06:40.290945Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Viewing data","metadata":{}},{"cell_type":"code","source":"import os\nimport math\nimport cv2\nimport random\nimport matplotlib.pyplot as plt","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:06:43.265873Z","iopub.execute_input":"2026-01-13T08:06:43.266556Z","iopub.status.idle":"2026-01-13T08:06:43.545118Z","shell.execute_reply.started":"2026-01-13T08:06:43.266526Z","shell.execute_reply":"2026-01-13T08:06:43.544571Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"IMAGES_DIR = \"/kaggle/input/geospatialdetection-dualityai/train_ALL/images\"\nLABELS_DIR = \"/kaggle/input/geospatialdetection-dualityai/train_ALL/labels\"\n\nSTART = 60\nEND = 69\nSTEP = 1         # step (1 = all)\nSHUFFLE = False  # True = random\n\nSHOW_CLASS_ID = True\nBOX_COLOR = (0, 255, 0)\nTHICKNESS = 2\nFONT_SCALE = 0.5\n\ndef load_yolo_labels(label_path):\n    boxes = []\n    if not os.path.exists(label_path):\n        return boxes\n\n    with open(label_path, \"r\") as f:\n        for line in f.readlines():\n            parts = line.strip().split()\n            if len(parts) != 5:\n                continue\n            cls, xc, yc, w, h = map(float, parts)\n            boxes.append((int(cls), xc, yc, w, h))\n    return boxes\n\n\ndef draw_boxes(img, boxes):\n    h, w, _ = img.shape\n    for cls, xc, yc, bw, bh in boxes:\n        x1 = int((xc - bw / 2) * w)\n        y1 = int((yc - bh / 2) * h)\n        x2 = int((xc + bw / 2) * w)\n        y2 = int((yc + bh / 2) * h)\n\n        cv2.rectangle(img, (x1, y1), (x2, y2), BOX_COLOR, THICKNESS)\n\n        if SHOW_CLASS_ID:\n            cv2.putText(img, f\"{cls}\", (x1, max(0, y1 - 5)), cv2.FONT_HERSHEY_SIMPLEX, FONT_SCALE, BOX_COLOR, 1, cv2.LINE_AA)\n    return img\n\nall_imgs = sorted([f for f in os.listdir(IMAGES_DIR) if f.lower().endswith((\".jpg\", \".jpeg\", \".png\"))])\n\nif SHUFFLE:\n    random.shuffle(all_imgs)\n\nsubset = all_imgs[START:END:STEP]\n\nMAX_IMAGES = 9\ngrid_imgs = subset[:MAX_IMAGES]\n\ncols = 3\nrows = math.ceil(len(grid_imgs) / cols)\n\nplt.figure(figsize=(15, 4 * rows))\nfor i, name in enumerate(grid_imgs):\n    img_path = os.path.join(IMAGES_DIR, name)\n    label_path = os.path.join(LABELS_DIR, os.path.splitext(name)[0] + \".txt\")\n\n    img = cv2.imread(img_path)\n    if img is None:\n        print(f\"Error to open - {name}\")\n        continue\n\n    boxes = load_yolo_labels(label_path)\n    img = draw_boxes(img, boxes)\n    img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n    plt.subplot(rows, cols, i + 1)\n    plt.imshow(img_rgb)\n    plt.axis(\"off\")\n    plt.title(f\"{name}\\nboxes: {len(boxes)}\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T09:42:31.684161Z","iopub.execute_input":"2026-01-12T09:42:31.684904Z","iopub.status.idle":"2026-01-12T09:42:34.482515Z","shell.execute_reply.started":"2026-01-12T09:42:31.684858Z","shell.execute_reply":"2026-01-12T09:42:34.481664Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Dataset Preparation and Training Pipeline / Train YOLO detection model","metadata":{}},{"cell_type":"markdown","source":"This section describes the complete data preparation and training pipeline used for YOLOv8-based object detection.  \nThe process is divided into three main stages:\n\n1. Creation of empty label files for images without objects  \n2. Safe merging of multiple dataset versions  \n3. Final training configuration and execution  \n\nEach stage is critical for achieving stable training and reliable mAP50 metrics.\n\n## 🟢 Stage 1: Creating Empty Label Files for Negative Samples\n\nYOLO requires that **every image has a corresponding label file**, even if no objects are present in the image.\n\nFor images containing **no target objects**, the corresponding `.txt` file must exist and be **empty**.  \nIf such files are missing, YOLO may:\n- silently skip images,\n- incorrectly compute statistics,\n- produce lower precision and unstable mAP values.\n\n### Affected datasets (check Stage 2, there is a description of the files)\nThe following dataset versions contain images **without any objects**, and therefore require empty label files:\n- `train_V5`, `val_V5`\n- `train_V6`, `val_V6`\n\n### Solution\nFor each image in these datasets, an empty `.txt` file is automatically created in the `labels/` directory if it does not already exist.\n\nThis ensures:\n- proper handling of negative samples,\n- correct learning of background vs. object distinction,\n- reduced false positives during inference.","metadata":{}},{"cell_type":"code","source":"import os\nfrom pathlib import Path\n\ndef create_empty_labels(dataset_dir):\n    images_dir = Path(dataset_dir) / \"images\"\n    labels_dir = Path(dataset_dir) / \"labels\"\n\n    labels_dir.mkdir(parents=True, exist_ok=True)\n    image_extensions = {\".jpg\", \".jpeg\", \".png\", \".tif\", \".tiff\"}\n    images = [p for p in images_dir.iterdir() if p.suffix.lower() in image_extensions]\n\n    created = 0\n    for img_path in images:\n        label_path = labels_dir / (img_path.stem + \".txt\")\n        if not label_path.exists():\n            label_path.touch()\n            created += 1\n\n    print(f\"[OK] {dataset_dir} → created {created} empty label files\")\n\nbase = r\"Jupyer Notebook Projects/Geospatial_Detection/scripts\"\n\ndatasets_without_objects = [\n                            \"train_V5\", \"val_V5\",\n                            \"train_V6\", \"val_V6\",\n                           ]\n\nfor ds in datasets_without_objects:\n    create_empty_labels(os.path.join(base, ds))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:07:27.845638Z","iopub.execute_input":"2026-01-13T08:07:27.846289Z","iopub.status.idle":"2026-01-13T08:07:27.852385Z","shell.execute_reply.started":"2026-01-13T08:07:27.846258Z","shell.execute_reply":"2026-01-13T08:07:27.851640Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🟡 Stage 2: Safe Merging of Multiple Dataset Versions\n\nAt this stage, the following folders were created:\n- train_V1  (5:00 - 21:00, H = 450 m, Intensity = 0.1, Texel Size = 0.2, other = default)\n- train_V2  (5:00 - 19:00, H = 400 m, Intensity = 0.1, Texel Size = 0.2, other = default)\n- train_V3  (5:00 - 17:00, H = 400 m, Intensity = 0.1, Texel Size = 0.2, other = default)\n- train_V4  (5:00 - 15:00, H = 400 m, Intensity = 0.1, Texel Size = 0.2, other = default)\n- train_V5  (oceanData without *.txt data) - this data was used as examples for which there are no targets\n- train_V6  (road_and_shore without *.txt data) - this data was used as examples for which there are no targets\n- train_V7  (5:00 - 21:00, H = 200 m, Intensity = 0.2, Texel Size = 0.4, other = default)\n- train_V8  (5:00 - 19:00, H = 200 m, Intensity = 0.2, Texel Size = 0.4, other = default)\n- train_V9  (5:00 - 17:00, H = 200 m, Intensity = 0.2, Texel Size = 0.4, other = default)\n- train_V10 (5:00 - 15:00, H = 200 m, Intensity = 0.2, Texel Size = 0.4, other = default)\n- train_V11 (5:00 - 13:00, H = 200 m, Intensity = 0.2, Texel Size = 0.4, other = default)\n- train_V12 (5:00 - 11:00, H = 200 m, Intensity = 0.2, Texel Size = 0.4, other = default)\n\nMax LB score on test data:\n- train_V1, val_V1 - **LB(max)=0.32** - around 300 pictures in train data - 1 session in Falcon\n- train_V1,...,trainV6,  val_V1,...,val_V6  - **LB(max)=0.53** - 2517 pictures in train data - 4 session in Falcon\n- train_V1,...,trainV12, val_V1,...,val_V12 - **LB(max)=0.52** - 4142 pictures in train data, 456 pictures in validation data - 10 sessions in Falcon\n\nValidation data was randomly taken from each folder in a random amount (10%)\n\nMultiple dataset versions (`train_V1`–`train_V12`, `val_V1`–`val_V12`) are combined into unified training and validation sets.\n\n### ⚠️ Important Issue: Filename Collisions\nSome datasets contain **images with identical filenames** but different content.\n\nIf merged directly:\n- files may be overwritten,\n- image–label pairs may become mismatched,\n- training quality may degrade significantly.\n\n### Solution: Dataset-Aware Renaming\nTo prevent conflicts, each file is renamed using a **dataset prefix** during merging.\n\nExample:\n- train_V2/images/000078.jpg → new file name: train_V2_000078.jpg\n- train_V5/images/000078.jpg → new file name: train_V5_000078.jpg\n\nThe same renaming logic is applied to label files, ensuring perfect correspondence between images and annotations.\n\n### Final dataset structure\nAfter merging, the unified dataset has the following structure:\n\n```\n\ntrain_ALL/\n├── images/\n└── labels/\n\nval_ALL/\n├── images/\n└── labels/\n\n````\n\nThis approach guarantees:\n- zero filename conflicts,\n- correct image–label alignment,\n- traceability of data origin,\n- full compatibility with YOLO training pipelines.","metadata":{}},{"cell_type":"code","source":"import shutil\n\ndef merge_datasets_with_prefix(source_dirs, target_dir):\n    target_images = Path(target_dir) / \"images\"\n    target_labels = Path(target_dir) / \"labels\"\n\n    target_images.mkdir(parents = True, exist_ok = True)\n    target_labels.mkdir(parents = True, exist_ok = True)\n\n    for src in source_dirs:\n        src = Path(src)\n        prefix = src.name  # train_V2, train_V3, etc.\n\n        src_images = src / \"images\"\n        src_labels = src / \"labels\"\n\n        for img in src_images.iterdir():\n            new_img_name = f\"{prefix}_{img.name}\"\n            shutil.copy(img, target_images / new_img_name)\n\n            label_src = src_labels / f\"{img.stem}.txt\"\n            label_dst = target_labels / f\"{prefix}_{img.stem}.txt\"\n\n            if label_src.exists():\n                shutil.copy(label_src, label_dst)\n            else:\n                label_dst.touch()\n\n        print(f\"[MERGED SAFELY] {src}\")\n\n# APPLY\nbase = r\"Jupyer Notebook Projects/Geospatial_Detection/scripts\"\n\ntrain_sources = [\"train_V1\",\n                 \"train_V2\",\n                 \"train_V3\",\n                 \"train_V4\",\n                 \"train_V5\",\n                 \"train_V6\",\n                 \"train_V7\",\n                 \"train_V8\",\n                 \"train_V9\",\n                 \"train_V10\",\n                 \"train_V11\",\n                 \"train_V12\",\n                ]\n\nval_sources = [\"val_V1\",\n               \"val_V2\",\n               \"val_V3\",\n               \"val_V4\",\n               \"val_V5\",\n               \"val_V6\",\n               \"val_V7\",\n               \"val_V8\",\n               \"val_V9\",\n               \"val_V10\",\n               \"val_V11\",\n               \"val_V12\",\n              ]\n\nmerge_datasets_with_prefix([Path(base) / d for d in train_sources], Path(base) / \"train_ALL\")\nmerge_datasets_with_prefix([Path(base) / d for d in val_sources], Path(base) / \"val_ALL\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:08:30.401143Z","iopub.execute_input":"2026-01-13T08:08:30.401435Z","iopub.status.idle":"2026-01-13T08:08:30.408363Z","shell.execute_reply.started":"2026-01-13T08:08:30.401410Z","shell.execute_reply":"2026-01-13T08:08:30.407689Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔵 Stage 3: Final Training Configuration\n\n### Dataset configuration (YAML)\n\nThe unified dataset is referenced through a single YAML file:\n\n```yaml\ntrain: \"/scripts/train_ALL\"\nval:   \"/scripts/val_ALL\"\ntest:  \"/testImages/images\"\n\nnc: 2\nnames: [\"building\", \"vehicle\"]\n````\n\n### Training setup\n\nThe YOLOv8x model is trained using:\n\n* **SGD (AdamW) optimizer** for improved generalization,\n* fixed random seed for reproducibility,\n* online data augmentation,\n* dropout regularization,\n* validation monitoring\n\nKey training parameters:\n\n* Model: `YOLOv8x,v10l,12x`\n* Classes: 2 (`building`, `vehicle`)\n* Training on merged datasets (`train_ALL`)\n* Validation on merged datasets (`val_ALL`)","metadata":{}},{"cell_type":"code","source":"! pip install ultralytics","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:16:27.440595Z","iopub.execute_input":"2026-01-13T08:16:27.441168Z","iopub.status.idle":"2026-01-13T08:16:32.913491Z","shell.execute_reply.started":"2026-01-13T08:16:27.441137Z","shell.execute_reply":"2026-01-13T08:16:32.912588Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🟤 YOLO V8 (optimizer = SGD) {data: V1, V2, V3, V4, V5, V6}","metadata":{}},{"cell_type":"code","source":"# ===================================================\n# ===================================================\n# I variant (yolov8x - optimizator - SGD, data = {train_V1, train_V2, ..., train_V5, train_V6})\n# Initially, I tested on small models of version 8 - yolov8n, yolov8s and increased the maximum LB to no more than 0.32, then I tried a larger model - yolov8x, where the LB improved to 0.50807.\n# I was also interested to know how the optimizer affects learning. For this task, I chose the AdamW and SGD optimizers.\n# I tested each optimizer on the yolov8x, yolov10l, and yolo12x models for a small number of epochs (10-15 epochs)\n# Both optimizers show a good score on the validation metrics map@50 and map@50-95 (>0.85)\n# But when using the AdamW optimizer, the LB score on the test data decreases by an average of 0.5-0.7.\n# The resulting optimizer was chosen - SGD or stochastic gradient descent\n# For reproducibility of the results, I fixed the value of seed = 42\n# The rationale for choosing data augmentation methods is given above\n# LB = map@50 (on test data) = 0.50807\n# ===================================================\n# ===================================================\n#augment         = True\n#mosaic          = 0.4\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.05\n#hsv_s           = 1.0\n#hsv_v           = 0.75\n\n#flipud          = 0.1\n#fliplr          = 0.6\n\n#translate       = 0.1\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 5\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================\n# ===================================================","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:46:32.595118Z","iopub.execute_input":"2026-01-13T08:46:32.595579Z","iopub.status.idle":"2026-01-13T08:46:32.619675Z","shell.execute_reply.started":"2026-01-13T08:46:32.595545Z","shell.execute_reply":"2026-01-13T08:46:32.618354Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport sys\nimport numpy as np\nimport random\nimport argparse\nimport yaml\nfrom pathlib import Path\nfrom ultralytics import YOLO","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:16:32.915263Z","iopub.execute_input":"2026-01-13T08:16:32.915535Z","iopub.status.idle":"2026-01-13T08:16:36.234004Z","shell.execute_reply.started":"2026-01-13T08:16:32.915509Z","shell.execute_reply":"2026-01-13T08:16:36.233155Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\n# =========================\n# =========================\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n# =========================\n# =========================\nparser = argparse.ArgumentParser()\n# =========================\n# Reproducibility\n# =========================\nparser.add_argument('--seed', type=int, default = 42, help='Random seed')\n# =========================\n# Training\n# =========================\nparser.add_argument('--epochs', type=int, default         = 50)\nparser.add_argument('--optimizer', type=str, default      = 'SGD')\nparser.add_argument('--lr0', type=float, default          = 0.002,  help = 'Initial learning rate')\nparser.add_argument('--lrf', type=float, default          = 0.0001, help = 'Final learning rate')\nparser.add_argument('--momentum', type=float, default     = 0.9)\nparser.add_argument('--weight_decay', type=float, default = 0.0001)\nparser.add_argument('--dropout', type=float, default      = 0.3)\nparser.add_argument('--cos_lr', action = 'store_true', help = 'Use cosine LR schedule', default = True)\nparser.add_argument('--patience', type=int, default       = 100)\nparser.add_argument('--save_period', type=int, default    = 10)\nparser.add_argument('--exist_ok', action = 'store_true', help = 'Overwrite existing experiment', default = True)\n\n# =========================\n# Dataset / Task\n# =========================\nparser.add_argument('--single_cls', action = 'store_true', help = 'Train as single-class', default = False)\nparser.add_argument('--workers', type = int, default = 4)\n\n# =========================\n# Augmentations (YOLO-compatible)\n# =========================\nparser.add_argument('--augment', action = 'store_true', default = True)\nparser.add_argument('--mosaic',       type=float, default = 0.4)\nparser.add_argument('--mixup',        type=float, default = 0.25)\nparser.add_argument('--copy_paste',   type=float, default = 0.05)\nparser.add_argument('--close_mosaic', type=int,   default = 10)\n\nparser.add_argument('--hsv_h', type=float, default        = 0.05)\nparser.add_argument('--hsv_s', type=float, default        = 1.0)\nparser.add_argument('--hsv_v', type=float, default        = 0.75)\n\nparser.add_argument('--flipud', type=float, default       = 0.1)\nparser.add_argument('--fliplr', type=float, default       = 0.6)\n\nparser.add_argument('--translate', type=float, default    = 0.1)\nparser.add_argument('--scale',     type=float, default    = 0.6)\nparser.add_argument('--shear',     type=float, default    = 0.02)\n\n# =========================\n# Warmup\n# =========================\nparser.add_argument('--warmup_epochs',   type=int,   default = 5)\nparser.add_argument('--warmup_momentum', type=float, default = 1.0)\n\n# =========================\n# Inference / Evaluation\n# =========================\nparser.add_argument('--conf', type=float, default = 0.25, help = 'Confidence threshold')\nparser.add_argument('--iou',  type=float, default = 0.5,  help = 'IoU threshold')\n\n# =========================\n# Logging / Visualization\n# =========================\nparser.add_argument('--plots', action = 'store_true', help = 'Enable plots', default = True)\n\nargs, _ = parser.parse_known_args()\nos.chdir(\"/kaggle/working\")\n\nmodel = YOLO(\"/kaggle/input/geospatialdetection-dualityai/yolov8x.pt\")\n\nresults = model.train(data         = \"/kaggle/input/geospatialdetection-dualityai/yolo_params_all.yaml\",\n                      device       = [0,1],\n                      epochs       = args.epochs,\n                      #batch        = 16,# auto\n                      #imgsz        = 512,# auto\n                      optimizer    = args.optimizer,\n                      lr0          = args.lr0,\n                      lrf          = args.lrf,\n                      weight_decay = args.weight_decay,\n                      momentum     = args.momentum,\n                      dropout      = args.dropout,\n                      #dfl          = 1.00,# auto\n                      cos_lr       = args.cos_lr,\n                      patience     = args.patience,\n                      save_period  = args.save_period,\n            \n                      seed         = args.seed,\n                      project      = \"/kaggle/working/runs\",\n                      name         = \"train_V8_SGD\",\n                      val          = True,\n                      exist_ok     = args.exist_ok,\n                      single_cls   = args.single_cls,\n                      plots        = args.plots,\n            \n                      augment      = args.augment,\n                      mosaic       = args.mosaic,\n                      mixup        = args.mixup,\n                      copy_paste   = args.copy_paste,\n                      close_mosaic = args.close_mosaic,\n                      hsv_h        = args.hsv_h, \n                      hsv_s        = args.hsv_s, \n                      hsv_v        = args.hsv_v,\n                      flipud       = args.flipud, \n                      fliplr       = args.fliplr,\n                      translate    = args.translate, \n                      scale        = args.scale, \n                      shear        = args.shear,\n            \n                      warmup_epochs   = args.warmup_epochs, \n                      warmup_momentum = args.warmup_momentum,\n                      workers         = args.workers,\n            \n                      conf = args.conf,\n                      iou  = args.iou,\n                     )\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"✅ TRAINING COMPLETE!\")\nprint(\"=\" * 70)\n\nprint(\"\\n📁 Model Weights Saved:\")\nprint(f\"   Best model: runs/detect/train/weights/best.pt\")\nprint(f\"   Last model: runs/detect/train/weights/last.pt\")\nprint(\"\\n Use 'best.pt' for predictions and submissions (highest validation mAP)\")\n\"\"\";","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-19T12:57:04.171439Z","iopub.execute_input":"2025-12-19T12:57:04.172229Z","iopub.status.idle":"2025-12-19T15:43:23.912947Z","shell.execute_reply.started":"2025-12-19T12:57:04.172193Z","shell.execute_reply":"2025-12-19T15:43:23.912132Z"},"jupyter":{"source_hidden":true},"_kg_hide-output":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🔴 YOLO V8 (optimizer = AdamW) {data: V1, V2, V3, V4, V5, V6}","metadata":{}},{"cell_type":"code","source":"# ===================================================\n# II variant (yolov10l - optimizator - SGD, data = {train_V1, train_V2, ..., train_V5, train_V6})\n# LB = 0.45554\n# ===================================================\n# ===================================================\n#augment         = True\n#mosaic          = 0.4\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.05\n#hsv_s           = 1.0\n#hsv_v           = 0.75\n\n#flipud          = 0.1\n#fliplr          = 0.6\n\n#translate       = 0.1\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 5\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================\n# ===================================================","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:13:33.013082Z","iopub.execute_input":"2026-01-13T08:13:33.013607Z","iopub.status.idle":"2026-01-13T08:13:33.017337Z","shell.execute_reply.started":"2026-01-13T08:13:33.013576Z","shell.execute_reply":"2026-01-13T08:13:33.016613Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\n# =========================\n# =========================\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n# =========================\n# =========================\nparser = argparse.ArgumentParser()\n\n# =========================\n# Reproducibility\n# =========================\nparser.add_argument('--seed', type=int, default = 42, help='Random seed')\n\n# =========================\n# Training\n# =========================\nparser.add_argument('--epochs', type=int, default         = 80)\nparser.add_argument('--optimizer', type=str, default      = 'AdamW')\nparser.add_argument('--lr0', type=float, default          = 0.005,  help = 'Initial learning rate')\nparser.add_argument('--lrf', type=float, default          = 0.00001, help = 'Final learning rate')\nparser.add_argument('--momentum', type=float, default     = 0.9)\nparser.add_argument('--weight_decay', type=float, default = 0.0001)\nparser.add_argument('--dropout', type=float, default      = 0.15)\nparser.add_argument('--cos_lr', action = 'store_true', help = 'Use cosine LR schedule', default = True)\nparser.add_argument('--patience', type=int, default       = 100)\nparser.add_argument('--save_period', type=int, default    = 10)\nparser.add_argument('--exist_ok', action = 'store_true', help = 'Overwrite existing experiment', default = True)\n\n# =========================\n# Dataset / Task\n# =========================\nparser.add_argument('--single_cls', action = 'store_true', help = 'Train as single-class', default = False)\nparser.add_argument('--workers', type = int, default = 4)\n\n# =========================\n# Augmentations (YOLO-compatible)\n# =========================\nparser.add_argument('--augment', action = 'store_true', default = True)\nparser.add_argument('--mosaic', type=float, default     = 0.4)\nparser.add_argument('--mixup', type=float, default      = 0.25)\nparser.add_argument('--cutmix', type=float, default     = 0.05)\nparser.add_argument('--copy_paste', type=float, default = 0.05)\nparser.add_argument('--close_mosaic', type=int, default = 10)\n\nparser.add_argument('--hsv_h', type=float, default      = 0.35)\nparser.add_argument('--hsv_s', type=float, default      = 0.69)\nparser.add_argument('--hsv_v', type=float, default      = 0.75)\n\nparser.add_argument('--flipud', type=float, default     = 0.3)\nparser.add_argument('--fliplr', type=float, default     = 0.3)\n\nparser.add_argument('--translate', type=float, default  = 0.1)\nparser.add_argument('--scale', type=float, default      = 0.6)\nparser.add_argument('--shear', type=float, default      = 0.02)\n\n# =========================\n# Warmup\n# =========================\nparser.add_argument('--warmup_epochs', type=int, default     = 5)\nparser.add_argument('--warmup_momentum', type=float, default = 1.0)\n\n# =========================\n# Inference / Evaluation\n# =========================\nparser.add_argument('--conf', type=float, default = 0.35, help = 'Confidence threshold')\nparser.add_argument('--iou',  type=float, default  = 0.50, help = 'IoU threshold')\n\n# =========================\n# Logging / Visualization\n# =========================\nparser.add_argument('--plots', action = 'store_true', help = 'Enable plots')\n\nargs, _ = parser.parse_known_args()\nos.chdir(\"/kaggle/working\")\n\nmodel = YOLO(\"/kaggle/input/geospatialdetection-dualityai/yolov8x.pt\")\n\nresults = model.train(data         = \"/kaggle/input/geospatialdetection-dualityai/yolo_params_all.yaml\",\n                      device       = [0,1],\n                      epochs       = args.epochs,\n                      #batch        = 16,# auto\n                      #imgsz        = 512,# auto\n                      optimizer    = args.optimizer,\n                      lr0          = args.lr0,\n                      lrf          = args.lrf,\n                      weight_decay = args.weight_decay,\n                      momentum     = args.momentum,\n                      dropout      = args.dropout,\n                      #dfl          = 1.00,# auto\n                      cos_lr       = args.cos_lr,\n                      patience     = args.patience,\n                      save_period  = args.save_period,\n            \n                      seed         = args.seed,\n                      project      = \"/kaggle/working/runs\",\n                      name         = \"train_V8_AdamW\",\n                      val          = True,\n                      exist_ok     = args.exist_ok,\n                      single_cls   = args.single_cls,\n                      plots        = args.plots,\n            \n                      augment      = args.augment,\n                      mosaic       = args.mosaic,\n                      mixup        = args.mixup,\n                      cutmix       = args.cutmix,\n                      copy_paste   = args.copy_paste,\n                      close_mosaic = args.close_mosaic,\n                      hsv_h        = args.hsv_h, \n                      hsv_s        = args.hsv_s, \n                      hsv_v        = args.hsv_v,\n                      flipud       = args.flipud, \n                      fliplr       = args.fliplr,\n                      translate    = args.translate, \n                      scale        = args.scale, \n                      shear        = args.shear,\n            \n                      warmup_epochs   = args.warmup_epochs, \n                      warmup_momentum = args.warmup_momentum,\n                      workers         = args.workers,\n            \n                      conf = args.conf,\n                      iou  = args.iou,\n                     )\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"✅ TRAINING COMPLETE!\")\nprint(\"=\" * 70)\n\nprint(\"\\n📁 Model Weights Saved:\")\nprint(f\"   Best model: runs/detect/train/weights/best.pt\")\nprint(f\"   Last model: runs/detect/train/weights/last.pt\")\nprint(\"\\n Use 'best.pt' for predictions and submissions (highest validation mAP)\")\n\"\"\";","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-22T08:56:43.312149Z","iopub.execute_input":"2025-12-22T08:56:43.312388Z","iopub.status.idle":"2025-12-22T13:26:37.531900Z","shell.execute_reply.started":"2025-12-22T08:56:43.312362Z","shell.execute_reply":"2025-12-22T13:26:37.530981Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true,"_kg_hide-output":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🟣 YOLO V10 (optimizer = SGD) {data: V1, V2, V3, V4, V5, V6}","metadata":{}},{"cell_type":"code","source":"# ===================================================\n# ===================================================\n# III variant (yolov10l - optimizator - SGD, data = {train_V1, train_V2, ..., train_V5, train_V6})\n# I chose the large version because there is a lack of GPU memory for the x version\n# LB = 0.49528\n# ===================================================\n# ===================================================\n#augment         = True\n#mosaic          = 0.4\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.05\n#hsv_s           = 1.0\n#hsv_v           = 0.75\n\n#flipud          = 0.1\n#fliplr          = 0.6\n\n#translate       = 0.1\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 5\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================\n# ===================================================","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:13:35.219742Z","iopub.execute_input":"2026-01-13T08:13:35.220477Z","iopub.status.idle":"2026-01-13T08:13:35.224010Z","shell.execute_reply.started":"2026-01-13T08:13:35.220450Z","shell.execute_reply":"2026-01-13T08:13:35.223245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\n# =========================\n# =========================\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n# =========================\n# =========================\n\nparser = argparse.ArgumentParser()\n\n# =========================\n# Reproducibility\n# =========================\nparser.add_argument('--seed', type=int, default = 42, help='Random seed')\n\n# =========================\n# Training\n# =========================\nparser.add_argument('--epochs', type=int, default         = 80)\nparser.add_argument('--optimizer', type=str, default      = 'SGD')\nparser.add_argument('--lr0', type=float, default          = 0.002,  help = 'Initial learning rate')\nparser.add_argument('--lrf', type=float, default          = 0.0001, help = 'Final learning rate')\nparser.add_argument('--momentum', type=float, default     = 0.9)\nparser.add_argument('--weight_decay', type=float, default = 0.0001)\nparser.add_argument('--dropout', type=float, default      = 0.3)\nparser.add_argument('--cos_lr', action = 'store_true', help = 'Use cosine LR schedule')\nparser.add_argument('--patience', type=int, default       = 100)\nparser.add_argument('--save_period', type=int, default    = 10)\nparser.add_argument('--exist_ok', action = 'store_true', help = 'Overwrite existing experiment')\n\n# =========================\n# Dataset / Task\n# =========================\nparser.add_argument('--single_cls', action = 'store_true', help = 'Train as single-class')\nparser.add_argument('--val', action = 'store_true', help = 'Val data analysis')\nparser.add_argument('--workers', type = int, default = 4)\n\n# =========================\n# Augmentations (YOLO-compatible)\n# =========================\nparser.add_argument('--augment', action = 'store_true')\nparser.add_argument('--mosaic', type=float, default     = 0.4)\nparser.add_argument('--mixup', type=float, default      = 0.25)\nparser.add_argument('--copy_paste', type=float, default = 0.05)\nparser.add_argument('--close_mosaic', type=int, default = 10)\n\nparser.add_argument('--hsv_h', type=float, default      = 0.05)\nparser.add_argument('--hsv_s', type=float, default      = 1.0)\nparser.add_argument('--hsv_v', type=float, default      = 0.75)\n\nparser.add_argument('--flipud', type=float, default     = 0.1)\nparser.add_argument('--fliplr', type=float, default     = 0.6)\n\nparser.add_argument('--translate', type=float, default  = 0.1)\nparser.add_argument('--scale', type=float, default      = 0.6)\nparser.add_argument('--shear', type=float, default      = 0.02)\n\n# =========================\n# Warmup\n# =========================\nparser.add_argument('--warmup_epochs', type=int, default     = 5)\nparser.add_argument('--warmup_momentum', type=float, default = 1.0)\n\n# =========================\n# Inference / Evaluation\n# =========================\nparser.add_argument('--conf', type=float, default = 0.25, help = 'Confidence threshold')\nparser.add_argument('--iou', type=float, default  = 0.5, help = 'IoU threshold')\n\n# =========================\n# Logging / Visualization\n# =========================\nparser.add_argument('--plots', action = 'store_true', help = 'Enable plots')\n\nargs, _ = parser.parse_known_args()\nos.chdir(\"/kaggle/working\")\n\nmodel = YOLO(\"/kaggle/input/geospatialdetection-dualityai-weights/yolov10l.pt\")\n\nresults = model.train(data         = \"/kaggle/input/geospatialdetection-dualityai/yolo_params_all.yaml\",\n                      device       = 0,\n                      epochs       = args.epochs,\n                      #batch        = 16,# auto\n                      #imgsz        = 512,# auto\n                      optimizer    = args.optimizer,\n                      lr0          = args.lr0,\n                      lrf          = args.lrf,\n                      weight_decay = args.weight_decay,\n                      momentum     = args.momentum,\n                      dropout      = args.dropout,\n                      #dfl          = 1.00,# auto\n                      cos_lr       = args.cos_lr,\n                      patience     = args.patience,\n                      save_period  = args.save_period,\n            \n                      seed         = args.seed,\n                      project      = \"/kaggle/working/runs\",\n                      name         = \"train_V7\",\n                      val          = args.val,\n                      exist_ok     = args.exist_ok,\n                      single_cls   = args.single_cls,\n                      plots        = args.plots,\n            \n                      augment      = args.augment,\n                      mosaic       = args.mosaic,\n                      mixup        = args.mixup,\n                      #cutmix       = args.cutmix, лучше поулчается без него\n                      copy_paste   = args.copy_paste,\n                      close_mosaic = args.close_mosaic,\n                      hsv_h        = args.hsv_h, \n                      hsv_s        = args.hsv_s, \n                      hsv_v        = args.hsv_v,\n                      flipud       = args.flipud, \n                      fliplr       = args.fliplr,\n                      translate    = args.translate, \n                      scale        = args.scale, \n                      shear        = args.shear,\n            \n                      warmup_epochs   = args.warmup_epochs, \n                      warmup_momentum = args.warmup_momentum,\n                      workers         = args.workers,\n            \n                      conf = args.conf,\n                      iou  = args.iou,\n                     )\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"✅ TRAINING COMPLETE!\")\nprint(\"=\" * 70)\n\nprint(\"\\n📁 Model Weights Saved:\")\nprint(f\"   Best model: runs/detect/train/weights/best.pt\")\nprint(f\"   Last model: runs/detect/train/weights/last.pt\")\nprint(\"\\n Use 'best.pt' for predictions and submissions (highest validation mAP)\")\n\n# 🟠 YOLO V12 (optimizer = SGD) {data: V1, V2, V3, V4, V5, V6}\n\"\"\";","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-22T14:35:21.727550Z","iopub.execute_input":"2025-12-22T14:35:21.727836Z","iopub.status.idle":"2025-12-22T18:19:52.594729Z","shell.execute_reply.started":"2025-12-22T14:35:21.727809Z","shell.execute_reply":"2025-12-22T18:19:52.593664Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"_kg_hide-output":true,"collapsed":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🟡 YOLO V12 (optimizer = SGD)","metadata":{}},{"cell_type":"code","source":"# ===================================================\n# imgsz = 640 (with higher quality (1024, 2048) - Error - Out of memory)\n# ===================================================\n# IV variant (yolov12x - optimizator - SGD, data = {train_V1, train_V2, ..., train_V5, train_V6})\n# LB = 0.52315 (the best on data V1-V6 - first augmentation)\n# LB = 0.53399 (the best on data V1-V6 - second augmentation)\n# ===================================================\n# ===================================================\n# **First augmentation version:** {LB = 0.52315}\n#coslr = True\n#dropout = 0.15\n#lr  = 0.0001\n#lr0 = 0.002\n#augment         = True\n#mosaic          = 0.4\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.05\n#hsv_s           = 1.0\n#hsv_v           = 0.75\n\n#flipud          = 0.1\n#fliplr          = 0.6\n\n#translate       = 0.1\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 5\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================\n# ===================================================\n# **Second augmentation version:** {LB = 0.47692}\n#coslr = False\n#dropout = 0.15\n#lr  = 0.0001\n#lr0 = 0.001\n#augment         = True\n#mosaic          = 0.6\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.55\n#hsv_s           = 0.65\n#hsv_v           = 0.75\n\n#flipud          = 0.6\n#fliplr          = 0.6\n\n#translate       = 0.5\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 10\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:13:37.191116Z","iopub.execute_input":"2026-01-13T08:13:37.191411Z","iopub.status.idle":"2026-01-13T08:13:37.195373Z","shell.execute_reply.started":"2026-01-13T08:13:37.191387Z","shell.execute_reply":"2026-01-13T08:13:37.194689Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\n# =========================\n# =========================\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n# =========================\n# =========================\nparser = argparse.ArgumentParser()\n\n# =========================\n# Reproducibility\n# =========================\nparser.add_argument('--seed', type=int, default = 42, help='Random seed')\n\n# =========================\n# Training\n# =========================\nparser.add_argument('--epochs',       type = int,   default = 80)\nparser.add_argument('--optimizer',    type = str,   default = 'SGD')\nparser.add_argument('--lr0',          type = float, default = 0.002,  help = 'Initial learning rate')\nparser.add_argument('--lrf',          type = float, default = 0.0001, help = 'Final learning rate')\nparser.add_argument('--momentum',     type = float, default = 0.9)\nparser.add_argument('--weight_decay', type = float, default = 0.0001)\nparser.add_argument('--dropout',      type = float, default = 0.3)\nparser.add_argument('--cos_lr',   action = 'store_true', help = 'Use cosine LR schedule',        default = True)\nparser.add_argument('--patience',     type = int,   default = 100)\nparser.add_argument('--save_period',  type = int,   default = 10)\nparser.add_argument('--exist_ok', action = 'store_true', help = 'Overwrite existing experiment', default = True)\n\n# =========================\n# Dataset / Task\n# =========================\nparser.add_argument('--single_cls', action = 'store_true', help = 'Train as single-class', default = False)\nparser.add_argument('--val',        action = 'store_true', help = 'Val data analysis',     default = True)\nparser.add_argument('--workers',    type = int,                                            default = 4)\n\n# =========================\n# Augmentations (YOLO-compatible)\n# =========================\nparser.add_argument('--augment', action = 'store_true', default = True)\nparser.add_argument('--mosaic',       type = float,     default = 0.4)\nparser.add_argument('--mixup',        type = float,     default = 0.25)\nparser.add_argument('--copy_paste',   type = float,     default = 0.05)\nparser.add_argument('--close_mosaic', type = int,       default = 10)\n\nparser.add_argument('--hsv_h', type = float, default = 0.05)\nparser.add_argument('--hsv_s', type = float, default = 1.0)\nparser.add_argument('--hsv_v', type = float, default = 0.75)\n\nparser.add_argument('--flipud', type = float, default  = 0.1)\nparser.add_argument('--fliplr', type = float, default  = 0.6)\n\nparser.add_argument('--translate', type = float, default = 0.1)\nparser.add_argument('--scale',     type = float, default = 0.6)\nparser.add_argument('--shear',     type = float, default = 0.02)\n\n# =========================\n# Warmup\n# =========================\nparser.add_argument('--warmup_epochs',   type = int,   default = 5)\nparser.add_argument('--warmup_momentum', type = float, default = 1.0)\n\n# =========================\n# Inference / Evaluation\n# =========================\nparser.add_argument('--conf', type = float, default = 0.25, help = 'Confidence threshold')\nparser.add_argument('--iou',  type = float, default = 0.50, help = 'IoU threshold')\n\n# =========================\n# Logging / Visualization\n# =========================\nparser.add_argument('--plots', action = 'store_true', help = 'Enable plots', default = True)\n\nargs, _ = parser.parse_known_args()\nos.chdir(\"/kaggle/working\")\n\nmodel = YOLO(\"/kaggle/input/geospatialdetection-dualityai-weights/yolo12x.pt\")\n\nresults = model.train(data         = \"/kaggle/input/geospatialdetection-dualityai/yolo_params_all.yaml\",\n                      device       = [0,1],\n                      epochs       = args.epochs,\n                      #batch        = 16,# auto 16\n                      #imgsz        = 512,# auto 640\n                      optimizer    = args.optimizer,\n                      lr0          = args.lr0,\n                      lrf          = args.lrf,\n                      weight_decay = args.weight_decay,\n                      momentum     = args.momentum,\n                      dropout      = args.dropout,\n                      #dfl          = 1.00,# auto\n                      cos_lr       = True,\n                      patience     = args.patience,\n                      save_period  = args.save_period,\n            \n                      seed         = args.seed,\n                      project      = \"/kaggle/working/runs\",\n                      name         = \"train_V12\",\n                      val          = True,\n                      exist_ok     = True,\n                      single_cls   = False,\n                      plots        = True,\n            \n                      augment      = True,\n                      mosaic       = args.mosaic,\n                      mixup        = args.mixup,\n                      copy_paste   = args.copy_paste,\n                      close_mosaic = args.close_mosaic,\n                      hsv_h        = args.hsv_h, \n                      hsv_s        = args.hsv_s, \n                      hsv_v        = args.hsv_v,\n                      flipud       = args.flipud, \n                      fliplr       = args.fliplr,\n                      translate    = args.translate, \n                      scale        = args.scale, \n                      shear        = args.shear,\n            \n                      warmup_epochs   = args.warmup_epochs, \n                      warmup_momentum = args.warmup_momentum,\n                      workers         = args.workers,\n            \n                      conf = args.conf,\n                      iou  = args.iou,\n                     )\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"✅ TRAINING COMPLETE!\")\nprint(\"=\" * 70)\n\nprint(\"\\n📁 Model Weights Saved:\")\nprint(f\"   Best model: runs/detect/train/weights/best.pt\")\nprint(f\"   Last model: runs/detect/train/weights/last.pt\")\nprint(\"\\n Use 'best.pt' for predictions and submissions (highest validation mAP)\")\n\"\"\";","metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ===================================================\n# imgsz = 640 (with higher quality (1024, 2048) - Error - Out of memory)\n# ===================================================\n# V variant (yolov12x - optimizator - SGD, data = {train_V1, train_V2, ..., train_V11, train_V12})\n# LB = 0.49023 (the best on data V1-V12 - first augmentation)\n# LB = 0.47547 (the best on data V1-V12 - second augmentation)\n# ===================================================\n# ===================================================\n# **First augmentation version:** {LB = 0.49023}\n#augment         = True\n#mosaic          = 0.4\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.05\n#hsv_s           = 1.0\n#hsv_v           = 0.75\n\n#flipud          = 0.1\n#fliplr          = 0.6\n\n#translate       = 0.1\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 5\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================\n# ===================================================\n# **Second augmentation version:** {LB = 0.47547}\n#augment         = True\n#mosaic          = 0.1\n#mixup           = 0.25\n#copy_paste      = 0.05\n#close_mosaic    = 10\n\n#hsv_h           = 0.55\n#hsv_s           = 0.65\n#hsv_v           = 0.75\n\n#flipud          = 0.6\n#fliplr          = 0.6\n\n#translate       = 0.1\n#scale           = 0.6\n#shear           = 0.02\n\n#warmup_epochs   = 5\n#warmup_momentum = 1.0\n\n#conf            = 0.25\n#iou             = 0.5\n# ===================================================\n# ===================================================\n# Small CONF → higher than recall → often higher than mAP@50\n# Large CONF → the model looks “stricter\", but loses recall\n\n# IoU in train = affects how many takes are left\n# With high IoU, NMS is softer → higher recall\n# At low IoU, NMS is more aggressive → less FP, but more FN\n# ===================================================\n# ===================================================","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:14:33.533027Z","iopub.execute_input":"2026-01-13T08:14:33.533672Z","iopub.status.idle":"2026-01-13T08:14:33.537635Z","shell.execute_reply.started":"2026-01-13T08:14:33.533644Z","shell.execute_reply":"2026-01-13T08:14:33.536928Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# =========================\n# =========================\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\n# =========================\n# =========================\nparser = argparse.ArgumentParser()\n\n# =========================\n# Reproducibility\n# =========================\nparser.add_argument('--seed',    type=int,   default = SEED,                   help = 'Random seed')\nparser.add_argument('--project', type = str, default = \"/kaggle/working/runs\", help = 'Project_Path')\nparser.add_argument('--name',    type = str, default = \"train_V12\",            help = 'Project_Name')\n\n# =========================\n# Training\n# =========================\nparser.add_argument('--epochs',       type = int,   default = 100)\nparser.add_argument('--batch',        type = int,   default = 4)\nparser.add_argument('--imgsz',        type = int,   default = 640)\nparser.add_argument('--optimizer',    type = str,   default = 'SGD')\nparser.add_argument('--lr0',          type = float, default = 0.001,  help = 'Initial learning rate')\nparser.add_argument('--lrf',          type = float, default = 0.0001, help = 'Final learning rate')\nparser.add_argument('--momentum',     type = float, default = 0.9)\nparser.add_argument('--weight_decay', type = float, default = 0.0001)\nparser.add_argument('--dropout',      type = float, default = 0.15)\nparser.add_argument('--cos_lr',   action = 'store_true', help = 'Use cosine LR schedule', default = False)\nparser.add_argument('--dfl',          type = float, default = 1.2)\nparser.add_argument('--patience',     type = int,   default = 100)\nparser.add_argument('--save_period',  type = int,   default = 10)\n\n# =========================\n# Dataset / Task\n# =========================\nparser.add_argument('--val',        action = 'store_true', help = 'Val data analysis',           default = True)\nparser.add_argument('--show',       action = 'store_true', help = 'show res',                    default = True)\nparser.add_argument('--exist_ok', action = 'store_true', help = 'Overwrite existing experiment', default = True)\nparser.add_argument('--single_cls', action = 'store_true', help = 'Train as single-class',       default = False)\nparser.add_argument('--plots',      action = 'store_true', help = 'final grap',                  default = True)\nparser.add_argument('--workers',    type = int,                                                  default = 8)\n\n# =========================\n# Augmentations (YOLO-compatible)\n# =========================\nparser.add_argument('--augment', action = 'store_true', default = True)\nparser.add_argument('--mosaic', type=float, default     = 0.1)\nparser.add_argument('--mixup', type=float, default      = 0.25)\nparser.add_argument('--copy_paste', type=float, default = 0.05)\nparser.add_argument('--close_mosaic', type=int, default = 10)\n\nparser.add_argument('--hsv_h', type=float, default      = 0.05)\nparser.add_argument('--hsv_s', type=float, default      = 1.00)\nparser.add_argument('--hsv_v', type=float, default      = 0.75)\n\nparser.add_argument('--flipud', type=float, default     = 0.1)\nparser.add_argument('--fliplr', type=float, default     = 0.6)\n\nparser.add_argument('--translate', type=float, default  = 0.1)\nparser.add_argument('--scale', type=float, default      = 0.6)\nparser.add_argument('--shear', type=float, default      = 0.02)\n\n# =========================\n# Warmup\n# =========================\nparser.add_argument('--warmup_epochs',   type = int,   default = 5)\nparser.add_argument('--warmup_momentum', type = float, default = 1.0)\n\n# =========================\n# Inference / Evaluation\n# =========================\nparser.add_argument('--conf', type = float, default = 0.25, help = 'Confidence threshold')\nparser.add_argument('--iou',  type = float, default = 0.50, help = 'IoU threshold')\n\nargs, _ = parser.parse_known_args()\nos.chdir(\"/kaggle/working\")\n\nmodel = YOLO(\"/kaggle/input/geospatialdetection-dualityai-weights/yolo12x.pt\")\n\n# Let's make a fine-tune of the model weights calculated from V1-V6 (V1-V12) data\n# This approach worsened the detection results, we are trying to refine the model from scratch\n\"\"\"\nargs_path = Path(\"/kaggle/working/runs/train_V12/args.yaml\")\nwith open(args_path) as f:\n    args = yaml.safe_load(f)\n\nargs[\"epochs\"] = 100\n\nwith open(args_path, \"w\") as f:\n    yaml.safe_dump(args, f)\n\nprint(\"Updated epochs to:\", args[\"epochs\"])\n\nmodel = YOLO(\"/kaggle/working/runs/train_V12/weights/last.pt\",)\nresults = model.train(resume = True)\n\"\"\";\n\nresults = model.train(data         = \"/kaggle/input/geospatialdetection-dualityai/yolo_params_all.yaml\",\n                      device       = [0,1],\n                      epochs       = args.epochs,\n                      batch        = args.batch,         # auto 16\n                      imgsz        = args.imgsz,         # auto 640\n                      optimizer    = args.optimizer,\n                      lr0          = args.lr0,\n                      lrf          = args.lrf,\n                      #weight_decay = args.weight_decay,  # l2 regularization term\n                      #momentum     = args.momentum,      # momentum factor for SGD or beta1 for Adam optimizers,\n                      dropout      = args.dropout,\n                      #dfl          = args.dfl,           # auto\n                      cos_lr       = args.cos_lr,        # utilizes a cosine learning rate scheduler \n                      patience     = args.patience,\n                      save_period  = args.save_period,\n            \n                      seed         = args.seed,\n                      project      = args.project,\n                      name         = args.name,\n                      val          = args.val,\n                      show         = args.show,\n                      exist_ok     = args.exist_ok,\n                      single_cls   = args.single_cls,\n                      plots        = args.plots,\n            \n                      augment      = args.augment,     # validation augmentation\n                      mosaic       = args.mosaic,      # combines 4 training images into one\n                      mixup        = args.mixup,       # mixes 2 images and their labels, creating a composite image\n                      copy_paste   = args.copy_paste,  # complements the training photo with additional objects from other photos\n                      close_mosaic = args.close_mosaic,# disables mosaic data augmentation in the last N epochs \n                      hsv_h        = args.hsv_h,       # allows you to adjust the hue of the image by a small fraction of the color circle\n                      hsv_s        = args.hsv_s,       # slightly changes the saturation of the image, affecting the intensity of colors\n                      hsv_v        = args.hsv_v,       # changes the value (brightness) of the image by an insignificant amount.\n                      flipud       = args.flipud,      # up down flip\n                      fliplr       = args.fliplr,      # left right flip\n                      translate    = args.translate,   # converts an image horizontally and vertically by a fraction of the image size\n                      scale        = args.scale,       # scales the image by a gain factor\n                      shear        = args.shear,       # shifts the image by a preset degree, viewing objects from different angles\n            \n                      warmup_epochs   = args.warmup_epochs,  # gradually increasing the learning rate from a low value to the initial learning rate\n                      warmup_momentum = args.warmup_momentum,# initial momentum for warmup phase\n                      workers         = args.workers,        # number of worker threads for data loading\n            \n                      conf = args.conf,                      # sets the minimum confidence threshold for detections\n                      iou  = args.iou,                       # sets the Intersection Over Union threshold for Non-Maximum Suppression (NMS)\n                     )\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"✅ TRAINING COMPLETE!\")\nprint(\"=\" * 70)\n\nprint(\"\\n📁 Model Weights Saved:\")\nprint(f\"   Best model: runs/detect/train/weights/best.pt\")\nprint(f\"   Last model: runs/detect/train/weights/last.pt\")\nprint(\"\\n Use 'best.pt' for predictions and submissions (highest validation mAP)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:16:40.647160Z","iopub.execute_input":"2026-01-13T08:16:40.647852Z","iopub.status.idle":"2026-01-13T08:16:42.462445Z","shell.execute_reply.started":"2026-01-13T08:16:40.647817Z","shell.execute_reply":"2026-01-13T08:16:42.461810Z"},"_kg_hide-output":false},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(\"/kaggle/input/geospatialdetection-dualityai-images/Tables_YOLO_Models.png\")","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-01-13T08:16:58.772103Z","iopub.execute_input":"2026-01-13T08:16:58.772424Z","iopub.status.idle":"2026-01-13T08:16:58.788842Z","shell.execute_reply.started":"2026-01-13T08:16:58.772395Z","shell.execute_reply":"2026-01-13T08:16:58.788135Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Key Findings of the Study:**\n\n1.  The utilization of the AdamW **optimizer** adversely impacts the precision of bounding box predictions on the test set, even though the mAP@50 (>0.94) and mAP@50-95 (>0.820) metrics exhibit only marginal fluctuations.\n2.  The **`yolo12x.pt`** model yielded superior performance and was identified as optimal for the final prediction task.\n3.  Employing pre-trained models (**transfer learning**) results in a degradation of bounding box prediction accuracy.\n4.  The following variables were found to have a negligible impact on the model's final performance: the composition of the training **dataset**, the altitude of image capture, as well as the intensity and size of the texel.\n5.  Among the data augmentations tested, **mosaic** proved to be the most impactful. Its selective application alone achieved an LB Score of 0.40.","metadata":{}},{"cell_type":"markdown","source":"# Predict on Test data with yolo12x model","metadata":{}},{"cell_type":"code","source":"from ultralytics import YOLO\nfrom pathlib import Path\nimport cv2\nimport os\nimport yaml\n\n# -----------------------------\n# Prediction function\n# -----------------------------\ndef predict_and_save(model, image_path, output_path_img, output_path_txt):\n    # Small conf = 0.0001, because Infuence of Recall more important then Infuence of Precision\n    # It doesn't matter how many false boxes we have, it's more important whether we hit the object with them or not.\n    # R - Of all the real objects, how many did the model find? \n    # The metric depends on the number of missed boxes.\n    # Box(P) - Of all the boxes that the model predicted, how many were superfluous?\n    # Here, on the contrary, the metric depends on the number of superfluous boxes.\n    # default:\n    # conf = 0.25\n    # iou  = 0.7\n    results = model.predict(source = image_path, \n                            imgsz  = 1024, \n                            conf   = 0.0001,\n                           )\n    result = results[0]\n\n    img = result.plot()\n    cv2.imwrite(str(output_path_img), img)\n\n    with open(output_path_txt, 'w') as f:\n        if result.boxes is None:\n            return\n        for box in result.boxes:\n            cls_id = int(box.cls)\n            x_center, y_center, width, height = box.xywhn[0].tolist()\n            conf = float(box.conf[0])\n            f.write(f\"{cls_id} {conf:.6f} {x_center} {y_center} {width} {height}\\n\")\n\n# -----------------------------\n# Paths and config\n# -----------------------------\nos.chdir(\"/kaggle/working\")\n\n#yaml_path = \"/kaggle/input/geospatialdetection-final-dataset/yolo_params_all.yaml\"\nyaml_path  = \"/kaggle/input/geospatialdetection-dualityai/yolo_params_all.yaml\"\nwith open(yaml_path, 'r') as f:\n    data = yaml.safe_load(f)\n\nif 'test' not in data or data['test'] is None:\n    raise ValueError(\"❌ No 'test' field in YAML\")\n\ntest_path = Path(data['test'])\n\n# ✅ robust test images path handling\nif (test_path / \"images\").exists():\n    images_dir = test_path / \"images\"\nelse:\n    images_dir = test_path\n\nif not images_dir.exists() or not images_dir.is_dir():\n    raise FileNotFoundError(f\"❌ Test images directory not found: {images_dir}\")\n\nif not any(images_dir.iterdir()):\n    raise ValueError(f\"❌ Test images directory is empty: {images_dir}\")\n\n# -----------------------------\n# Load model\n# -----------------------------\n# Change optimizator\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolov8x_SGD.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolov8x_AdamW.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolov10l_SGD.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolov10l_AdamW.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo12x_SGD.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo12x_AdamW.pt\"\n\n# Pre-trained(PT) or None pre-trained (NonePT) model with first augmentation (1)\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo8x_V1-V6_SGD_PT.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo10l_V1-V6_SGD_PT.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo12x_V1-V6_SGD_PT.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo12x_V1-V12_SGD_PT.pt\"\n#model_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_yolo12x_V1-V12_NonePT.pt\"\n\n# Final Weights\nmodel_path = \"/kaggle/input/geospatialdetection-dualityai-weights/best_Submission12.pt\"\nmodel = YOLO(model_path)\n\n# -----------------------------\n# Output directories\n# -----------------------------\noutput_dir = Path(\"/kaggle/working/predictions\")\nimages_output_dir = output_dir / \"images\"\nlabels_output_dir = output_dir / \"labels\"\n\nimages_output_dir.mkdir(parents = True, exist_ok = True)\nlabels_output_dir.mkdir(parents = True, exist_ok = True)\n\n# -----------------------------\n# Run inference\n# -----------------------------\nfor img_path in images_dir.glob(\"*\"):\n    if img_path.suffix.lower() not in [\".jpg\", \".png\", \".jpeg\"]:\n        continue\n\n    output_img = images_output_dir / img_path.name\n    output_txt = labels_output_dir / (img_path.stem + \".txt\")\n\n    predict_and_save(model, img_path, output_img, output_txt)\n\nprint(f\"✅ Predicted images saved in: {images_output_dir}\")\nprint(f\"✅ Bounding box labels saved in: {labels_output_dir}\")\n\n# -----------------------------\n# Evaluation\n# -----------------------------\nmetrics = model.val(data = yaml_path, split = \"test\")\nprint(metrics)","metadata":{"scrolled":true,"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T18:26:18.419213Z","iopub.execute_input":"2026-01-12T18:26:18.419776Z","iopub.status.idle":"2026-01-12T18:52:09.837896Z","shell.execute_reply.started":"2026-01-12T18:26:18.419746Z","shell.execute_reply":"2026-01-12T18:52:09.837193Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Loss and eval metrics on train data (YOLO V12x - dataset V0 - V6)","metadata":{}},{"cell_type":"code","source":"# Results of best (LB = 0.52) model\nfrom IPython.display import Image\nImage(\"/kaggle/input/geospatialdetection-dualityai-images/BoxF1_curve.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:38.979835Z","iopub.execute_input":"2026-01-13T08:24:38.980180Z","iopub.status.idle":"2026-01-13T08:24:38.995097Z","shell.execute_reply.started":"2026-01-13T08:24:38.980152Z","shell.execute_reply":"2026-01-13T08:24:38.994502Z"},"_kg_hide-output":false},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(\"/kaggle/input/geospatialdetection-dualityai-images/BoxPR_curve.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:39.277882Z","iopub.execute_input":"2026-01-13T08:24:39.278148Z","iopub.status.idle":"2026-01-13T08:24:39.288228Z","shell.execute_reply.started":"2026-01-13T08:24:39.278123Z","shell.execute_reply":"2026-01-13T08:24:39.287586Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(\"/kaggle/input/geospatialdetection-dualityai-images/BoxP_curve.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:44.241899Z","iopub.execute_input":"2026-01-13T08:24:44.242231Z","iopub.status.idle":"2026-01-13T08:24:44.253137Z","shell.execute_reply.started":"2026-01-13T08:24:44.242200Z","shell.execute_reply":"2026-01-13T08:24:44.252525Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(\"/kaggle/input/geospatialdetection-dualityai-images/BoxR_curve.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:44.435479Z","iopub.execute_input":"2026-01-13T08:24:44.435732Z","iopub.status.idle":"2026-01-13T08:24:44.447607Z","shell.execute_reply.started":"2026-01-13T08:24:44.435707Z","shell.execute_reply":"2026-01-13T08:24:44.446955Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(\"/kaggle/input/geospatialdetection-dualityai-images/confusion_matrix_normalized.png\")","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:47.890464Z","iopub.execute_input":"2026-01-13T08:24:47.891079Z","iopub.status.idle":"2026-01-13T08:24:47.902618Z","shell.execute_reply.started":"2026-01-13T08:24:47.891041Z","shell.execute_reply":"2026-01-13T08:24:47.901917Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(\"/kaggle/input/geospatialdetection-dualityai-images/results.png\")","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:48.071286Z","iopub.execute_input":"2026-01-13T08:24:48.071512Z","iopub.status.idle":"2026-01-13T08:24:48.084965Z","shell.execute_reply.started":"2026-01-13T08:24:48.071489Z","shell.execute_reply":"2026-01-13T08:24:48.084287Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Validation and Pred batch labels\nImage(\"/kaggle/input/geospatialdetection-dualityai-images/Val_Pred_Batch_0.png\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:24:55.969518Z","iopub.execute_input":"2026-01-13T08:24:55.969850Z","iopub.status.idle":"2026-01-13T08:24:56.067842Z","shell.execute_reply.started":"2026-01-13T08:24:55.969820Z","shell.execute_reply":"2026-01-13T08:24:56.067090Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Save predictions to *.csv file","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nimport pandas as pd\nimport csv\nimport sys\n\ndef predictions_to_csv(preds_folder: str = \"/kaggle/working/predictions/labels\", \n                       output_csv: str = \"/kaggle/working/submission_yolov12x_(16)_augmentation_V1-V6.csv\", \n                       test_images_folder: str = \"/kaggle/input/geospatialdetection-final-dataset/testImages/images\",\n                       allowed_extensions: tuple = (\".jpg\", \".png\", \".jpeg\")\n                      ):\n\n    # Convert YOLO prediction files to Kaggle submission CSV format with strict validation.\n    # Validate inno boxputs\n    preds_path = Path(preds_folder)\n    if not preds_path.exists():\n        print(f\"ERROR: Prediction folder '{preds_folder}' does not exist\")\n        sys.exit(1)\n\n    # Get test image IDs (without extensions)\n    test_images_path = Path(test_images_folder)\n    if not test_images_path.exists():\n        print(f\"ERROR: Test images folder '{test_images_folder}' not found\")\n        sys.exit(1)\n        \n    test_images = {p.stem: True \n                   for p in test_images_path.glob(\"*\")\n                   if p.suffix.lower() in allowed_extensions\n                  }\n    print(f\"Found {len(test_images)} test images\")\n\n    # Collect predictions with validation\n    predictions = []\n    error_count = 0\n    \n    for txt_file in preds_path.glob(\"*.txt\"):\n        image_id = txt_file.stem\n        \n        # Validate image_id\n        if image_id not in test_images:\n            print(f\"Skipping non-test image prediction: {txt_file.name}\")\n            continue\n            \n        with open(txt_file, \"r\") as f:\n            valid_lines = []\n            for line_num, line in enumerate(f, 1):\n                line = line.strip()\n                if not line:\n                    continue  # Skip empty lines\n                    \n                parts = line.split()\n                # Validate YOLO format: 6 values per line\n                if len(parts) != 6:\n                    print(f\"Invalid prediction in {txt_file.name} line {line_num}: {line}\")\n                    error_count += 1\n                    continue\n                    \n                try:\n                    # Validate numerical values\n                    [float(x) for x in parts]\n                    valid_lines.append(line)\n                except ValueError:\n                    print(f\"Non-numeric values in {txt_file.name} line {line_num}: {line}\")\n                    error_count += 1\n                    continue\n\n        pred_str = \" \".join(valid_lines) if valid_lines else \"no box\"\n        predictions.append({\"image_id\": image_id, \"prediction_string\": pred_str})\n\n    # Create submission dataframe\n    submission_df = pd.DataFrame({\"image_id\": list(test_images.keys())})\n    \n    if predictions:\n        preds_df = pd.DataFrame(predictions)\n        final_df = submission_df.merge(preds_df, on=\"image_id\", how=\"left\").fillna(\"no boxes\")\n    else:\n        final_df = submission_df\n        final_df[\"prediction_string\"] = \"no boxes\"\n\n    # Save with CSV quoting rules\n    final_df.to_csv(output_csv, index=False, quoting=csv.QUOTE_NONNUMERIC)\n    \n    print(f\"\\n Success! Submission saved to {output_csv}\")\n    print(f\"   Total predictions: {len(predictions)}\")\n    print(f\"   Validation errors: {error_count}\")\n\nimport argparse\nparser = argparse.ArgumentParser(description = \"Convert YOLO predictions to Kaggle CSV\")\nparser.add_argument(\"--preds_folder\",\n                    default = \"/kaggle/working/predictions/labels\",\n                    help    = \"Folder with prediction .txt files\")\nparser.add_argument(\"--output_csv\",\n                    default = \"/kaggle/working/submission_yolov12x_(16)_augmentation_V1-V6.csv\",\n                    help    = \"Output CSV filename\")\nparser.add_argument(\"--test_images_folder\", \n                    default = \"/kaggle/input/geospatialdetection-final-dataset/testImages/images\",\n                    help    = \"Path to test images directory\")\nargs, _ = parser.parse_known_args()\npredictions_to_csv(preds_folder       = args.preds_folder,\n                   output_csv         = args.output_csv,\n                   test_images_folder = args.test_images_folder\n                  )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T18:52:09.839607Z","iopub.execute_input":"2026-01-12T18:52:09.839898Z","iopub.status.idle":"2026-01-12T18:52:11.122053Z","shell.execute_reply.started":"2026-01-12T18:52:09.839869Z","shell.execute_reply":"2026-01-12T18:52:11.121451Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🤖 Detectron2 prediction (YOLO -> COCO format)","metadata":{}},{"cell_type":"markdown","source":"## 1️⃣ Imports + paths\n\n* import the necessary libraries for working with files, images and JSON (`pathlib`, `os`, `json`, `cv2`, `numpy`);\n* connect Detectron2 (datasets, configs, trainer, inference);\n* set the paths to the `train/val/test` folders in structure (`images` and `labels');\n* we set a list of classes (`building`, `vehicle`) and a working directory where COCO annotations and learning outcomes will be saved.\n\nThe goal: to bring all the surroundings and paths to a single view","metadata":{}},{"cell_type":"code","source":"!pip -q install -U \"git+https://github.com/facebookresearch/detectron2.git\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:25:17.986311Z","iopub.execute_input":"2026-01-13T08:25:17.987232Z","iopub.status.idle":"2026-01-13T08:27:21.307052Z","shell.execute_reply.started":"2026-01-13T08:25:17.987185Z","shell.execute_reply":"2026-01-13T08:27:21.306252Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport json\nimport torch\nimport cv2\nimport copy\nimport csv\n\n# Detectron2\nfrom detectron2.data import detection_utils as utils\nfrom detectron2.data.datasets import register_coco_instances\nfrom detectron2.engine import DefaultTrainer, DefaultPredictor\nfrom detectron2.config import get_cfg\nimport detectron2.data.transforms as T\nfrom detectron2 import model_zoo\nfrom detectron2.evaluation import COCOEvaluator, inference_on_dataset\nfrom detectron2.data import build_detection_train_loader, build_detection_test_loader\nfrom detectron2.utils.logger import setup_logger","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:27:21.308768Z","iopub.execute_input":"2026-01-13T08:27:21.309035Z","iopub.status.idle":"2026-01-13T08:27:22.163572Z","shell.execute_reply.started":"2026-01-13T08:27:21.309002Z","shell.execute_reply":"2026-01-13T08:27:22.162927Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"setup_logger()\n#logging.getLogger(\"detectron2\").setLevel(logging.WARNING)\n\n# Paths\nROOT = Path(\"/kaggle/input/geospatialdetection-final-dataset\")\n\nTRAIN_DIR = ROOT / \"train_ALL\"\nVAL_DIR   = ROOT / \"val_ALL\"\nTEST_DIR  = ROOT / \"testImages\"\n\n# If there is an images subfolder in testImages:\nTEST_IMAGES_DIR = (TEST_DIR / \"images\") if (TEST_DIR / \"images\").exists() else TEST_DIR\n\n# Class names\nCLASS_NAMES = [\"building\", \"vehicle\"]\nNC = len(CLASS_NAMES)\n\nWORK_DIR = Path(\"/kaggle/working/detectron2_frcnn\")\nWORK_DIR.mkdir(parents = True, exist_ok = True)\n\nprint(\"Train:\", TRAIN_DIR)\nprint(\"Val  :\", VAL_DIR)\nprint(\"Test :\", TEST_IMAGES_DIR)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:29:28.732958Z","iopub.execute_input":"2026-01-13T08:29:28.733992Z","iopub.status.idle":"2026-01-13T08:29:28.752506Z","shell.execute_reply.started":"2026-01-13T08:29:28.733956Z","shell.execute_reply":"2026-01-13T08:29:28.751929Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2️⃣ YOLO → COCO JSON Converter\n\nDetectron2 is most convenient to train on datasets in the **COCO** format.\nThe data is now in YOLO format: for each image there is a `.txt`, where each row is an object:\n\n`class_id, x_center, y_center, width, height` (all coordinates are normalized in the range 0..1)\n\nThis converter does the following:\n\n* reads all images from `images/` and their dimensions (width, height);\n* for each image, it searches for the corresponding `.txt` in `labels/`;\n* converts YOLO coordinates (center/width/height) to COCO bbox (x_min, y_min, width, height) in pixels;\n* generates the final `*.json` in the COCO structure: `images`, `annotations`, `categories'.\n\nImportant: if the `.txt` is missing or empty, this is considered the correct case of “no objects in the image\".","metadata":{}},{"cell_type":"code","source":"def yolo_to_coco_json(images_dir: Path, labels_dir: Path,\n                      out_json: Path,\n                      class_names,\n                      img_exts=(\".jpg\", \".jpeg\", \".png\")):\n    \n    images_dir = Path(images_dir)\n    labels_dir = Path(labels_dir)\n    out_json = Path(out_json)\n    out_json.parent.mkdir(parents=True, exist_ok=True)\n\n    image_files = sorted([p for p in images_dir.iterdir() if p.suffix.lower() in img_exts])\n    if len(image_files) == 0:\n        raise FileNotFoundError(f\"No images found in {images_dir}\")\n\n    coco = {\"images\": [],\n            \"annotations\": [],\n            \"categories\": [{\"id\": i+1, \"name\": n} for i, n in enumerate(class_names)]\n           }\n\n    ann_id = 1\n    for img_id, img_path in enumerate(image_files, start=1):\n        img = cv2.imread(str(img_path))\n        if img is None:\n            print(f\"WARNING: can't read {img_path}, skipping\")\n            continue\n\n        h, w = img.shape[:2]\n\n        coco[\"images\"].append({\"id\": img_id,\n                               \"file_name\": img_path.name,\n                               \"width\": w,\n                               \"height\": h\n                              })\n\n        # label file name matches image stem\n        label_path = labels_dir / f\"{img_path.stem}.txt\"\n\n        # empty txt or no txt => just \"no objects\"\n        if not label_path.exists():\n            continue\n\n        lines = [ln.strip() for ln in label_path.read_text().splitlines() if ln.strip()]\n        for ln in lines:\n            parts = ln.split()\n            # waiting for 5 columns: cls xc yc bw bh\n            if len(parts) < 5:\n                continue\n\n            cls = int(float(parts[0]))\n            xc  = float(parts[1])\n            yc  = float(parts[2])\n            bw  = float(parts[3])\n            bh  = float(parts[4])\n\n            # YOLO normalized -> absolute\n            box_w = bw * w\n            box_h = bh * h\n            x_min = (xc * w) - box_w / 2\n            y_min = (yc * h) - box_h / 2\n\n            # clip on borders\n            x_min = max(0.0, min(x_min, w - 1.0))\n            y_min = max(0.0, min(y_min, h - 1.0))\n            box_w = max(1.0, min(box_w, w - x_min))\n            box_h = max(1.0, min(box_h, h - y_min))\n\n            coco[\"annotations\"].append({\"id\": ann_id,\n                                        \"image_id\": img_id,\n                                        \"category_id\": cls + 1, # COCO category\n                                        \"bbox\": [x_min, y_min, box_w, box_h],\n                                        \"area\": float(box_w * box_h),\n                                        \"iscrowd\": 0\n                                       })\n            ann_id += 1\n\n    out_json.write_text(json.dumps(coco))\n    print(f\"Saved COCO JSON: {out_json} | images = {len(coco['images'])} anns = {len(coco['annotations'])}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:29:51.489460Z","iopub.execute_input":"2026-01-13T08:29:51.490160Z","iopub.status.idle":"2026-01-13T08:29:51.500826Z","shell.execute_reply.started":"2026-01-13T08:29:51.490128Z","shell.execute_reply":"2026-01-13T08:29:51.500059Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3️⃣ Сreate COCO annotations for train/val and register them in Detectron2\n\nHere we are:\n\n* we call the converter twice: separately for `train` and for `val`;\n* we get two files: `train_coco.json` and 'val_coco.json`;\n* we register them in Detectron2 via `register_coco_instances()`under the names, for example, `geo_train` and `geo_val'.\n\nRegistration means that Detectron2 now “knows” about these datasets and can:\n\n* upload images;\n* Read annotations;\n* Use them in training and validation.","metadata":{}},{"cell_type":"code","source":"TRAIN_COCO = WORK_DIR / \"train_coco.json\"\nVAL_COCO   = WORK_DIR / \"val_coco.json\"\n\nyolo_to_coco_json(TRAIN_DIR / \"images\", TRAIN_DIR / \"labels\", TRAIN_COCO, CLASS_NAMES)\nyolo_to_coco_json(VAL_DIR   / \"images\", VAL_DIR   / \"labels\", VAL_COCO,   CLASS_NAMES)\n\n# Dataset register\nregister_coco_instances(\"geo_train\", {}, str(TRAIN_COCO), str(TRAIN_DIR / \"images\"))\nregister_coco_instances(\"geo_val\",   {}, str(VAL_COCO),   str(VAL_DIR   / \"images\"))\n\nprint(\"Datasets registered: geo_train, geo_val\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:29:56.384011Z","iopub.execute_input":"2026-01-13T08:29:56.384584Z","iopub.status.idle":"2026-01-13T08:38:56.456350Z","shell.execute_reply.started":"2026-01-13T08:29:56.384550Z","shell.execute_reply":"2026-01-13T08:38:56.455627Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4️⃣ Faster R-CNN Configuration + training\n\nIn this block, we:\n\n* we take the ready-made basic Faster R-CNN config from the Detectron2 Model Zoo (`faster_rcnn_R_101_FPN_3x`);\n* changing key parameters:\n\n  * number of classes `ROI_HEADS.NUM_CLASSES = 2`;\n  * paths to train/val datasets;\n  * batch size, LR, number of iterations;\n  * Image size parameters (multi-scale train);\n* we specify pretrained weights (COCO) as a starting point to accelerate and stabilize learning.\n\nNext, we run the `DefaultTrainer`, which:\n\n* collects dataloader;\n* trains the model;\n* Saves checkpoints and logs in `OUTPUT_DIR'.","metadata":{}},{"cell_type":"code","source":"def patch_coco_json(path: str): # else error after fitting\n    p = Path(path)\n    coco = json.loads(p.read_text())\n\n    coco.setdefault(\"info\", {\"description\": \"patched\", \"version\": \"1.0\"})\n    coco.setdefault(\"licenses\", [{\"id\": 1, \"name\": \"Unknown\", \"url\": \"\"}])\n\n    # если у image нет license — добавляем\n    for img in coco.get(\"images\", []):\n        img.setdefault(\"license\", 1)\n\n    p.write_text(json.dumps(coco))\n    print(\"Patched:\", p)\n\npatch_coco_json(\"/kaggle/working/detectron2_frcnn/train_coco.json\")\npatch_coco_json(\"/kaggle/working/detectron2_frcnn/val_coco.json\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:38:56.457794Z","iopub.execute_input":"2026-01-13T08:38:56.458034Z","iopub.status.idle":"2026-01-13T08:38:56.558001Z","shell.execute_reply.started":"2026-01-13T08:38:56.458009Z","shell.execute_reply":"2026-01-13T08:38:56.557403Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cfg = get_cfg()\ncfg.merge_from_file(model_zoo.get_config_file(\"COCO-Detection/faster_rcnn_R_101_FPN_3x.yaml\"))\ncfg.DATASETS.TRAIN = (\"geo_train\",)\ncfg.DATASETS.TEST  = (\"geo_val\",)\ncfg.DATALOADER.NUM_WORKERS = 4\n\n# Weights from model zoo (pretrained)\ncfg.MODEL.WEIGHTS = model_zoo.get_checkpoint_url(\"COCO-Detection/faster_rcnn_R_101_FPN_3x.yaml\")\n\n# Number of classes (2)\ncfg.MODEL.ROI_HEADS.NUM_CLASSES = NC\n\n# Threshold for the test/inference\ncfg.MODEL.ROI_HEADS.SCORE_THRESH_TEST = 0.01\n\n# Solver\ncfg.SOLVER.IMS_PER_BATCH = 8\ncfg.SOLVER.BASE_LR = 0.01\ncfg.SOLVER.MAX_ITER = 19000    # Epochs = (IMS_PER_BATCH * MAX_ITER)/(N_train_images) = 15 epochs\ncfg.SOLVER.STEPS = (15000, 17600)\ncfg.SOLVER.GAMMA = 0.1\ncfg.SOLVER.WARMUP_ITERS = 1000\n\n# Size of FPN\ncfg.INPUT.MIN_SIZE_TRAIN = (640, 704, 768, 800)\ncfg.INPUT.MAX_SIZE_TRAIN = 1333\ncfg.INPUT.MIN_SIZE_TEST  = 800\ncfg.INPUT.MAX_SIZE_TEST  = 1333\n\ncfg.OUTPUT_DIR = str(WORK_DIR / \"output\")\nos.makedirs(cfg.OUTPUT_DIR, exist_ok = True)\n\nclass AugmentedMapper:\n    def __init__(self, cfg, is_train=True):\n        self.is_train = is_train\n        self.img_format = cfg.INPUT.FORMAT\n\n        if is_train:\n            self.augmentations = [\n                # Multi-scale resize\n                T.ResizeShortestEdge(short_edge_length = cfg.INPUT.MIN_SIZE_TRAIN,\n                                     max_size          = cfg.INPUT.MAX_SIZE_TRAIN,\n                                     sample_style      = \"choice\",\n                                    ),\n\n                # Flips\n                T.RandomFlip(prob = 0.5, horizontal = True, vertical = False),\n                T.RandomFlip(prob = 0.2, horizontal = False, vertical = True),\n\n                # Color jitter\n                T.RandomBrightness(0.85, 1.15),\n                T.RandomContrast(0.85, 1.15),\n                T.RandomSaturation(0.85, 1.15),\n\n                # Random rotation\n                T.RandomRotation(angle = [-10, 10], expand=False, center=None, sample_style=\"range\"),\n\n                # Crop — it can help for small objects,\n                # but it can worsen if it cuts the boxes too much.:\n            ]\n        else:\n            # For val/test — only resize\n            self.augmentations = [T.ResizeShortestEdge(short_edge_length = cfg.INPUT.MIN_SIZE_TEST,\n                                                       max_size          = cfg.INPUT.MAX_SIZE_TEST,\n                                                       sample_style      = \"choice\",\n                                                      )\n                                 ]\n\n    def __call__(self, dataset_dict):\n        dataset_dict = copy.deepcopy(dataset_dict)\n\n        image = utils.read_image(dataset_dict[\"file_name\"], format = self.img_format)\n        aug_input = T.AugInput(image)\n        transforms = T.AugmentationList(self.augmentations)(aug_input)\n        image = aug_input.image\n\n        dataset_dict[\"image\"] = torch.as_tensor(image.transpose(2, 0, 1).astype(\"float32\"))\n\n        if \"annotations\" in dataset_dict:\n            annos = [utils.transform_instance_annotations(obj, transforms, image.shape[:2])\n                     for obj in dataset_dict.pop(\"annotations\")\n                     if obj.get(\"iscrowd\", 0) == 0\n                    ]\n            instances = utils.annotations_to_instances(annos, image.shape[:2])\n            dataset_dict[\"instances\"] = utils.filter_empty_instances(instances)\n\n        return dataset_dict\n\nclass Trainer(DefaultTrainer):\n    @classmethod\n    def build_train_loader(cls, cfg):\n        return build_detection_train_loader(cfg, mapper = AugmentedMapper(cfg, is_train = True))\n\n    @classmethod\n    def build_evaluator(cls, cfg, dataset_name, output_folder=None):\n        return COCOEvaluator(dataset_name, cfg, False, output_folder or os.path.join(cfg.OUTPUT_DIR, \"inference\"))\n\n#trainer = Trainer(cfg)\n#trainer.resume_or_load(resume = False) # fine-tune from COCO-pretrained\n#trainer.train()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:38:56.558791Z","iopub.execute_input":"2026-01-13T08:38:56.559196Z","iopub.status.idle":"2026-01-13T08:38:56.580475Z","shell.execute_reply.started":"2026-01-13T08:38:56.559170Z","shell.execute_reply":"2026-01-13T08:38:56.579768Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5️⃣ Evaluation based on the validation sample\n\nAfter training, it is important to evaluate the quality on the `val` in order to:\n\n* check that the model is really learning and not degrading;\n* see standard COCO metrics: `AP`, `AP50`, `AP75`, `AP_small/medium/large'.\n\nDetectron2 does this via `COCOEvaluator` and 'inference_on_dataset()`:\n\n* runs the model through validation images;\n* compares predicted boxes with true annotations;\n* outputs metrics.","metadata":{}},{"cell_type":"code","source":"evaluator = COCOEvaluator(\"geo_val\", cfg, False, output_dir = os.path.join(cfg.OUTPUT_DIR, \"inference\"))\nval_loader = build_detection_test_loader(cfg, \"geo_val\")\nmetrics = inference_on_dataset(trainer.model, val_loader, evaluator)\nmetrics","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:38:56.581874Z","iopub.execute_input":"2026-01-13T08:38:56.582146Z","iopub.status.idle":"2026-01-13T08:38:56.585182Z","shell.execute_reply.started":"2026-01-13T08:38:56.582112Z","shell.execute_reply":"2026-01-13T08:38:56.584589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# =========================\n# Path to Detectron2 metrics.json\n# =========================\nmetrics_path = \"/kaggle/input/geospatialdetection-dualityai-weights/metrics (2).json\"\nprint(\"Reading:\", metrics_path)\n\nrows = []\nwith open(metrics_path, \"r\") as f:\n    for line in f:\n        line = line.strip()\n        if not line:\n            continue\n        try:\n            rows.append(json.loads(line))\n        except Exception:\n            pass\n\ndf = pd.DataFrame(rows)\nif \"iteration\" not in df.columns:\n    raise ValueError(\"No 'iteration' column in metrics.json. Something is wrong with logging.\")\n\ndf = df.sort_values(\"iteration\").reset_index(drop=True)\nprint(\"Columns:\", sorted(df.columns.tolist()))\nprint(\"Rows:\", len(df))\n\ndef plot_metric(df, col, title=None, rolling=None):\n    if col not in df.columns:\n        print(f\"Skipping '{col}' (not found in metrics.json)\")\n        return\n\n    d = df[[\"iteration\", col]].dropna()\n    if d.empty:\n        print(f\"Skipping '{col}' (all NaN)\")\n        return\n\n    x = d[\"iteration\"].values\n    y = d[col].values\n\n    plt.figure(figsize = (10, 4))\n    plt.plot(x, y, marker=\".\", linewidth=1)\n\n    if rolling is not None and len(d) >= rolling:\n        y_s = d[col].rolling(rolling, min_periods=max(1, rolling // 3)).mean()\n        plt.plot(d[\"iteration\"].values, y_s.values, linewidth=2)\n\n    plt.xlabel(\"iteration\")\n    plt.ylabel(col)\n    plt.title(title or col)\n    plt.grid(True)\n    plt.show()\n\n# =========================\n# Losses\n# =========================\nplot_metric(df, \"loss_cls\",      title=\"loss_cls vs iteration\",      rolling=50)\nplot_metric(df, \"loss_box_reg\",  title=\"loss_box_reg vs iteration\",  rolling=50)\nplot_metric(df, \"loss_rpn_cls\",  title=\"loss_rpn_cls vs iteration\",  rolling=50)\nplot_metric(df, \"loss_rpn_loc\",  title=\"loss_rpn_loc vs iteration\",  rolling=50)\n\n# =========================\n# Learning rate\n# =========================\nplot_metric(df, \"lr\", title = \"learning rate (lr) vs iteration\", rolling = 50)\n\n# =========================\nplot_metric(df, \"bbox/AP50\", title = \"mAP@50 (bbox/AP50) vs iteration\",  rolling = None)\nplot_metric(df, \"bbox/AP\",   title = \"mAP@50-95 (bbox/AP) vs iteration\", rolling = None)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:39:07.878530Z","iopub.execute_input":"2026-01-13T08:39:07.879188Z","iopub.status.idle":"2026-01-13T08:39:08.631960Z","shell.execute_reply.started":"2026-01-13T08:39:07.879156Z","shell.execute_reply":"2026-01-13T08:39:08.631335Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6️⃣ Inference on the test + saving in YOLO format\n\n* loading the trained model (`model_final.pth`);\n* run it through all the test images;\n* for each image, we save the predictions in `.txt` in a format compatible with your old pipeline:\n\n`class_id, confidence, x_center, y_center, width, height`\n\nwhere the coordinates are **normalized** (0..1), as in YOLO.\nThis is convenient, because then you can reuse your ready-made converter `.txt → submission.csv`.","metadata":{}},{"cell_type":"code","source":"predict_cfg = cfg.clone()\nWORK_DIR = Path(\"/kaggle/working/detectron2_frcnn\")\npredict_cfg.MODEL.WEIGHTS = \"/kaggle/input/geospatialdetection-dualityai-weights/model_final_R101.pth\"\n#os.path.join(WORK_DIR, \"output\", \"model_final.pth\")\npredict_cfg.MODEL.ROI_HEADS.SCORE_THRESH_TEST = 0.25\npredictor = DefaultPredictor(predict_cfg)\n\nout_dir = WORK_DIR / \"test_predictions_yolo\"\nlabels_out = out_dir / \"labels\"\nlabels_out.mkdir(parents=True, exist_ok=True)\n\nimg_exts = {\".jpg\", \".jpeg\", \".png\"}\n\nfor img_path in sorted(TEST_IMAGES_DIR.iterdir()):\n    if img_path.suffix.lower() not in img_exts:\n        continue\n\n    img = cv2.imread(str(img_path))\n    h, w = img.shape[:2]\n\n    outputs = predictor(img)\n    instances = outputs[\"instances\"].to(\"cpu\")\n\n    boxes = instances.pred_boxes.tensor.numpy() # xyxy\n    scores = instances.scores.numpy()\n    classes = instances.pred_classes.numpy()\n\n    txt_path = labels_out / f\"{img_path.stem}.txt\"\n    with open(txt_path, \"w\") as f:\n        for (x1, y1, x2, y2), sc, cl in zip(boxes, scores, classes):\n            bw = max(0.0, x2 - x1)\n            bh = max(0.0, y2 - y1)\n            xc = x1 + bw / 2.0\n            yc = y1 + bh / 2.0\n\n            f.write(\n                f\"{int(cl)} {float(sc):.6f} \"\n                f\"{xc / w:.6f} {yc / h:.6f} {bw / w:.6f} {bh / h:.6f}\\n\"\n            )\n\nprint(\"Saved YOLO-style predictions to:\", labels_out)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:41:01.464064Z","iopub.execute_input":"2026-01-13T08:41:01.464919Z","iopub.status.idle":"2026-01-13T08:41:01.469211Z","shell.execute_reply.started":"2026-01-13T08:41:01.464886Z","shell.execute_reply":"2026-01-13T08:41:01.468571Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7️⃣ Converter .txt → submission .csv (from yolo with some changes)\n\nIn this block, we take the prediction folder (`labels/*.txt`) and create the final Kaggle submission.:\n\n* A CSV string is created for each image from the test folder:\n\n  * `image_id`\n  * `prediction_string`\n* if there are boxes in `.txt`, the strings are combined into one long `prediction_string`\n* if there are no boxes or the file is empty, put `no boxes`\n\nAdditionally, the code performs validation:\n\n* skips the extra `.txt`, which are not among the test `ids`;\n* checks that each line in `.txt` has 6 numbers;\n* counts the number of format errors.\n\nThe result is a ready—made `submission.csv`, which can be sent for evaluation.","metadata":{}},{"cell_type":"code","source":"def predictions_to_csv(preds_folder: str,\n                       output_csv: str,\n                       test_images_folder: str,\n                       allowed_extensions: tuple = (\".jpg\", \".png\", \".jpeg\")):\n\n    preds_path = Path(preds_folder)\n    if not preds_path.exists():\n        print(f\"ERROR: Prediction folder '{preds_folder}' does not exist\")\n        sys.exit(1)\n\n    test_images_path = Path(test_images_folder)\n    if not test_images_path.exists():\n        print(f\"ERROR: Test images folder '{test_images_folder}' not found\")\n        sys.exit(1)\n\n    test_images = {p.stem: True for p in test_images_path.glob(\"*\") if p.suffix.lower() in allowed_extensions}\n    print(f\"Found {len(test_images)} test images\")\n\n    predictions = []\n    error_count = 0\n\n    for txt_file in preds_path.glob(\"*.txt\"):\n        image_id = txt_file.stem\n\n        if image_id not in test_images:\n            continue\n\n        with open(txt_file, \"r\") as f:\n            valid_lines = []\n            for line_num, line in enumerate(f, 1):\n                line = line.strip()\n                if not line:\n                    continue\n\n                parts = line.split()\n                if len(parts) != 6:\n                    error_count += 1\n                    continue\n\n                try:\n                    [float(x) for x in parts]\n                    valid_lines.append(line)\n                except ValueError:\n                    error_count += 1\n                    continue\n\n        pred_str = \" \".join(valid_lines) if valid_lines else \"no boxes\"\n        predictions.append({\"image_id\": image_id, \"prediction_string\": pred_str})\n\n    submission_df = pd.DataFrame({\"image_id\": list(test_images.keys())})\n\n    if predictions:\n        preds_df = pd.DataFrame(predictions)\n        final_df = submission_df.merge(preds_df, on=\"image_id\", how=\"left\").fillna(\"no boxes\")\n    else:\n        final_df = submission_df\n        final_df[\"prediction_string\"] = \"no boxes\"\n\n    final_df.to_csv(output_csv, index=False, quoting = csv.QUOTE_NONNUMERIC)\n\n    print(f\"\\n✅ Submission saved to {output_csv}\")\n    print(f\"   Total predictions (txt files read): {len(predictions)}\")\n    print(f\"   Validation errors: {error_count}\")\n\n# defaults\npredictions_to_csv(preds_folder       = \"/kaggle/working/detectron2_frcnn/test_predictions_yolo/labels\",\n                   output_csv         = \"/kaggle/working/submission_detectron_frcnn_R101_0.25.csv\",\n                   test_images_folder = \"/kaggle/input/geospatialdetection-final-dataset/testImages/images\"\n                  )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-13T08:41:22.714429Z","iopub.execute_input":"2026-01-13T08:41:22.715220Z","iopub.status.idle":"2026-01-13T08:41:22.720237Z","shell.execute_reply.started":"2026-01-13T08:41:22.715187Z","shell.execute_reply":"2026-01-13T08:41:22.719424Z"}},"outputs":[],"execution_count":null}]}