{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.12.12"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":117682,"databundleVersionId":15062069},{"sourceType":"datasetVersion","sourceId":14892898,"datasetId":9528952,"databundleVersionId":15756936},{"sourceType":"datasetVersion","sourceId":14882335,"datasetId":9327392,"databundleVersionId":15745501},{"sourceType":"datasetVersion","sourceId":14808160,"datasetId":9177003,"databundleVersionId":15663894}],"dockerImageVersionId":31260,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":533.974301,"end_time":"2026-02-24T01:07:03.070729","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2026-02-24T00:58:09.096428","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# WNet3D + nnUNet Ensemble — Vesuvius Challenge Surface Detection\n\n**125th Place Solution · 0.582 Private LB**\n\nThis notebook runs our full inference pipeline:\n\n1. **WNet3D preprocessing** — self-supervised feature extraction ([CellSeg3D, Achard et al.](https://elifesciences.org/reviewed-preprints/99848v1))\n2. **nnUNet ensemble** — two models with different training data and preprocessors\n3. **Postprocessing** — hysteresis thresholding + morphological closing + dust removal\n\n### Models\n\n| Model | Training Data | WNet Preprocessor | GT Version | Loss |\n|-------|--------------|-------------------|------------|------|\n| **A** (Dataset324) | All | WNet3D (20 epochs) | Original | DC+CE |\n| **B** (Dataset325) | All | WNet3D (100 epochs) | Cleaned | BoundaryDoULoss |\n\n### Key Idea\n\nRather than feeding raw CT to nnUNet, we first run WNet3D (a self-supervised dual U-Net from microscopy) to extract learned structural features. The 4-channel input (raw CT + 3 WNet feature channels) gives nnUNet richer information about sheet boundaries and tissue separation.\n\nSee our [solution writeup](link) for details.\n\nThis notebook uses a P100 GPU, but this could have been optimized by using 2 x T4 GPUs and having each WNET+nnUNET combo in each GPU. This would have made it faster","metadata":{}},{"cell_type":"markdown","source":"## 1. Install Dependencies\n\nOffline install from pre-packaged wheels (no internet required on Kaggle).","metadata":{}},{"cell_type":"code","source":"!uv pip install --system --no-index \\\n    --find-links=/kaggle/input/datasets/pr4deepr/cellseg3d-nnunetv2-wheels/wheels \\\n    torch==2.8.0 \\\n    torchvision==0.23.0 \\\n    nnunetv2 \\\n    napari_cellseg3d\n\n!uv cache clean\n!rm -rf /root/.cache/uv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:47:26.333866Z","iopub.execute_input":"2026-03-02T09:47:26.334393Z","iopub.status.idle":"2026-03-02T09:47:39.199895Z","shell.execute_reply.started":"2026-03-02T09:47:26.334366Z","shell.execute_reply":"2026-03-02T09:47:39.199137Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Configuration\n\nAll paths, model weights, ensemble weights, and postprocessing parameters.","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nimport os\n\n# ── Paths ──\nINPUT_DIR = Path(\"/kaggle/input/vesuvius-challenge-surface-detection\")\nWORKING_DIR = Path(\"/kaggle/temp\")\nWORKING_DIR.mkdir(parents=True, exist_ok=True)\nTEST_IMAGES_DIR = INPUT_DIR / \"test_images\"\nSUBMISSION_DIR = Path(\"/kaggle/working\")\n\n# ── WNet3D model checkpoints ──\nWNET_MODEL_PATH_A = Path(\"/kaggle/input/datasets/pr4deepr/wnet-latest-model/wnet_model/WNet3D_ep20/wnet.pth\")\nWNET_MODEL_PATH_B = Path(\"/kaggle/input/datasets/pr4deepr/wnet-latest-model/wnet_model/WNet3D_ep100/wnet.pth\")\n\n# ── nnUNet model folders ──\nNNUNET_MODEL_A = \"/kaggle/input/datasets/pr4deepr/nnunet/Dataset324_wnet/Dataset324_wnet/nnUNetTrainer__nnUNetResEncUNetXLPlans__3d_fullres_bs18\"\nNNUNET_MODEL_B = \"/kaggle/input/datasets/pr4deepr/nnunet/Dataset325_cleaned/Dataset325_cleaned/nnUNetTrainerBoundaryDoU__nnUNetResEncUNetXLPlans__3d_fullres_bs\"\n\n# ── Ensemble weights ──\nWEIGHT_A = 0.5  # Dataset324\nWEIGHT_B = 0.5  # Dataset325\n\n# ── nnUNet inference settings ──\nCHUNK_SIZE = 4           # images per batch (memory management)\nTILE_STEP_SIZE = 0.5     # sliding window overlap\nUSE_MIRRORING = True     # test-time augmentation\nMIRROR_AXES_OVERRIDE = (0, 1)  # 4-pass TTA (reduced from default 8)\n\n# ── Postprocessing ──\nT_LOW = 0.30\nT_HIGH = 0.90\nZ_RADIUS = 1\nXY_RADIUS = 0\nDUST_MIN_SIZE = 100\n\n# ── Derived paths ──\nos.environ[\"nnUNet_results\"] = \"/kaggle/input/nnunet/\"\nCHECKPOINT_NAME = \"checkpoint_final.pth\"\nOUTPUT_DIR = Path(\"/kaggle/temp/predictions\")\nOUTPUT_DIR.mkdir(exist_ok=True)\nWNET_TEMP_A = WORKING_DIR / \"wnet_a\"\nWNET_TEMP_B = WORKING_DIR / \"wnet_b\"\n\n# ── Sanity checks ──\nassert WNET_MODEL_PATH_A.exists(), f\"WNet A not found: {WNET_MODEL_PATH_A}\"\nassert WNET_MODEL_PATH_B.exists(), f\"WNet B not found: {WNET_MODEL_PATH_B}\"\nassert Path(NNUNET_MODEL_A).exists(), f\"nnUNet A not found: {NNUNET_MODEL_A}\"\nassert Path(NNUNET_MODEL_B).exists(), f\"nnUNet B not found: {NNUNET_MODEL_B}\"\nprint(\"All model paths verified.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:47:39.201841Z","iopub.execute_input":"2026-03-02T09:47:39.202444Z","iopub.status.idle":"2026-03-02T09:47:39.280420Z","shell.execute_reply.started":"2026-03-02T09:47:39.202413Z","shell.execute_reply":"2026-03-02T09:47:39.279783Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Imports","metadata":{}},{"cell_type":"code","source":"import gc\nimport sys\nimport shutil\nimport contextlib\nimport subprocess\nimport torch\nimport numpy as np\nimport tifffile\nfrom tqdm import tqdm\nimport zipfile\nfrom copy import deepcopy\n\nfrom napari_cellseg3d.dev_scripts import remote_inference as cs3d\nfrom napari_cellseg3d.config import ModelInfo, WeightsInfo, SlidingWindowConfig\nfrom scipy import ndimage as ndi\nfrom skimage.morphology import remove_small_objects\n\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Device: {DEVICE}\")\nif DEVICE.type == 'cuda':\n    print(f\"GPU: {torch.cuda.get_device_name(0)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:47:39.281395Z","iopub.execute_input":"2026-03-02T09:47:39.281677Z","iopub.status.idle":"2026-03-02T09:48:24.278699Z","shell.execute_reply.started":"2026-03-02T09:47:39.281649Z","shell.execute_reply":"2026-03-02T09:48:24.278024Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Custom Loss for Dataset325\n\nDataset325 was trained with BoundaryDoULoss. We need to register the custom trainer\nand loss so nnUNet can load the checkpoint.","metadata":{}},{"cell_type":"code","source":"import nnunetv2\n\nnnunet_trainer_dir = Path(nnunetv2.__path__[0]) / \"training\" / \"nnUNetTrainer\" / \"variants\" / \"loss\"\nnnunet_loss_dir = Path(nnunetv2.__path__[0]) / \"training\" / \"loss\"\nCUSTOM_FILES = Path(\"/kaggle/input/datasets/pr4deepr/nnunet/boundaryloss_Dataset325/boundaryloss_Dataset325\")\n\nshutil.copy(CUSTOM_FILES / \"nnUNetTrainerBoundaryDoU.py\", nnunet_trainer_dir / \"nnUNetTrainerBoundaryDoU.py\")\nshutil.copy(CUSTOM_FILES / \"boundary_dou_loss_3d.py\", nnunet_loss_dir / \"boundary_dou_loss_3d.py\")\nprint(f\"Registered BoundaryDoU trainer and loss.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:48:24.279506Z","iopub.execute_input":"2026-03-02T09:48:24.280303Z","iopub.status.idle":"2026-03-02T09:48:24.330189Z","shell.execute_reply.started":"2026-03-02T09:48:24.280276Z","shell.execute_reply":"2026-03-02T09:48:24.329482Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Utility Functions\n\nHelper to suppress verbose output (prevents notebook size blowup) and clean up\ntemporary files left by CellSeg3D.","metadata":{}},{"cell_type":"code","source":"#Used this initially to suppr\n@contextlib.contextmanager\ndef suppress_stdout():\n    \"\"\"Temporarily redirect stdout/stderr to /dev/null.\"\"\"\n    devnull = open(os.devnull, 'w')\n    old_stdout, old_stderr = sys.stdout, sys.stderr\n    sys.stdout, sys.stderr = devnull, devnull\n    try:\n        yield\n    finally:\n        sys.stdout, sys.stderr = old_stdout, old_stderr\n        devnull.close()\n\n\ndef nuke_cs3d_leftovers():\n    \"\"\"Delete temporary files created by CellSeg3D inference.\"\"\"\n    for search_root in [Path.home(), Path(\"/kaggle/working\"), Path(\"/kaggle/temp\"), Path(\"/tmp\")]:\n        if not search_root.exists():\n            continue\n        for pattern in [\"volume_WNet3D*\", \"*_pred_*\"]:\n            for f in search_root.rglob(pattern):\n                try:\n                    if f.is_file():\n                        f.unlink()\n                    elif f.is_dir():\n                        shutil.rmtree(f)\n                except Exception:\n                    pass","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:48:24.332317Z","iopub.execute_input":"2026-03-02T09:48:24.332543Z","iopub.status.idle":"2026-03-02T09:48:24.352865Z","shell.execute_reply.started":"2026-03-02T09:48:24.332521Z","shell.execute_reply":"2026-03-02T09:48:24.352172Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. WNet3D Setup\n\nTwo WNet3D configurations — one per nnUNet model:\n- **WNet A** (20 epochs, patch=64) → feeds nnUNet Model A (Dataset324)\n- **WNet B** (100 epochs, patch=128) → feeds nnUNet Model B (Dataset325)","metadata":{}},{"cell_type":"code","source":"WNET_PATCH_SIZE_A = 64\nWNET_PATCH_SIZE_B = 128\nWNET_NUM_CLASSES = 3\nWNET_OVERLAP = 0.25\n\nwnet_model_info = ModelInfo(\n    name=\"WNet3D\",\n    num_classes=WNET_NUM_CLASSES,\n)\n\nwnet_config_A = deepcopy(cs3d.CONFIG)\nwnet_config_A.model_info = wnet_model_info\nwnet_config_A.weights_config = WeightsInfo(path=str(WNET_MODEL_PATH_A), use_custom=True)\nwnet_config_A.sliding_window_config = SlidingWindowConfig(window_size=WNET_PATCH_SIZE_A, window_overlap=WNET_OVERLAP)\n\nwnet_config_B = deepcopy(cs3d.CONFIG)\nwnet_config_B.model_info = wnet_model_info\nwnet_config_B.weights_config = WeightsInfo(path=str(WNET_MODEL_PATH_B), use_custom=True)\nwnet_config_B.sliding_window_config = SlidingWindowConfig(window_size=WNET_PATCH_SIZE_B, window_overlap=WNET_OVERLAP)\n\nprint(\"WNet3D configs ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:48:24.353653Z","iopub.execute_input":"2026-03-02T09:48:24.353909Z","iopub.status.idle":"2026-03-02T09:48:24.367346Z","shell.execute_reply.started":"2026-03-02T09:48:24.353876Z","shell.execute_reply":"2026-03-02T09:48:24.366697Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. nnUNet Setup\n\nInitialize both nnUNet predictors with reduced TTA (4 passes instead of 8).","metadata":{}},{"cell_type":"code","source":"from nnunetv2.inference.predict_from_raw_data import nnUNetPredictor\n\n\ndef init_nnunet_predictor(model_folder, checkpoint_name='checkpoint_best.pth'):\n    predictor = nnUNetPredictor(\n        tile_step_size=TILE_STEP_SIZE,\n        use_gaussian=True,\n        use_mirroring=USE_MIRRORING,\n        perform_everything_on_device=True,\n        device=torch.device('cuda'), #With multiple GPUs, specify GPU device\n        verbose=False,\n        verbose_preprocessing=False,\n        allow_tqdm=False\n    )\n    predictor.initialize_from_trained_model_folder(\n        model_folder,\n        use_folds=(0,),\n        checkpoint_name=checkpoint_name\n    )\n    if MIRROR_AXES_OVERRIDE is not None and USE_MIRRORING:\n        original = predictor.allowed_mirroring_axes\n        predictor.allowed_mirroring_axes = MIRROR_AXES_OVERRIDE\n        print(f\"  TTA: {original} → {MIRROR_AXES_OVERRIDE} \"\n              f\"({2**len(MIRROR_AXES_OVERRIDE)} passes instead of {2**len(original)})\")\n    return predictor\n\n\nprint(\"Initializing nnUNet predictors...\")\npredictor_a = init_nnunet_predictor(NNUNET_MODEL_A, 'checkpoint_latest.pth')\npredictor_b = init_nnunet_predictor(NNUNET_MODEL_B, CHECKPOINT_NAME)\nprint(\"Done.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:53:57.140999Z","iopub.execute_input":"2026-03-02T09:53:57.141537Z","iopub.status.idle":"2026-03-02T09:54:18.741968Z","shell.execute_reply.started":"2026-03-02T09:53:57.141507Z","shell.execute_reply":"2026-03-02T09:54:18.741067Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Postprocessing\n\nHysteresis thresholding with anisotropic morphological closing and small component removal.\nThis simple recipe consistently outperformed Optuna-tuned and topology-aware alternatives on the leaderboard.","metadata":{}},{"cell_type":"code","source":"def build_anisotropic_struct(z_radius: int, xy_radius: int):\n    \"\"\"Build a 3D structuring element with independent Z and XY radii.\"\"\"\n    z, r = z_radius, xy_radius\n    if z == 0 and r == 0:\n        return None\n    if z == 0 and r > 0:\n        size = 2 * r + 1\n        struct = np.zeros((1, size, size), dtype=bool)\n        for dy in range(-r, r + 1):\n            for dx in range(-r, r + 1):\n                if dy * dy + dx * dx <= r * r:\n                    struct[0, r + dy, r + dx] = True\n        return struct\n    if z > 0 and r == 0:\n        struct = np.zeros((2 * z + 1, 1, 1), dtype=bool)\n        struct[:, 0, 0] = True\n        return struct\n    depth, size = 2 * z + 1, 2 * r + 1\n    struct = np.zeros((depth, size, size), dtype=bool)\n    for dz in range(-z, z + 1):\n        for dy in range(-r, r + 1):\n            for dx in range(-r, r + 1):\n                if dy * dy + dx * dx <= r * r:\n                    struct[z + dz, r + dy, r + dx] = True\n    return struct\n\n\ndef topo_postprocess(probs, T_low=T_LOW, T_high=T_HIGH, z_radius=Z_RADIUS,\n                     xy_radius=XY_RADIUS, dust_min_size=DUST_MIN_SIZE):\n    \"\"\"Hysteresis threshold → morphological closing → dust removal.\"\"\"\n    strong = probs >= T_high\n    weak = probs >= T_low\n    if not strong.any():\n        return np.zeros_like(probs, dtype=np.uint8)\n\n    struct_hyst = ndi.generate_binary_structure(3, 3)\n    mask = ndi.binary_propagation(strong, mask=weak, structure=struct_hyst)\n    if not mask.any():\n        return np.zeros_like(probs, dtype=np.uint8)\n\n    if z_radius > 0 or xy_radius > 0:\n        struct_close = build_anisotropic_struct(z_radius, xy_radius)\n        if struct_close is not None:\n            mask = ndi.binary_closing(mask, structure=struct_close)\n\n    if dust_min_size > 0:\n        mask = remove_small_objects(mask.astype(bool), min_size=dust_min_size)\n\n    return mask.astype(np.uint8)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:54:31.163333Z","iopub.execute_input":"2026-03-02T09:54:31.164196Z","iopub.status.idle":"2026-03-02T09:54:31.173621Z","shell.execute_reply.started":"2026-03-02T09:54:31.164160Z","shell.execute_reply":"2026-03-02T09:54:31.172869Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9. Inference Pipeline\n\nFor each chunk of images:\n\n1. **WNet3D** — run both checkpoints, concatenate features with raw CT → 4-channel arrays\n2. **nnUNet A** — predict on WNet-A enriched input\n3. **nnUNet B** — predict on WNet-B enriched input\n4. **Ensemble** — weighted average of probability maps → postprocess → save\n\nProcessing in chunks keeps memory usage manageable on Kaggle's P100 GPUs.","metadata":{}},{"cell_type":"code","source":"test_images = sorted(list(TEST_IMAGES_DIR.glob(\"*.tif\")))\nprint(f\"Found {len(test_images)} test images\")\n\nchunks = [test_images[i:i + CHUNK_SIZE] for i in range(0, len(test_images), CHUNK_SIZE)]\nprint(f\"Processing in {len(chunks)} chunk(s) of up to {CHUNK_SIZE} images\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:54:31.174948Z","iopub.execute_input":"2026-03-02T09:54:31.175204Z","iopub.status.idle":"2026-03-02T09:54:31.196684Z","shell.execute_reply.started":"2026-03-02T09:54:31.175170Z","shell.execute_reply":"2026-03-02T09:54:31.195965Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for chunk_idx, chunk in enumerate(chunks):\n    print(f\"\\n{'='*60}\")\n    print(f\"CHUNK {chunk_idx+1}/{len(chunks)} — {len(chunk)} images\")\n    print(f\"{'='*60}\")\n\n    wnet_buffers_a = []\n    wnet_buffers_b = []\n    chunk_stems = []\n\n    # ── Step 1: WNet3D feature extraction ──\n    for img_idx, img_path in enumerate(chunk):\n        print(f\"  [{img_idx+1}/{len(chunk)}] WNet preprocessing: {img_path.stem}\")\n\n        img = tifffile.imread(img_path)\n        if img.ndim == 4 and img.shape[0] == 1:\n            img_3d = img[0]\n        elif img.ndim == 3:\n            img_3d = img\n        else:\n            img_3d = img\n\n        raw = img_3d[np.newaxis, ...] if img_3d.ndim == 3 else img_3d\n\n        # WNet A (20 epochs, patch=64)\n        with suppress_stdout():\n            result_a = cs3d.inference_on_images(img_3d, config=wnet_config_A)\n        pred_a = result_a[0].semantic_segmentation\n        pred_c_a = pred_a[np.newaxis, ...] if pred_a.ndim == 3 else pred_a\n        wnet_buffers_a.append(np.concatenate([raw, pred_c_a], axis=0).astype(np.float32))\n        del result_a, pred_a, pred_c_a\n        nuke_cs3d_leftovers()\n        torch.cuda.empty_cache()\n\n        # WNet B (100 epochs, patch=128)\n        with suppress_stdout():\n            result_b = cs3d.inference_on_images(img_3d, config=wnet_config_B)\n        pred_b = result_b[0].semantic_segmentation\n        pred_c_b = pred_b[np.newaxis, ...] if pred_b.ndim == 3 else pred_b\n        wnet_buffers_b.append(np.concatenate([raw, pred_c_b], axis=0).astype(np.float32))\n        del result_b, pred_b, pred_c_b, img, img_3d, raw\n        nuke_cs3d_leftovers()\n        torch.cuda.empty_cache()\n\n        chunk_stems.append(img_path.stem)\n\n    # ── Step 2: nnUNet Model A (Dataset324) ──\n    print(f\"  Running nnUNet Model A (Dataset324)...\")\n    props_list = [{'spacing': (1.0, 1.0, 1.0)} for _ in wnet_buffers_a]\n    dummy = [None] * len(wnet_buffers_a)\n\n    with suppress_stdout():\n        results_a = predictor_a.predict_from_list_of_npy_arrays(\n            wnet_buffers_a, None, props_list, dummy,\n            save_probabilities=True, num_processes=1\n        )\n    del wnet_buffers_a\n    gc.collect()\n    torch.cuda.empty_cache()\n\n    # ── Step 3: nnUNet Model B (Dataset325) ──\n    print(f\"  Running nnUNet Model B (Dataset325)...\")\n    with suppress_stdout():\n        results_b = predictor_b.predict_from_list_of_npy_arrays(\n            wnet_buffers_b, None, props_list, dummy,\n            save_probabilities=True, num_processes=1\n        )\n    del wnet_buffers_b\n    gc.collect()\n    torch.cuda.empty_cache()\n\n    # ── Step 4: Ensemble + postprocess + save ──\n    print(f\"  Ensembling and postprocessing...\")\n    for i, (res_a, res_b) in enumerate(zip(results_a, results_b)):\n        prob_a = res_a[1][1] if isinstance(res_a, (tuple, list)) else res_a[1]\n        prob_b = res_b[1][1] if isinstance(res_b, (tuple, list)) else res_b[1]\n\n        merged = (prob_a * WEIGHT_A) + (prob_b * WEIGHT_B)\n        total_w = WEIGHT_A + WEIGHT_B\n        if total_w != 1.0 and total_w != 0:\n            merged /= total_w\n\n        pred_mask = topo_postprocess(merged)\n\n        save_path = f\"{OUTPUT_DIR}/{chunk_stems[i]}.tif\"\n        tifffile.imwrite(save_path, pred_mask.astype(np.uint8))\n        print(f\"    {chunk_stems[i]}: {pred_mask.sum():,} foreground voxels\")\n        del prob_a, prob_b, merged, pred_mask\n\n    del results_a, results_b\n    gc.collect()\n    torch.cuda.empty_cache()\n\n    print(f\"  Chunk {chunk_idx+1}/{len(chunks)} complete.\")\n\nprint(\"\\nAll inference complete!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T09:54:31.197700Z","iopub.execute_input":"2026-03-02T09:54:31.197983Z","iopub.status.idle":"2026-03-02T10:01:22.626279Z","shell.execute_reply.started":"2026-03-02T09:54:31.197961Z","shell.execute_reply":"2026-03-02T10:01:22.625597Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 10. Generate Submission\n\nZip all prediction TIFFs into `submission.zip` for Kaggle submission.","metadata":{}},{"cell_type":"code","source":"def generate_zip_submission(\n    predictions_dir: Path,\n    output_zip: Path = Path(\"/kaggle/working/submission.zip\"),\n    delete_after_zip: bool = True\n) -> Path:\n    tiff_files = sorted(predictions_dir.glob(\"*.tif\"))\n    if not tiff_files:\n        raise ValueError(f\"No TIFF files in {predictions_dir}\")\n\n    print(f\"Zipping {len(tiff_files)} files to {output_zip}...\")\n    try:\n        with zipfile.ZipFile(output_zip, 'w', zipfile.ZIP_DEFLATED) as zipf:\n            for tiff_path in tqdm(tiff_files, desc=\"Zipping\"):\n                zipf.write(tiff_path, arcname=tiff_path.name)\n        if delete_after_zip:\n            for tiff_path in tiff_files:\n                tiff_path.unlink()\n    except Exception as e:\n        print(f\"Error during zipping: {e}\")\n        return None\n\n    size_mb = output_zip.stat().st_size / (1024 * 1024)\n    print(f\"Submission saved: {output_zip} ({size_mb:.2f} MB)\")\n    return output_zip\n\n\nsubmission_path = generate_zip_submission(OUTPUT_DIR)\nprint(f\"\\nDone! Submit: {submission_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-02T10:01:22.627744Z","iopub.execute_input":"2026-03-02T10:01:22.628903Z","iopub.status.idle":"2026-03-02T10:01:22.991733Z","shell.execute_reply.started":"2026-03-02T10:01:22.628872Z","shell.execute_reply":"2026-03-02T10:01:22.991138Z"}},"outputs":[],"execution_count":null}]}