{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":25563,"databundleVersionId":2094376},{"sourceType":"datasetVersion","sourceId":15616925,"datasetId":9994860,"databundleVersionId":16551011},{"sourceType":"datasetVersion","sourceId":15740249,"datasetId":10085887,"databundleVersionId":16682554}],"dockerImageVersionId":31329,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"title","cell_type":"markdown","source":"# Assignment: Kaggle Competition — Plant Pathology 2021 (B3 5-Fold Inference)\n\n**Competition:** [Plant Pathology 2021 - FGVC8](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8)  \n**Task:** Classify diseases in apple leaf images (multi-label classification)  \n**Metric:** Mean F1-Score (higher = better)\n\n---\n\n## Before you start — complete these steps in order\n\n**Step 1 — Create a Kaggle account** (free): [https://www.kaggle.com](https://www.kaggle.com)\n\n**Step 2 — Accept the competition rules:**  \nGo to → [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules)  \nScroll to the bottom → click **\"I Understand and Accept\"**  \n⚠️ You cannot access the data or submit without accepting first.\n\n**Step 3 — Open the competition notebook:**  \nGo to the competition page → click **\"Code\"** tab → click **\"New Notebook\"**. Then, Click **\"File\"** and **\"Import\"** this notebook.\nThe dataset attaches automatically. No download needed.\n\n---\n\n## ⚠️ Important: This competition requires Internet OFF\n\nKaggle re-runs your notebook on a hidden test set to generate your score.  \nThis competition **does not allow internet access** during that re-run.\n\n**What this means for you:**\n- The notebook must run completely **without downloading anything from the internet**\n- `weights=\"imagenet\"` will **fail** because it tries to download the weights\n- You must upload the model weights as a Kaggle Dataset and load them offline\n\n---\n\n## How to upload offline model weights (do this once)\n\n**Part A — Get the weights file**\n\nOption 1 — Download directly (paste this URL in your browser):\n```\nhttps://storage.googleapis.com/tensorflow/keras-applications/efficientnet_v2/efficientnetv2-b0_notop.h5\n```\nThis downloads a file called `efficientnetv2-b0_notop.h5` (~29 MB).\n\nOption 2 — Run this on your local machine or in a Kaggle notebook with internet ON:\n```python\nimport tensorflow as tf\nm = tf.keras.applications.EfficientNetV2B0(include_top=False, weights=\"imagenet\")\nm.save_weights(\"efficientnetv2b0_notop.h5\")\n```\n\n**Part B — Upload as a Kaggle Dataset**\n1. Click **\"Upload\"** → select your `.h5` file\n2. Give it a name, e.g. `efficientnetv2b0`\n3. Click **\"Create\"** → wait for upload to finish\n\n---\n\n## How to submit\n\nThis competition uses **Notebook submission** — you do NOT upload a CSV file manually.\n\n1. Make sure **Internet is OFF** in Settings (right panel → Settings → Internet → OFF)\n2. Click **\"Save & Run All (Commit)\"** — top right button\n3. Wait for the notebook to finish running (~15–25 min with GPU)\n4. Once done → click **\"Submit to Competition\"** button that appears (at the bottom of right panel)\n5. Kaggle re-runs your notebook privately and extracts `submission.csv` from the output\n\n---\n\n## What to submit on Canvas\n\n1. Your Kaggle notebook **public URL** (Settings → Sharing → Public)\n2. **Screenshot** of the leaderboard showing your username and score\n3. Your **public Mean F1 score**","metadata":{}},{"id":"s1","cell_type":"markdown","source":"## Step 1: Imports","metadata":{}},{"id":"c1","cell_type":"code","source":"import os\nimport json\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"2\"\n\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\n\ntf.get_logger().setLevel(\"ERROR\")\n\nprint(\"TensorFlow:\", tf.__version__)\nprint(\"GPU available:\", len(tf.config.list_physical_devices(\"GPU\")) > 0)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:53:59.511910Z","iopub.execute_input":"2026-04-15T00:53:59.512299Z","iopub.status.idle":"2026-04-15T00:54:25.327195Z","shell.execute_reply.started":"2026-04-15T00:53:59.512262Z","shell.execute_reply":"2026-04-15T00:54:25.326176Z"}},"outputs":[],"execution_count":null},{"id":"s2","cell_type":"markdown","source":"## Step 2: Set Paths","metadata":{}},{"id":"c2","cell_type":"code","source":"# Path setup for Kaggle inference-only submission\nIS_KAGGLE = os.path.exists(\"/kaggle/input\")\nBACKBONE_NAME = \"b3\"\nBACKBONE_CONFIGS = {\n    \"b3\": {\n        \"constructor\": \"EfficientNetV2B3\",\n        \"img_size\": 300,\n    },\n}\nBACKBONE_CONFIG = BACKBONE_CONFIGS[BACKBONE_NAME]\nBACKBONE_FACTORY = getattr(tf.keras.applications, BACKBONE_CONFIG[\"constructor\"])\nIMG_SIZE = BACKBONE_CONFIG[\"img_size\"]\n\nFINAL_MODEL_DATASET = \"plantpathology2021-b3-final-model\"\nFOLD_WEIGHT_FILENAMES = [f\"fold{i}_best.weights.h5\" for i in range(1, 6)]\nTHRESHOLD_FILENAMES = [\"b3_5fold_thresholds.json\", \"b3_3fold_thresholds.json\", \"thresholds.json\", \"b3_thresholds.json\"]\nTEST_PROBS_FILENAMES = [\"b3_5fold_test_probs.npy\", \"b3_3fold_test_probs.npy\", \"test_probs.npy\"]\n\ndef find_first_file(search_root, candidate_names):\n    for root, _, files in os.walk(search_root):\n        for filename in candidate_names:\n            if filename in files:\n                return os.path.join(root, filename)\n    return None\n\ndef find_all_files(search_root, candidate_names):\n    found = []\n    for root, _, files in os.walk(search_root):\n        for filename in files:\n            if filename in candidate_names:\n                found.append(os.path.join(root, filename))\n    return sorted(found)\n\nif IS_KAGGLE:\n    BASE_DIR = \"/kaggle/input/competitions/plant-pathology-2021-fgvc8\"\n    MODEL_WEIGHTS_PATHS = find_all_files(\"/kaggle/input\", FOLD_WEIGHT_FILENAMES)\n    THRESHOLDS_PATH = find_first_file(\"/kaggle/input\", THRESHOLD_FILENAMES)\n    TEST_PROBS_PATH = find_first_file(\"/kaggle/input\", TEST_PROBS_FILENAMES)\nelse:\n    BASE_DIR = next((path for path in [os.environ.get(\"PLANT_PATHOLOGY_BASE_DIR\"), \"../data/plant-pathology-2021-fgvc8\", \"./data/plant-pathology-2021-fgvc8\"] if path and os.path.exists(path)), \"../data/plant-pathology-2021-fgvc8\")\n    MODEL_WEIGHTS_PATHS = [\n        path for path in [\n            os.environ.get(\"PLANT_PATHOLOGY_FINAL_MODEL_PATH_FOLD1\"),\n            os.environ.get(\"PLANT_PATHOLOGY_FINAL_MODEL_PATH_FOLD2\"),\n            os.environ.get(\"PLANT_PATHOLOGY_FINAL_MODEL_PATH_FOLD3\"),\n            os.environ.get(\"PLANT_PATHOLOGY_FINAL_MODEL_PATH_FOLD4\"),\n            os.environ.get(\"PLANT_PATHOLOGY_FINAL_MODEL_PATH_FOLD5\"),\n            \"./artifacts/fold1_best.weights.h5\",\n            \"./artifacts/fold2_best.weights.h5\",\n            \"./artifacts/fold3_best.weights.h5\",\n            \"./artifacts/fold4_best.weights.h5\",\n            \"./artifacts/fold5_best.weights.h5\",\n            \"../artifacts/fold1_best.weights.h5\",\n            \"../artifacts/fold2_best.weights.h5\",\n            \"../artifacts/fold3_best.weights.h5\",\n            \"../artifacts/fold4_best.weights.h5\",\n            \"../artifacts/fold5_best.weights.h5\",\n        ] if path and os.path.exists(path)\n    ]\n    THRESHOLDS_PATH = next((path for path in [os.environ.get(\"PLANT_PATHOLOGY_THRESHOLDS_PATH\"), \"./artifacts/b3_5fold_thresholds.json\", \"../artifacts/b3_5fold_thresholds.json\", \"./artifacts/b3_3fold_thresholds.json\", \"../artifacts/b3_3fold_thresholds.json\"] if path and os.path.exists(path)), None)\n    TEST_PROBS_PATH = next((path for path in [os.environ.get(\"PLANT_PATHOLOGY_TEST_PROBS_PATH\"), \"./artifacts/b3_5fold_test_probs.npy\", \"../artifacts/b3_5fold_test_probs.npy\", \"./artifacts/b3_3fold_test_probs.npy\", \"../artifacts/b3_3fold_test_probs.npy\"] if path and os.path.exists(path)), None)\n\nTRAIN_DIR = os.path.join(BASE_DIR, \"train_images\")\nTEST_DIR = os.path.join(BASE_DIR, \"test_images\")\nOUTPUT_DIR = os.environ.get(\"PLANT_PATHOLOGY_OUTPUT_DIR\", \"./artifacts\")\nos.makedirs(OUTPUT_DIR, exist_ok=True)\n\nprint(\"Environment:\", \"Kaggle\" if IS_KAGGLE else \"Local\")\nprint(\"Backbone:\", BACKBONE_NAME, \"->\", BACKBONE_CONFIG[\"constructor\"])\nprint(\"BASE_DIR:\", BASE_DIR)\nprint(\"TEST_DIR exists:\", os.path.exists(TEST_DIR))\nprint(\"MODEL_WEIGHTS_PATHS:\", MODEL_WEIGHTS_PATHS)\nprint(\"MODEL_WEIGHTS_COUNT:\", len(MODEL_WEIGHTS_PATHS))\nprint(\"MODEL_WEIGHTS all exist:\", all(os.path.exists(path) for path in MODEL_WEIGHTS_PATHS) if MODEL_WEIGHTS_PATHS else False)\nprint(\"THRESHOLDS_PATH:\", THRESHOLDS_PATH)\nprint(\"THRESHOLDS_PATH exists:\", bool(THRESHOLDS_PATH and os.path.exists(THRESHOLDS_PATH)))\nprint(\"OUTPUT_DIR:\", OUTPUT_DIR)\n\nif len(MODEL_WEIGHTS_PATHS) < 5:\n    raise FileNotFoundError(f\"Could not find all 5 fold weights. Expected files: {FOLD_WEIGHT_FILENAMES}\")\nif THRESHOLDS_PATH is None:\n    raise FileNotFoundError(f\"Could not find threshold JSON. Upload one of {THRESHOLD_FILENAMES} as a Kaggle Dataset.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:54:27.310717Z","iopub.execute_input":"2026-04-15T00:54:27.311533Z","iopub.status.idle":"2026-04-15T00:55:11.895788Z","shell.execute_reply.started":"2026-04-15T00:54:27.311492Z","shell.execute_reply":"2026-04-15T00:55:11.894998Z"}},"outputs":[],"execution_count":null},{"id":"s3","cell_type":"markdown","source":"## Step 3: Load Submission Template\n\nThis inference-only notebook reads `sample_submission.csv` to get the test image order.\n","metadata":{}},{"id":"c3","cell_type":"code","source":"sample_df = pd.read_csv(os.path.join(BASE_DIR, \"sample_submission.csv\"))\nprint(\"Sample submission shape:\", sample_df.shape)\nprint(sample_df.head())\n\nALL_LABELS = [\"complex\", \"frog_eye_leaf_spot\", \"healthy\", \"powdery_mildew\", \"rust\", \"scab\"]\nNUM_CLASSES = len(ALL_LABELS)\nprint(\"\\nClasses:\", ALL_LABELS)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:21.443594Z","iopub.execute_input":"2026-04-15T00:55:21.443870Z","iopub.status.idle":"2026-04-15T00:55:21.469062Z","shell.execute_reply.started":"2026-04-15T00:55:21.443847Z","shell.execute_reply":"2026-04-15T00:55:21.468395Z"}},"outputs":[],"execution_count":null},{"id":"c3b","cell_type":"code","source":"print(\"Inference-only notebook: training label distribution is skipped.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:23.987194Z","iopub.execute_input":"2026-04-15T00:55:23.987528Z","iopub.status.idle":"2026-04-15T00:55:23.991851Z","shell.execute_reply.started":"2026-04-15T00:55:23.987496Z","shell.execute_reply":"2026-04-15T00:55:23.991258Z"}},"outputs":[],"execution_count":null},{"id":"s4","cell_type":"markdown","source":"## Step 4: Skip Training Visualizations\n\nTraining-image previews are omitted in the inference-only submission notebook.\n","metadata":{}},{"id":"c4","cell_type":"code","source":"print(\"Inference-only notebook: sample training image visualization is skipped.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:24.412256Z","iopub.execute_input":"2026-04-15T00:55:24.412861Z","iopub.status.idle":"2026-04-15T00:55:24.417095Z","shell.execute_reply.started":"2026-04-15T00:55:24.412833Z","shell.execute_reply":"2026-04-15T00:55:24.416277Z"}},"outputs":[],"execution_count":null},{"id":"s5","cell_type":"markdown","source":"## Step 5: Label Setup\n\nWe keep only the class-name list needed to decode predictions back into submission strings.\n","metadata":{}},{"id":"c5","cell_type":"code","source":"print(\"Inference-only notebook: no train/val split is created.\")\nprint(\"This notebook expects a fully trained model checkpoint plus a threshold JSON file.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:26.412136Z","iopub.execute_input":"2026-04-15T00:55:26.412784Z","iopub.status.idle":"2026-04-15T00:55:26.416490Z","shell.execute_reply.started":"2026-04-15T00:55:26.412754Z","shell.execute_reply":"2026-04-15T00:55:26.415788Z"}},"outputs":[],"execution_count":null},{"id":"s6","cell_type":"markdown","source":"## Step 6: Build the tf.data Pipeline","metadata":{}},{"id":"c6","cell_type":"code","source":"def load_thresholds(path):\n    with open(path, \"r\", encoding=\"utf-8\") as f:\n        payload = json.load(f)\n    threshold = np.asarray(payload[\"thresholds\"], dtype=np.float32)\n    print(\"Loaded threshold payload:\", payload)\n    return threshold\n\nTHRESHOLD = load_thresholds(THRESHOLDS_PATH)\nprint(\"Threshold vector shape:\", THRESHOLD.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:26.819048Z","iopub.execute_input":"2026-04-15T00:55:26.819817Z","iopub.status.idle":"2026-04-15T00:55:26.828373Z","shell.execute_reply.started":"2026-04-15T00:55:26.819783Z","shell.execute_reply":"2026-04-15T00:55:26.827662Z"}},"outputs":[],"execution_count":null},{"id":"s7","cell_type":"markdown","source":"## Step 7: Skip Training Augmentation\n\nNo training-time augmentation is needed for inference-only submission. TTA is applied later at prediction time.\n","metadata":{}},{"id":"c7","cell_type":"code","source":"print(\"Inference-only notebook: training augmentation blocks are skipped.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:29.691774Z","iopub.execute_input":"2026-04-15T00:55:29.692499Z","iopub.status.idle":"2026-04-15T00:55:29.696211Z","shell.execute_reply.started":"2026-04-15T00:55:29.692469Z","shell.execute_reply":"2026-04-15T00:55:29.695381Z"}},"outputs":[],"execution_count":null},{"id":"s8","cell_type":"markdown","source":"## Step 8: Build the Model\n\n**Key difference from previous notebooks:**  \nThis is **multi-label** classification:\n- Output activation: `sigmoid` — each of the 6 classes is predicted independently  \n- Loss: `binary_crossentropy` — loss is computed separately per class  \n\n**Offline weights:**  \nWe use `weights=None` and load the `.h5` file manually.  \nThis is required because **internet is OFF** in this competition.","metadata":{}},{"id":"c8","cell_type":"code","source":"def build_submission_model():\n    inputs = tf.keras.Input(shape=(IMG_SIZE, IMG_SIZE, 3), name=\"input_image\")\n    x = tf.keras.applications.efficientnet_v2.preprocess_input(inputs * 255.0)\n    base_model = BACKBONE_FACTORY(include_top=False, weights=None)\n    x = base_model(x, training=False)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dropout(0.4)(x)\n    outputs = tf.keras.layers.Dense(NUM_CLASSES, activation=\"sigmoid\", dtype=\"float32\", name=\"predictions\")(x)\n    return tf.keras.Model(inputs, outputs, name=\"plant_pathology_model\")\n\npreview_model = build_submission_model()\npreview_model.summary()\n\ndef probs_to_multihot(probabilities, threshold):\n    threshold_array = np.asarray(threshold, dtype=np.float32)\n    if threshold_array.ndim == 0:\n        threshold_array = np.full(probabilities.shape[1], float(threshold_array), dtype=np.float32)\n    binary = (probabilities >= threshold_array.reshape(1, -1)).astype(\"int32\")\n    empty_rows = binary.sum(axis=1) == 0\n    if np.any(empty_rows):\n        best_idx = np.argmax(probabilities[empty_rows], axis=1)\n        binary[empty_rows] = 0\n        binary[empty_rows, best_idx] = 1\n    return binary\n\ndef labels_from_probabilities(probabilities, threshold):\n    binary = probs_to_multihot(probabilities, threshold)\n    decoded = []\n    for row in binary:\n        selected = [ALL_LABELS[i] for i, flag in enumerate(row) if flag == 1]\n        decoded.append(\" \".join(selected))\n    return decoded\n\nTEST_BATCH_SIZE = 32\nAUTOTUNE = tf.data.AUTOTUNE\n\ndef _load_test_image(image_name, tta_mode=\"original\"):\n    image_path = tf.strings.join([TEST_DIR, \"/\", image_name])\n    image = tf.io.read_file(image_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.convert_image_dtype(image, tf.float32)\n    image = tf.image.resize(image, [IMG_SIZE, IMG_SIZE])\n    if tta_mode == \"flip_lr\":\n        image = tf.image.flip_left_right(image)\n    return image\n\ndef make_test_dataset(paths, tta_mode=\"original\"):\n    ds = tf.data.Dataset.from_tensor_slices(paths)\n    ds = ds.map(lambda x: _load_test_image(x, tta_mode=tta_mode), num_parallel_calls=AUTOTUNE)\n    ds = ds.batch(TEST_BATCH_SIZE)\n    ds = ds.prefetch(AUTOTUNE)\n    return ds\n\ndef predict_probabilities(model, paths, use_tta=False, tta_modes=None):\n    if use_tta:\n        modes = tta_modes or [\"original\", \"flip_lr\"]\n        probs = []\n        for mode in modes:\n            test_ds = make_test_dataset(paths, tta_mode=mode)\n            print(f\"Running TTA mode: {mode}\")\n            probs.append(model.predict(test_ds, verbose=1))\n        return np.mean(probs, axis=0)\n    test_ds = make_test_dataset(paths, tta_mode=\"original\")\n    return model.predict(test_ds, verbose=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:30.105162Z","iopub.execute_input":"2026-04-15T00:55:30.105748Z","iopub.status.idle":"2026-04-15T00:55:33.187060Z","shell.execute_reply.started":"2026-04-15T00:55:30.105715Z","shell.execute_reply":"2026-04-15T00:55:33.186494Z"}},"outputs":[],"execution_count":null},{"id":"s9","cell_type":"markdown","source":"## Step 9: Compile and Train (Frozen Backbone)","metadata":{}},{"id":"c9","cell_type":"code","source":"print(\"Inference-only submission notebook: frozen-backbone training is skipped.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:35.931808Z","iopub.execute_input":"2026-04-15T00:55:35.932592Z","iopub.status.idle":"2026-04-15T00:55:35.936939Z","shell.execute_reply.started":"2026-04-15T00:55:35.932551Z","shell.execute_reply":"2026-04-15T00:55:35.936035Z"}},"outputs":[],"execution_count":null},{"id":"s10","cell_type":"markdown","source":"## Step 10: Fine-Tune the Backbone (improves score)\n\nUnfreeze the backbone and retrain with a smaller learning rate.  \nThis usually gives a significant improvement in F1 score.","metadata":{}},{"id":"c10","cell_type":"code","source":"print(\"Inference-only submission notebook: fine-tuning is skipped.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:36.331556Z","iopub.execute_input":"2026-04-15T00:55:36.331825Z","iopub.status.idle":"2026-04-15T00:55:36.336025Z","shell.execute_reply.started":"2026-04-15T00:55:36.331803Z","shell.execute_reply":"2026-04-15T00:55:36.335228Z"}},"outputs":[],"execution_count":null},{"id":"s11","cell_type":"markdown","source":"## Step 11: Make Predictions on Test Set\n\nUse the best threshold configuration found on the local validation split.  \nIf no class passes the threshold, we still force one label by taking the highest probability class.","metadata":{}},{"id":"c11","cell_type":"code","source":"sample_submission_path = os.path.join(BASE_DIR, \"sample_submission.csv\")\nsample_df_for_submission = globals().get(\"sample_df\")\nif sample_df_for_submission is None or \"image\" not in sample_df_for_submission.columns:\n    sample_df_for_submission = pd.read_csv(sample_submission_path)\n    print(\"Reloaded sample_submission.csv for Step 11 from:\", sample_submission_path)\nelse:\n    print(\"Reusing sample_submission.csv already loaded in memory.\")\nsample_df = sample_df_for_submission\ntest_paths = sample_df_for_submission[\"image\"].astype(str).tolist()\nglobals()[\"test_paths\"] = test_paths\nprint(\"Test image count:\", len(test_paths))\nUSE_TTA = globals().get(\"USE_TTA\", True)\nTTA_MODES = globals().get(\"TTA_MODES\", [\"original\", \"flip_lr\"])\nprint(\"Using TTA:\", USE_TTA, \"| modes:\", TTA_MODES if USE_TTA else [\"original\"])\n\nuse_weight_ensemble = len(MODEL_WEIGHTS_PATHS) >= 5\n\nif use_weight_ensemble:\n    fold_predictions = []\n    for weight_path in MODEL_WEIGHTS_PATHS:\n        tf.keras.backend.clear_session()\n        model = build_submission_model()\n        model.load_weights(weight_path)\n        print(\"Loaded fold weights:\", weight_path)\n        fold_predictions.append(\n            predict_probabilities(model, test_paths, use_tta=USE_TTA, tta_modes=TTA_MODES).astype(np.float32)\n        )\n    preds = np.mean(fold_predictions, axis=0)\nelif TEST_PROBS_PATH and os.path.exists(TEST_PROBS_PATH):\n    print(\"Using precomputed ensemble test probabilities from:\", TEST_PROBS_PATH)\n    preds = np.load(TEST_PROBS_PATH).astype(np.float32)\nelse:\n    raise FileNotFoundError(\"Need either five fold weights or a precomputed test probabilities file.\")\n\nif len(preds) != len(test_paths):\n    raise ValueError(f\"Prediction count {len(preds)} does not match test image count {len(test_paths)}\")\n\npredicted_labels = labels_from_probabilities(preds, THRESHOLD)\nprint(\"Using threshold:\", THRESHOLD)\nprint(\"Fold ensemble size:\", len(MODEL_WEIGHTS_PATHS))\nprint(\"Sample predictions:\")\nfor i in range(min(3, len(test_paths))):\n    print(f\"  {test_paths[i]}  →  {predicted_labels[i]}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:55:38.664010Z","iopub.execute_input":"2026-04-15T00:55:38.664703Z","iopub.status.idle":"2026-04-15T00:56:57.772629Z","shell.execute_reply.started":"2026-04-15T00:55:38.664671Z","shell.execute_reply":"2026-04-15T00:56:57.771917Z"}},"outputs":[],"execution_count":null},{"id":"s12","cell_type":"markdown","source":"## Step 12: Save submission.csv","metadata":{}},{"id":"c12","cell_type":"code","source":"sample_submission_path = os.path.join(BASE_DIR, \"sample_submission.csv\")\nsample_df_for_submission = globals().get(\"sample_df\")\nif sample_df_for_submission is None or \"image\" not in sample_df_for_submission.columns:\n    sample_df_for_submission = pd.read_csv(sample_submission_path)\n    print(\"Reloaded sample_submission.csv for Step 12 from:\", sample_submission_path)\nsubmission_images = sample_df_for_submission[\"image\"].astype(str).tolist()\nif \"predicted_labels\" not in globals():\n    raise NameError(\"predicted_labels is not defined. Re-run Step 11 before Step 12.\")\nif len(predicted_labels) != len(submission_images):\n    raise ValueError(f\"Predicted label count {len(predicted_labels)} does not match submission image count {len(submission_images)}\")\n\nsubmission = pd.DataFrame({\n    \"image\": submission_images,\n    \"labels\": predicted_labels,\n})\n\nsubmission.to_csv(\"submission.csv\", index=False)\nprint(\"submission.csv saved!\")\nprint(submission.head())\nprint(\"Shape:\", submission.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T00:57:08.874995Z","iopub.execute_input":"2026-04-15T00:57:08.875474Z","iopub.status.idle":"2026-04-15T00:57:08.887588Z","shell.execute_reply.started":"2026-04-15T00:57:08.875437Z","shell.execute_reply":"2026-04-15T00:57:08.886692Z"}},"outputs":[],"execution_count":null},{"id":"s13","cell_type":"markdown","source":"## Step 13: Submit to Kaggle\n\nThis version is **inference-only**. Kaggle runs the notebook, loads your final trained B3 checkpoint, applies TTA, writes `submission.csv`, and then scores it.\n\n### Checklist before committing\n\n| Item | Check |\n|---|---|\n| Internet toggle is OFF | ☐ |\n| Final B3 model dataset is attached | ☐ |\n| Threshold JSON is attached | ☐ |\n| `MODEL_WEIGHTS_PATH exists: True` in Step 2 | ☐ |\n| `THRESHOLDS_PATH exists: True` in Step 2 | ☐ |\n| `submission.csv` is saved in the last cell | ☐ |\n\nThen click **Save & Run All (Commit)** and, once it finishes, **Submit to Competition**.\n","metadata":{}},{"id":"s14","cell_type":"markdown","source":"## Step 14: Share on Canvas\n\n**Make your notebook public:**  \nIn your Kaggle notebook → **Settings** → **Sharing** → set to **Public**\n\n**Submit on Canvas:**\n1. Your Kaggle notebook public URL\n2. Screenshot of the leaderboard showing your username and score\n3. Your public Mean F1 score\n\n---\n\n## How to improve your score\n\n| Idea | Expected gain |\n|---|---|\n| Fine-tune the backbone (Step 10) | +5–10% |\n| Use a larger model (EfficientNetV2B2 or B3) | +3–5% |\n| Tune the threshold (try 0.3, 0.4, 0.5) | +1–3% |\n| Train more epochs | +2–5% |\n| Add stronger augmentation | +1–3% |\n\nBaseline score with this notebook: **~0.75–0.82 Mean F1**","metadata":{}},{"id":"e735e030-fff3-4014-b090-07783331f178","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}