{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.12.12"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":25563,"databundleVersionId":2094376},{"sourceType":"datasetVersion","sourceId":15616925,"datasetId":9994860,"databundleVersionId":16551011},{"sourceType":"datasetVersion","sourceId":15621361,"datasetId":9997984,"databundleVersionId":16555672}],"dockerImageVersionId":31329,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Assignment: Kaggle Competition — Plant Pathology 2021\n\n**Competition:** [Plant Pathology 2021 - FGVC8](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8)  \n**Task:** Classify diseases in apple leaf images (multi-label classification)  \n**Metric:** Mean F1-Score (higher = better)\n\n---\n\n## Before you start — complete these steps in order\n\n**Step 1 — Create a Kaggle account** (free): [https://www.kaggle.com](https://www.kaggle.com)\n\n**Step 2 — Accept the competition rules:**  \nGo to → [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules)  \nScroll to the bottom → click **\"I Understand and Accept\"**  \n⚠️ You cannot access the data or submit without accepting first.\n\n**Step 3 — Open the competition notebook:**  \nGo to the competition page → click **\"Code\"** tab → click **\"New Notebook\"**. Then, Click **\"File\"** and **\"Import\"** this notebook.\nThe dataset attaches automatically. No download needed.\n\n---\n\n## ⚠️ Important: This competition requires Internet OFF\n\nKaggle re-runs your notebook on a hidden test set to generate your score.  \nThis competition **does not allow internet access** during that re-run.\n\n**What this means for you:**\n- The notebook must run completely **without downloading anything from the internet**\n- `weights=\"imagenet\"` will **fail** because it tries to download the weights\n- You must upload the model weights as a Kaggle Dataset and load them offline\n\n---\n\n## How to upload offline model weights (do this once)\n\n**Part A — Get the weights file**\n\nOption 1 — Download directly (paste this URL in your browser):\n```\nhttps://storage.googleapis.com/tensorflow/keras-applications/efficientnet_v2/efficientnetv2-b0_notop.h5\n```\nThis downloads a file called `efficientnetv2-b0_notop.h5` (~29 MB).\n\nOption 2 — Run this on your local machine or in a Kaggle notebook with internet ON:\n```python\nimport tensorflow as tf\nm = tf.keras.applications.EfficientNetV2B0(include_top=False, weights=\"imagenet\")\nm.save_weights(\"efficientnetv2b0_notop.h5\")\n```\n\n**Part B — Upload as a Kaggle Dataset**\n1. Click **\"Upload\"** → select your `.h5` file\n2. Give it a name, e.g. `efficientnetv2b0`\n3. Click **\"Create\"** → wait for upload to finish\n\n---\n\n## How to submit\n\nThis competition uses **Notebook submission** — you do NOT upload a CSV file manually.\n\n1. Make sure **Internet is OFF** in Settings (right panel → Settings → Internet → OFF)\n2. Click **\"Save & Run All (Commit)\"** — top right button\n3. Wait for the notebook to finish running (~15–25 min with GPU)\n4. Once done → click **\"Submit to Competition\"** button that appears (at the bottom of right panel)\n5. Kaggle re-runs your notebook privately and extracts `submission.csv` from the output\n\n---\n\n## What to submit on Canvas\n\n1. Your Kaggle notebook **public URL** (Settings → Sharing → Public)\n2. **Screenshot** of the leaderboard showing your username and score\n3. Your **public Mean F1 score**","metadata":{}},{"cell_type":"markdown","source":"## Step 1: Imports","metadata":{}},{"cell_type":"code","source":"import os\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"2\"\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom sklearn.model_selection import train_test_split\n\ntf.get_logger().setLevel('ERROR')\n\nprint(\"TensorFlow:\", tf.__version__)\nprint(\"GPU available:\", len(tf.config.list_physical_devices('GPU')) > 0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:30.614205Z","iopub.execute_input":"2026-04-15T12:34:30.614530Z","iopub.status.idle":"2026-04-15T12:34:59.499992Z","shell.execute_reply.started":"2026-04-15T12:34:30.614455Z","shell.execute_reply":"2026-04-15T12:34:59.499186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Set Paths","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom sklearn.metrics import f1_score\n\nSEED = 42\nnp.random.seed(SEED)\ntf.random.set_seed(SEED)\n\nBASE_DIR = \"/kaggle/input/competitions/plant-pathology-2021-fgvc8\"\nTRAIN_DIR = os.path.join(BASE_DIR, \"train_images\")\nTEST_DIR  = os.path.join(BASE_DIR, \"test_images\")\n\nWEIGHTS_PATH = \"/kaggle/input/datasets/shiyuchen2005/efficientnetv2b0-weights/efficientnetv2-b0_notop.h5\"\n\nprint(\"Train images:\", TRAIN_DIR)\nprint(\"Test images :\", TEST_DIR)\nprint(\"Weights path:\", WEIGHTS_PATH)\nprint(\"Train dir exists:\", os.path.exists(TRAIN_DIR))\nprint(\"Test dir exists :\", os.path.exists(TEST_DIR))\nprint(\"Weights exist  :\", os.path.exists(WEIGHTS_PATH))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:59.501916Z","iopub.execute_input":"2026-04-15T12:34:59.502360Z","iopub.status.idle":"2026-04-15T12:34:59.520791Z","shell.execute_reply.started":"2026-04-15T12:34:59.502335Z","shell.execute_reply":"2026-04-15T12:34:59.519938Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3: Load and Explore the Data\n\nThis is a **multi-label** problem — one leaf can have multiple diseases at the same time.  \nLabels are stored as space-separated strings, e.g. `\"scab frog_eye_leaf_spot\"`.","metadata":{}},{"cell_type":"code","source":"# Load CSV files\ntrain_df = pd.read_csv(os.path.join(BASE_DIR, \"train.csv\"))\nsample_df = pd.read_csv(os.path.join(BASE_DIR, \"sample_submission.csv\"))\n\nprint(\"Train shape:\", train_df.shape)\nprint(\"Sample submission shape:\", sample_df.shape)\n\nprint(\"\\nFirst 5 rows of train_df:\")\nprint(train_df.head())\n\n# Define class labels\nALL_LABELS = ['complex', 'frog_eye_leaf_spot', 'healthy', 'powdery_mildew', 'rust', 'scab']\nNUM_CLASSES = len(ALL_LABELS)\n\nprint(\"\\nClasses:\", ALL_LABELS)\nprint(\"Number of classes:\", NUM_CLASSES)\n\n# Check missing values\nprint(\"\\nMissing values in train_df:\")\nprint(train_df.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:59.521977Z","iopub.execute_input":"2026-04-15T12:34:59.522648Z","iopub.status.idle":"2026-04-15T12:34:59.577329Z","shell.execute_reply.started":"2026-04-15T12:34:59.522624Z","shell.execute_reply":"2026-04-15T12:34:59.576586Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Show label distribution (multi-label friendly)\n\nall_labels_flat = []\n\nfor label_str in train_df['labels']:\n    all_labels_flat.extend(label_str.split())\n\nlabel_counts = pd.Series(all_labels_flat).value_counts()\n\nplt.figure(figsize=(10, 4))\nlabel_counts.plot(kind='bar')\nplt.title(\"Label Frequency (Individual Classes)\")\nplt.xticks(rotation=30, ha='right')\nplt.tight_layout()\nplt.show()\n\nprint(\"\\nLabel frequency:\")\nprint(label_counts)\n\nprint(\"\\nUnique label combinations:\", train_df['labels'].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:59.578378Z","iopub.execute_input":"2026-04-15T12:34:59.578627Z","iopub.status.idle":"2026-04-15T12:34:59.845322Z","shell.execute_reply.started":"2026-04-15T12:34:59.578607Z","shell.execute_reply":"2026-04-15T12:34:59.844602Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 4: Show Sample Images","metadata":{}},{"cell_type":"code","source":"ALL_CLASSES = ['complex', 'frog_eye_leaf_spot', 'healthy', 'powdery_mildew', 'rust', 'scab']\n\ntrain_df[\"label_list\"] = train_df[\"labels\"].apply(lambda x: x.split())\n\nmlb = MultiLabelBinarizer(classes=ALL_CLASSES)\ny = mlb.fit_transform(train_df[\"label_list\"])\n\nprint(\"Classes:\", mlb.classes_)\nprint(\"Encoded label shape:\", y.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:59.846239Z","iopub.execute_input":"2026-04-15T12:34:59.846660Z","iopub.status.idle":"2026-04-15T12:34:59.877627Z","shell.execute_reply.started":"2026-04-15T12:34:59.846632Z","shell.execute_reply":"2026-04-15T12:34:59.876931Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 5: Encode Labels\n\nWe convert space-separated label strings into **binary vectors** of length 6.  \nExample: `\"scab frog_eye_leaf_spot\"` → `[0, 1, 0, 0, 0, 1]`  \nThis is called **multi-hot encoding** — multiple positions can be 1 at the same time.","metadata":{}},{"cell_type":"code","source":"# Use number of labels per image for a simple stratification signal\nlabel_count = y.sum(axis=1)\n\ntrain_paths, val_paths, y_train, y_val = train_test_split(\n    train_df[\"image\"].values,\n    y,\n    test_size=0.2,\n    random_state=SEED,\n    stratify=label_count\n)\n\nprint(\"Training samples  :\", len(train_paths))\nprint(\"Validation samples:\", len(val_paths))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:59.878684Z","iopub.execute_input":"2026-04-15T12:34:59.879019Z","iopub.status.idle":"2026-04-15T12:34:59.901715Z","shell.execute_reply.started":"2026-04-15T12:34:59.878996Z","shell.execute_reply":"2026-04-15T12:34:59.900922Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 6: Build the tf.data Pipeline","metadata":{}},{"cell_type":"code","source":"IMG_SIZE = 260\nBATCH_SIZE = 32\nAUTOTUNE = tf.data.AUTOTUNE\n\naugment = tf.keras.Sequential([\n    tf.keras.layers.RandomFlip(\"horizontal\"),\n    tf.keras.layers.RandomRotation(0.08),\n    tf.keras.layers.RandomZoom(0.15),\n    tf.keras.layers.RandomTranslation(0.10, 0.10),\n    tf.keras.layers.RandomContrast(0.15),\n], name=\"augmentation\")\n\ndef decode_train_image(path, label):\n    image_path = tf.strings.join([TRAIN_DIR, \"/\", path])\n    image = tf.io.read_file(image_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) / 255.0\n    image = augment(image, training=True)\n    image = image * 255.0\n    image = tf.keras.applications.efficientnet_v2.preprocess_input(image)\n    return image, tf.cast(label, tf.float32)\n\ndef decode_val_image(path, label):\n    image_path = tf.strings.join([TRAIN_DIR, \"/\", path])\n    image = tf.io.read_file(image_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) / 255.0\n    image = image * 255.0\n    image = tf.keras.applications.efficientnet_v2.preprocess_input(image)\n    return image, tf.cast(label, tf.float32)\n\ntrain_ds = tf.data.Dataset.from_tensor_slices((train_paths, y_train))\ntrain_ds = (\n    train_ds\n    .shuffle(2048, seed=SEED)\n    .map(decode_train_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nval_ds = tf.data.Dataset.from_tensor_slices((val_paths, y_val))\nval_ds = (\n    val_ds\n    .map(decode_val_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nprint(\"train_ds and val_ds ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:34:59.903679Z","iopub.execute_input":"2026-04-15T12:34:59.904282Z","iopub.status.idle":"2026-04-15T12:35:02.373820Z","shell.execute_reply.started":"2026-04-15T12:34:59.904250Z","shell.execute_reply":"2026-04-15T12:35:02.373080Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 7: Define Augmentation Blocks\n\nCombined geometric + color augmentation from Lectures 11a and 11b.","metadata":{}},{"cell_type":"code","source":"# Step 7: Prepare Test Dataset\n\ntest_paths = sample_df[\"image\"].values\n\ndef decode_test_image(path):\n    image_path = tf.strings.join([TEST_DIR, \"/\", path])\n    image = tf.io.read_file(image_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) / 255.0\n    image = image * 255.0\n    image = tf.keras.applications.efficientnet_v2.preprocess_input(image)\n    return image\n\ntest_ds = tf.data.Dataset.from_tensor_slices(test_paths)\ntest_ds = (\n    test_ds\n    .map(decode_test_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nprint(\"test_ds ready.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:36:48.676072Z","iopub.execute_input":"2026-04-15T12:36:48.676755Z","iopub.status.idle":"2026-04-15T12:36:48.733731Z","shell.execute_reply.started":"2026-04-15T12:36:48.676722Z","shell.execute_reply":"2026-04-15T12:36:48.733128Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 8: Build the Model\n\n**Key difference from previous notebooks:**  \nThis is **multi-label** classification:\n- Output activation: `sigmoid` — each of the 6 classes is predicted independently  \n- Loss: `binary_crossentropy` — loss is computed separately per class  \n\n**Offline weights:**  \nWe use `weights=None` and load the `.h5` file manually.  \nThis is required because **internet is OFF** in this competition.","metadata":{}},{"cell_type":"code","source":"inputs = tf.keras.Input(shape=(IMG_SIZE, IMG_SIZE, 3))\n\nbase_model = tf.keras.applications.EfficientNetV2B0(\n    include_top=False,\n    weights=None,\n    input_tensor=inputs\n)\n\nbase_model.load_weights(WEIGHTS_PATH)\nprint(\"Weights loaded from:\", WEIGHTS_PATH)\n\nbase_model.trainable = False\n\nx = base_model.output\nx = tf.keras.layers.GlobalAveragePooling2D()(x)\nx = tf.keras.layers.Dropout(0.35)(x)\noutputs = tf.keras.layers.Dense(len(ALL_CLASSES), activation=\"sigmoid\")(x)\n\nmodel = tf.keras.Model(inputs, outputs)\n\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:36:52.418200Z","iopub.execute_input":"2026-04-15T12:36:52.418922Z","iopub.status.idle":"2026-04-15T12:36:55.903740Z","shell.execute_reply.started":"2026-04-15T12:36:52.418845Z","shell.execute_reply":"2026-04-15T12:36:55.903145Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 9: Compile and Train (Frozen Backbone)","metadata":{}},{"cell_type":"code","source":"# Step 9: Compile and Train (Frozen Backbone)\n\ncheckpoint_path = \"/kaggle/working/best_model.keras\"\n\ncallbacks = [\n    tf.keras.callbacks.EarlyStopping(\n        monitor=\"val_loss\",\n        patience=2,\n        restore_best_weights=True\n    ),\n    tf.keras.callbacks.ReduceLROnPlateau(\n        monitor=\"val_loss\",\n        factor=0.5,\n        patience=1,\n        verbose=1\n    ),\n    tf.keras.callbacks.ModelCheckpoint(\n        checkpoint_path,\n        monitor=\"val_loss\",\n        save_best_only=True,\n        verbose=1\n    )\n]\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n    loss=\"binary_crossentropy\",\n    metrics=[tf.keras.metrics.BinaryAccuracy(name=\"binary_accuracy\")]\n)\n\nprint(\"Training frozen backbone (3 epochs)...\")\nhistory_frozen = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=3,\n    callbacks=callbacks\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T12:48:52.152437Z","iopub.execute_input":"2026-04-15T12:48:52.153153Z","iopub.status.idle":"2026-04-15T13:08:17.693280Z","shell.execute_reply.started":"2026-04-15T12:48:52.153119Z","shell.execute_reply":"2026-04-15T13:08:17.692602Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 10: Fine-Tune the Backbone (improves score)\n\nUnfreeze the backbone and retrain with a smaller learning rate.  \nThis usually gives a significant improvement in F1 score.","metadata":{}},{"cell_type":"code","source":"# Step 10: Fine-Tune the Model\n\nbase_model.trainable = True\n\nfor layer in base_model.layers[:-30]:\n    layer.trainable = False\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-5),\n    loss=\"binary_crossentropy\",\n    metrics=[tf.keras.metrics.BinaryAccuracy(name=\"binary_accuracy\")]\n)\n\nprint(\"Fine-tuning top layers (5 epochs)...\")\nhistory_finetune = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=5,\n    callbacks=callbacks\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T13:09:39.964520Z","iopub.execute_input":"2026-04-15T13:09:39.965351Z","iopub.status.idle":"2026-04-15T13:41:19.300246Z","shell.execute_reply.started":"2026-04-15T13:09:39.965319Z","shell.execute_reply":"2026-04-15T13:41:19.299615Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 11: Make Predictions on Test Set\n\n**Threshold = 0.5** — if probability > 0.5 we predict that class.  \nIf no class passes the threshold, we pick the highest probability class.","metadata":{}},{"cell_type":"code","source":"# Step 11: Validation Prediction and Threshold Search\n\nprint(\"Predicting on validation set...\")\nval_probs = model.predict(val_ds, verbose=1)\n\nthresholds = np.arange(0.20, 0.81, 0.05)\nbest_threshold = 0.5\nbest_f1 = -1\n\nfor t in thresholds:\n    val_pred = (val_probs > t).astype(int)\n    score = f1_score(y_val, val_pred, average=\"samples\")\n\n    print(f\"Threshold = {t:.2f} | Validation Sample F1 = {score:.5f}\")\n\n    if score > best_f1:\n        best_f1 = score\n        best_threshold = t\n\nprint(\"\\nBest threshold:\", best_threshold)\nprint(\"Best validation sample F1:\", best_f1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T13:42:24.182694Z","iopub.execute_input":"2026-04-15T13:42:24.183070Z","iopub.status.idle":"2026-04-15T13:43:26.144404Z","shell.execute_reply.started":"2026-04-15T13:42:24.183040Z","shell.execute_reply":"2026-04-15T13:43:26.143783Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 12: Save submission.csv","metadata":{}},{"cell_type":"code","source":"# Step 12: Save submission.csv\n\nprint(\"Predicting on test set...\")\ntest_probs = model.predict(test_ds, verbose=1)\n\ntest_pred = (test_probs > best_threshold).astype(int)\n\npred_labels = []\nfor idx, row in enumerate(test_pred):\n    labels = [ALL_LABELS[i] for i in range(NUM_CLASSES) if row[i] == 1]\n\n    # If nothing is predicted, choose the highest-probability label\n    if len(labels) == 0:\n        top_idx = np.argmax(test_probs[idx])\n        labels = [ALL_LABELS[top_idx]]\n\n    pred_labels.append(\" \".join(labels))\n\nsubmission = pd.DataFrame({\n    \"image\": sample_df[\"image\"],\n    \"labels\": pred_labels\n})\n\nsubmission.to_csv(\"submission.csv\", index=False)\n\nprint(\"submission.csv saved!\")\nprint(\"Shape:\", submission.shape)\nprint(submission.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-15T13:43:46.405302Z","iopub.execute_input":"2026-04-15T13:43:46.406135Z","iopub.status.idle":"2026-04-15T13:43:58.183912Z","shell.execute_reply.started":"2026-04-15T13:43:46.406104Z","shell.execute_reply":"2026-04-15T13:43:58.183022Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 13: Submit to Kaggle\n\nThis competition uses **Notebook submission** — you do NOT upload the CSV file manually.\n\n### How to submit\n\n1. Make sure **Internet is OFF**  \n   → Right panel → **Settings** → **Internet** → toggle OFF\n\n2. Click **\"Save & Run All (Commit)\"** — top right button  \n   → This runs the full notebook from top to bottom  \n   → Wait ~15–25 minutes for it to finish\n\n3. Once the commit is done → click **\"Submit to Competition\"**  \n   → Kaggle re-runs your notebook privately on the hidden test set  \n   → It reads `submission.csv` from the output and scores it\n\n4. Check your score at the leaderboard:  \n   [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/leaderboard](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/leaderboard)\n\n---\n\n### Checklist before committing\n\n| Item | Check |\n|---|---|\n| Internet toggle is OFF | ☐ |\n| Weights dataset is attached | ☐ |\n| `weights=None` in model code | ☐ |\n| `load_weights(WEIGHTS_PATH)` is in model code | ☐ |\n| `submission.csv` is saved in the last cell | ☐ |\n| All cells run without errors | ☐ |","metadata":{}},{"cell_type":"markdown","source":"## Step 14: Share on Canvas\n\n**Make your notebook public:**  \nIn your Kaggle notebook → **Settings** → **Sharing** → set to **Public**\n\n**Submit on Canvas:**\n1. Your Kaggle notebook public URL\n2. Screenshot of the leaderboard showing your username and score\n3. Your public Mean F1 score\n\n---\n\n## How to improve your score\n\n| Idea | Expected gain |\n|---|---|\n| Fine-tune the backbone (Step 10) | +5–10% |\n| Use a larger model (EfficientNetV2B2 or B3) | +3–5% |\n| Tune the threshold (try 0.3, 0.4, 0.5) | +1–3% |\n| Train more epochs | +2–5% |\n| Add stronger augmentation | +1–3% |\n\nBaseline score with this notebook: **~0.75–0.82 Mean F1**","metadata":{}}]}