{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.12.12"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":25563,"databundleVersionId":2094376},{"sourceType":"datasetVersion","sourceId":7643626,"datasetId":4453002,"databundleVersionId":7740213}],"dockerImageVersionId":31329,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":5428.430883,"end_time":"2026-04-10T01:37:47.932342+00:00","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2026-04-10T00:07:19.501459+00:00","version":"2.7.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"00c94f6f","cell_type":"markdown","source":"# Assignment: Kaggle Competition — Plant Pathology 2021\n\n**Competition:** [Plant Pathology 2021 - FGVC8](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8)  \n**Task:** Classify diseases in apple leaf images (multi-label classification)  \n**Metric:** Mean F1-Score (higher = better)  \n\n---\n\n## Before you start — 4 required steps\n\n**Step 1 — Create a Kaggle account** (free): [https://www.kaggle.com](https://www.kaggle.com)\n\n**Step 2 — Accept the competition rules:**  \nGo to → [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules)  \nScroll to the bottom → click **\"I Understand and Accept\"**  \n⚠️ You cannot download data or submit without accepting.\n\n**Step 3 — Download the data:**  \nGo to → [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/data](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/data)  \nClick **\"Download All\"** — this gives you a zip file with:\n```\ntrain_images/       ← ~18,000 apple leaf images\ntest_images/        ← ~2,700 images to predict\ntrain.csv           ← image names + labels\nsample_submission.csv  ← format for your submission\n```\n\n**Step 4 — Choose your environment:**\n\n| Option | Recommended? | Notes |\n|---|---|---|\n| **Kaggle Notebook** | ✅ Best | Free GPU, data already loaded, easy submit |\n| Local machine | ⚠️ OK | You must download ~15 GB data manually |\n\n> **Kaggle Notebook path:**  \n> Go to the competition page → click **\"Code\"** → **\"New Notebook\"**  \n> The dataset is automatically attached — no download needed.\n\n---\n\n## What you submit to me (Canvas)\n\n1. Link to your **Kaggle notebook** (make it public: Settings → Sharing → Public)\n2. Screenshot of your **leaderboard rank**\n3. Your **public score** (Mean F1)","metadata":{"papermill":{"duration":0.004214,"end_time":"2026-04-10T00:07:21.858418+00:00","exception":false,"start_time":"2026-04-10T00:07:21.854204+00:00","status":"completed"},"tags":[]}},{"id":"bf87f800","cell_type":"markdown","source":"## Step 1: Imports","metadata":{"papermill":{"duration":0.003068,"end_time":"2026-04-10T00:07:21.864965+00:00","exception":false,"start_time":"2026-04-10T00:07:21.861897+00:00","status":"completed"},"tags":[]}},{"id":"51294398","cell_type":"code","source":"import os\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"2\"\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom sklearn.model_selection import train_test_split\n\ntf.get_logger().setLevel('ERROR')\n\nprint(\"TensorFlow:\", tf.__version__)\nprint(\"GPU available:\", len(tf.config.list_physical_devices('GPU')) > 0)","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:55:26.933563Z","iopub.execute_input":"2026-04-12T23:55:26.933791Z","iopub.status.idle":"2026-04-12T23:55:56.060188Z","shell.execute_reply.started":"2026-04-12T23:55:26.933767Z","shell.execute_reply":"2026-04-12T23:55:56.059195Z"},"papermill":{"duration":26.48466,"end_time":"2026-04-10T00:07:48.352769+00:00","exception":false,"start_time":"2026-04-10T00:07:21.868109+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"f857737d","cell_type":"markdown","source":"## Step 2: Set Paths\n\nChange `BASE_DIR` depending on where you run this notebook.","metadata":{"papermill":{"duration":0.003416,"end_time":"2026-04-10T00:07:48.359760+00:00","exception":false,"start_time":"2026-04-10T00:07:48.356344+00:00","status":"completed"},"tags":[]}},{"id":"8d2bd211","cell_type":"code","source":"# ── Kaggle Notebook (recommended) ──────────────────────────────────────────\nBASE_DIR    = \"/kaggle/input/plant-pathology-2021-fgvc8\"\n\n# ── Local machine ── uncomment and set your own path ───────────────────────\n# BASE_DIR  = \"./plant-pathology-2021-fgvc8\"  \n\nBASE_DIR   = \"/kaggle/input/competitions/plant-pathology-2021-fgvc8\"\nTRAIN_DIR  = os.path.join(BASE_DIR, \"train_images\")\nTEST_DIR   = os.path.join(BASE_DIR, \"test_images\")\n\n# Confirm images exist\nprint(os.listdir(TRAIN_DIR)[:5])\n\nprint(\"Train images:\", TRAIN_DIR)\nprint(\"Test  images:\", TEST_DIR)\n","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:56:19.984900Z","iopub.execute_input":"2026-04-12T23:56:19.985341Z","iopub.status.idle":"2026-04-12T23:56:20.202548Z","shell.execute_reply.started":"2026-04-12T23:56:19.985309Z","shell.execute_reply":"2026-04-12T23:56:20.201733Z"},"papermill":{"duration":0.195667,"end_time":"2026-04-10T00:07:48.558690+00:00","exception":false,"start_time":"2026-04-10T00:07:48.363023+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"19c9dc1f","cell_type":"markdown","source":"## Step 3: Load and Explore the Data\n\nThis is a **multi-label** problem. One leaf can have multiple diseases at the same time.  \nLabels are stored as space-separated strings, e.g. `\"scab frog_eye_leaf_spot\"`.","metadata":{"papermill":{"duration":0.00331,"end_time":"2026-04-10T00:07:48.565642+00:00","exception":false,"start_time":"2026-04-10T00:07:48.562332+00:00","status":"completed"},"tags":[]}},{"id":"bfe04d59","cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/competitions/plant-pathology-2021-fgvc8/train.csv\")\nsample_df = pd.read_csv(\"/kaggle/input/competitions/plant-pathology-2021-fgvc8/sample_submission.csv\")\n\nprint(\"Train shape:\", train_df.shape)\nprint(\"\\nFirst 5 rows:\")\nprint(train_df.head())\n\n# 6 unique disease classes\nALL_LABELS = ['complex', 'frog_eye_leaf_spot', 'healthy',\n              'powdery_mildew', 'rust', 'scab']\nNUM_CLASSES = len(ALL_LABELS)\nprint(\"\\nClasses:\", ALL_LABELS)","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:56:25.903349Z","iopub.execute_input":"2026-04-12T23:56:25.904148Z","iopub.status.idle":"2026-04-12T23:56:25.955287Z","shell.execute_reply.started":"2026-04-12T23:56:25.904096Z","shell.execute_reply":"2026-04-12T23:56:25.954442Z"},"papermill":{"duration":0.050365,"end_time":"2026-04-10T00:07:48.619305+00:00","exception":false,"start_time":"2026-04-10T00:07:48.568940+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"305aaa7b","cell_type":"code","source":"# Show label distribution\nlabel_counts = train_df['labels'].value_counts().head(10)\n\nplt.figure(figsize=(10, 4))\nlabel_counts.plot(kind='bar')\nplt.title(\"Top 10 Label Combinations\")\nplt.xticks(rotation=30, ha='right')\nplt.tight_layout()\nplt.show()\n\nprint(\"\\nUnique label combinations:\", train_df['labels'].nunique())","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:00.998541Z","iopub.execute_input":"2026-04-12T23:57:00.999323Z","iopub.status.idle":"2026-04-12T23:57:01.176217Z","shell.execute_reply.started":"2026-04-12T23:57:00.999290Z","shell.execute_reply":"2026-04-12T23:57:01.175563Z"},"papermill":{"duration":0.270453,"end_time":"2026-04-10T00:07:48.893429+00:00","exception":false,"start_time":"2026-04-10T00:07:48.622976+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"55f9214d","cell_type":"markdown","source":"## Step 4: Show Sample Images","metadata":{"papermill":{"duration":0.004129,"end_time":"2026-04-10T00:07:48.902049+00:00","exception":false,"start_time":"2026-04-10T00:07:48.897920+00:00","status":"completed"},"tags":[]}},{"id":"37a478ac","cell_type":"code","source":"plt.figure(figsize=(16, 4))\nfor i in range(8):\n    row   = train_df.iloc[i]\n    path  = os.path.join(TRAIN_DIR, row['image'])\n    image = tf.keras.utils.load_img(path, target_size=(128, 128))\n    image = tf.keras.utils.img_to_array(image).astype(\"uint8\")\n    plt.subplot(1, 8, i + 1)\n    plt.imshow(image)\n    plt.title(row['labels'], fontsize=6)\n    plt.axis(\"off\")\n\nplt.suptitle(\"Sample Training Images\", fontsize=13)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:04.692034Z","iopub.execute_input":"2026-04-12T23:57:04.692766Z","iopub.status.idle":"2026-04-12T23:57:06.176638Z","shell.execute_reply.started":"2026-04-12T23:57:04.692735Z","shell.execute_reply":"2026-04-12T23:57:06.175758Z"},"papermill":{"duration":1.302266,"end_time":"2026-04-10T00:07:50.208735+00:00","exception":false,"start_time":"2026-04-10T00:07:48.906469+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"dceb12e6","cell_type":"markdown","source":"## Step 5: Encode Labels\n\nWe convert space-separated label strings into **binary vectors** of length 6.  \nExample: `\"scab healthy\"` → `[0, 0, 1, 0, 0, 1]`  \nThis is called **multi-hot encoding**.","metadata":{"papermill":{"duration":0.010181,"end_time":"2026-04-10T00:07:50.229727+00:00","exception":false,"start_time":"2026-04-10T00:07:50.219546+00:00","status":"completed"},"tags":[]}},{"id":"8a28d1a4","cell_type":"code","source":"# Split label strings into lists\ntrain_df['label_list'] = train_df['labels'].apply(lambda x: x.split())\n\n# Multi-hot encode\nmlb = MultiLabelBinarizer(classes=ALL_LABELS)\ny_all = mlb.fit_transform(train_df['label_list']).astype('float32')\n\nprint(\"Label matrix shape:\", y_all.shape)   # (18632, 6)\nprint(\"Classes:\", mlb.classes_)\nprint(\"\\nExample — first row label:\", train_df['labels'].iloc[0])\nprint(\"Encoded:\", y_all[0])","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:11.209878Z","iopub.execute_input":"2026-04-12T23:57:11.210319Z","iopub.status.idle":"2026-04-12T23:57:11.462233Z","shell.execute_reply.started":"2026-04-12T23:57:11.210287Z","shell.execute_reply":"2026-04-12T23:57:11.461323Z"},"papermill":{"duration":0.2454,"end_time":"2026-04-10T00:07:50.484742+00:00","exception":false,"start_time":"2026-04-10T00:07:50.239342+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"f414508b","cell_type":"markdown","source":"## Step 6: Build the tf.data Pipeline\n\nWe use **both geometric and color augmentation** (from Lectures 11a and 11b).  \nAugmentation is applied inside the model — only active during training.","metadata":{"papermill":{"duration":0.009534,"end_time":"2026-04-10T00:07:50.504754+00:00","exception":false,"start_time":"2026-04-10T00:07:50.495220+00:00","status":"completed"},"tags":[]}},{"id":"071a9d85","cell_type":"code","source":"IMG_SIZE   = 224\nBATCH_SIZE = 32\nAUTOTUNE   = tf.data.AUTOTUNE\n\n# Train / validation split\nX_train, X_val, y_train, y_val = train_test_split(\n    train_df['image'].values, y_all,\n    test_size=0.2, random_state=42\n)\nprint(f\"Train: {len(X_train)} | Val: {len(X_val)}\")\n\n\ndef load_image(path, label):\n    \"\"\"Read JPEG, resize, normalize to [0,1].\"\"\"\n    full_path = tf.strings.join([TRAIN_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) / 255.0\n    return image, label\n\ndef load_test_image(path):\n    \"\"\"Read test image — no label.\"\"\"\n    full_path = tf.strings.join([TEST_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) / 255.0\n    return image\n\n\ntrain_ds = (\n    tf.data.Dataset.from_tensor_slices((X_train, y_train))\n    .shuffle(2000)\n    .map(load_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nval_ds = (\n    tf.data.Dataset.from_tensor_slices((X_val, y_val))\n    .map(load_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\ntest_paths = sample_df['image'].values\ntest_ds = (\n    tf.data.Dataset.from_tensor_slices(test_paths)\n    .map(load_test_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nprint(\"Pipelines ready!\")","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:15.134042Z","iopub.execute_input":"2026-04-12T23:57:15.134828Z","iopub.status.idle":"2026-04-12T23:57:15.415936Z","shell.execute_reply.started":"2026-04-12T23:57:15.134793Z","shell.execute_reply":"2026-04-12T23:57:15.415198Z"},"papermill":{"duration":0.276836,"end_time":"2026-04-10T00:07:50.792361+00:00","exception":false,"start_time":"2026-04-10T00:07:50.515525+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"e93e2159","cell_type":"markdown","source":"## Step 7: Define Augmentation Blocks\n\nCombined geometric + color augmentation from Lectures 11a and 11b.","metadata":{"papermill":{"duration":0.010289,"end_time":"2026-04-10T00:07:50.812824+00:00","exception":false,"start_time":"2026-04-10T00:07:50.802535+00:00","status":"completed"},"tags":[]}},{"id":"844acf74","cell_type":"code","source":"# ── Geometric augmentation ──────────────────────────────────────────────────\ngeometric_augmentation = tf.keras.Sequential([\n    tf.keras.layers.RandomFlip(\"horizontal\"),\n    tf.keras.layers.RandomRotation(factor=0.05),\n    tf.keras.layers.RandomZoom(height_factor=(-0.1, 0.1), width_factor=(-0.1, 0.1)),\n    tf.keras.layers.RandomTranslation(height_factor=0.1, width_factor=0.1),\n], name=\"geometric_augmentation\")\n\n\n# ── Color augmentation custom layers ───────────────────────────────────────\nclass RandomBrightnessLayer(tf.keras.layers.Layer):\n    def __init__(self, max_delta=0.2, **kwargs):\n        super().__init__(**kwargs)\n        self.max_delta = max_delta\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_brightness(x, self.max_delta)\n            x = tf.clip_by_value(x, 0.0, 1.0)\n        return x\n\nclass RandomContrastLayer(tf.keras.layers.Layer):\n    def __init__(self, lower=0.7, upper=1.3, **kwargs):\n        super().__init__(**kwargs)\n        self.lower, self.upper = lower, upper\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_contrast(x, self.lower, self.upper)\n            x = tf.clip_by_value(x, 0.0, 1.0)\n        return x\n\nclass RandomSaturationLayer(tf.keras.layers.Layer):\n    def __init__(self, lower=0.6, upper=1.4, **kwargs):\n        super().__init__(**kwargs)\n        self.lower, self.upper = lower, upper\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_saturation(x, self.lower, self.upper)\n            x = tf.clip_by_value(x, 0.0, 1.0)\n        return x\n\nclass RandomHueLayer(tf.keras.layers.Layer):\n    def __init__(self, max_delta=0.1, **kwargs):\n        super().__init__(**kwargs)\n        self.max_delta = max_delta\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_hue(x, self.max_delta)\n            x = tf.clip_by_value(x, 0.0, 1.0)\n        return x\n\ncolor_augmentation = tf.keras.Sequential([\n    RandomBrightnessLayer(max_delta=0.2),\n    RandomContrastLayer(lower=0.7, upper=1.3),\n    RandomSaturationLayer(lower=0.6, upper=1.4),\n    RandomHueLayer(max_delta=0.1),\n], name=\"color_augmentation\")\n\nprint(\"Augmentation blocks ready!\")","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:26.619566Z","iopub.execute_input":"2026-04-12T23:57:26.620242Z","iopub.status.idle":"2026-04-12T23:57:26.648502Z","shell.execute_reply.started":"2026-04-12T23:57:26.620210Z","shell.execute_reply":"2026-04-12T23:57:26.647775Z"},"papermill":{"duration":0.042651,"end_time":"2026-04-10T00:07:50.865455+00:00","exception":false,"start_time":"2026-04-10T00:07:50.822804+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"5a001430","cell_type":"markdown","source":"## Step 8: Build the Model\n\n**Key difference from previous notebooks:**  \nThis is **multi-label**, not multi-class.  \n- Output activation: `sigmoid` (not softmax) — each class is independent  \n- Loss: `binary_crossentropy` (not sparse categorical)  \n- Each output is a probability 0–1 for each of the 6 classes","metadata":{"papermill":{"duration":0.010012,"end_time":"2026-04-10T00:07:50.885812+00:00","exception":false,"start_time":"2026-04-10T00:07:50.875800+00:00","status":"completed"},"tags":[]}},{"id":"4a0ef679","cell_type":"code","source":"inputs = tf.keras.Input(shape=(IMG_SIZE, IMG_SIZE, 3), name=\"input_image\")\n\n# 1) Geometric augmentation\nx = geometric_augmentation(inputs)\n\n# 2) Color augmentation (expects [0,1])\nx = color_augmentation(x)\n\n# 3) Scale back to [0,255] for EfficientNet\nx = x * 255.0\nx = tf.keras.applications.efficientnet_v2.preprocess_input(x)\n\n# 4) Backbone — EfficientNetV2B0 pretrained on ImageNet\nbase_model = tf.keras.applications.EfficientNetV2M(\n    include_top=False,\n    weights=\"/kaggle/input/datasets/shubhamcodez/keras-weights/efficientnetv2-m_notop.h5\"\n)\nbase_model.trainable = False   # freeze backbone first\nx = base_model(x, training=False)\n\n# 5) Classification head — sigmoid for multi-label\nx = tf.keras.layers.GlobalAveragePooling2D()(x)\nx = tf.keras.layers.Dropout(0.3)(x)\noutputs = tf.keras.layers.Dense(\n    NUM_CLASSES, activation=\"sigmoid\", name=\"predictions\"\n)(x)\n\nmodel = tf.keras.Model(inputs, outputs, name=\"plant_pathology_model\")\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:30.121437Z","iopub.execute_input":"2026-04-12T23:57:30.121863Z","iopub.status.idle":"2026-04-12T23:57:38.935635Z","shell.execute_reply.started":"2026-04-12T23:57:30.121831Z","shell.execute_reply":"2026-04-12T23:57:38.934950Z"},"papermill":{"duration":7.808981,"end_time":"2026-04-10T00:07:58.704464+00:00","exception":false,"start_time":"2026-04-10T00:07:50.895483+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"d7744f01","cell_type":"markdown","source":"## Step 9: Compile and Train","metadata":{"papermill":{"duration":0.010038,"end_time":"2026-04-10T00:07:58.725358+00:00","exception":false,"start_time":"2026-04-10T00:07:58.715320+00:00","status":"completed"},"tags":[]}},{"id":"4f1ccf50","cell_type":"code","source":"model.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),\n    loss=\"binary_crossentropy\",       # multi-label loss\n    metrics=[\"accuracy\"]\n)\n\nprint(\"Training frozen backbone (5 epochs)...\")\nhistory = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=5\n)","metadata":{"execution":{"iopub.status.busy":"2026-04-12T23:57:45.905276Z","iopub.execute_input":"2026-04-12T23:57:45.906000Z","iopub.status.idle":"2026-04-13T00:02:40.472702Z","shell.execute_reply.started":"2026-04-12T23:57:45.905966Z","shell.execute_reply":"2026-04-13T00:02:40.470391Z"},"papermill":{"duration":1474.131485,"end_time":"2026-04-10T00:32:32.867056+00:00","exception":false,"start_time":"2026-04-10T00:07:58.735571+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"4535a2af","cell_type":"markdown","source":"## Step 10: Fine-Tune the Backbone (Optional — improves score)\n\nAfter training the head, we **unfreeze the backbone** and train with a lower learning rate.  \nThis usually improves the score significantly.","metadata":{"papermill":{"duration":0.095791,"end_time":"2026-04-10T00:32:33.059614+00:00","exception":false,"start_time":"2026-04-10T00:32:32.963823+00:00","status":"completed"},"tags":[]}},{"id":"409e803c","cell_type":"code","source":"# Unfreeze backbone\nbase_model.trainable = True\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-4),  # lower LR for fine-tuning\n    loss=\"binary_crossentropy\",\n    metrics=[\"accuracy\"]\n)\n\nprint(\"Fine-tuning full model (10 epochs)...\")\nhistory_ft = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=10\n)","metadata":{"execution":{"iopub.execute_input":"2026-04-10T00:32:33.256661Z","iopub.status.busy":"2026-04-10T00:32:33.251998Z","iopub.status.idle":"2026-04-10T01:37:33.429955Z","shell.execute_reply":"2026-04-10T01:37:33.429170Z"},"papermill":{"duration":3900.552536,"end_time":"2026-04-10T01:37:33.708227+00:00","exception":false,"start_time":"2026-04-10T00:32:33.155691+00:00","status":"completed"},"tags":[]},"outputs":[],"execution_count":null},{"id":"a01760db","cell_type":"markdown","source":"## Step 11: Make Predictions on Test Set\n\nWe need to convert probabilities → label strings to match the submission format.\n\n**Threshold = 0.5** — if probability > 0.5 we predict that class.  \nYou can tune this threshold to improve your F1 score.","metadata":{"papermill":{"duration":0.271279,"end_time":"2026-04-10T01:37:34.266154+00:00","exception":false,"start_time":"2026-04-10T01:37:33.994875+00:00","status":"completed"},"tags":[]}},{"id":"9046c2a9","cell_type":"code","source":"# Get probabilities for all test images\npreds = model.predict(test_ds, verbose=1)   # shape: (N, 6)\n\nTHRESHOLD = 0.35\n\ndef probs_to_label(prob_row, threshold=THRESHOLD):\n    \"\"\"Convert probability array to space-separated label string.\"\"\"\n    selected = [ALL_LABELS[i] for i, p in enumerate(prob_row) if p >= threshold]\n    # If nothing passes threshold, pick the class with the highest probability\n    if len(selected) == 0:\n        selected = [ALL_LABELS[np.argmax(prob_row)]]\n    return \" \".join(selected)\n\npredicted_labels = [probs_to_label(row) for row in preds]\n\nprint(\"Sample predictions:\")\nfor i in range(min(5, len(test_paths))):\n    print(f\"  {test_paths[i]}  →  {predicted_labels[i]}\")","metadata":{"execution":{"iopub.execute_input":"2026-04-10T01:37:34.914120Z","iopub.status.busy":"2026-04-10T01:37:34.913799Z","iopub.status.idle":"2026-04-10T01:37:41.611501Z","shell.execute_reply":"2026-04-10T01:37:41.610794Z"},"papermill":{"duration":7.06822,"end_time":"2026-04-10T01:37:41.613065+00:00","exception":false,"start_time":"2026-04-10T01:37:34.544845+00:00","status":"completed"},"tags":[]},"outputs":[],"execution_count":null},{"id":"3525431c","cell_type":"markdown","source":"## Step 12: Create the Submission File\n\nThe submission must match `sample_submission.csv` format exactly:  \n- Column 1: `image` (filename)  \n- Column 2: `labels` (space-separated predicted classes)","metadata":{"papermill":{"duration":0.269194,"end_time":"2026-04-10T01:37:42.151424+00:00","exception":false,"start_time":"2026-04-10T01:37:41.882230+00:00","status":"completed"},"tags":[]}},{"id":"a5e5d20e","cell_type":"code","source":"submission = pd.DataFrame({\n    'image':  test_paths,\n    'labels': predicted_labels\n})\n\nsubmission.to_csv('submission.csv', index=False)\n\nprint(\"submission.csv saved!\")\nprint(\"\\nFirst 5 rows:\")\nprint(submission.head())\nprint(\"\\nShape:\", submission.shape)","metadata":{"execution":{"iopub.execute_input":"2026-04-10T01:37:42.693353Z","iopub.status.busy":"2026-04-10T01:37:42.692624Z","iopub.status.idle":"2026-04-10T01:37:42.704498Z","shell.execute_reply":"2026-04-10T01:37:42.703677Z"},"papermill":{"duration":0.286444,"end_time":"2026-04-10T01:37:42.705973+00:00","exception":false,"start_time":"2026-04-10T01:37:42.419529+00:00","status":"completed"},"tags":[]},"outputs":[],"execution_count":null},{"id":"dfefd004","cell_type":"markdown","source":"## Step 13: Submit to Kaggle\n\n### Option A — Kaggle Notebook (easiest)\n\n1. Click **\"Save & Run All (Commit)\"** (top right button)\n2. Wait for the notebook to finish running\n3. Go to **Output** tab on the right panel → find `submission.csv`\n4. Click **\"Submit to Competition\"** button\n\n### Option B — Upload manually\n\n1. Go to → [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/submit](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/submit)\n2. Click **\"Upload Submission File\"**\n3. Select your `submission.csv`\n4. Add a short description (e.g. \"EfficientNetV2B0 baseline\")\n5. Click **\"Make Submission\"**\n\n### Option C — Kaggle API (command line)\n\n```bash\npip install kaggle\nkaggle competitions submit -c plant-pathology-2021-fgvc8 -f submission.csv -m \"my submission\"\n```\n\n---\n\nAfter submitting, go to the **Leaderboard** tab to see your rank:  \n[https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/leaderboard](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/leaderboard)","metadata":{"papermill":{"duration":0.268573,"end_time":"2026-04-10T01:37:43.244887+00:00","exception":false,"start_time":"2026-04-10T01:37:42.976314+00:00","status":"completed"},"tags":[]}},{"id":"a18c6c13","cell_type":"markdown","source":"## Step 14: Share on Canvas\n\n**Make your notebook public:**  \nIn your Kaggle notebook → top right → **Settings** → **Sharing** → set to **Public**\n\n**What to submit on Canvas:**\n1. Your Kaggle notebook public URL\n2. Screenshot of the leaderboard showing your username and score\n3. Your public Mean F1 score\n\n---\n\n## How to improve your score\n\n| Idea | Expected gain |\n|---|---|\n| Fine-tune the backbone (Step 10) | +5–10% |\n| Use a larger model (EfficientNetV2M or B3) | +3–5% |\n| Tune the threshold (try 0.3, 0.4, 0.5) | +1–3% |\n| Train more epochs | +2–5% |\n| Add stronger augmentation | +1–3% |\n\nA good baseline score with this notebook: **~0.75–0.82 Mean F1**","metadata":{"papermill":{"duration":0.268028,"end_time":"2026-04-10T01:37:43.861719+00:00","exception":false,"start_time":"2026-04-10T01:37:43.593691+00:00","status":"completed"},"tags":[]}}]}