{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.12.12"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":25563,"databundleVersionId":2094376},{"sourceType":"datasetVersion","sourceId":7643626,"datasetId":4453002,"databundleVersionId":7740213}],"dockerImageVersionId":31329,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"papermill":{"default_parameters":{},"duration":3217.18409,"end_time":"2026-04-10T07:35:19.914390+00:00","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2026-04-10T06:41:42.730300+00:00","version":"2.7.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Assignment: Kaggle Competition — Plant Pathology 2021\n\n**Competition:** [Plant Pathology 2021 - FGVC8](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8)  \n**Task:** Classify diseases in apple leaf images (multi-label classification)  \n**Metric:** Mean F1-Score (higher = better)\n\n---\n\n## Before you start — complete these steps in order\n\n**Step 1 — Create a Kaggle account** (free): [https://www.kaggle.com](https://www.kaggle.com)\n\n**Step 2 — Accept the competition rules:**  \nGo to → [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/rules)  \nScroll to the bottom → click **\"I Understand and Accept\"**  \n⚠️ You cannot access the data or submit without accepting first.\n\n**Step 3 — Open the competition notebook:**  \nGo to the competition page → click **\"Code\"** tab → click **\"New Notebook\"**. Then, Click **\"File\"** and **\"Import\"** this notebook.\nThe dataset attaches automatically. No download needed.\n\n---\n\n## ⚠️ Important: This competition requires Internet OFF\n\nKaggle re-runs your notebook on a hidden test set to generate your score.  \nThis competition **does not allow internet access** during that re-run.\n\n**What this means for you:**\n- The notebook must run completely **without downloading anything from the internet**\n- `weights=\"imagenet\"` will **fail** because it tries to download the weights\n- You must upload the model weights as a Kaggle Dataset and load them offline\n\n---\n\n## How to upload offline model weights (do this once)\n\n**Part A — Get the weights file**\n\nOption 1 — Download directly (paste this URL in your browser):\n```\nhttps://storage.googleapis.com/tensorflow/keras-applications/efficientnet_v2/efficientnetv2-b0_notop.h5\n```\nThis downloads a file called `efficientnetv2-b0_notop.h5` (~29 MB).\n\nOption 2 — Run this on your local machine or in a Kaggle notebook with internet ON:\n```python\nimport tensorflow as tf\nm = tf.keras.applications.EfficientNetV2B0(include_top=False, weights=\"imagenet\")\nm.save_weights(\"efficientnetv2b0_notop.h5\")\n```\n\n**Part B — Upload as a Kaggle Dataset**\n1. Click **\"Upload\"** → select your `.h5` file\n2. Give it a name, e.g. `efficientnetv2b0`\n3. Click **\"Create\"** → wait for upload to finish\n\n---\n\n## How to submit\n\nThis competition uses **Notebook submission** — you do NOT upload a CSV file manually.\n\n1. Make sure **Internet is OFF** in Settings (right panel → Settings → Internet → OFF)\n2. Click **\"Save & Run All (Commit)\"** — top right button\n3. Wait for the notebook to finish running (~15–25 min with GPU)\n4. Once done → click **\"Submit to Competition\"** button that appears (at the bottom of right panel)\n5. Kaggle re-runs your notebook privately and extracts `submission.csv` from the output\n\n---\n\n## What to submit on Canvas\n\n1. Your Kaggle notebook **public URL** (Settings → Sharing → Public)\n2. **Screenshot** of the leaderboard showing your username and score\n3. Your **public Mean F1 score**","metadata":{"papermill":{"duration":0.004439,"end_time":"2026-04-10T06:41:45.281496+00:00","exception":false,"start_time":"2026-04-10T06:41:45.277057+00:00","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Step 1: Imports","metadata":{"papermill":{"duration":0.005158,"end_time":"2026-04-10T06:42:02.615624+00:00","exception":false,"start_time":"2026-04-10T06:42:02.610466+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import os\nos.environ[\"TF_CPP_MIN_LOG_LEVEL\"] = \"2\"\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau\n\ntf.get_logger().setLevel('ERROR')\n\nprint(\"TensorFlow:\", tf.__version__)\nprint(\"GPU available:\", len(tf.config.list_physical_devices('GPU')) > 0)","metadata":{"papermill":{"duration":29.886762,"end_time":"2026-04-10T06:42:32.505818+00:00","exception":false,"start_time":"2026-04-10T06:42:02.619056+00:00","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:01:52.750138Z","iopub.execute_input":"2026-05-20T07:01:52.750464Z","iopub.status.idle":"2026-05-20T07:01:55.389279Z","shell.execute_reply.started":"2026-05-20T07:01:52.750438Z","shell.execute_reply":"2026-05-20T07:01:55.388570Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Set Paths","metadata":{"papermill":{"duration":0.003682,"end_time":"2026-04-10T06:42:32.513691+00:00","exception":false,"start_time":"2026-04-10T06:42:32.510009+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import os\n\nBASE_DIR   = \"/kaggle/input/competitions/plant-pathology-2021-fgvc8\"\nTRAIN_DIR  = os.path.join(BASE_DIR, \"train_images\")\nTEST_DIR   = os.path.join(BASE_DIR, \"test_images\")\n\nprint(\"Train images:\", TRAIN_DIR)\nprint(\"Test images:\", TEST_DIR)\n\nprint(\"Train dir exists:\", os.path.exists(TRAIN_DIR))\nprint(\"Test dir exists:\", os.path.exists(TEST_DIR))\n\nif os.path.exists(TRAIN_DIR):\n    print(\"Sample train files:\", os.listdir(TRAIN_DIR)[:3])\n\nif os.path.exists(TEST_DIR):\n    print(\"Sample test files:\", os.listdir(TEST_DIR)[:3])","metadata":{"papermill":{"duration":0.021964,"end_time":"2026-04-10T06:42:32.539333+00:00","exception":false,"start_time":"2026-04-10T06:42:32.517369+00:00","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:01:55.391014Z","iopub.execute_input":"2026-05-20T07:01:55.391546Z","iopub.status.idle":"2026-05-20T07:01:55.711721Z","shell.execute_reply.started":"2026-05-20T07:01:55.391519Z","shell.execute_reply":"2026-05-20T07:01:55.711120Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3: Load and Explore the Data\n\nThis is a **multi-label** problem — one leaf can have multiple diseases at the same time.  \nLabels are stored as space-separated strings, e.g. `\"scab frog_eye_leaf_spot\"`.","metadata":{"papermill":{"duration":0.003738,"end_time":"2026-04-10T06:42:32.563008+00:00","exception":false,"start_time":"2026-04-10T06:42:32.559270+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_df  = pd.read_csv(os.path.join(BASE_DIR, \"train.csv\"))\nsample_df = pd.read_csv(os.path.join(BASE_DIR, \"sample_submission.csv\"))\n\nprint(\"Train shape:\", train_df.shape)\nprint(train_df.head())\n\nALL_LABELS = ['complex', 'frog_eye_leaf_spot', 'healthy',\n              'powdery_mildew', 'rust', 'scab']\n\nNUM_CLASSES = len(ALL_LABELS)\n\nassert set(\" \".join(train_df['labels']).split()) == set(ALL_LABELS)\n\nprint(\"\\nClasses:\", ALL_LABELS)\nprint(train_df['labels'].value_counts().head(10))","metadata":{"papermill":{"duration":0.051881,"end_time":"2026-04-10T06:42:32.618743+00:00","exception":false,"start_time":"2026-04-10T06:42:32.566862+00:00","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:01:55.712515Z","iopub.execute_input":"2026-05-20T07:01:55.713202Z","iopub.status.idle":"2026-05-20T07:01:55.805784Z","shell.execute_reply.started":"2026-05-20T07:01:55.713176Z","shell.execute_reply":"2026-05-20T07:01:55.805164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from collections import Counter\n\nall_labels = train_df['labels'].str.split().sum()\nlabel_counts = Counter(all_labels)\n\nlabel_counts = pd.Series(label_counts).sort_values(ascending=False)\n\nplt.figure(figsize=(10, 4))\nlabel_counts.plot(kind='bar')\nplt.title(\"Label Distribution\")\nplt.xticks(rotation=30, ha='right')\nplt.tight_layout()\nplt.show()\n\nprint(\"\\nNumber of classes:\", len(label_counts))","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:01:55.806594Z","iopub.execute_input":"2026-05-20T07:01:55.806898Z","iopub.status.idle":"2026-05-20T07:01:56.555945Z","shell.execute_reply.started":"2026-05-20T07:01:55.806864Z","shell.execute_reply":"2026-05-20T07:01:56.555318Z"},"papermill":{"duration":0.331898,"end_time":"2026-04-10T06:42:32.954890+00:00","exception":false,"start_time":"2026-04-10T06:42:32.622992+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 4: Show Sample Images","metadata":{"papermill":{"duration":0.004569,"end_time":"2026-04-10T06:42:32.964481+00:00","exception":false,"start_time":"2026-04-10T06:42:32.959912+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport tensorflow as tf\nimport os\n\nplt.figure(figsize=(18, 4))\n\nfor i in range(8):\n    row = train_df.iloc[i]\n    path = os.path.join(BASE_DIR, 'train_images', row['image'])\n    \n    image = tf.keras.utils.load_img(path, target_size=(128, 128))\n    image = tf.keras.utils.img_to_array(image).astype(\"uint8\")\n    \n    plt.subplot(1, 8, i + 1)\n    plt.imshow(image)\n    \n    labels = row['labels']\n    if isinstance(labels, str):\n        label_text = \"\\n\".join(labels.split())  \n    else:\n        label_text = \"\\n\".join(labels)         \n        \n    plt.title(label_text, fontsize=7)\n    plt.axis(\"off\")\n\nplt.suptitle(\"Sample Training Images\", fontsize=14)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:01:56.557500Z","iopub.execute_input":"2026-05-20T07:01:56.557791Z","iopub.status.idle":"2026-05-20T07:01:57.860508Z","shell.execute_reply.started":"2026-05-20T07:01:56.557768Z","shell.execute_reply":"2026-05-20T07:01:57.859675Z"},"papermill":{"duration":1.38077,"end_time":"2026-04-10T06:42:34.350046+00:00","exception":false,"start_time":"2026-04-10T06:42:32.969276+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 5: Encode Labels\n\nWe convert space-separated label strings into **binary vectors** of length 6.  \nExample: `\"scab frog_eye_leaf_spot\"` → `[0, 1, 0, 0, 0, 1]`  \nThis is called **multi-hot encoding** — multiple positions can be 1 at the same time.","metadata":{"papermill":{"duration":0.01188,"end_time":"2026-04-10T06:42:34.374431+00:00","exception":false,"start_time":"2026-04-10T06:42:34.362551+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_df['label_list'] = train_df['labels'].apply(lambda x: x.split())\n\nmlb = MultiLabelBinarizer(classes=ALL_LABELS)\nmlb.fit(train_df['label_list'])\n\ny_all = mlb.transform(train_df['label_list']).astype('float32')\n\nprint(\"Label matrix shape:\", y_all.shape)\nprint(\"Classes:\", mlb.classes_)\nprint(\"Example:\", train_df['labels'].iloc[0], \"->\", y_all[0])","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:01:57.861751Z","iopub.execute_input":"2026-05-20T07:01:57.862104Z","iopub.status.idle":"2026-05-20T07:01:58.099373Z","shell.execute_reply.started":"2026-05-20T07:01:57.862077Z","shell.execute_reply":"2026-05-20T07:01:58.098637Z"},"papermill":{"duration":0.274218,"end_time":"2026-04-10T06:42:34.659635+00:00","exception":false,"start_time":"2026-04-10T06:42:34.385417+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nSEED = 42\nIMAGE_SIZE = (224, 224)\nBATCH_SIZE = 32\nBASE_DIR = '/kaggle/input/plant-pathology-2021-fgvc8/'\n\nfor i, class_name in enumerate(mlb.classes_):\n    train_df[class_name] = y_all[:, i]\n\ntrain_data, val_data = train_test_split(train_df, test_size=0.2, random_state=SEED, shuffle=True)\n\nprint(\"Train data shape:\", train_data.shape)\nprint(\"Val data shape:\", val_data.shape)\n\nif 'TRAIN_DIR' in globals() and os.path.exists(TRAIN_DIR):\n    ACTUAL_IMAGE_DIR = TRAIN_DIR\nelif os.path.exists(os.path.join(BASE_DIR, 'train_images')):\n    ACTUAL_IMAGE_DIR = os.path.join(BASE_DIR, 'train_images')\nelse:\n    ACTUAL_IMAGE_DIR = BASE_DIR\n\nprint(\"Image directory found at:\", ACTUAL_IMAGE_DIR)\n\ntrain_datagen = ImageDataGenerator(\n    rescale=1./255,\n    rotation_range=20,\n    horizontal_flip=True,\n    vertical_flip=True\n)\n\nval_datagen = ImageDataGenerator(rescale=1./255)\n\ntrain_generator = train_datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=ACTUAL_IMAGE_DIR,\n    x_col='image',\n    y_col=list(mlb.classes_), \n    target_size=IMAGE_SIZE,\n    batch_size=BATCH_SIZE,\n    class_mode='raw',        \n    seed=SEED\n)\n\nval_generator = val_datagen.flow_from_dataframe(\n    dataframe=val_data,\n    directory=ACTUAL_IMAGE_DIR,\n    x_col='image',\n    y_col=list(mlb.classes_),\n    target_size=IMAGE_SIZE,\n    batch_size=BATCH_SIZE,\n    class_mode='raw',\n    shuffle=False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:01:58.100434Z","iopub.execute_input":"2026-05-20T07:01:58.100755Z","iopub.status.idle":"2026-05-20T07:03:06.976341Z","shell.execute_reply.started":"2026-05-20T07:01:58.100729Z","shell.execute_reply":"2026-05-20T07:03:06.975507Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models, optimizers\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau\n\nLOCAL_WEIGHTS_PATH = '/kaggle/input/datasets/shubhamcodez/keras-weights/efficientnetb0_notop.h5'\n\nbase_model = tf.keras.applications.EfficientNetB0(\n    include_top=False, \n    weights=None,      \n    input_shape=(IMAGE_SIZE[0], IMAGE_SIZE[1], 3)\n)\n\nbase_model.load_weights(LOCAL_WEIGHTS_PATH)\n\nx = base_model.output\nx = layers.GlobalAveragePooling2D()(x)\nx = layers.Dropout(0.5)(x)\npredictions = layers.Dense(len(mlb.classes_), activation='sigmoid')(x)\n\nmodel = models.Model(inputs=base_model.input, outputs=predictions)\n\nmodel.compile(\n    optimizer=optimizers.Adam(learning_rate=1e-4),\n    loss='binary_crossentropy',                    \n    metrics=[\n        tf.keras.metrics.BinaryAccuracy(name='accuracy'),\n        tf.keras.metrics.AUC(multi_label=True, name='auc')\n    ]\n)\n\ncallbacks = [\n        EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True),\n        ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=2, min_lr=1e-6)\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:03:06.977336Z","iopub.execute_input":"2026-05-20T07:03:06.977637Z","iopub.status.idle":"2026-05-20T07:03:11.128354Z","shell.execute_reply.started":"2026-05-20T07:03:06.977614Z","shell.execute_reply":"2026-05-20T07:03:11.127753Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 6: Build the tf.data Pipeline","metadata":{"papermill":{"duration":0.010478,"end_time":"2026-04-10T06:42:34.681266+00:00","exception":false,"start_time":"2026-04-10T06:42:34.670788+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import os\nIMG_SIZE = 224\nBATCH_SIZE = 32\nAUTOTUNE = tf.data.AUTOTUNE\n\nif os.path.exists('/kaggle/input/competitions/plant-pathology-2021-fgvc8/train_images'):\n    TRAIN_DIR = '/kaggle/input/competitions/plant-pathology-2021-fgvc8/train_images'\n    TEST_DIR = '/kaggle/input/competitions/plant-pathology-2021-fgvc8/test_images'\nelse:\n    TRAIN_DIR = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n    TEST_DIR = '/kaggle/input/plant-pathology-2021-fgvc8/test_images'\n\nprint(\"Using TRAIN_DIR:\", TRAIN_DIR)\nX_train, X_val, y_train, y_val = train_test_split(\n    train_df['image'].values, y_all,\n    test_size=0.15, random_state=42, shuffle=True\n)\n\nprint(f\"Train: {len(X_train)} | Val: {len(X_val)}\")\n\ndef load_train_image(path, label):\n    full_path = tf.strings.join([TRAIN_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) \n    return image, label\n\ndef load_val_image(path, label):\n    full_path = tf.strings.join([TRAIN_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) \n    return image, label\n\ndef load_test_image(path):\n    full_path = tf.strings.join([TEST_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) \n    return image\n\ntrain_ds = (\n    tf.data.Dataset.from_tensor_slices((X_train, y_train))\n    .shuffle(len(X_train))\n    .map(load_train_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nval_ds = (\n    tf.data.Dataset.from_tensor_slices((X_val, y_val))\n    .map(load_val_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\ntest_paths = sample_df['image'].values\ntest_ds = (\n    tf.data.Dataset.from_tensor_slices(test_paths)\n    .map(load_test_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nprint(\"Pipelines ready!\")","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:03:11.129379Z","iopub.execute_input":"2026-05-20T07:03:11.129671Z","iopub.status.idle":"2026-05-20T07:03:11.289468Z","shell.execute_reply.started":"2026-05-20T07:03:11.129640Z","shell.execute_reply":"2026-05-20T07:03:11.288887Z"},"papermill":{"duration":0.302911,"end_time":"2026-04-10T06:42:34.994800+00:00","exception":false,"start_time":"2026-04-10T06:42:34.691889+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 7: Define Augmentation Blocks\n\nCombined geometric + color augmentation from Lectures 11a and 11b.","metadata":{"papermill":{"duration":0.011765,"end_time":"2026-04-10T06:42:35.018334+00:00","exception":false,"start_time":"2026-04-10T06:42:35.006569+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"geometric_augmentation = tf.keras.Sequential([\n    tf.keras.layers.RandomFlip(\"horizontal\"),\n    tf.keras.layers.RandomRotation(factor=0.05),\n    tf.keras.layers.RandomZoom(height_factor=(-0.1, 0.1),\n                               width_factor=(-0.1, 0.1)),\n    tf.keras.layers.RandomTranslation(height_factor=0.1,\n                                      width_factor=0.1),\n], name=\"geometric_augmentation\")\n\n\n\nclass RandomBrightnessLayer(tf.keras.layers.Layer):\n    def __init__(self, max_delta=0.2, **kwargs):\n        super().__init__(**kwargs)\n        self.max_delta = max_delta\n\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_brightness(x, self.max_delta)\n        return x\n\n\nclass RandomContrastLayer(tf.keras.layers.Layer):\n    def __init__(self, lower=0.7, upper=1.3, **kwargs):\n        super().__init__(**kwargs)\n        self.lower = lower\n        self.upper = upper\n\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_contrast(x, self.lower, self.upper)\n        return x\n\n\nclass RandomSaturationLayer(tf.keras.layers.Layer):\n    def __init__(self, lower=0.6, upper=1.4, **kwargs):\n        super().__init__(**kwargs)\n        self.lower = lower\n        self.upper = upper\n\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_saturation(x, self.lower, self.upper)\n        return x\n\n\nclass RandomHueLayer(tf.keras.layers.Layer):\n    def __init__(self, max_delta=0.1, **kwargs):\n        super().__init__(**kwargs)\n        self.max_delta = max_delta\n\n    def call(self, x, training=False):\n        if training:\n            x = tf.image.random_hue(x, self.max_delta)\n        return x\n\n\ncolor_augmentation = tf.keras.Sequential([\n    RandomBrightnessLayer(max_delta=0.2),\n    RandomContrastLayer(lower=0.7, upper=1.3),\n    RandomSaturationLayer(lower=0.6, upper=1.4),\n    RandomHueLayer(max_delta=0.1),\n], name=\"color_augmentation\")\n\n\nprint(\"Augmentation blocks ready!\")","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:03:11.290223Z","iopub.execute_input":"2026-05-20T07:03:11.290416Z","iopub.status.idle":"2026-05-20T07:03:11.312384Z","shell.execute_reply.started":"2026-05-20T07:03:11.290395Z","shell.execute_reply":"2026-05-20T07:03:11.311639Z"},"papermill":{"duration":0.049648,"end_time":"2026-04-10T06:42:35.079279+00:00","exception":false,"start_time":"2026-04-10T06:42:35.029631+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 8: Build the Model\n\n**Key difference from previous notebooks:**  \nThis is **multi-label** classification:\n- Output activation: `sigmoid` — each of the 6 classes is predicted independently  \n- Loss: `binary_crossentropy` — loss is computed separately per class  \n\n**Offline weights:**  \nWe use `weights=None` and load the `.h5` file manually.  \nThis is required because **internet is OFF** in this competition.","metadata":{"papermill":{"duration":0.011063,"end_time":"2026-04-10T06:42:35.102013+00:00","exception":false,"start_time":"2026-04-10T06:42:35.090950+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"WEIGHTS_PATH = \"/kaggle/input/datasets/shubhamcodez/keras-weights/efficientnetb0_notop.h5\"\n\ninputs = tf.keras.Input(shape=(IMG_SIZE, IMG_SIZE, 3), name=\"input_image\")\n\nx = geometric_augmentation(inputs)\nx = color_augmentation(x)\n\nbase_model = tf.keras.applications.EfficientNetB0(\n    include_top=False,\n    weights=None,   \n    input_shape=(IMG_SIZE, IMG_SIZE, 3)\n)\n\nbase_model.load_weights(\n    WEIGHTS_PATH,\n    by_name=True,\n    skip_mismatch=True\n)\n\nbase_model.trainable = False\n\nx = base_model(x, training=False)\n\nx = tf.keras.layers.GlobalAveragePooling2D()(x)\nx = tf.keras.layers.Dropout(0.4)(x)\n\noutputs = tf.keras.layers.Dense(\n    NUM_CLASSES,\n    activation=\"sigmoid\",\n    name=\"predictions\"\n)(x)\n\nmodel = tf.keras.Model(inputs, outputs, name=\"plant_pathology_model\")\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:03:11.313369Z","iopub.execute_input":"2026-05-20T07:03:11.313710Z","iopub.status.idle":"2026-05-20T07:03:12.354956Z","shell.execute_reply.started":"2026-05-20T07:03:11.313686Z","shell.execute_reply":"2026-05-20T07:03:12.354334Z"},"papermill":{"duration":3.519775,"end_time":"2026-04-10T06:42:38.633043+00:00","exception":false,"start_time":"2026-04-10T06:42:35.113268+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\nif os.path.exists('/kaggle/input/competitions/plant-pathology-2021-fgvc8/train_images'):\n    TRAIN_DIR = '/kaggle/input/competitions/plant-pathology-2021-fgvc8/train_images'\n    TEST_DIR = '/kaggle/input/competitions/plant-pathology-2021-fgvc8/test_images'\nelse:\n    TRAIN_DIR = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n    TEST_DIR = '/kaggle/input/plant-pathology-2021-fgvc8/test_images'\n\nprint(\"Using TRAIN_DIR:\", TRAIN_DIR)\n\ndef load_train_image(path, label):\n    full_path = tf.strings.join([TRAIN_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) \n    return image, label\n\ndef load_val_image(path, label):\n    full_path = tf.strings.join([TRAIN_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) \n    return image, label\n\ndef load_test_image(path):\n    full_path = tf.strings.join([TEST_DIR + \"/\", path])\n    image = tf.io.read_file(full_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = tf.cast(image, tf.float32) \n    return image\n\ntrain_ds = (\n    tf.data.Dataset.from_tensor_slices((X_train, y_train))\n    .shuffle(len(X_train))\n    .map(load_train_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nval_ds = (\n    tf.data.Dataset.from_tensor_slices((X_val, y_val))\n    .map(load_val_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\ntest_paths = sample_df['image'].values\ntest_ds = (\n    tf.data.Dataset.from_tensor_slices(test_paths)\n    .map(load_test_image, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n\nprint(\"Pipelines ready!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:03:12.355899Z","iopub.execute_input":"2026-05-20T07:03:12.356148Z","iopub.status.idle":"2026-05-20T07:03:12.495780Z","shell.execute_reply.started":"2026-05-20T07:03:12.356126Z","shell.execute_reply":"2026-05-20T07:03:12.495124Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 9: Compile and Train (Frozen Backbone)","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\n\ncallbacks = [\n    tf.keras.callbacks.EarlyStopping(\n        monitor='val_loss', \n        patience=3, \n        restore_best_weights=True,\n        verbose=1\n    ),\n    tf.keras.callbacks.ReduceLROnPlateau(\n        monitor='val_loss', \n        factor=0.5, \n        patience=2, \n        min_lr=1e-6,\n        verbose=1\n    )\n]\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-4),  \n    loss=tf.keras.losses.BinaryCrossentropy(),                \n    metrics=[\n        tf.keras.metrics.BinaryAccuracy(name='accuracy'),\n        tf.keras.metrics.AUC(multi_label=True, name='auc') \n    ]\n)\n\nprint(\"Training frozen backbone...\")\n\nhistory = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=8,\n    callbacks=callbacks   \n)","metadata":{"execution":{"iopub.status.busy":"2026-05-20T07:03:12.496645Z","iopub.execute_input":"2026-05-20T07:03:12.496952Z","iopub.status.idle":"2026-05-20T07:40:58.766046Z","shell.execute_reply.started":"2026-05-20T07:03:12.496928Z","shell.execute_reply":"2026-05-20T07:40:58.765375Z"},"papermill":{"duration":1540.329906,"end_time":"2026-04-10T07:08:18.997881+00:00","exception":false,"start_time":"2026-04-10T06:42:38.667975+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 10: Fine-Tune the Backbone (improves score)\n\nUnfreeze the backbone and retrain with a smaller learning rate.  \nThis usually gives a significant improvement in F1 score.","metadata":{"papermill":{"duration":0.108395,"end_time":"2026-04-10T07:08:19.214672+00:00","exception":false,"start_time":"2026-04-10T07:08:19.106277+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import tensorflow as tf\n\nbase_model.trainable = True\n\nfor layer in base_model.layers[:-50]:\n    layer.trainable = False\n\nfor layer in base_model.layers[-50:]:\n    if isinstance(layer, tf.keras.layers.BatchNormalization):\n        layer.trainable = False\n    else:\n        layer.trainable = True\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=5e-6),  \n    loss=tf.keras.losses.BinaryCrossentropy(),                \n    metrics=[\n        tf.keras.metrics.BinaryAccuracy(name='accuracy'),\n        tf.keras.metrics.AUC(multi_label=True, name='auc')\n    ]\n)\n\ncallbacks = [\n    tf.keras.callbacks.EarlyStopping(\n        monitor='val_auc',\n        patience=3,\n        mode='max',\n        restore_best_weights=True,\n        verbose=1\n    ),\n    tf.keras.callbacks.ReduceLROnPlateau(\n        monitor='val_auc',\n        factor=0.3,\n        patience=2,\n        mode='max',\n        min_lr=1e-7,\n        verbose=1\n    )\n]\n\nprint(\"Fine-tuning model...\")\n\nhistory_ft = model.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=10,\n    callbacks=callbacks\n)","metadata":{"papermill":{"duration":1611.699296,"end_time":"2026-04-10T07:35:11.029987+00:00","exception":false,"start_time":"2026-04-10T07:08:19.330691+00:00","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2026-05-20T07:40:58.768277Z","iopub.execute_input":"2026-05-20T07:40:58.769008Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 11: Make Predictions on Test Set\n\n**Threshold = 0.5** — if probability > 0.5 we predict that class.  \nIf no class passes the threshold, we pick the highest probability class.","metadata":{"papermill":{"duration":0.215915,"end_time":"2026-04-10T07:35:11.465058+00:00","exception":false,"start_time":"2026-04-10T07:35:11.249143+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def tta_predict(model, dataset, tta=5):\n    preds = []\n    for _ in range(tta):\n        preds.append(model.predict(dataset, verbose=0))\n    return np.mean(preds, axis=0)\n\npreds = tta_predict(model, test_ds, tta=5)\n\nTHRESHOLD = 0.5\n\ndef to_label(p):\n    labels = [ALL_LABELS[i] for i, v in enumerate(p) if v >= THRESHOLD]\n    if not labels:\n        labels = [ALL_LABELS[np.argmax(p)]]\n    return \" \".join(labels)\n\npred_labels = [to_label(p) for p in preds]","metadata":{"papermill":{"duration":3.144046,"end_time":"2026-04-10T07:35:14.808539+00:00","exception":false,"start_time":"2026-04-10T07:35:11.664493+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 12: Save submission.csv","metadata":{"papermill":{"duration":0.19449,"end_time":"2026-04-10T07:35:15.204974+00:00","exception":false,"start_time":"2026-04-10T07:35:15.010484+00:00","status":"completed"},"tags":[]}},{"cell_type":"code","source":"submission = pd.DataFrame({\n    \"image\": sample_df[\"image\"],\n    \"labels\": pred_labels\n})\n\nsubmission.to_csv(\"submission.csv\", index=False)\nprint(submission.head())","metadata":{"papermill":{"duration":0.206897,"end_time":"2026-04-10T07:35:15.607657+00:00","exception":false,"start_time":"2026-04-10T07:35:15.400760+00:00","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 13: Submit to Kaggle\n\nThis competition uses **Notebook submission** — you do NOT upload the CSV file manually.\n\n### How to submit\n\n1. Make sure **Internet is OFF**  \n   → Right panel → **Settings** → **Internet** → toggle OFF\n\n2. Click **\"Save & Run All (Commit)\"** — top right button  \n   → This runs the full notebook from top to bottom  \n   → Wait ~15–25 minutes for it to finish\n\n3. Once the commit is done → click **\"Submit to Competition\"**  \n   → Kaggle re-runs your notebook privately on the hidden test set  \n   → It reads `submission.csv` from the output and scores it\n\n4. Check your score at the leaderboard:  \n   [https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/leaderboard](https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8/leaderboard)\n\n---\n\n### Checklist before committing\n\n| Item | Check |\n|---|---|\n| Internet toggle is OFF | ☐ |\n| Weights dataset is attached | ☐ |\n| `weights=None` in model code | ☐ |\n| `load_weights(WEIGHTS_PATH)` is in model code | ☐ |\n| `submission.csv` is saved in the last cell | ☐ |\n| All cells run without errors | ☐ |","metadata":{"papermill":{"duration":0.199164,"end_time":"2026-04-10T07:35:16.094937+00:00","exception":false,"start_time":"2026-04-10T07:35:15.895773+00:00","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Step 14: Share on Canvas\n\n**Make your notebook public:**  \nIn your Kaggle notebook → **Settings** → **Sharing** → set to **Public**\n\n**Submit on Canvas:**\n1. Your Kaggle notebook public URL\n2. Screenshot of the leaderboard showing your username and score\n3. Your public Mean F1 score\n\n---\n\n## How to improve your score\n\n| Idea | Expected gain |\n|---|---|\n| Fine-tune the backbone (Step 10) | +5–10% |\n| Use a larger model (EfficientNetV2B2 or B3) | +3–5% |\n| Tune the threshold (try 0.3, 0.4, 0.5) | +1–3% |\n| Train more epochs | +2–5% |\n| Add stronger augmentation | +1–3% |\n\nBaseline score with this notebook: **~0.75–0.82 Mean F1**","metadata":{"papermill":{"duration":0.199804,"end_time":"2026-04-10T07:35:16.506827+00:00","exception":false,"start_time":"2026-04-10T07:35:16.307023+00:00","status":"completed"},"tags":[]}}]}