{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":13836,"databundleVersionId":1718836,"sourceType":"competition"}],"dockerImageVersionId":31154,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🧠 Deep Data Discovery & Research Mapping — Cassava Leaf Disease\n**Week 01 – Discover & Map | Author: MD. tanvir Arif Siddiqui**\n\n**Objective.** Forensically inspect the dataset, map risks (imbalance, duplicates, leakage, label noise), and define a clear PRD + 6–10 research questions to guide modeling.\n\n**Outcome.** 1‑page PRD • research questions • visual + textual data map • robust CV split plan.\n\n**Data.** Kaggle Competition: Cassava Leaf Disease Classification (Makerere AI Lab / NaCRRI)  \nFacts: 21,367 labeled images • 5 classes • field images • potential label noise.  \n","metadata":{}},{"cell_type":"code","source":"# Core\nimport os, sys, json, glob, math, random, textwrap, zipfile, gc\nfrom pathlib import Path\nfrom collections import Counter, defaultdict\n\n# Data\nimport numpy as np\nimport pandas as pd\n\n# Imaging\nfrom PIL import Image, ImageStat, ImageOps\nimport cv2\n\n# Viz\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# ML helpers\nfrom sklearn.model_selection import StratifiedKFold, GroupKFold\nfrom sklearn.metrics import classification_report, confusion_matrix\n\n# Pretty plotting defaults\nplt.rcParams[\"figure.dpi\"] = 120\nsns.set_style(\"whitegrid\")\n\n# Paths (Kaggle)\nINPUT = Path(\"/kaggle/input\")\nWORK  = Path(\"/kaggle/working\")\nROOTS = [p for p in INPUT.glob(\"*cassava*\")] or [INPUT]  # be resilient to dataset aliasing\nROOTS\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-10-09T23:19:49.996113Z","iopub.execute_input":"2025-10-09T23:19:49.996325Z","iopub.status.idle":"2025-10-09T23:19:53.228143Z","shell.execute_reply.started":"2025-10-09T23:19:49.996308Z","shell.execute_reply":"2025-10-09T23:19:53.227235Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def find_first(*patterns):\n    for base in ROOTS:\n        for pat in patterns:\n            hits = list(base.glob(pat))\n            if hits:\n                return hits[0]\n    return None\n\n# Unzip helper (if images are packed)\ndef unzip_to_work(zpath: Path, outdir: Path):\n    outdir.mkdir(parents=True, exist_ok=True)\n    with zipfile.ZipFile(zpath, \"r\") as zf:\n        zf.extractall(outdir)\n\n# Locate canonical files/folders (covers both zipped/unzipped layouts)\ncsv_path  = find_first(\"**/train.csv\")\njson_path = find_first(\"**/label_num_to_disease_map.json\")\ntrain_dir = find_first(\"**/train_images\")  # typical competition layout\ntest_dir  = find_first(\"**/test_images\")\n\n# If zipped images, unzip to /kaggle/working\nif train_dir is None:\n    z_train = find_first(\"**/train_images.zip\")\n    if z_train: \n        train_dir = WORK/\"train_images\"\n        unzip_to_work(z_train, WORK)\nif test_dir is None:\n    z_test = find_first(\"**/test_images.zip\")\n    if z_test:\n        test_dir = WORK/\"test_images\"\n        unzip_to_work(z_test, WORK)\n\nprint(\"csv:\", csv_path)\nprint(\"json:\", json_path)\nprint(\"train_dir:\", train_dir)\nprint(\"test_dir:\", test_dir)\n\n# Load metadata\ndf = pd.read_csv(csv_path)\nwith open(json_path) as f:\n    label_map = json.load(f)\n\ninv_label_map = {v:k for k,v in label_map.items()}  # names -> ids\ndf[\"label_name\"] = df[\"label\"].map(label_map)\n\ndisplay(df.head())\nprint(\"n_train:\", len(df), \"n_classes:\", df['label'].nunique(), \"label_map:\", label_map)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-09T23:19:53.228857Z","iopub.execute_input":"2025-10-09T23:19:53.229234Z","iopub.status.idle":"2025-10-09T23:20:56.307926Z","shell.execute_reply.started":"2025-10-09T23:19:53.229207Z","shell.execute_reply":"2025-10-09T23:20:56.307229Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Ensure label_name exists robustly\n# (handles label_map keys as str or int)\nif list(label_map.keys()) and isinstance(next(iter(label_map.keys())), str):\n    df[\"label_name\"] = df[\"label\"].map(lambda i: label_map[str(i)])\nelse:\n    df[\"label_name\"] = df[\"label\"].map(label_map)\n\n# Basic sanity\nassert train_dir is not None and train_dir.exists(), \"Missing train_images\"\nmissing_files = [im for im in df.image_id if not (train_dir/im).exists()]\nprint(\"Missing files:\", len(missing_files))\n\n# Class balance — FIXED\ncls_counts = df['label_name'].value_counts()\ncls_counts = cls_counts.sort_index(key=lambda idx: idx.map(inv_label_map))  # <— use idx, not idx.index\n\nplt.figure(figsize=(7,3))\nsns.barplot(x=cls_counts.index, y=cls_counts.values)\nplt.title(\"Class Distribution (train)\")\nplt.xticks(rotation=20)\nplt.ylabel(\"count\"); plt.xlabel(\"\")\nplt.show()\n\nfor k, v in cls_counts.items():\n    print(f\"{k:30s} : {v:5d} ({100*v/len(df):.1f}%)\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-09T23:20:56.308686Z","iopub.execute_input":"2025-10-09T23:20:56.308924Z","iopub.status.idle":"2025-10-09T23:20:56.814657Z","shell.execute_reply.started":"2025-10-09T23:20:56.308896Z","shell.execute_reply":"2025-10-09T23:20:56.813989Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def show_grid(rows=2, cols=5):\n    fig, axes = plt.subplots(rows, cols, figsize=(cols*2.4, rows*2.4))\n    axes = axes.flatten()\n    for i, cls in enumerate(sorted(label_map.values(), key=lambda x: inv_label_map[x])[:rows*cols]):\n        sample = df[df.label_name==cls].sample(1).iloc[0]\n        img = Image.open(train_dir/sample.image_id).convert(\"RGB\")\n        axes[i].imshow(img); axes[i].set_title(cls, fontsize=9); axes[i].axis(\"off\")\n    plt.tight_layout(); plt.show()\n\nshow_grid()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-09T23:20:56.816331Z","iopub.execute_input":"2025-10-09T23:20:56.816908Z","iopub.status.idle":"2025-10-09T23:20:58.415578Z","shell.execute_reply.started":"2025-10-09T23:20:56.816880Z","shell.execute_reply":"2025-10-09T23:20:58.414664Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def image_stats(path: Path):\n    # returns dict of width, height, size_kb, brightness, blur, aspect\n    with Image.open(path) as im:\n        w, h = im.size\n        gray = ImageOps.grayscale(im)\n        stat = ImageStat.Stat(gray)\n        brightness = stat.mean[0]  # 0..255\n        # Blur via Laplacian variance\n        gray_cv = np.array(gray)\n        blur = cv2.Laplacian(gray_cv, cv2.CV_64F).var()\n    size_kb = path.stat().st_size/1024\n    aspect = w/h\n    return dict(w=w, h=h, size_kb=size_kb, brightness=brightness, blur=blur, aspect=aspect)\n\n# Sample to keep runtime reasonable (use all if you want full audit)\nSAMPLE_N = min(4000, len(df))\nsample_ids = df.sample(SAMPLE_N, random_state=42).image_id.values\n\nstats = []\nfor iid in sample_ids:\n    stats.append(image_stats(train_dir/iid))\nstats_df = pd.DataFrame(stats)\n\nfig, ax = plt.subplots(1,3, figsize=(12,3))\nsns.histplot(stats_df['w'], ax=ax[0]); ax[0].set_title(\"Width\")\nsns.histplot(stats_df['h'], ax=ax[1]); ax[1].set_title(\"Height\")\nsns.histplot(stats_df['brightness'], ax=ax[2]); ax[2].set_title(\"Brightness\")\nplt.show()\n\nplt.figure(figsize=(4,3))\nsns.histplot(stats_df['blur'])\nplt.title(\"Laplacian Variance (Blur ~ low = blurrier)\")\nplt.show()\n\nstats_df.describe().T\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-10T00:21:48.123118Z","iopub.execute_input":"2025-10-10T00:21:48.123826Z","iopub.status.idle":"2025-10-10T00:22:18.657918Z","shell.execute_reply.started":"2025-10-10T00:21:48.123803Z","shell.execute_reply":"2025-10-10T00:22:18.657344Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from collections import Counter\n\ndef get_exif_model(path: Path):\n    try:\n        with Image.open(path) as im:\n            ex = im.getexif()\n            # 272 (0x110) = Model; 271 = Make\n            model = ex.get(272, None)\n            make  = ex.get(271, None)\n            return f\"{make} {model}\".strip() if (make or model) else None\n    except Exception:\n        return None\n\ndevice_counts = Counter()\nfor iid in random.sample(list(sample_ids), min(1000, len(sample_ids))):\n    device = get_exif_model(train_dir/iid)\n    if device: device_counts[device]+=1\n\ndevice_counts.most_common(10)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-10T00:22:23.666821Z","iopub.execute_input":"2025-10-10T00:22:23.667441Z","iopub.status.idle":"2025-10-10T00:22:24.254896Z","shell.execute_reply.started":"2025-10-10T00:22:23.667415Z","shell.execute_reply":"2025-10-10T00:22:24.254273Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Lightweight pHash (no external deps)\ndef phash(im: Image.Image, hash_size=8, highfreq_fact=4):\n    # https://www.hackerfactor.com/blog/index.php?/archives/432-Looks-Like-It.html\n    img = ImageOps.grayscale(im).resize((hash_size*highfreq_fact, hash_size*highfreq_fact), Image.Resampling.LANCZOS)\n    pixels = np.asarray(img, dtype=np.float32)\n    dct = cv2.dct(pixels)\n    dctlow = dct[:hash_size, :hash_size]\n    med = np.median(dctlow)\n    diff = dctlow > med\n    return \"\".join(\"1\" if v else \"0\" for v in diff.flatten())\n\ndef hamming(a,b): return sum(c1!=c2 for c1,c2 in zip(a,b))\n\n# Compute hashes for a (stratified) sample\nN_HASH = min(5000, len(df))\nids_to_hash = df.groupby(\"label_name\", group_keys=False).apply(lambda g: g.sample(min(len(g), math.ceil(N_HASH/df['label_name'].nunique())), random_state=0)).image_id.tolist()\n\nhashes = {}\nfor iid in ids_to_hash:\n    with Image.open(train_dir/iid) as im:\n        hashes[iid] = phash(im)\n\n# Find near-duplicates (Hamming distance <= 5 as heuristic)\npairs = []\nids = list(hashes.keys())\nfor i in range(len(ids)):\n    for j in range(i+1, len(ids)):\n        d = hamming(hashes[ids[i]], hashes[ids[j]])\n        if d <= 5:\n            pairs.append((ids[i], ids[j], d))\nlen(pairs), pairs[:5]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-10T00:22:41.174156Z","iopub.execute_input":"2025-10-10T00:22:41.174700Z","iopub.status.idle":"2025-10-10T00:24:34.315479Z","shell.execute_reply.started":"2025-10-10T00:22:41.174679Z","shell.execute_reply":"2025-10-10T00:24:34.314649Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Build grouping by hash (exact or near-dup clusters). For simplicity, use exact hash groups here.\nhash_to_group = {}\nfor iid in df.image_id:\n    hp = None\n    try:\n        with Image.open(train_dir/iid) as im:\n            hp = phash(im)\n    except:\n        pass\n    hash_to_group[iid] = hp or iid  # fallback: unique\n\ndf[\"group\"] = df[\"image_id\"].map(hash_to_group)\n\n# GroupKFold split blueprint (store folds to CSV for modeling notebook)\ngkf = GroupKFold(n_splits=5)\ndf[\"fold\"] = -1\nfor fold, (tr, va) in enumerate(gkf.split(df, y=df[\"label\"], groups=df[\"group\"])):\n    df.loc[df.index[va], \"fold\"] = fold\n\ndf.fold.value_counts().sort_index(), df.head()\ndf.to_csv(WORK/\"folds_groupkfold.csv\", index=False)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-10T00:25:23.658035Z","iopub.execute_input":"2025-10-10T00:25:23.658741Z","iopub.status.idle":"2025-10-10T00:28:58.969086Z","shell.execute_reply.started":"2025-10-10T00:25:23.658714Z","shell.execute_reply":"2025-10-10T00:28:58.968438Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🗺️ Data Map (features → outcomes)\n\n**Inputs.** RGB pixel data from `train_images/*.jpg` (field photos; varying resolution/lighting).  \n**Target.** Multi‑class label in `train.csv` (`0..4`) mapped via `label_num_to_disease_map.json` →  \n0: CBB, 1: CBSD, 2: CGM, 3: CMD, 4: Healthy. :contentReference[oaicite:5]{index=5}\n\n**Derived signals (for EDA only).**  \n- Image stats: width/height, aspect, file size, brightness, blur (Laplacian var).  \n- EXIF: camera make/model (if present) → device bias check.  \n- pHash clusters: duplicates/near‑duplicates → leakage risk, grouping for CV.\n\n**Potential leakage/bias.**  \n- Background or device artifacts correlated with labels (mitigate via leaf‑centric crops/augmentations).  \n- Duplicate shots (same leaf) across folds → group‑aware CV.\n","metadata":{"execution":{"iopub.status.busy":"2025-10-09T23:21:52.876192Z","iopub.execute_input":"2025-10-09T23:21:52.876910Z","iopub.status.idle":"2025-10-09T23:22:07.644471Z","shell.execute_reply.started":"2025-10-09T23:21:52.876892Z","shell.execute_reply":"2025-10-09T23:22:07.643746Z"}}},{"cell_type":"markdown","source":"## 📄 1‑Page Product Requirement Document (PRD)\n\n**Title.** Cassava Leaf Disease Detector — Field‑Robust Image Classifier\n\n**Objective.** Classify a cassava leaf photo into one of **5 classes** (4 diseases + healthy) with **field‑robust** performance and leakage‑resistant validation. :contentReference[oaicite:6]{index=6}\n\n**User / Value.** Smallholder farmers & agronomists get **fast triage** without needing experts on‑site; earlier detection reduces losses and improves food security.\n\n**Scope.**  \n- **In:** Forensic EDA (integrity, imbalance, duplicates), leakage‑resistant CV, augmentation plan, baseline CNN plan.  \n- **Out:** Mobile app build, geographic spread modeling, treatment recommendation.\n\n**Success Metrics.**  \n- **Primary:** Macro‑F1 ≥ **0.80** on leakage‑resistant CV (GroupKFold by duplicate clusters).  \n- **Secondary:** Per‑class recall ≥ 0.75 (esp. minority classes); stable performance across brightness/blur strata; runtime ≤ 1s/image on GPU.  \n- **Aspirational:** Accuracy ≥ **0.90** based on strong community solutions. :contentReference[oaicite:7]{index=7}\n\n**Constraints.**  \n- **Data:** Imbalanced labels; possible label noise; varied lighting/blur; device bias (EXIF). :contentReference[oaicite:8]{index=8}  \n- **Technical:** Kaggle GPU/TPU; no internet; storage/time limits; heavy image IO.  \n- **Ethical:** Communicate uncertainty; avoid false reassurance on severe diseases; be transparent about limitations.\n\n**Deliverables.**  \n1) Forensic EDA report & data map  2) CV split file (`folds_groupkfold.csv`)  \n3) Augmentation preview  4) Baseline model plan  5) Research Qs  6) Publishable notebook\n","metadata":{"execution":{"iopub.status.busy":"2025-10-09T23:22:07.645336Z","iopub.execute_input":"2025-10-09T23:22:07.645610Z","iopub.status.idle":"2025-10-09T23:22:07.893474Z","shell.execute_reply.started":"2025-10-09T23:22:07.645586Z","shell.execute_reply":"2025-10-09T23:22:07.892635Z"}}},{"cell_type":"markdown","source":"## 🧩 Research Questions\n\n### Descriptive\n1) How imbalanced are the 5 classes (counts & %), and how does imbalance vary across devices/brightness/blur strata?  \n2) What is the distribution of resolutions, brightness, and blur, and do these differ by class?\n\n### Diagnostic\n3) Are there **duplicates/near‑duplicates** within/between classes? How would they bias naive splits (accuracy inflation)?  \n4) Do EXIF camera models correlate with specific labels → risk of spurious device cues?\n\n### Predictive\n5) Which **augmentations** (flip/rotate/jitter/cutout) most improve minority‑class recall without hurting others?  \n6) Does **leaf‑centric cropping** (HSV green mask + largest contour) improve signal over raw full‑frame images?  \n7) Do models trained on **brightness/blur‑balanced** batches generalize better than naive sampling?\n\n### Prescriptive\n8) What **CV strategy** gives the most stable estimates (StratifiedKFold vs **GroupKFold by pHash clusters**)?  \n9) What is a **safe operating threshold** per class to minimize dangerous false negatives (e.g., CMD/CBSD)?  \n10) If label noise exists, does **co‑teaching or label smoothing** improve robustness relative to standard CE loss?\n","metadata":{}}]}