{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"b0923aea-02b2-4230-9440-33099bfa44c1","cell_type":"markdown","source":"# 1. Prétraitement et mise en cache des images\n\n## Objectif de ce notebook\n\nCe notebook constitue la **première étape** du pipeline : il prépare une seule fois les images (recadrage de la rétine + amélioration de contraste selon la méthode de Ben Graham, redimensionnement à 512×512) et les sauvegarde sous forme d'archives réutilisables.\n\n## Pourquoi un notebook séparé pour cette étape ?\n\nLe prétraitement de l'ensemble des images (environ 35 000 en entraînement et 53 000 en test) est l'étape la plus coûteuse en temps du pipeline, alors qu'elle **ne dépend d'aucun choix de modèle** : le résultat est strictement identique quel que soit le modèle entraîné ensuite. Il aurait été inefficace de refaire ce travail à chaque nouvelle expérimentation.\n\nEn isolant cette étape dans un notebook dédié dont la sortie (les archives `.zip` des images prétraitées) est réutilisée par tous les notebooks suivants, chaque nouvelle expérimentation (nouveau modèle, nouvelle architecture) peut **ignorer complètement le prétraitement** et démarrer directement l'entraînement — ce qui a permis de réduire le temps de chaque nouvelle itération de plusieurs heures à quelques minutes.\n\n## Contenu\n\n1. Chargement des métadonnées (labels d'entraînement, liste des images de test)\n2. Recadrage de la rétine et amélioration de contraste (Ben Graham) sur toutes les images\n3. Sauvegarde du résultat sous forme d'archives `.zip` réutilisables par les notebooks 2 et 3\n","metadata":{}},{"id":"907722b4-ed4c-4ecc-87dd-41d5988ee557","cell_type":"code","source":"# ============================================================\n# CONFIGURATION\n# ============================================================\n\nimport os\nimport time\nimport zipfile\nimport warnings\nfrom pathlib import Path\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport cv2\nimport numpy as np\nimport pandas as pd\nfrom tqdm.auto import tqdm\n\nwarnings.filterwarnings(\"ignore\")\n\nSEED = 42  # identique a v11_secure.ipynb, ne pas changer\n\nCOMPETITION_DIR = Path(\"/kaggle/input/competitions/diabetic-retinopathy-detection\")\n\nTRAIN_IMAGE_DIR = Path(\n    \"/kaggle/input/datasets/josephrynkiewicz/\"\n    \"diabetic-retinopathy-train-unzipped/train\"\n)\n\nTEST_IMAGE_DIR = Path(\n    \"/kaggle/input/datasets/josephrynkiewicz/\"\n    \"diabetic-retinopathy-test-unzipped/test\"\n)\n\nTRAIN_LABELS_PATH = COMPETITION_DIR / \"trainLabels.csv.zip\"\nSAMPLE_SUBMISSION_PATH = COMPETITION_DIR / \"sampleSubmission.csv.zip\"\n\nassert TRAIN_LABELS_PATH.exists(), TRAIN_LABELS_PATH\nassert TRAIN_IMAGE_DIR.exists(), TRAIN_IMAGE_DIR\nassert TEST_IMAGE_DIR.exists(), TEST_IMAGE_DIR\nassert SAMPLE_SUBMISSION_PATH.exists(), SAMPLE_SUBMISSION_PATH\n\nWORK_DIR = Path(\"/kaggle/working/dr_v11_t4\")\nTRAIN_GRAHAM_DIR = WORK_DIR / \"graham_train_512\"\nTEST_GRAHAM_DIR = WORK_DIR / \"graham_test_512\"\n\nfor d in [WORK_DIR, TRAIN_GRAHAM_DIR, TEST_GRAHAM_DIR]:\n    d.mkdir(parents=True, exist_ok=True)\n\nTRAIN_CACHE_ARCHIVE = Path(\"/kaggle/working/v11_graham_train_512.zip\")\nTEST_CACHE_ARCHIVE = Path(\"/kaggle/working/v11_graham_test_512.zip\")\n\nIMAGE_SIZE_CACHE = 512\nJPEG_QUALITY = 92\n\nprint(\"Chemins OK.\")\nprint(\"CPU disponibles :\", os.cpu_count())\n","metadata":{},"outputs":[],"execution_count":null},{"id":"4059ec93-9bfc-43ba-b0ff-39f3dbebdfeb","cell_type":"markdown","source":"## 1. Chargement des métadonnées (mêmes labels/split que V11)","metadata":{}},{"id":"8946d2a3-5d4f-423e-b784-a91871d4d699","cell_type":"code","source":"# ============================================================\n# CHARGEMENT DES LABELS TRAIN + METADATA TEST\n# ============================================================\n\ndf = pd.read_csv(TRAIN_LABELS_PATH)\ndf[\"patient_id\"] = df[\"image\"].str.rsplit(\"_\", n=1).str[0]\ndf[\"eye\"] = df[\"image\"].str.rsplit(\"_\", n=1).str[1]\n\nprint(\"Images train :\", len(df))\nprint(df[\"level\"].value_counts().sort_index())\n\nsample_sub = pd.read_csv(SAMPLE_SUBMISSION_PATH)\ntest_df = sample_sub[[\"image\"]].copy()\n\nprint(\"\\nImages test :\", len(test_df))\n","metadata":{},"outputs":[],"execution_count":null},{"id":"67b807d3-5f42-4264-bc68-9c0b7810be83","cell_type":"markdown","source":"## 2. Vérification d'un cache déjà existant\n\nAvant de relancer un prétraitement coûteux, on vérifie si une archive de cache a déjà été générée lors d'une exécution précédente et rendue disponible en entrée de ce notebook. Si c'est le cas, l'étape de prétraitement est simplement ignorée.\n","metadata":{}},{"id":"6ee81f8d-b424-4e14-a6c5-38e72a57df51","cell_type":"code","source":"# ============================================================\n# RESTAURER UN CACHE EXISTANT (evite de refaire le travail si deja fait)\n# ============================================================\n\ndef find_attached_archive(filename):\n    matches = list(Path(\"/kaggle/input\").rglob(filename))\n    return matches[0] if matches else None\n\n\ndef restore_cache(archive_name, destination_dir, expected_min_files):\n    destination_dir.mkdir(parents=True, exist_ok=True)\n    current = len(list(destination_dir.glob(\"*.jpg\")))\n\n    if current >= expected_min_files:\n        print(f\"Cache deja present : {current} fichiers\")\n        return True\n\n    archive = find_attached_archive(archive_name)\n    if archive is None:\n        print(f\"{archive_name} non trouve dans /kaggle/input.\")\n        return False\n\n    print(\"Cache trouve :\", archive)\n    start = time.time()\n    with zipfile.ZipFile(archive, \"r\") as zf:\n        zf.extractall(destination_dir)\n\n    count = len(list(destination_dir.glob(\"*.jpg\")))\n    print(f\"Cache restaure : {count} images en {(time.time()-start)/60:.1f} min\")\n    return count >= expected_min_files\n\n\nTRAIN_CACHE_RESTORED = restore_cache(\"v11_graham_train_512.zip\", TRAIN_GRAHAM_DIR, 35000)\nTEST_CACHE_RESTORED = restore_cache(\"v11_graham_test_512.zip\", TEST_GRAHAM_DIR, 53000)\n\nprint(\"\\nTRAIN_CACHE_RESTORED =\", TRAIN_CACHE_RESTORED)\nprint(\"TEST_CACHE_RESTORED  =\", TEST_CACHE_RESTORED)\n","metadata":{},"outputs":[],"execution_count":null},{"id":"2a939f84-8ccf-47bc-9b37-e1281d1e0149","cell_type":"markdown","source":"## 3. Fonctions de prétraitement (crop rétine + Ben Graham, identiques à V11)","metadata":{}},{"id":"26409209-0811-487e-b23d-cbc92057da70","cell_type":"code","source":"# ============================================================\n# FONCTIONS DE PRÉTRAITEMENT\n# ============================================================\n\ndef crop_retina(image):\n    intensity = image.max(axis=2)\n    mask = (intensity > 7).astype(np.uint8)\n\n    if mask.sum() == 0:\n        return image\n\n    kernel = np.ones((5, 5), np.uint8)\n    mask = cv2.morphologyEx(mask, cv2.MORPH_CLOSE, kernel)\n\n    n_labels, labels, stats, _ = cv2.connectedComponentsWithStats(mask, connectivity=8)\n\n    if n_labels <= 1:\n        return image\n\n    component = 1 + np.argmax(stats[1:, cv2.CC_STAT_AREA])\n\n    x = stats[component, cv2.CC_STAT_LEFT]\n    y = stats[component, cv2.CC_STAT_TOP]\n    w = stats[component, cv2.CC_STAT_WIDTH]\n    h = stats[component, cv2.CC_STAT_HEIGHT]\n\n    mx = int(w * 0.015)\n    my = int(h * 0.015)\n\n    x1 = max(0, x - mx)\n    y1 = max(0, y - my)\n    x2 = min(image.shape[1], x + w + mx)\n    y2 = min(image.shape[0], y + h + my)\n\n    crop = image[y1:y2, x1:x2]\n    return crop if crop.size else image\n\n\ndef pad_square(image):\n    h, w = image.shape[:2]\n    side = max(h, w)\n    result = np.zeros((side, side, 3), dtype=np.uint8)\n    top = (side - h) // 2\n    left = (side - w) // 2\n    result[top:top+h, left:left+w] = image\n    return result\n\n\ndef create_natural(image):\n    image = crop_retina(image)\n    image = pad_square(image)\n    return cv2.resize(image, (IMAGE_SIZE_CACHE, IMAGE_SIZE_CACHE), interpolation=cv2.INTER_AREA)\n\n\ndef create_graham(image):\n    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    retina_mask = gray > 5\n\n    blurred = cv2.GaussianBlur(image, (0, 0), sigmaX=IMAGE_SIZE_CACHE / 30)\n    processed = cv2.addWeighted(image, 4.0, blurred, -4.0, 128)\n    processed = np.clip(processed, 0, 255).astype(np.uint8)\n    processed[~retina_mask] = 0\n\n    return processed\n\n\ndef process_train_image(image_id):\n    dst = TRAIN_GRAHAM_DIR / f\"{image_id}.jpg\"\n    if dst.exists():\n        return True\n\n    src = TRAIN_IMAGE_DIR / f\"{image_id}.jpeg\"\n    image = cv2.imread(str(src), cv2.IMREAD_COLOR)\n    if image is None:\n        return False\n\n    graham = create_graham(create_natural(image))\n    return bool(cv2.imwrite(str(dst), graham, [cv2.IMWRITE_JPEG_QUALITY, JPEG_QUALITY]))\n\n\ndef process_test_image(image_id):\n    dst = TEST_GRAHAM_DIR / f\"{image_id}.jpg\"\n    if dst.exists():\n        return True\n\n    src = TEST_IMAGE_DIR / f\"{image_id}.jpeg\"\n    image = cv2.imread(str(src), cv2.IMREAD_COLOR)\n    if image is None:\n        return False\n\n    graham = create_graham(create_natural(image))\n    return bool(cv2.imwrite(str(dst), graham, [cv2.IMWRITE_JPEG_QUALITY, JPEG_QUALITY]))\n","metadata":{},"outputs":[],"execution_count":null},{"id":"636c5100-754c-4c45-9a09-0d586a107c97","cell_type":"markdown","source":"## 4. Prétraitement TRAIN","metadata":{}},{"id":"0eec17f8-1215-40e3-be29-caabfc3ed640","cell_type":"code","source":"# ============================================================\n# PRÉTRAITEMENT TRAIN\n# ============================================================\n\nworkers = min(8, os.cpu_count() or 4)\n\nif TRAIN_CACHE_RESTORED:\n    print(\"Cache train deja restaure : preprocessing ignore.\")\nelse:\n    print(\"Creation cache TRAIN 512 - workers :\", workers)\n    start = time.time()\n\n    with ThreadPoolExecutor(max_workers=workers) as executor:\n        results = list(tqdm(\n            executor.map(process_train_image, df[\"image\"].astype(str).tolist()),\n            total=len(df),\n            desc=\"Graham train\"\n        ))\n\n    print(f\"Cache train termine en {(time.time()-start)/60:.1f} min\")\n    print(\"Images cache :\", len(list(TRAIN_GRAHAM_DIR.glob(\"*.jpg\"))))\n\n    if not all(results):\n        n_failed = sum(1 for r in results if not r)\n        print(f\"Attention : {n_failed} images n'ont pas pu etre traitees.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"7f95ea78-f02a-4463-9b09-f2845ae3e513","cell_type":"markdown","source":"## 5. Prétraitement TEST","metadata":{}},{"id":"d875f21f-a1c1-4544-a374-672915c6911c","cell_type":"code","source":"# ============================================================\n# PRÉTRAITEMENT TEST\n# ============================================================\n\nif TEST_CACHE_RESTORED:\n    print(\"Cache test deja restaure : preprocessing ignore.\")\nelse:\n    print(\"Creation cache TEST 512 - workers :\", workers)\n    start = time.time()\n\n    with ThreadPoolExecutor(max_workers=workers) as executor:\n        test_results = list(tqdm(\n            executor.map(process_test_image, test_df[\"image\"].astype(str).tolist()),\n            total=len(test_df),\n            desc=\"Graham test\"\n        ))\n\n    print(f\"Cache test termine en {(time.time()-start)/60:.1f} min\")\n    print(\"Images cache test :\", len(list(TEST_GRAHAM_DIR.glob(\"*.jpg\"))))\n\n    if not all(test_results):\n        n_failed = sum(1 for r in test_results if not r)\n        print(f\"Attention : {n_failed} images test n'ont pas pu etre traitees.\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"62bac072-9a6c-4908-86e6-fb483dc09139","cell_type":"markdown","source":"## 6. Création des archives .zip réutilisables","metadata":{}},{"id":"433bff41-c6e3-4674-93a9-7a1658c5a7d8","cell_type":"code","source":"# ============================================================\n# CRÉATION DES ARCHIVES ZIP\n# ============================================================\n\ndef create_store_zip(source_dir, archive_path, label):\n    files = sorted(source_dir.glob(\"*.jpg\"))\n\n    if not files:\n        raise RuntimeError(f\"Aucun fichier dans {source_dir}\")\n\n    if archive_path.exists():\n        print(f\"Archive {label} deja presente :\", archive_path)\n        return\n\n    print(f\"Creation de {archive_path.name} avec {len(files)} images ({label})...\")\n    start = time.time()\n\n    with zipfile.ZipFile(archive_path, \"w\", compression=zipfile.ZIP_STORED, allowZip64=True) as zf:\n        for path in tqdm(files, desc=f\"Archive {label}\"):\n            zf.write(path, arcname=path.name)\n\n    size_gb = archive_path.stat().st_size / 1024**3\n    print(f\"Archive {label} creee : {size_gb:.2f} GB en {(time.time()-start)/60:.1f} min\")\n\n\ncreate_store_zip(TRAIN_GRAHAM_DIR, TRAIN_CACHE_ARCHIVE, \"train\")\ncreate_store_zip(TEST_GRAHAM_DIR, TEST_CACHE_ARCHIVE, \"test\")\n","metadata":{},"outputs":[],"execution_count":null},{"id":"5142849c-9778-4357-89d7-981237274753","cell_type":"markdown","source":"## 7. Résumé\n\nVérification que les deux archives de cache (train et test) ont bien été générées avec succès. Ces deux fichiers constituent la sortie de ce notebook et sont réutilisés directement par les notebooks suivants (`02_modele_base_efficientnet_b4` et `03_ensembling_efficientnet_b5`), qui n'ont donc pas besoin de refaire ce prétraitement.\n","metadata":{}},{"id":"968decae-1db9-4283-866c-c9ba896da755","cell_type":"code","source":"required_files = {\n    \"Archive cache TRAIN\": TRAIN_CACHE_ARCHIVE,\n    \"Archive cache TEST\": TEST_CACHE_ARCHIVE,\n}\n\nprint(\"=\" * 60)\nprint(\"RESUME\")\nprint(\"=\" * 60)\n\nfor label, path in required_files.items():\n    ok = Path(path).exists()\n    status = \"OK\" if ok else \"MANQUANT\"\n    size = f\"({Path(path).stat().st_size / 1024**3:.2f} GB)\" if ok else \"\"\n    print(f\"[{status:8}] {label:25} {path} {size}\")\n","metadata":{},"outputs":[],"execution_count":null}]}