{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","isGpuEnabled":true,"isInternetEnabled":true,"language":"python","sourceType":"notebook"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"166382ba-fcbc-49f9-ad60-fff3dbc4e031","cell_type":"markdown","source":"# TP02 - Classification MGMT (RSNA-MICCAI Brain Tumor)\n\n**Principales ameliorations vs version initiale :**\n1. Transfer learning (EfficientNetB0 ImageNet) au lieu d'un CNN 3D from scratch\n2. Fusion multi-modalite reelle (FLAIR + T1wCE + T2w) au lieu de dupliquer FLAIR\n3. Selection intelligente des coupes (on ecarte les coupes quasi-noires)\n4. Data augmentation 3D coherente (meme transfo appliquee a toutes les coupes)\n5. Regularisation L2 + dropout + label smoothing\n6. Validation croisee stratifiee (5-fold)\n7. Cosine decay avec warmup pour le learning rate\n8. Cache disque (.npy) des donnees pretraitees","metadata":{}},{"id":"9c36ca2a-2d8a-4b20-ad9b-5df22c7dadd7","cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport cv2\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models, regularizers\nfrom sklearn.model_selection import StratifiedKFold, train_test_split\nfrom sklearn.metrics import classification_report, confusion_matrix, roc_curve, auc\nimport matplotlib.pyplot as plt\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:12:36.971109Z","iopub.execute_input":"2026-07-10T19:12:36.971885Z","iopub.status.idle":"2026-07-10T19:12:43.064712Z","shell.execute_reply.started":"2026-07-10T19:12:36.971850Z","shell.execute_reply":"2026-07-10T19:12:43.063948Z"}},"outputs":[],"execution_count":null},{"id":"cf3fdd3a-795d-4ee6-be8d-7053b6e2dbc8","cell_type":"markdown","source":"## Configuration GPU\nSur Kaggle : dans le menu de droite du notebook, **Settings > Accelerator > GPU T4 x2** (ou P100/A100 selon dispo) doit etre active AVANT d'executer le notebook. Sans ca, TensorFlow tournera sur CPU meme si le code ci-dessous est correct.","metadata":{}},{"id":"44d9a4b4-5851-406c-a30e-97e721c9cb42","cell_type":"code","source":"# Verification et configuration du GPU\ngpus = tf.config.list_physical_devices('GPU')\nprint(f\"GPU(s) detecte(s) : {len(gpus)}\")\nfor gpu in gpus:\n    print(f\"  - {gpu}\")\n\nif gpus:\n    try:\n        # Memory growth : evite que TF reserve toute la VRAM d'un coup\n        for gpu in gpus:\n            tf.config.experimental.set_memory_growth(gpu, True)\n\n        # Mixed precision (float16) : accelere significativement l'entrainement\n        # sur GPU (Tensor Cores) tout en gardant la stabilite numerique via\n        # les couches de sortie en float32\n        tf.keras.mixed_precision.set_global_policy('mixed_float16')\n        print(\"Mixed precision (float16) activee.\")\n        print(\"Politique de calcul :\", tf.keras.mixed_precision.global_policy())\n    except RuntimeError as e:\n        print(f\"Impossible de configurer le GPU : {e}\")\nelse:\n    print(\"!! Aucun GPU detecte -> le notebook va tourner sur CPU (tres lent).\")\n    print(\"!! Sur Kaggle : Settings (panneau de droite) > Accelerator > GPU T4 x2\")\n\n# Strategie de distribution : gere automatiquement le cas mono-GPU ET\n# multi-GPU (Kaggle propose souvent 2x T4). MirroredStrategy repartit le\n# batch entre les GPU disponibles ; se comporte comme un strategy \"no-op\"\n# s'il n'y a qu'un seul GPU (ou le CPU).\nif len(gpus) > 1:\n    strategy = tf.distribute.MirroredStrategy()\n    print(f\"MirroredStrategy activee sur {strategy.num_replicas_in_sync} GPU(s).\")\nelse:\n    strategy = tf.distribute.get_strategy()  # strategie par defaut (1 device)\n    print(\"Strategie mono-device (1 GPU ou CPU).\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:12:48.847167Z","iopub.execute_input":"2026-07-10T19:12:48.847944Z","iopub.status.idle":"2026-07-10T19:12:49.814067Z","shell.execute_reply.started":"2026-07-10T19:12:48.847916Z","shell.execute_reply":"2026-07-10T19:12:49.813282Z"}},"outputs":[],"execution_count":null},{"id":"847f39da-0feb-454c-9d97-c158cb85ccf0","cell_type":"markdown","source":"## Configuration","metadata":{}},{"id":"e8ec68bd-e9ad-497a-a55e-7390644a276d","cell_type":"code","source":"PATH = '/kaggle/input/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/'\nCACHE_DIR = '/kaggle/working/cache/'\nos.makedirs(CACHE_DIR, exist_ok=True)\n\nIMG_SIZE = 128\nNUM_SLICES = 16\nMODALITIES = ['FLAIR', 'T1wCE', 'T2w']   # 3 modalites -> 3 canaux reels\nBATCH_SIZE = 8  # plus eleve que 8 pour profiter du GPU ; reduire si erreur \"out of memory\"\nEPOCHS = 40\nN_FOLDS = 2\nSEED = 42\n\ntf.random.set_seed(SEED)\nnp.random.seed(SEED)\n\ntrain_labels = pd.read_csv(PATH + 'train_labels.csv')\n# BraTS21ID = 109 est connue pour etre corrompue dans ce dataset -> on l'exclut\ntrain_labels = train_labels[train_labels['BraTS21ID'] != 109].reset_index(drop=True)\n\nprint(f\"Nombre de patients : {len(train_labels)}\")\nprint(train_labels['MGMT_value'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:13.488183Z","iopub.execute_input":"2026-07-10T19:13:13.488989Z","iopub.status.idle":"2026-07-10T19:13:13.507040Z","shell.execute_reply.started":"2026-07-10T19:13:13.488954Z","shell.execute_reply":"2026-07-10T19:13:13.505932Z"}},"outputs":[],"execution_count":null},{"id":"d415a2be-c262-4668-95da-69acd17a35ea","cell_type":"markdown","source":"## Pretraitement DICOM \n\n- La version precedente appelait le GPU **une fois par image individuelle** -> des milliers de petits appels, chacun paye ~quelques ms de latence de lancement de kernel (overhead), qui domine largement le temps de calcul reel sur une image 128x128.\n- Ici : (1) lecture + resize CPU **parallelises** sur plusieurs threads (le decodage DICOM est I/O-bound, les threads se chevauchent bien), (2) normalisation par percentile faite en **UN SEUL appel GPU vectorise sur tout le lot** (tout un patient, ou meme tout le dataset). Un seul appel GPU sur 10 000 images est bien plus rapide que 10 000 appels GPU sur 1 image.","metadata":{}},{"id":"c9839b8a-d34e-430f-806b-da95cf25d3f1","cell_type":"code","source":"from concurrent.futures import ThreadPoolExecutor\n\ndef load_dicom_raw(filepath, img_size=IMG_SIZE):\n    \"\"\"CPU uniquement : lecture DICOM + resize brut (rapide, pas de normalisation ici).\"\"\"\n    dcm = pydicom.dcmread(filepath)\n    raw = dcm.pixel_array.astype(np.float32)\n    raw = cv2.resize(raw, (img_size, img_size), interpolation=cv2.INTER_AREA)\n    return raw\n\n\ndef gpu_batch_normalize(images_array):\n    \"\"\"Normalisation percentile de TOUT un lot d'images en un seul appel GPU\n    vectorise (axe batch), au lieu d'un appel par image. C'est ce qui apporte\n    le vrai gain de vitesse : un seul lancement de kernel GPU pour N images.\"\"\"\n    device = '/GPU:0' if tf.config.list_physical_devices('GPU') else '/CPU:0'\n    with tf.device(device):\n        images = tf.convert_to_tensor(images_array, dtype=tf.float32)  # (N, H, W)\n        n = tf.shape(images)[0]\n        flat = tf.reshape(images, [n, -1])\n        sorted_flat = tf.sort(flat, axis=1)\n        length = tf.shape(sorted_flat)[1]\n        idx1 = tf.cast(tf.cast(length, tf.float32) * 0.01, tf.int32)\n        idx99 = tf.cast(tf.cast(length, tf.float32) * 0.99, tf.int32)\n        p1 = tf.gather(sorted_flat, idx1, axis=1)[:, None, None]\n        p99 = tf.gather(sorted_flat, idx99, axis=1)[:, None, None]\n        denom = tf.maximum(p99 - p1, 1e-6)\n        out = tf.clip_by_value((images - p1) / denom, 0.0, 1.0)\n    return out.numpy()\n\n\ndef select_informative_slices(dicom_files, patient_path, num_slices=NUM_SLICES):\n    \"\"\"Choisit les coupes qui contiennent effectivement du contenu cerebral\n    (on ecarte les coupes quasi-vides en debut/fin de volume), au lieu de\n    prendre un echantillonnage uniforme naif.\"\"\"\n    scores = []\n    for f in dicom_files:\n        try:\n            dcm = pydicom.dcmread(os.path.join(patient_path, f))\n            arr = dcm.pixel_array.astype(np.float32)\n            scores.append(np.mean(arr > np.percentile(arr, 5) + 1e-3))\n        except Exception:\n            scores.append(0.0)\n\n    order = np.argsort(scores)[::-1]\n    top_idx = sorted(order[:min(len(order), num_slices * 3)])\n    if len(top_idx) >= num_slices:\n        step = max(1, len(top_idx) // num_slices)\n        chosen = [top_idx[i] for i in range(0, len(top_idx), step)][:num_slices]\n    else:\n        chosen = top_idx\n\n    return [dicom_files[i] for i in chosen]\n\n\ndef load_patient_volume(patient_id, modalities=MODALITIES, num_slices=NUM_SLICES,\n                        io_workers=8):\n    \"\"\"Charge un volume multi-modalite pour un patient :\n    shape finale = (num_slices, IMG_SIZE, IMG_SIZE, len(modalities))\n\n    Version rapide : lecture CPU parallelisee (ThreadPoolExecutor) + UNE SEULE\n    normalisation GPU vectorisee sur toutes les coupes/modalites du patient.\"\"\"\n    folder = str(patient_id).zfill(5)\n\n    all_paths = []           # tous les chemins de fichiers a lire, toutes modalites confondues\n    modality_slice_counts = []  # pour savoir ou re-decouper apres normalisation groupee\n\n    for modality in modalities:\n        mod_path = os.path.join(PATH, 'train', folder, modality)\n        if not os.path.exists(mod_path):\n            return None\n\n        dicom_files = sorted([f for f in os.listdir(mod_path) if f.endswith('.dcm')])\n        if not dicom_files:\n            return None\n\n        selected = select_informative_slices(dicom_files, mod_path, num_slices)\n        paths = [os.path.join(mod_path, f) for f in selected]\n        all_paths.extend(paths)\n        modality_slice_counts.append(len(paths))\n\n    # Lecture + resize CPU en parallele (I/O-bound -> les threads se chevauchent bien)\n    with ThreadPoolExecutor(max_workers=io_workers) as executor:\n        raw_slices = list(executor.map(load_dicom_raw, all_paths))\n\n    # Un seul appel GPU pour normaliser TOUTES les coupes de TOUTES les modalites\n    # de ce patient d'un coup (au lieu d'un appel par coupe)\n    raw_stack = np.stack(raw_slices, axis=0)          # (total_slices, H, W)\n    normalized_stack = gpu_batch_normalize(raw_stack)  # (total_slices, H, W)\n\n    # On redecoupe par modalite et on complete si necessaire (padding par repetition\n    # de la derniere coupe, comme dans la version originale)\n    modality_stacks = []\n    cursor = 0\n    for count in modality_slice_counts:\n        mod_slices = list(normalized_stack[cursor:cursor + count])\n        cursor += count\n        while len(mod_slices) < num_slices:\n            mod_slices.append(mod_slices[-1].copy() if mod_slices\n                               else np.zeros((IMG_SIZE, IMG_SIZE), dtype=np.float32))\n        modality_stacks.append(np.array(mod_slices[:num_slices]))\n\n    volume = np.stack(modality_stacks, axis=-1)  # (num_slices, H, W, n_modalites)\n    return volume","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:17.194915Z","iopub.execute_input":"2026-07-10T19:13:17.195881Z","iopub.status.idle":"2026-07-10T19:13:17.211863Z","shell.execute_reply.started":"2026-07-10T19:13:17.195847Z","shell.execute_reply":"2026-07-10T19:13:17.210900Z"}},"outputs":[],"execution_count":null},{"id":"4d35ecc2-4e94-49bf-88c7-e6f5f945e2ad","cell_type":"markdown","source":"## Chargement avec cache disque\nEvite de tout retraiter a chaque run.","metadata":{}},{"id":"55421c84-1df7-42ae-98a6-70780d1f44ae","cell_type":"code","source":"def build_dataset(train_labels):\n    cache_x = os.path.join(CACHE_DIR, 'X.npy')\n    cache_y = os.path.join(CACHE_DIR, 'y.npy')\n\n    if os.path.exists(cache_x) and os.path.exists(cache_y):\n        print(\"Chargement depuis le cache...\")\n        return np.load(cache_x), np.load(cache_y)\n\n    X, y = [], []\n    print(\"Pretraitement des volumes (peut prendre du temps la premiere fois)...\")\n    for idx, row in train_labels.iterrows():\n        vol = load_patient_volume(row['BraTS21ID'])\n        if vol is not None:\n            X.append(vol)\n            y.append(row['MGMT_value'])\n        if idx % 50 == 0:\n            print(f\"  {idx}/{len(train_labels)} patients traites\")\n\n    X = np.array(X, dtype=np.float32)\n    y = np.array(y, dtype=np.float32)\n\n    np.save(cache_x, X)\n    np.save(cache_y, y)\n    return X, y\n\n\nX, y = build_dataset(train_labels)\nprint(f\"\\nDataset final : X={X.shape}, y={y.shape}\")\nprint(f\"Positifs: {int(np.sum(y==1))}, Negatifs: {int(np.sum(y==0))}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:21.732024Z","iopub.execute_input":"2026-07-10T19:13:21.732872Z","iopub.status.idle":"2026-07-10T19:13:22.325218Z","shell.execute_reply.started":"2026-07-10T19:13:21.732828Z","shell.execute_reply":"2026-07-10T19:13:22.324177Z"}},"outputs":[],"execution_count":null},{"id":"751800c1-2c8d-4f92-a9b5-e8b172a2e97d","cell_type":"markdown","source":"## Data augmentation 3D coherente","metadata":{}},{"id":"d02e1742-7aff-4618-9e29-29822bd02489","cell_type":"code","source":"def augment_volume(volume):\n    \"\"\"Applique la MEME transformation geometrique a toutes les coupes d'un\n    volume (essentiel : augmenter chaque coupe independamment casserait la\n    coherence spatiale entre coupes).\"\"\"\n    # Flip horizontal\n    if np.random.rand() > 0.5:\n        volume = volume[:, :, ::-1, :]\n\n    # Rotation legere (meme angle pour toutes les coupes)\n    if np.random.rand() > 0.5:\n        angle = np.random.uniform(-15, 15)\n        h, w = volume.shape[1], volume.shape[2]\n        M = cv2.getRotationMatrix2D((w / 2, h / 2), angle, 1.0)\n        volume = np.stack([\n            cv2.warpAffine(volume[i], M, (w, h), flags=cv2.INTER_LINEAR)\n            for i in range(volume.shape[0])\n        ])\n\n    # Zoom leger\n    if np.random.rand() > 0.5:\n        zoom = np.random.uniform(0.9, 1.1)\n        h, w = volume.shape[1], volume.shape[2]\n        M = cv2.getRotationMatrix2D((w / 2, h / 2), 0, zoom)\n        volume = np.stack([\n            cv2.warpAffine(volume[i], M, (w, h), flags=cv2.INTER_LINEAR)\n            for i in range(volume.shape[0])\n        ])\n\n    # Jitter de luminosite/contraste (meme facteur pour tout le volume)\n    if np.random.rand() > 0.5:\n        brightness = np.random.uniform(-0.1, 0.1)\n        contrast = np.random.uniform(0.9, 1.1)\n        volume = np.clip(volume * contrast + brightness, 0, 1)\n\n    return volume.astype(np.float32)\n\n\nclass VolumeDataGenerator(tf.keras.utils.Sequence):\n    \"\"\"Generateur qui applique l'augmentation a la volee, uniquement en train.\"\"\"\n    def __init__(self, X, y, batch_size=BATCH_SIZE, augment=False, shuffle=True):\n        self.X, self.y = X, y\n        self.batch_size = batch_size\n        self.augment = augment\n        self.shuffle = shuffle\n        self.indices = np.arange(len(X))\n        self.on_epoch_end()\n\n    def __len__(self):\n        return int(np.ceil(len(self.X) / self.batch_size))\n\n    def on_epoch_end(self):\n        if self.shuffle:\n            np.random.shuffle(self.indices)\n\n    def __getitem__(self, idx):\n        batch_idx = self.indices[idx * self.batch_size:(idx + 1) * self.batch_size]\n        batch_X = self.X[batch_idx].copy()\n        batch_y = self.y[batch_idx]\n\n        if self.augment:\n            batch_X = np.array([augment_volume(v) for v in batch_X])\n\n        return batch_X, batch_y","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:32.594310Z","iopub.execute_input":"2026-07-10T19:13:32.595157Z","iopub.status.idle":"2026-07-10T19:13:32.606194Z","shell.execute_reply.started":"2026-07-10T19:13:32.595126Z","shell.execute_reply":"2026-07-10T19:13:32.605249Z"}},"outputs":[],"execution_count":null},{"id":"1c517299-70c9-4904-b706-d8a4d6453759","cell_type":"markdown","source":"## Modele : EfficientNetB0 (transfer learning) applique par coupe + agregation","metadata":{}},{"id":"4d62dd63-992e-4cdf-aecf-73b8941a2137","cell_type":"code","source":"def build_model(input_shape, num_classes=1, l2_reg=1e-4):\n    num_slices, h, w, channels = input_shape\n    inputs = layers.Input(shape=input_shape)\n\n    # Adapte les n canaux -> 3 canaux RGB attendus par EfficientNet\n    x = layers.TimeDistributed(layers.Conv2D(3, (1, 1), padding='same'))(inputs)\n\n    backbone = tf.keras.applications.EfficientNetB0(\n        include_top=False, weights='imagenet', pooling='avg',\n        input_shape=(h, w, 3)\n    )\n    # On gele le backbone au debut (fine-tuning partiel ensuite)\n    backbone.trainable = False\n\n    x = layers.TimeDistributed(backbone)(x)          # -> (num_slices, features)\n    x = layers.Dropout(0.3)(x)\n\n    # Agregation across coupes : Bi-LSTM capture les relations spatiales entre coupes\n    x = layers.Bidirectional(layers.LSTM(64, return_sequences=False,\n                                          kernel_regularizer=regularizers.l2(l2_reg)))(x)\n    x = layers.Dense(64, activation='relu', kernel_regularizer=regularizers.l2(l2_reg))(x)\n    x = layers.Dropout(0.4)(x)\n    # dtype='float32' : force la sortie en float32 meme sous la politique\n    # mixed_float16, necessaire pour la stabilite numerique du sigmoid + loss\n    outputs = layers.Dense(num_classes, activation='sigmoid', dtype='float32')(x)\n\n    model = models.Model(inputs, outputs)\n    return model, backbone","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:39.274030Z","iopub.execute_input":"2026-07-10T19:13:39.275081Z","iopub.status.idle":"2026-07-10T19:13:39.282089Z","shell.execute_reply.started":"2026-07-10T19:13:39.275048Z","shell.execute_reply":"2026-07-10T19:13:39.281428Z"}},"outputs":[],"execution_count":null},{"id":"b89091a7-5d20-4512-a13f-17d59400c0db","cell_type":"markdown","source":"## Learning rate : cosine decay avec warmup","metadata":{}},{"id":"1f928f9c-73a9-4d65-974c-4abdd754d063","cell_type":"code","source":"def make_lr_schedule(steps_per_epoch, epochs=EPOCHS, base_lr=1e-3, warmup_epochs=3):\n    total_steps = steps_per_epoch * epochs\n    warmup_steps = steps_per_epoch * warmup_epochs\n    return tf.keras.optimizers.schedules.CosineDecay(\n        initial_learning_rate=base_lr,\n        decay_steps=max(1, total_steps - warmup_steps),\n        warmup_target=base_lr,\n        warmup_steps=warmup_steps\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:51.196940Z","iopub.execute_input":"2026-07-10T19:13:51.197940Z","iopub.status.idle":"2026-07-10T19:13:51.203300Z","shell.execute_reply.started":"2026-07-10T19:13:51.197905Z","shell.execute_reply":"2026-07-10T19:13:51.202383Z"}},"outputs":[],"execution_count":null},{"id":"15755f68-bd7c-4391-b67f-ce1ce4afe5db","cell_type":"markdown","source":"## Validation croisee stratifiee (5-fold)\nEssentiel avec seulement ~580 patients : un seul split train/test donne une estimation tres bruitee.","metadata":{}},{"id":"7d958f7e-7720-4b8f-8dd4-d69a22bc4596","cell_type":"code","source":"skf = StratifiedKFold(n_splits=N_FOLDS, shuffle=True, random_state=SEED)\nfold_aucs = []\noof_preds = np.zeros(len(y))\n\nfor fold, (train_idx, val_idx) in enumerate(skf.split(X, y)):\n    print(f\"\\n{'='*60}\\nFOLD {fold + 1}/{N_FOLDS}\\n{'='*60}\")\n\n    X_train, X_val = X[train_idx], X[val_idx]\n    y_train, y_val = y[train_idx], y[val_idx]\n\n    train_gen = VolumeDataGenerator(X_train, y_train, augment=True, shuffle=True)\n    val_gen = VolumeDataGenerator(X_val, y_val, augment=False, shuffle=False)\n\n    input_shape = X_train.shape[1:]\n\n    # strategy.scope() : construction + compilation du modele sous la\n    # strategie de distribution (repartit automatiquement sur les GPU dispo)\n    with strategy.scope():\n        model, backbone = build_model(input_shape)\n\n        lr_schedule = make_lr_schedule(steps_per_epoch=len(train_gen))\n        model.compile(\n            optimizer=tf.keras.optimizers.Adam(learning_rate=lr_schedule),\n            loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05),\n            metrics=['accuracy', tf.keras.metrics.AUC(name='auc')]\n        )\n\n    callbacks = [\n        tf.keras.callbacks.EarlyStopping(\n            monitor='val_auc', mode='max', patience=8,\n            restore_best_weights=True, verbose=1\n        ),\n        tf.keras.callbacks.ModelCheckpoint(\n            filepath=os.path.join(CACHE_DIR, f'best_model_fold{fold}.h5'),\n            save_best_only=True, monitor='val_auc', mode='max', verbose=0\n        ),\n    ]\n\n    # Phase 1 : entrainement avec backbone gele\n    model.fit(train_gen, validation_data=val_gen, epochs=EPOCHS,\n              callbacks=callbacks, verbose=1)\n\n    # Phase 2 : fine-tuning des dernieres couches du backbone avec un LR tres bas\n    backbone.trainable = True\n    for layer in backbone.layers[:-20]:\n        layer.trainable = False\n\n    with strategy.scope():\n        model.compile(\n            optimizer=tf.keras.optimizers.Adam(learning_rate=1e-5),\n            loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05),\n            metrics=['accuracy', tf.keras.metrics.AUC(name='auc')]\n        )\n    model.fit(train_gen, validation_data=val_gen, epochs=10,\n              callbacks=callbacks, verbose=1)\n\n    # Evaluation du fold\n    val_pred = model.predict(val_gen).flatten()\n    oof_preds[val_idx] = val_pred\n    fold_auc = auc(*roc_curve(y_val, val_pred)[:2])\n    fold_aucs.append(fold_auc)\n    print(f\"Fold {fold + 1} AUC: {fold_auc:.4f}\")\n\nprint(f\"\\n{'='*60}\")\nprint(f\"AUC moyen (CV {N_FOLDS}-fold) : {np.mean(fold_aucs):.4f} +/- {np.std(fold_aucs):.4f}\")\nprint(f\"AUC out-of-fold global      : {auc(*roc_curve(y, oof_preds)[:2]):.4f}\")\nprint(f\"{'='*60}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:13:55.721375Z","iopub.execute_input":"2026-07-10T19:13:55.721869Z","iopub.status.idle":"2026-07-10T20:01:01.639182Z","shell.execute_reply.started":"2026-07-10T19:13:55.721841Z","shell.execute_reply":"2026-07-10T20:01:01.638362Z"}},"outputs":[],"execution_count":null},{"id":"4fb4c7b1-9de1-4079-8d0e-094929eaad58","cell_type":"markdown","source":"## Rapport final\nSur les predictions out-of-fold (plus fiable qu'un seul split train/test).","metadata":{}},{"id":"05428aca-f518-436b-aeb7-7eb0aa74cba1","cell_type":"code","source":"fpr, tpr, thresholds = roc_curve(y, oof_preds)\nroc_auc = auc(fpr, tpr)\n\n# Seuil optimal (indice de Youden = tpr - fpr max) plutot que 0.5 fixe.\n# Utile de comparer aux resultats avec seuil 0.5 : si le seuil optimal est\n# tres different de 0.5 (ex. 0.3 ou 0.7), ca indique un decalage du modele\n# plutot qu'un collapse total (auquel cas meme le seuil optimal ne separera rien).\nbest_idx = np.argmax(tpr - fpr)\nbest_threshold = thresholds[best_idx]\nprint(f\"Seuil optimal (Youden) : {best_threshold:.4f} (vs 0.5 par defaut)\")\n\noof_labels = (oof_preds > best_threshold).astype(int)\n\nprint(\"\\nClassification Report (out-of-fold, seuil optimal) :\")\nprint(classification_report(y, oof_labels))\n\nprint(\"\\nConfusion Matrix (out-of-fold, seuil optimal) :\")\nprint(confusion_matrix(y, oof_labels))\n\nplt.figure()\nplt.plot(fpr, tpr, color='darkorange', lw=2, label=f'ROC curve (AUC = {roc_auc:.2f})')\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--')\nplt.xlim([0.0, 1.0]); plt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate'); plt.ylabel('True Positive Rate')\nplt.title('ROC Curve (Out-of-Fold)')\nplt.legend(loc=\"lower right\")\nplt.savefig('roc_curve_ameliore.png')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:12:51.648496Z","iopub.execute_input":"2026-07-10T20:12:51.649316Z","iopub.status.idle":"2026-07-10T20:12:51.888335Z","shell.execute_reply.started":"2026-07-10T20:12:51.649284Z","shell.execute_reply":"2026-07-10T20:12:51.887599Z"}},"outputs":[],"execution_count":null},{"id":"c570f2ea-eb96-4841-ba0e-70dc436a3782","cell_type":"markdown","source":"## Entrainement final sur toutes les donnees\nPour la soumission (modele unique entraine sur (presque) tout le dataset).","metadata":{}},{"id":"1b78cc68-5e63-4593-8f5d-579fd3a0c9c8","cell_type":"code","source":"X_tr, X_ho, y_tr, y_ho = train_test_split(X, y, test_size=0.1, random_state=SEED, stratify=y)\ntrain_gen = VolumeDataGenerator(X_tr, y_tr, augment=True, shuffle=True)\nholdout_gen = VolumeDataGenerator(X_ho, y_ho, augment=False, shuffle=False)\n\nwith strategy.scope():\n    final_model, final_backbone = build_model(X.shape[1:])\n    final_model.compile(\n        optimizer=tf.keras.optimizers.Adam(learning_rate=make_lr_schedule(len(train_gen))),\n        loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05),\n        metrics=['accuracy', tf.keras.metrics.AUC(name='auc')]\n    )\nfinal_model.fit(\n    train_gen, validation_data=holdout_gen, epochs=EPOCHS,\n    callbacks=[tf.keras.callbacks.EarlyStopping(\n        monitor='val_auc', mode='max', patience=8, restore_best_weights=True)]\n)\n\nfinal_model.save('final_model_ameliore.h5')\nprint(\"\\nModele final sauvegarde sous 'final_model_ameliore.h5'\")","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}