{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"12880a3e-0b24-4399-ad8a-c7c1fab28d80","cell_type":"markdown","source":"# MAIRA-2 — Génération de rapport ancré, localisation et classification dérivée\n\n**Projet ARV-DS03 — approfondissement : modèle médical spécialisé MAIRA-2 (Microsoft)**\n\nMAIRA-2 est un modèle multimodal (encodeur RAD-DINO gelé + Vicuna-7B fine-tuné) qui **génère des comptes-rendus radiologiques ancrés** : chaque observation est accompagnée de boîtes englobantes localisant l'anomalie. Contrairement à Gemma 4, il s'utilise **par prompting** (aucun fine-tuning), dans son format d'entrée prévu.\n\nCe notebook :\n1. génère un rapport ancré par image (localisation native, vraies boîtes) ;\n2. **dérive une classe** (normal / suspected_opacity / uncertain) par post-traitement du texte du rapport ;\n3. évalue la **classification** (vs vérité RSNA) et la **localisation** (accord par quadrant, comparable à la partie B Gemma).\n\n> Accès : MAIRA-2 est un modèle *gated*. Il faut avoir demandé l'accès sur sa page Hugging Face et enregistré un secret Kaggle `HF_TOKEN` (Add-ons → Secrets).","metadata":{}},{"id":"a28c77ef-cfdf-40fa-93ad-de819fb076ca","cell_type":"markdown","source":"## 1. Environnement\n\nMAIRA-2 exige une version précise de `transformers` (testé en 4.51.3) et `trust_remote_code`. On épingle la version pour éviter les incompatibilités du code distant.","metadata":{}},{"id":"d615b9e8-c00b-4fbf-9de9-53c16b93c7f1","cell_type":"code","source":"!pip install -q \"transformers==4.51.3\" \"bitsandbytes>=0.43.0\" pillow protobuf sentencepiece 2>/dev/null\n\n# Diagnostic : verifier que bitsandbytes s'importe VRAIMENT (pas juste installe)\ntry:\n    import bitsandbytes as bnb\n    print(\"bitsandbytes importe :\", bnb.__version__)\nexcept Exception as e:\n    print(\"ECHEC import bitsandbytes :\", e)\n\nfrom transformers.utils import is_bitsandbytes_available\nprint(\"transformers voit bitsandbytes :\", is_bitsandbytes_available())","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:28:48.524415Z","iopub.execute_input":"2026-07-05T10:28:48.525210Z","iopub.status.idle":"2026-07-05T10:29:17.858176Z","shell.execute_reply.started":"2026-07-05T10:28:48.525164Z","shell.execute_reply":"2026-07-05T10:29:17.857175Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"cbeda403-8a6a-46e0-8863-2f6da6b34c26","cell_type":"code","source":"import os, json, re, torch, numpy as np, pandas as pd, cv2\nfrom PIL import Image\nimport pydicom\nfrom kaggle_secrets import UserSecretsClient\nfrom huggingface_hub import login\n\n# Authentification Hugging Face (secret Kaggle \"HF_TOKEN\" requis : modele gated)\nlogin(UserSecretsClient().get_secret(\"HF_TOKEN\"))\n\nimport transformers\nprint(\"transformers :\", transformers.__version__)\nprint(\"CUDA         :\", torch.cuda.get_device_name(0))","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:29:17.859474Z","iopub.execute_input":"2026-07-05T10:29:17.859965Z","iopub.status.idle":"2026-07-05T10:29:19.101766Z","shell.execute_reply.started":"2026-07-05T10:29:17.859937Z","shell.execute_reply":"2026-07-05T10:29:19.100960Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"47667562-01ba-4bd9-8a02-b153c321e3fe","cell_type":"markdown","source":"## 2. Dataset RSNA (labels + boîtes)\n\nMêmes données que les notebooks Gemma, pour que les comparaisons soient valides. On garde le mapping trois classes pour l'évaluation de classification, et les boîtes `Target=1` pour la localisation.","metadata":{}},{"id":"c554b331-3b47-4c02-bca5-a19a253bca3b","cell_type":"code","source":"import glob\nmatches = glob.glob('/kaggle/input/**/stage_2_detailed_class_info.csv', recursive=True)\nif not matches:\n    raise FileNotFoundError(\"Dataset RSNA introuvable (Add Input).\")\nBASE_DIR = os.path.dirname(matches[0]) + '/'\nTRAIN_IMAGES_DIR = os.path.join(BASE_DIR, 'stage_2_train_images')\n\ndf_class = pd.read_csv(os.path.join(BASE_DIR, 'stage_2_detailed_class_info.csv')).drop_duplicates('patientId').reset_index(drop=True)\nCLASS_MAP = {\n    \"Normal\": \"normal\",\n    \"Lung Opacity\": \"suspected_opacity\",\n    \"No Lung Opacity / Not Normal\": \"uncertain\",\n}\ndf_class[\"label\"] = df_class[\"class\"].map(CLASS_MAP)\ndf_boxes = pd.read_csv(os.path.join(BASE_DIR, 'stage_2_train_labels.csv'))\n\nfrom sklearn.model_selection import train_test_split\n# Pas d'entrainement : un seul echantillon d'evaluation suffit (prompting pur).\n_, eval_sub = train_test_split(df_class, test_size=0.20, stratify=df_class[\"label\"], random_state=42)\nN_EVAL = 100  # ajustable ; MAIRA-2 est lourd, on commence petit\neval_sub, _ = train_test_split(eval_sub, train_size=N_EVAL, stratify=eval_sub[\"label\"], random_state=42)\neval_sub = eval_sub.reset_index(drop=True)\nprint(\"Echantillon eval :\", len(eval_sub))\nprint(eval_sub[\"label\"].value_counts())","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:29:19.102881Z","iopub.execute_input":"2026-07-05T10:29:19.103347Z","iopub.status.idle":"2026-07-05T10:30:07.924845Z","shell.execute_reply.started":"2026-07-05T10:29:19.103321Z","shell.execute_reply":"2026-07-05T10:30:07.924021Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"b22b24ff-562c-49fa-a215-d39c6bb22d42","cell_type":"markdown","source":"## 3. Chargement des images\n\nMAIRA-2 attend une image PIL RGB. On lit le DICOM, on normalise et on applique CLAHE (rehaussement de contraste), comme pour Gemma. On ne redimensionne pas : le processeur de MAIRA-2 s'en charge (et fournit de quoi réajuster les boîtes).","metadata":{}},{"id":"ebe8dd8c-4f99-427a-b76b-b0afdd95d193","cell_type":"code","source":"def load_cxr(patient_id):\n    \"\"\"DICOM -> image PIL RGB (sans resize : le processeur MAIRA-2 recadre lui-meme).\"\"\"\n    path = os.path.join(TRAIN_IMAGES_DIR, f\"{patient_id}.dcm\")\n    if not os.path.exists(path):\n        return None\n    dcm = pydicom.dcmread(path)\n    img = dcm.pixel_array.astype(np.float32)\n    img = (img - img.min()) / (img.max() - img.min() + 1e-8) * 255.0\n    img = img.astype(np.uint8)\n    clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8))\n    img = clahe.apply(img)\n    return Image.fromarray(cv2.cvtColor(img, cv2.COLOR_GRAY2RGB))\n\n_t = load_cxr(eval_sub.iloc[0][\"patientId\"])\nprint(\"Image test :\", _t.size, _t.mode)","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:30:07.926527Z","iopub.execute_input":"2026-07-05T10:30:07.927051Z","iopub.status.idle":"2026-07-05T10:30:08.016712Z","shell.execute_reply.started":"2026-07-05T10:30:07.927027Z","shell.execute_reply":"2026-07-05T10:30:08.015906Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"8bfeae35-ba4a-4a56-980b-952e2667dee0","cell_type":"markdown","source":"## 4. Chargement de MAIRA-2 (4-bit pour tenir sur T4)\n\nLe modèle fait ~7 Md de paramètres. En précision native il ne tient pas sur un T4 (16 Go) avec la marge nécessaire à l'inférence. On le charge en **4-bit** (quantisation), comme pour Gemma.","metadata":{}},{"id":"470d08b6-2d9e-4691-b5b6-22dd1a3cbe80","cell_type":"code","source":"from transformers import AutoModelForCausalLM, AutoProcessor, BitsAndBytesConfig\n\nMODEL_ID = \"microsoft/maira-2\"\nbnb = BitsAndBytesConfig(\n    load_in_4bit=True,\n    bnb_4bit_use_double_quant=True,\n    bnb_4bit_quant_type=\"nf4\",\n    bnb_4bit_compute_dtype=torch.float16,\n)\n\nprocessor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)\nmodel = AutoModelForCausalLM.from_pretrained(\n    MODEL_ID,\n    trust_remote_code=True,\n    quantization_config=bnb,\n    device_map={\"\": 0},\n)\nmodel = model.eval()\nprint(\"MAIRA-2 charge en 4-bit.\")","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:30:08.017648Z","iopub.execute_input":"2026-07-05T10:30:08.018161Z","iopub.status.idle":"2026-07-05T10:34:39.989598Z","shell.execute_reply.started":"2026-07-05T10:30:08.018122Z","shell.execute_reply":"2026-07-05T10:34:39.988723Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"7fd968af-0178-46a8-a2fc-4e033b5a0dcc","cell_type":"markdown","source":"## 5. Génération d'un rapport ancré\n\nOn utilise l'API prévue par MAIRA-2 : `format_and_preprocess_reporting_input` avec `get_grounding=True`. On ne fournit que la vue frontale (RSNA n'a pas de vue latérale ni d'antériorité). On ne modifie **pas** le format d'instruction : le modèle y est très sensible.\n\nLa sortie est une liste de couples `(phrase, [boîtes])`, les boîtes étant en coordonnées relatives à l'image **recadrée** vue par le modèle.","metadata":{}},{"id":"0ae2d515-6a3d-4283-83c3-39c044ec7514","cell_type":"code","source":"def generate_grounded_report(image):\n    \"\"\"Retourne la sortie ancree parsee : liste de (phrase, boites|None).\"\"\"\n    processed = processor.format_and_preprocess_reporting_input(\n        current_frontal=image,\n        current_lateral=None,\n        prior_frontal=None,\n        indication=None,\n        technique=None,\n        comparison=None,\n        prior_report=None,\n        return_tensors=\"pt\",\n        get_grounding=True,   # rapport ANCRE (avec boites)\n    ).to(model.device)\n\n    with torch.no_grad():\n        out = model.generate(**processed, max_new_tokens=450, use_cache=True)\n    prompt_len = processed[\"input_ids\"].shape[-1]\n    text = processor.decode(out[0][prompt_len:], skip_special_tokens=True).lstrip()\n    # Parsing officiel MAIRA-2 -> sequence structuree (phrase, boites)\n    parsed = processor.convert_output_to_plaintext_or_grounded_sequence(text)\n    return parsed, text\n\n# Test rapide sur une image\ndemo_parsed, demo_raw = generate_grounded_report(load_cxr(eval_sub.iloc[0][\"patientId\"]))\nprint(\"Sortie brute :\\n\", demo_raw[:500], \"\\n\")\nprint(\"Sortie parsee :\\n\", demo_parsed)","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:34:39.990745Z","iopub.execute_input":"2026-07-05T10:34:39.991492Z","iopub.status.idle":"2026-07-05T10:34:46.711483Z","shell.execute_reply.started":"2026-07-05T10:34:39.991464Z","shell.execute_reply":"2026-07-05T10:34:46.710802Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"a6d7d160-18ca-420c-a3bd-b4f18554dbbf","cell_type":"markdown","source":"## 6. Dérivation de la classe à partir du rapport (post-traitement)\n\nMAIRA-2 ne produit pas de classe : il rédige un compte-rendu. On **dérive** la classe des trois catégories du projet à partir du texte, par détection de mots-clés radiologiques.\n\nLogique (prudente, alignée sur la priorité dépistage) :\n- présence de termes d'**opacité / consolidation / pneumonie / infiltrat** → `suspected_opacity` ;\n- rapport explicitement **normal / clair / sans anomalie** → `normal` ;\n- sinon (rapport non concluant ou ambigu) → `uncertain` (abstention).\n\nC'est une heuristique documentée, à assumer comme telle : elle traduit un texte libre en étiquette, et constitue une source d'erreur distincte du modèle lui-même.","metadata":{}},{"id":"0dc04aba-68bb-42c6-b9c9-ef587da8053a","cell_type":"code","source":"import re\n\n# --- Termes indiquant une ANOMALIE de type opacite (findings pulmonaires) ---\nOPACITY_TERMS = [\n    \"opacity\", \"opacities\", \"opacification\",\n    \"consolidation\", \"consolidations\",\n    \"pneumonia\", \"pneumonic\",\n    \"infiltrate\", \"infiltrates\", \"infiltration\",\n    \"airspace disease\", \"air-space disease\", \"air space disease\",\n    \"effusion\", \"effusions\",\n    \"atelectasis\", \"atelectatic\",\n    \"edema\", \"pulmonary edema\",\n    \"haziness\", \"hazy\",\n    \"density\", \"densities\",\n    \"mass\", \"nodule\", \"nodules\",\n]\n\n# --- Termes de NORMALITE explicite ---\nNORMAL_TERMS = [\n    \"clear\", \"unremarkable\", \"no acute\", \"no focal\", \"normal\",\n    \"no evidence\", \"within normal limits\", \"no abnormal\",\n]\n\n# --- Marqueurs de NEGATION (si presents juste avant un terme, on l'annule) ---\nNEGATION_CUES = [\n    \"no\", \"not\", \"without\", \"negative for\", \"free of\", \"clear of\",\n    \"resolution of\", \"resolved\", \"absence of\", \"rule out\", \"ruled out\",\n]\n\n# --- Phrases decrivant des DISPOSITIFS / structures NON pathologiques ---\n# (une boite sur ces phrases ne doit PAS compter comme opacite)\nDEVICE_TERMS = [\n    \"picc\", \"line\", \"tube\", \"catheter\", \"pacemaker\", \"lead\", \"wire\",\n    \"sternotomy\", \"clip\", \"device\", \"port\", \"valve\", \"stent\",\n    \"et tube\", \"ng tube\", \"tracheostomy\",\n]\n\ndef _negated(text, term):\n    \"\"\"Vrai si `term` apparait dans `text` precede (a <=4 mots) d'une negation.\"\"\"\n    for m in re.finditer(re.escape(term), text):\n        start = m.start()\n        # fenetre de contexte : ~40 caracteres avant le terme\n        window = text[max(0, start-40):start]\n        if any(cue in window for cue in NEGATION_CUES):\n            return True\n    return False\n\ndef _term_present_affirmative(text, terms):\n    \"\"\"Vrai s'il existe au moins un terme de la liste present ET non nie.\"\"\"\n    for t in terms:\n        if t in text and not _negated(text, t):\n            return True\n    return False\n\ndef phrase_is_pathological(phrase):\n    \"\"\"Une phrase decrit-elle une anomalie pulmonaire (vs dispositif/normalite) ?\"\"\"\n    p = phrase.lower()\n    if any(d in p for d in DEVICE_TERMS):\n        return False  # dispositif medical : pas une opacite\n    if _term_present_affirmative(p, OPACITY_TERMS):\n        return True\n    return False\n\ndef derive_class(parsed_output):\n    \"\"\"\n    Classe derivee en DEUX niveaux :\n      1) signal fort : une BOITE ancree sur une phrase pathologique => suspected_opacity\n      2) signal texte : termes d'opacite non nies => suspected_opacity\n                        sinon normalite explicite => normal\n                        sinon => uncertain\n    Retourne (classe, rapport_texte, a_une_boite_pathologique)\n    \"\"\"\n    if not isinstance(parsed_output, (list, tuple)) or len(parsed_output) == 0:\n        return \"uncertain\", \"\", False\n\n    phrases = []\n    box_on_pathology = False\n    for item in parsed_output:\n        if isinstance(item, (list, tuple)) and len(item) >= 1:\n            phrase = str(item[0])\n            boxes = item[1] if len(item) >= 2 else None\n            phrases.append(phrase)\n            # NIVEAU 1 : une boite posee sur une phrase pathologique\n            if boxes and phrase_is_pathological(phrase):\n                box_on_pathology = True\n        else:\n            phrases.append(str(item))\n\n    report = \" \".join(phrases).lower()\n\n    # Niveau 1 : signal fort par la localisation\n    if box_on_pathology:\n        return \"suspected_opacity\", report, True\n\n    # Niveau 2 : analyse texte avec gestion des negations\n    if _term_present_affirmative(report, OPACITY_TERMS):\n        return \"suspected_opacity\", report, False\n    if _term_present_affirmative(report, NORMAL_TERMS):\n        return \"normal\", report, False\n    return \"uncertain\", report, False\n\n# Test sur la demo\ncls_demo, rep_demo, boxpath_demo = derive_class(demo_parsed)\nprint(\"Classe derivee   :\", cls_demo)\nprint(\"Boite patho ?    :\", boxpath_demo)\nprint(\"Rapport (extrait):\", rep_demo[:300])\n","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:34:46.712490Z","iopub.execute_input":"2026-07-05T10:34:46.712836Z","iopub.status.idle":"2026-07-05T10:34:46.725942Z","shell.execute_reply.started":"2026-07-05T10:34:46.712774Z","shell.execute_reply":"2026-07-05T10:34:46.724705Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"e6e84dfe-9347-46ad-8aac-ef957e87dc97","cell_type":"markdown","source":"## 7. Extraction des boîtes pour la localisation\n\nPour comparer à la partie B de Gemma, on convertit les boîtes de MAIRA-2 en quadrant. Deux précautions :\n- les boîtes sont **relatives à l'image recadrée** vue par le modèle ; le processeur fournit `adjust_box_for_original_image_size` pour repasser aux dimensions d'origine ;\n- MAIRA-2 renvoie des coordonnées `(x_tl, y_tl, x_br, y_br)` normalisées (0–1).","metadata":{}},{"id":"6e7401bb-c1d6-4970-b648-8854e3a31bd8","cell_type":"code","source":"def box_center_to_quadrant(x_tl, y_tl, x_br, y_br):\n    \"\"\"Boite normalisee (0-1) -> quadrant qualitatif (comme la partie B Gemma).\"\"\"\n    cx = (x_tl + x_br) / 2\n    cy = (y_tl + y_br) / 2\n    horiz = \"left\" if cx < 0.5 else \"right\"\n    if cy < 1/3:\n        vert = \"upper\"\n    elif cy < 2/3:\n        vert = \"middle\"\n    else:\n        vert = \"lower\"\n    return f\"{vert}-{horiz}\"\n\ndef first_box_quadrant(parsed_output):\n    \"\"\"Retourne le quadrant de la premiere boite trouvee dans la sortie ancree, ou None.\"\"\"\n    if not isinstance(parsed_output, (list, tuple)):\n        return None\n    for item in parsed_output:\n        if isinstance(item, (list, tuple)) and len(item) >= 2 and item[1]:\n            boxes = item[1]\n            if boxes:\n                b = boxes[0]\n                if len(b) == 4:\n                    return box_center_to_quadrant(*b)\n    return None\n\n# Verite terrain quadrant (comme partie B Gemma)\ndef gt_quadrant(patient_id):\n    rows = df_boxes[(df_boxes[\"patientId\"] == patient_id) & (df_boxes[\"Target\"] == 1)]\n    if len(rows) == 0:\n        return None\n    r = rows.iloc[0]\n    dcm = pydicom.dcmread(os.path.join(TRAIN_IMAGES_DIR, f\"{patient_id}.dcm\"))\n    H, W = dcm.pixel_array.shape\n    cx, cy = r[\"x\"] + r[\"width\"]/2, r[\"y\"] + r[\"height\"]/2\n    horiz = \"left\" if cx < W/2 else \"right\"\n    if cy < H/3: vert = \"upper\"\n    elif cy < 2*H/3: vert = \"middle\"\n    else: vert = \"lower\"\n    return f\"{vert}-{horiz}\"","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:34:46.726962Z","iopub.execute_input":"2026-07-05T10:34:46.727246Z","iopub.status.idle":"2026-07-05T10:34:46.742646Z","shell.execute_reply.started":"2026-07-05T10:34:46.727221Z","shell.execute_reply":"2026-07-05T10:34:46.741798Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"58806534-0974-4dce-8c1c-6f1b66e26461","cell_type":"markdown","source":"## 8. Évaluation complète\n\nBoucle unique : pour chaque image d'éval, on génère le rapport ancré, on en dérive la classe et le quadrant. On calcule ensuite la classification (3 classes) et l'accord de localisation (sur les vraies opacités uniquement).","metadata":{}},{"id":"3dc9b98e-cd8b-42fb-be0d-a7bb2a5115f9","cell_type":"code","source":"from sklearn.metrics import accuracy_score, f1_score, confusion_matrix, classification_report\nfrom tqdm.auto import tqdm\n\npreds, trues, reports = [], [], []\nloc_pred_q, loc_true_q = [], []\nviz_records = []  # pour la visualisation (section suivante)\n\nfor _, row in tqdm(eval_sub.iterrows(), total=len(eval_sub)):\n    pid = row[\"patientId\"]\n    img = load_cxr(pid)\n    if img is None:\n        preds.append(\"uncertain\"); trues.append(row[\"label\"]); reports.append(\"\")\n        continue\n    try:\n        parsed, raw = generate_grounded_report(img)\n    except Exception:\n        parsed, raw = [], \"\"\n    cls, report, box_on_patho = derive_class(parsed)\n    preds.append(cls); trues.append(row[\"label\"]); reports.append(report)\n\n    # Boites predites (toutes) pour la visualisation\n    pred_boxes = []\n    for item in parsed if isinstance(parsed, (list, tuple)) else []:\n        if isinstance(item, (list, tuple)) and len(item) >= 2 and item[1]:\n            for b in item[1]:\n                if len(b) == 4:\n                    pred_boxes.append((item[0], b))  # (phrase, boite normalisee)\n\n    # Boites reelles (verite terrain RSNA)\n    gt_rows = df_boxes[(df_boxes[\"patientId\"] == pid) & (df_boxes[\"Target\"] == 1)]\n\n    viz_records.append({\n        \"pid\": pid, \"true\": row[\"label\"], \"pred\": cls,\n        \"report\": report, \"pred_boxes\": pred_boxes,\n        \"gt_rows\": gt_rows, \"box_on_patho\": box_on_patho,\n        \"raw\": raw, \"n_pred_boxes\": len(pred_boxes),\n    })\n\n    # Localisation (sur vraies opacites)\n    if row[\"label\"] == \"suspected_opacity\":\n        gtq = gt_quadrant(pid)\n        pq = first_box_quadrant(parsed)\n        if gtq is not None:\n            loc_true_q.append(gtq)\n            loc_pred_q.append(pq if pq is not None else \"none\")\n\nlabels = [\"normal\", \"suspected_opacity\", \"uncertain\"]\nacc = accuracy_score(trues, preds)\nf1  = f1_score(trues, preds, labels=labels, average=\"macro\")\nprint(\"=== CLASSIFICATION (derivee du rapport MAIRA-2, v2 avec negations + boites) ===\")\nprint(f\"Accuracy : {acc:.4f}   Macro-F1 : {f1:.4f}\")\nprint(classification_report(trues, preds, labels=labels, zero_division=0))\nprint(pd.DataFrame(confusion_matrix(trues, preds, labels=labels), index=labels, columns=labels))\n\ncm = confusion_matrix(trues, preds, labels=labels)\nop_i = labels.index(\"suspected_opacity\")\ntp = cm[op_i, op_i]; support = cm[op_i].sum()\nif support:\n    print(f\"\\n>>> Rappel suspected_opacity : {tp/support:.4f}\")\n\nif loc_true_q:\n    exact = np.mean([p == t for p, t in zip(loc_pred_q, loc_true_q)])\n    side  = np.mean([p.split(\"-\")[-1] == t.split(\"-\")[-1]\n                     for p, t in zip(loc_pred_q, loc_true_q) if p != \"none\"])\n    detected = np.mean([p != \"none\" for p in loc_pred_q])\n    print(\"\\n=== LOCALISATION ===\")\n    print(f\"Cas evalues      : {len(loc_true_q)}\")\n    print(f\"Taux detection   : {detected:.4f}\")\n    print(f\"Accord quadrant  : {exact:.4f}\")\n    print(f\"Accord cote (H)  : {side:.4f}\")\n\npd.DataFrame({\"patientId\": eval_sub[\"patientId\"], \"true\": trues, \"pred\": preds,\n              \"report\": reports}).to_csv(\"/kaggle/working/maira2_results.csv\", index=False)\nprint(\"\\nResultats sauvegardes.\")\n","metadata":{"execution":{"iopub.status.busy":"2026-07-05T10:34:46.743547Z","iopub.execute_input":"2026-07-05T10:34:46.744173Z","iopub.status.idle":"2026-07-05T10:46:15.413576Z","shell.execute_reply.started":"2026-07-05T10:34:46.744148Z","shell.execute_reply":"2026-07-05T10:46:15.412728Z"},"trusted":true},"outputs":[],"execution_count":null},{"id":"5d7ef904-3930-411a-97da-529f96ce9af1","cell_type":"markdown","source":"## 9. Notes pour le rapport\n\n- **MAIRA-2 vs Gemma / MedGemma.** MAIRA-2 est un modèle *spécialisé rapport radiologique*, utilisé par prompting (aucun fine-tuning). La classe est **dérivée** d'un texte libre, ce qui ajoute une source d'erreur (l'heuristique de mots-clés) distincte du modèle.\n- **Localisation native.** Contrairement à l'approche par quadrants demandée à Gemma, MAIRA-2 produit de **vraies boîtes**. On les a ramenées au quadrant uniquement pour comparer à la partie B — mais on pourrait aussi rapporter un IoU avec les boîtes RSNA, métrique plus fine (piste d'extension).\n- **Sensibilité au format.** On a utilisé strictement l'API `format_and_preprocess_reporting_input` sans modifier l'instruction : MAIRA-2 se dégrade si on change son format d'entrée.\n- **Limite honnête.** MAIRA-2 a été entraîné sur des rapports US/Espagne ; la distribution RSNA peut différer. Les résultats sont à interpréter comme un transfert hors distribution.","metadata":{}},{"id":"7a5096ac","cell_type":"markdown","source":"## 10. Visualisation : boîtes prédites (MAIRA-2) vs réelles (RSNA)\n\nOn affiche une dizaine d'images avec, superposées :\n- en **vert** : les boîtes de vérité terrain RSNA (opacités annotées) ;\n- en **rouge** : les boîtes prédites par MAIRA-2 (avec la phrase associée).\n\nOn privilégie les cas de vraies opacités (là où la comparaison est parlante).","metadata":{}},{"id":"ea4ef69a","cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.patches as mpatches\n\n# On priorise les vraies opacites (boites GT presentes), puis on complete.\nviz_sorted = sorted(viz_records, key=lambda r: 0 if r[\"true\"] == \"suspected_opacity\" else 1)\nto_show = viz_sorted[:10]\n\nn = len(to_show)\ncols = 2\nrows_n = (n + cols - 1) // cols\nfig, axes = plt.subplots(rows_n, cols, figsize=(13, 5.5 * rows_n))\naxes = np.array(axes).reshape(-1)\n\nfor ax, rec in zip(axes, to_show):\n    pid = rec[\"pid\"]\n    img = load_cxr(pid)\n    W, H = img.size  # PIL : (width, height)\n    ax.imshow(img, cmap=\"gray\")\n\n    # Boites reelles (RSNA) en vert — coordonnees en pixels de l'image d'origine\n    dcm = pydicom.dcmread(os.path.join(TRAIN_IMAGES_DIR, f\"{pid}.dcm\"))\n    oh, ow = dcm.pixel_array.shape\n    sx, sy = W / ow, H / oh  # facteur d'echelle origine -> image affichee\n    for _, gr in rec[\"gt_rows\"].iterrows():\n        rect = mpatches.Rectangle((gr[\"x\"]*sx, gr[\"y\"]*sy), gr[\"width\"]*sx, gr[\"height\"]*sy,\n                                  linewidth=2.2, edgecolor=\"lime\", facecolor=\"none\")\n        ax.add_patch(rect)\n\n    # Boites predites (MAIRA-2) en rouge — coordonnees normalisees 0-1\n    for phrase, b in rec[\"pred_boxes\"]:\n        x_tl, y_tl, x_br, y_br = b\n        rect = mpatches.Rectangle((x_tl*W, y_tl*H), (x_br-x_tl)*W, (y_br-y_tl)*H,\n                                  linewidth=2.0, edgecolor=\"red\", facecolor=\"none\", linestyle=\"--\")\n        ax.add_patch(rect)\n\n    ok = \"OK\" if rec[\"true\"] == rec[\"pred\"] else \"X\"\n    ax.set_title(f\"[{ok}] vrai={rec['true']} | pred={rec['pred']}\"\n                 + (\" | boite patho\" if rec[\"box_on_patho\"] else \"\"),\n                 fontsize=10)\n    ax.axis(\"off\")\n\n# Legende + masquer axes inutilises\nfor ax in axes[n:]:\n    ax.axis(\"off\")\ngreen_patch = mpatches.Patch(color=\"lime\", label=\"Verite RSNA (opacite reelle)\")\nred_patch   = mpatches.Patch(color=\"red\", label=\"Prediction MAIRA-2\")\nfig.legend(handles=[green_patch, red_patch], loc=\"upper center\", ncol=2, fontsize=11)\nplt.tight_layout(rect=[0, 0, 1, 0.98])\nplt.savefig(\"/kaggle/working/maira2_visualisation.png\", dpi=110, bbox_inches=\"tight\")\nplt.show()\nprint(\"Figure sauvegardee : /kaggle/working/maira2_visualisation.png\")\n\n# Afficher aussi les rapports texte des images montrees (utile pour l'analyse)\nprint(\"\\n--- Rapports des images affichees ---\")\nfor rec in to_show:\n    print(f\"\\n[{rec['pid']}] vrai={rec['true']} pred={rec['pred']}\")\n    print(\"  \", rec[\"report\"][:280])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-05T10:46:15.415648Z","iopub.execute_input":"2026-07-05T10:46:15.415970Z","iopub.status.idle":"2026-07-05T10:46:19.872442Z","shell.execute_reply.started":"2026-07-05T10:46:15.415945Z","shell.execute_reply":"2026-07-05T10:46:19.871611Z"}},"outputs":[],"execution_count":null},{"id":"0ba9b975-d04b-4249-a5e4-9dab385bc3bb","cell_type":"markdown","source":"## 11. Export Excel pour analyse externe (par une autre IA)\n\nOn exporte, pour les 100 images de test, l'intégralité de ce que MAIRA-2 a produit : rapport brut, rapport nettoyé, chaque phrase avec sa/ses boîte(s), la classe dérivée, la vérité terrain, et les boîtes réelles RSNA. Le fichier comporte deux feuilles :\n\n- **Synthèse** : une ligne par image (vue d'ensemble, idéale pour un premier tri) ;\n- **Détail_boîtes** : une ligne par boîte prédite (phrase + coordonnées), pour l'analyse fine de la localisation.\n\nCe format est pensé pour être fourni à un autre modèle qui analysera les diagnostics.","metadata":{}},{"id":"9e3927eb-eefe-404f-8000-042e3fedd4eb","cell_type":"code","source":"import pandas as pd\n\n# ---------- Feuille 1 : synthese (1 ligne / image) ----------\nsynth_rows = []\nfor rec in viz_records:\n    pid = rec[\"pid\"]\n    # vérité terrain : boîtes réelles en texte\n    gt_boxes = []\n    for _, gr in rec[\"gt_rows\"].iterrows():\n        gt_boxes.append(f\"(x={int(gr['x'])},y={int(gr['y'])},w={int(gr['width'])},h={int(gr['height'])})\")\n    gt_boxes_str = \" ; \".join(gt_boxes) if gt_boxes else \"\"\n\n    # boîtes prédites en texte (phrase -> coords normalisées)\n    pred_boxes_str = \" ; \".join(\n        f\"[{phrase}] ({x1:.2f},{y1:.2f},{x2:.2f},{y2:.2f})\"\n        for phrase, (x1, y1, x2, y2) in rec[\"pred_boxes\"]\n    )\n\n    synth_rows.append({\n        \"patientId\": pid,\n        \"verite_terrain\": rec[\"true\"],\n        \"classe_predite\": rec[\"pred\"],\n        \"correct\": rec[\"true\"] == rec[\"pred\"],\n        \"boite_sur_pathologie\": rec[\"box_on_patho\"],\n        \"nb_boites_predites\": rec[\"n_pred_boxes\"],\n        \"nb_boites_reelles\": len(rec[\"gt_rows\"]),\n        \"rapport_maira\": rec[\"report\"],\n        \"rapport_brut\": rec[\"raw\"],\n        \"boites_predites\": pred_boxes_str,\n        \"boites_reelles_RSNA\": gt_boxes_str,\n    })\ndf_synth = pd.DataFrame(synth_rows)\n\n# ---------- Feuille 2 : detail des boites (1 ligne / boite predite) ----------\nbox_rows = []\nfor rec in viz_records:\n    if not rec[\"pred_boxes\"]:\n        box_rows.append({\n            \"patientId\": rec[\"pid\"], \"verite_terrain\": rec[\"true\"],\n            \"classe_predite\": rec[\"pred\"], \"phrase\": \"(aucune boite)\",\n            \"x_tl\": None, \"y_tl\": None, \"x_br\": None, \"y_br\": None,\n            \"quadrant\": None,\n        })\n        continue\n    for phrase, (x1, y1, x2, y2) in rec[\"pred_boxes\"]:\n        cx, cy = (x1+x2)/2, (y1+y2)/2\n        horiz = \"left\" if cx < 0.5 else \"right\"\n        vert = \"upper\" if cy < 1/3 else (\"middle\" if cy < 2/3 else \"lower\")\n        box_rows.append({\n            \"patientId\": rec[\"pid\"], \"verite_terrain\": rec[\"true\"],\n            \"classe_predite\": rec[\"pred\"], \"phrase\": phrase,\n            \"x_tl\": round(x1,3), \"y_tl\": round(y1,3),\n            \"x_br\": round(x2,3), \"y_br\": round(y2,3),\n            \"quadrant\": f\"{vert}-{horiz}\",\n        })\ndf_boxes_detail = pd.DataFrame(box_rows)\n\n# ---------- Feuille 3 : analyse d'erreurs (cas mal classes) ----------\ndf_err = df_synth[df_synth[\"correct\"] == False][\n    [\"patientId\", \"verite_terrain\", \"classe_predite\", \"boite_sur_pathologie\",\n     \"nb_boites_predites\", \"rapport_maira\", \"boites_predites\"]\n].copy()\n# Focus particulier : les uncertain versés à tort dans suspected_opacity\ndf_err[\"type_erreur\"] = df_err[\"verite_terrain\"] + \" -> \" + df_err[\"classe_predite\"]\ndf_err = df_err.sort_values(\"type_erreur\").reset_index(drop=True)\n\n# ---------- Ecriture du fichier Excel ----------\nout_path = \"/kaggle/working/maira2_analyse_complete.xlsx\"\nwith pd.ExcelWriter(out_path, engine=\"openpyxl\") as writer:\n    df_synth.to_excel(writer, sheet_name=\"Synthese\", index=False)\n    df_boxes_detail.to_excel(writer, sheet_name=\"Detail_boites\", index=False)\n    df_err.to_excel(writer, sheet_name=\"Analyse_erreurs\", index=False)\n\n    # Mise en forme legere : largeurs de colonnes + en-tetes en gras\n    from openpyxl.styles import Font, PatternFill, Alignment\n    header_fill = PatternFill(\"solid\", fgColor=\"1F3864\")\n    header_font = Font(bold=True, color=\"FFFFFF\")\n    for sheet_name, df in [(\"Synthese\", df_synth), (\"Detail_boites\", df_boxes_detail), (\"Analyse_erreurs\", df_err)]:\n        ws = writer.sheets[sheet_name]\n        for col_idx, col in enumerate(df.columns, start=1):\n            cell = ws.cell(row=1, column=col_idx)\n            cell.fill = header_fill\n            cell.font = header_font\n            cell.alignment = Alignment(horizontal=\"center\", vertical=\"center\")\n            # largeur : longue pour les colonnes de texte\n            width = 60 if col in (\"rapport_maira\", \"rapport_brut\", \"boites_predites\", \"boites_reelles_RSNA\", \"phrase\") else 18\n            from openpyxl.utils import get_column_letter\n            ws.column_dimensions[get_column_letter(col_idx)].width = width\n        ws.freeze_panes = \"A2\"  # fige la ligne d'en-tete\n\nprint(f\"Fichier Excel cree : {out_path}\")\nprint(f\"  Feuille 'Synthese'      : {len(df_synth)} lignes (1 par image)\")\nprint(f\"  Feuille 'Detail_boites' : {len(df_boxes_detail)} lignes (1 par boite)\")\nprint(f\"  Feuille 'Analyse_erreurs': {len(df_err)} lignes (cas mal classes)\")\nprint(\"\\nApercu synthese :\")\nprint(df_synth[[\"patientId\",\"verite_terrain\",\"classe_predite\",\"nb_boites_predites\"]].head(10).to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-05T10:46:19.873886Z","iopub.execute_input":"2026-07-05T10:46:19.874229Z","iopub.status.idle":"2026-07-05T10:46:20.378168Z","shell.execute_reply.started":"2026-07-05T10:46:19.874186Z","shell.execute_reply":"2026-07-05T10:46:20.377467Z"}},"outputs":[],"execution_count":null},{"id":"e2b3370a-08c8-42dc-bd18-a43a3ba4b1c9","cell_type":"markdown","source":"## 12. Analyse des termes déclencheurs (diagnostic de la sur-détection)\n\nLe principal défaut observé est que des cas `uncertain` sont classés `suspected_opacity`. Cette cellule identifie **quels termes d'opacité** déclenchent le plus ces erreurs, afin de savoir si un terme est trop large (candidat à retrait) ou si le problème est ailleurs (MAIRA-2 décrit réellement des findings sur ces images ambiguës).","metadata":{}},{"id":"73b7e51f-24cc-4d74-a645-87774333e11f","cell_type":"code","source":"from collections import Counter\n\n# Cas uncertain classes a tort en suspected_opacity\nmis = [rec for rec in viz_records\n       if rec[\"true\"] == \"uncertain\" and rec[\"pred\"] == \"suspected_opacity\"]\n\nprint(f\"Cas 'uncertain' -> 'suspected_opacity' : {len(mis)}\\n\")\n\ntrigger_counter = Counter()\nfor rec in mis:\n    report = rec[\"report\"]\n    for term in OPACITY_TERMS:\n        if term in report and not _negated(report, term):\n            trigger_counter[term] += 1\n\nprint(\"Termes d'opacite ayant declenche le faux positif (frequence) :\")\nfor term, cnt in trigger_counter.most_common():\n    print(f\"  {term:<25} : {cnt}\")\n\nprint(\"\\nExemples de rapports mal classes (5 premiers) :\")\nfor rec in mis[:5]:\n    print(f\"\\n[{rec['pid']}]\")\n    print(\"  \", rec[\"report\"][:260])\n\n# Part des erreurs venant d'une BOITE sur pathologie vs du TEXTE seul\nby_box = sum(1 for rec in mis if rec[\"box_on_patho\"])\nprint(f\"\\nParmi ces {len(mis)} erreurs : {by_box} declenchees par une BOITE sur phrase pathologique, \"\n      f\"{len(mis)-by_box} par le TEXTE seul.\")\nprint(\"=> Si beaucoup sont dues aux boites, durcir le filtre de boites ; \"\n      \"si dues au texte, MAIRA-2 decrit de vrais findings sur ces cas ambigus (limite intrinseque).\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-05T10:46:20.378970Z","iopub.execute_input":"2026-07-05T10:46:20.379526Z","iopub.status.idle":"2026-07-05T10:46:20.387936Z","shell.execute_reply.started":"2026-07-05T10:46:20.379500Z","shell.execute_reply":"2026-07-05T10:46:20.386908Z"}},"outputs":[],"execution_count":null},{"id":"bf4d53b3-b595-4d4d-a3b8-072288c7625c","cell_type":"markdown","source":"## 13. Classification secondaire par LLM (GPT-OSS-120B via OpenRouter)\n\nLa classification par règles atteint ses limites sur les cas ambigus. On délègue ici l'interprétation du rapport MAIRA-2 à un LLM de raisonnement médical, **GPT-OSS-120B** (fort sur HealthBench), via OpenRouter.\n\n**Principe.** Pour chaque image, on envoie au LLM le rapport ancré de MAIRA-2 (phrases + boîtes) et on lui demande de classer en `normal` / `suspected_opacity` / `uncertain`, en gérant les négations et les dispositifs médicaux. On compare ensuite à la vérité terrain, comme pour la classification par règles.\n\n> Nécessite un secret Kaggle `OPENROUTER_API_KEY`. On réutilise `viz_records` (déjà calculé) : **pas besoin de relancer MAIRA-2**, on ne fait qu'analyser ses sorties.","metadata":{}},{"id":"f3da239c-b027-4294-98b7-91ed9c7133d7","cell_type":"code","source":"# OpenRouter s'utilise via le SDK OpenAI standard (base_url personnalisee).\n!pip install -q openai 2>/dev/null\n\nfrom openai import OpenAI\nfrom kaggle_secrets import UserSecretsClient\n\nOPENROUTER_KEY = UserSecretsClient().get_secret(\"OPENROUTER_API_KEY\")\nclient = OpenAI(\n    base_url=\"https://openrouter.ai/api/v1\",\n    api_key=OPENROUTER_KEY,\n)\n\nLLM_MODEL = \"openai/gpt-oss-120b:free\"   # choix : fort raisonnement medical (HealthBench)\n\nSYSTEM_PROMPT = \"\"\"You are a radiology expert assisting an educational triage tool.\nYou receive the grounded findings of a frontal chest X-ray, as a list of\n(finding_description, bounding_boxes) produced by an upstream model.\n\nClassify the image into EXACTLY ONE of three classes:\n- \"suspected_opacity\": the findings describe a pulmonary opacity, consolidation,\n  infiltrate, pneumonia, effusion or similar airspace abnormality that is ASSERTED\n  (not negated).\n- \"normal\": the findings explicitly indicate clear lungs / no acute abnormality,\n  OR describe only medical devices (PICC line, ET tube, NG tube, catheter) without\n  any pulmonary abnormality.\n- \"uncertain\": the findings are ambiguous, non-specific, or you cannot confidently\n  decide between normal and opacity.\n\nImportant rules:\n- Treat negated findings (\"no focal infiltrate\", \"no effusion\") as NORMAL evidence,\n  not as opacity.\n- A bounding box on a device (line, tube, catheter) is NOT an opacity.\n- Prioritise patient safety: if a genuine airspace abnormality is asserted, prefer\n  \"suspected_opacity\" over \"uncertain\".\n\nReturn ONLY a valid JSON object, no prose:\n{\"predicted_class\": \"...\", \"confidence\": 0.0-1.0, \"reason\": \"short justification\"}\"\"\"\n\ndef parsed_to_prompt_text(parsed):\n    \"\"\"Formate la sortie MAIRA-2 (liste (phrase, boites)) en texte pour le LLM.\"\"\"\n    lines = []\n    for item in parsed if isinstance(parsed, (list, tuple)) else []:\n        if isinstance(item, (list, tuple)) and len(item) >= 1:\n            phrase = str(item[0])\n            boxes = item[1] if len(item) >= 2 else None\n            lines.append(f\"- {phrase}  boxes={boxes}\")\n    return \"\\n\".join(lines) if lines else \"(no findings)\"\n\nprint(\"Client OpenRouter pret. Modele :\", LLM_MODEL)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-05T10:46:20.388872Z","iopub.execute_input":"2026-07-05T10:46:20.389637Z","iopub.status.idle":"2026-07-05T10:46:25.809648Z","shell.execute_reply.started":"2026-07-05T10:46:20.389607Z","shell.execute_reply":"2026-07-05T10:46:25.808858Z"}},"outputs":[],"execution_count":null},{"id":"70022c62-5d55-4cde-b11c-1109deb4bb8e","cell_type":"code","source":"import json as _json\nimport time\n\ndef classify_with_llm(parsed, max_retries=5):\n    \"\"\"Envoie le rapport MAIRA-2 au LLM. Retourne (classe, confiance, raison, success).\n    - success=True  : le LLM a bien repondu (prediction fiable)\n    - success=False : echec (429 apres retries, ou autre erreur) -> A EXCLURE des metriques\n    Gestion du rate limit : retry avec back-off exponentiel (3, 6, 12, 24, 48 s).\"\"\"\n    user_text = parsed_to_prompt_text(parsed)\n    for attempt in range(max_retries):\n        try:\n            resp = client.chat.completions.create(\n                model=LLM_MODEL,\n                messages=[\n                    {\"role\": \"system\", \"content\": SYSTEM_PROMPT},\n                    {\"role\": \"user\", \"content\": user_text},\n                ],\n                temperature=0,\n                max_tokens=300,\n            )\n            content = resp.choices[0].message.content.strip()\n            content = content.replace(\"```json\", \"\").replace(\"```\", \"\").strip()\n            s, e = content.find(\"{\"), content.rfind(\"}\") + 1\n            obj = _json.loads(content[s:e])\n            cls = obj.get(\"predicted_class\", \"uncertain\")\n            if cls not in {\"normal\", \"suspected_opacity\", \"uncertain\"}:\n                cls = \"uncertain\"\n            return cls, obj.get(\"confidence\", 0.0), obj.get(\"reason\", \"\"), True\n        except Exception as ex:\n            msg = str(ex)\n            if \"429\" in msg or \"rate\" in msg.lower():\n                wait = (2 ** attempt) * 3   # 3, 6, 12, 24, 48 s\n                time.sleep(wait)\n                continue\n            # erreur non liee au rate limit : inutile d'insister\n            return \"uncertain\", 0.0, f\"error: {ex}\", False\n    # rate limit persistant apres tous les retries\n    return \"uncertain\", 0.0, \"error: rate limit after retries\", False\n\n# --- Test sur un rapport avant de lancer les 100 ---\ndemo_rec = viz_records[0]\ncls_llm, conf_llm, reason_llm, ok_llm = classify_with_llm(\n    [(phrase, box) for phrase, box in demo_rec[\"pred_boxes\"]]\n)\nprint(\"Image      :\", demo_rec[\"pid\"])\nprint(\"Verite     :\", demo_rec[\"true\"])\nprint(\"Regles     :\", demo_rec[\"pred\"])\nprint(\"LLM        :\", cls_llm, f\"(conf={conf_llm}, success={ok_llm})\")\nprint(\"Raison LLM :\", reason_llm)\nif not ok_llm:\n    print(\"\\\\n[!] Le test a echoue (rate limit probable). Si ca persiste sur la boucle,\")\n    print(\"    ajoute une cle OpenRouter payante (quelques $) et relance.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-05T10:46:25.810877Z","iopub.execute_input":"2026-07-05T10:46:25.811103Z","iopub.status.idle":"2026-07-05T10:46:30.563953Z","shell.execute_reply.started":"2026-07-05T10:46:25.811075Z","shell.execute_reply":"2026-07-05T10:46:30.563260Z"}},"outputs":[],"execution_count":null},{"id":"1ef21c36-24d0-4a09-b6d9-dc5c7314ab92","cell_type":"code","source":"from sklearn.metrics import accuracy_score, f1_score, confusion_matrix, classification_report\nfrom tqdm.auto import tqdm\nimport time\n\nllm_preds, llm_trues, llm_reasons, llm_ok = [], [], [], []\n\nfor rec in tqdm(viz_records, total=len(viz_records)):\n    parsed_like = [(phrase, box) for phrase, box in rec[\"pred_boxes\"]]\n    cls, conf, reason, ok = classify_with_llm(\n        parsed_like if parsed_like else [(rec[\"report\"], None)]\n    )\n    llm_preds.append(cls); llm_trues.append(rec[\"true\"])\n    llm_reasons.append(reason); llm_ok.append(ok)\n    time.sleep(2.0)   # tier gratuit : espacement prudent (Levier 2 partiel)\n\n# --- Compteur d'echecs : ESSENTIEL pour juger la fiabilite ---\nmask = np.array(llm_ok)\nn_ok, n_fail = int(mask.sum()), int((~mask).sum())\nprint(f\"Appels reussis : {n_ok} / {len(mask)}   |   echecs (exclus) : {n_fail}\")\nif n_fail > 5:\n    print(\"[!] Beaucoup d'echecs : metriques calculees sur un sous-ensemble reduit.\")\n    print(\"    Pour un run fiable, ajoute une cle OpenRouter payante et relance.\\\\n\")\n\n# --- Metriques UNIQUEMENT sur les appels reussis (les echecs ne sont PAS des predictions) ---\nvalid_preds = [p for p, o in zip(llm_preds, llm_ok) if o]\nvalid_trues = [t for t, o in zip(llm_trues, llm_ok) if o]\n\nlabels = [\"normal\", \"suspected_opacity\", \"uncertain\"]\nif len(valid_preds) > 0:\n    acc = accuracy_score(valid_trues, valid_preds)\n    f1  = f1_score(valid_trues, valid_preds, labels=labels, average=\"macro\")\n    print(\"=== CLASSIFICATION SECONDAIRE (GPT-OSS-120B) — sur appels reussis uniquement ===\")\n    print(f\"Base : {len(valid_preds)} images   Accuracy : {acc:.4f}   Macro-F1 : {f1:.4f}\")\n    print(classification_report(valid_trues, valid_preds, labels=labels, zero_division=0))\n    cmn = confusion_matrix(valid_trues, valid_preds, labels=labels)\n    print(pd.DataFrame(cmn, index=labels, columns=labels))\n    op_i = labels.index(\"suspected_opacity\")\n    if cmn[op_i].sum():\n        print(f\"\\\\n>>> Rappel suspected_opacity (LLM) : {cmn[op_i,op_i]/cmn[op_i].sum():.4f}\")\n\n    # Comparaison regles vs LLM sur les MEMES images reussies\n    rules_valid = [rec[\"pred\"] for rec, o in zip(viz_records, llm_ok) if o]\n    agree = np.mean([a == b for a, b in zip(rules_valid, valid_preds)])\n    print(f\"\\\\nAccord regles vs LLM (sur {len(valid_preds)} images) : {agree:.2%}\")\nelse:\n    print(\"Aucun appel reussi : impossible de calculer les metriques. Ajoute une cle payante.\")\n\n# --- Sauvegarde complete (avec flag succes pour tracabilite) ---\npd.DataFrame({\n    \"patientId\": [r[\"pid\"] for r in viz_records],\n    \"verite\": llm_trues,\n    \"pred_regles\": [r[\"pred\"] for r in viz_records],\n    \"pred_LLM\": llm_preds,\n    \"appel_reussi\": llm_ok,\n    \"raison_LLM\": llm_reasons,\n    \"rapport\": [r[\"report\"] for r in viz_records],\n}).to_csv(\"/kaggle/working/maira2_llm_comparison.csv\", index=False)\nprint(\"\\\\nSauvegarde : /kaggle/working/maira2_llm_comparison.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-05T10:46:30.565046Z","iopub.execute_input":"2026-07-05T10:46:30.565360Z","execution_failed":"2026-07-05T16:37:12.951Z"}},"outputs":[],"execution_count":null},{"id":"bb20256d-8eaa-402c-ae27-392bfa78618b","cell_type":"markdown","source":"## 14. Export final pour l'interface web (CSV complet avec boîtes)\n\nCette cellule produit `maira2_interface.csv`, pensé pour l'application web de secours.\nIl contient, pour chaque image de la banque :\n- la **classe finale** (celle du LLM GPT-OSS si l'appel a réussi, sinon celle des règles) ;\n- les deux classes séparées (règles et LLM) pour référence ;\n- les **boîtes prédites** par MAIRA-2, en pixels sur une image 1024×1024 (format standard RSNA) ;\n- les **boîtes réelles** RSNA (vérité terrain), au même format ;\n- le rapport MAIRA-2 et la vérité terrain.\n\nLes boîtes sont stockées en JSON dans les cellules, pour être relues facilement par l'interface.","metadata":{}},{"id":"2e6aee57-c416-4f53-895f-18c5bc560c97","cell_type":"code","source":"import json as _json\n\n# Dimension de reference pour l'interface (RSNA = 1024x1024)\nREF = 1024\n\ndef pred_boxes_to_pixels(pred_boxes):\n    \"\"\"(phrase, boite normalisee 0-1) -> liste de dict pixels {x1,y1,x2,y2,phrase}.\"\"\"\n    out = []\n    for phrase, (x1, y1, x2, y2) in pred_boxes:\n        out.append({\n            \"x1\": round(x1 * REF), \"y1\": round(y1 * REF),\n            \"x2\": round(x2 * REF), \"y2\": round(y2 * REF),\n            \"phrase\": phrase,\n        })\n    return out\n\ndef gt_rows_to_pixels(gt_rows):\n    \"\"\"Boites RSNA (x,y,w,h en pixels 1024) -> liste de dict {x1,y1,x2,y2}.\"\"\"\n    out = []\n    for _, r in gt_rows.iterrows():\n        out.append({\n            \"x1\": int(r[\"x\"]), \"y1\": int(r[\"y\"]),\n            \"x2\": int(r[\"x\"] + r[\"width\"]), \"y2\": int(r[\"y\"] + r[\"height\"]),\n        })\n    return out\n\n# Construire le mapping patientId -> classe LLM (si dispo et appel reussi)\nllm_by_pid = {}\ntry:\n    for pid, pred, ok in zip([r[\"pid\"] for r in viz_records], llm_preds, llm_ok):\n        llm_by_pid[pid] = pred if ok else None\nexcept NameError:\n    # Si la section 13 (LLM) n'a pas ete executee, on se rabat sur les regles seules\n    print(\"[i] Section LLM non executee : classe finale = regles uniquement.\")\n\nrows = []\nfor rec in viz_records:\n    pid = rec[\"pid\"]\n    cls_rules = rec[\"pred\"]\n    cls_llm = llm_by_pid.get(pid)  # None si echec ou section LLM non lancee\n    # Classe finale : priorite au LLM s'il a repondu, sinon regles\n    cls_final = cls_llm if cls_llm is not None else cls_rules\n\n    rows.append({\n        \"patientId\": pid,\n        \"classe_finale\": cls_final,\n        \"classe_regles\": cls_rules,\n        \"classe_LLM\": cls_llm if cls_llm is not None else \"\",\n        \"verite_terrain\": rec[\"true\"],\n        \"boites_predites\": _json.dumps(pred_boxes_to_pixels(rec[\"pred_boxes\"])),\n        \"boites_reelles\": _json.dumps(gt_rows_to_pixels(rec[\"gt_rows\"])),\n        \"rapport_maira\": rec[\"report\"],\n        \"ref_size\": REF,\n    })\n\ndf_interface = pd.DataFrame(rows)\nout_csv = \"/kaggle/working/maira2_interface.csv\"\ndf_interface.to_csv(out_csv, index=False)\nprint(f\"Export interface : {out_csv}  ({len(df_interface)} images)\")\nprint(\"\\nColonnes :\", list(df_interface.columns))\nprint(\"\\nApercu (2 premieres lignes) :\")\nfor _, r in df_interface.head(2).iterrows():\n    print(f\"\\n{r['patientId']} | finale={r['classe_finale']} | verite={r['verite_terrain']}\")\n    print(\"  boites predites :\", r[\"boites_predites\"][:150])\n    print(\"  boites reelles  :\", r[\"boites_reelles\"][:150])","metadata":{"trusted":true,"execution":{"execution_failed":"2026-07-05T16:37:12.958Z"}},"outputs":[],"execution_count":null}]}