{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🧠 PIPELINE D'ANALYSE RADIOGÉNOMIQUE DES TUMEURS CÉRÉBRALES\n\n## 📌 Contexte du projet\n\nCe notebook s'inscrit dans le cadre de la compétition **RSNA-MICCAI Brain Tumor Radiogenomic Classification**, dont l'objectif est de prédire le **statut de méthylation du promoteur MGMT** chez les patients atteints de glioblastome, à partir d'images IRM multiparamétriques.\n\n### 🎯 Objectif médical\n- **Prédire** si le gène MGMT est méthylé (1) ou non (0)\n- **Éviter** les biopsies invasives\n- **Aider** les médecins à personnaliser les traitements\n\n### 📊 Données utilisées\n- **585 patients** en entraînement\n- **87 patients** en test\n- **4 séquences IRM** par patient :\n  - **FLAIR** : œdème, infiltration tumorale\n  - **T1w** : anatomie cérébrale\n  - **T1wCE** : prise de contraste, tumeur active\n  - **T2w** : liquide, nécrose\n- **Dossiers DICOM** contenant les volumes 3D\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"# 🧠 Radiogenomics Pipeline Design\n## RSNA-MICCAI Brain Tumor Radiogenomic Classification\n\n```mermaid\nflowchart TD\n\nA[📥 MRI Dataset<br/>RSNA-MICCAI] --> B[Patient Data]\n\nB --> C1[FLAIR]\nB --> C2[T1w]\nB --> C3[T1wCE]\nB --> C4[T2w]\n\nC1 --> D\nC2 --> D\nC3 --> D\nC4 --> D\n\nsubgraph Preprocessing\nD[Load DICOM Images]\nD --> E[Resize & Resampling]\nE --> F[Intensity Normalization]\nF --> G[Slice Selection]\nG --> H[Data Augmentation]\nend\n\nH --> I[2.5D Dataset]\n\nsubgraph Deep Learning Model\nI --> J[EfficientNet-B0 Backbone]\nJ --> K[Feature Extraction]\nK --> L[Attention Pooling]\nL --> M[Fully Connected Layer]\nM --> N[MGMT Prediction]\nend\n\nsubgraph Training\nN --> O[Cross Validation]\nO --> P[Loss Function]\nP --> Q[Optimizer]\nQ --> R[Early Stopping]\nend\n\nR --> S[Best Model]\n\nsubgraph Evaluation\nS --> T[AUC ROC]\nS --> U[Accuracy]\nS --> V[Confusion Matrix]\nend\n```\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"# 1-INSTALLATION DES PACKAGES\n#### Cette cellule installe tous les packages Python nécessaires au bon fonctionnement du pipeline. Elle s'exécute en premier et ne doit être exécutée qu'une seule fois.","metadata":{}},{"cell_type":"code","source":"!pip install -q timm scikit-image medpy opencv-python-headless\n\nimport os\nimport gc\nimport glob\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm import tqdm\nimport warnings\nimport timm\nimport time\nwarnings.filterwarnings('ignore')\n\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import DataLoader, Dataset\nimport torch.nn.functional as F\nfrom torch.cuda.amp import autocast, GradScaler\n\nimport pydicom\nfrom scipy.ndimage import zoom\nfrom skimage.transform import resize\nimport cv2\n\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import roc_auc_score, roc_curve\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:27:06.572524Z","iopub.execute_input":"2026-07-10T20:27:06.57281Z","iopub.status.idle":"2026-07-10T20:27:10.041842Z","shell.execute_reply.started":"2026-07-10T20:27:06.572782Z","shell.execute_reply":"2026-07-10T20:27:10.040778Z"}},"outputs":[],"execution_count":25},{"cell_type":"markdown","source":"# 2-CHARGEMENT DES DONNÉES\n#### Cette cellule charge les données de la compétition RSNA-MICCAI à partir du Kaggle. Elle vérifie que les fichiers sont présents et affiche un aperçu des données.","metadata":{}},{"cell_type":"code","source":"DATA_PATH = '/kaggle/input/competitions/rsna-miccai-brain-tumor-radiogenomic-classification'\n\nprint(\"=\"*60)\nprint(\"📂 CHARGEMENT DES DONNÉES\")\nprint(\"=\"*60)\n\nif os.path.exists(DATA_PATH):\n    print(f\"✅ Chemin trouvé: {DATA_PATH}\")\n    print(f\"📁 Contenu: {os.listdir(DATA_PATH)}\")\nelse:\n    print(f\"❌ Chemin non trouvé: {DATA_PATH}\")\n    DATA_PATH = '/kaggle/input/rsna-miccai-brain-tumor-radiogenomic-classification'\n    if os.path.exists(DATA_PATH):\n        print(f\"✅ Chemin alternatif trouvé: {DATA_PATH}\")\n        print(f\"📁 Contenu: {os.listdir(DATA_PATH)}\")\n    else:\n        raise FileNotFoundError(\"❌ Impossible de trouver les données\")\n\nprint(\"\\n Chargement des fichiers CSV...\")\ntrain_df = pd.read_csv(os.path.join(DATA_PATH, 'train_labels.csv'))\ntest_df = pd.read_csv(os.path.join(DATA_PATH, 'sample_submission.csv'))\nprint(f\"✅ Train: {len(train_df)} patients\")\nprint(f\"✅ Test: {len(test_df)} patients\")\n\nprint(\"\\n📊 Aperçu des données d'entraînement:\")\nprint(train_df.head())\n\nprint(\"\\n📊 Distribution MGMT:\")\nprint(train_df['MGMT_value'].value_counts())\nprint(f\"Ratio de méthylation: {train_df['MGMT_value'].mean():.3f}\")\n\nprint(\"\\n🔍 Vérification des IDs:\")\nprint(f\"IDs uniques dans train: {train_df['BraTS21ID'].nunique()}\")\nprint(f\"IDs uniques dans test: {test_df['BraTS21ID'].nunique()}\")\n\ncorrupt_ids = ['00109', '00123', '00709']\ntrain_df['is_corrupt'] = train_df['BraTS21ID'].astype(str).isin(corrupt_ids)\n\nprint(\"\\n📊 Statistiques rapides:\")\nprint(f\"MGMT=0 (non-méthylé): {sum(train_df['MGMT_value'] == 0)} patients\")\nprint(f\"MGMT=1 (méthylé): {sum(train_df['MGMT_value'] == 1)} patients\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:27:12.349681Z","iopub.execute_input":"2026-07-10T20:27:12.350519Z","iopub.status.idle":"2026-07-10T20:27:12.375038Z","shell.execute_reply.started":"2026-07-10T20:27:12.350481Z","shell.execute_reply":"2026-07-10T20:27:12.37427Z"}},"outputs":[{"name":"stdout","text":"============================================================\n📂 CHARGEMENT DES DONNÉES\n============================================================\n✅ Chemin trouvé: /kaggle/input/competitions/rsna-miccai-brain-tumor-radiogenomic-classification\n📁 Contenu: ['sample_submission.csv', 'train_labels.csv', 'test', 'train']\n\n Chargement des fichiers CSV...\n✅ Train: 585 patients\n✅ Test: 87 patients\n\n📊 Aperçu des données d'entraînement:\n   BraTS21ID  MGMT_value\n0          0           1\n1          2           1\n2          3           0\n3          5           1\n4          6           1\n\n📊 Distribution MGMT:\nMGMT_value\n1    307\n0    278\nName: count, dtype: int64\nRatio de méthylation: 0.525\n\n🔍 Vérification des IDs:\nIDs uniques dans train: 585\nIDs uniques dans test: 87\n\n📊 Statistiques rapides:\nMGMT=0 (non-méthylé): 278 patients\nMGMT=1 (méthylé): 307 patients\n","output_type":"stream"}],"execution_count":26},{"cell_type":"markdown","source":"# 3-CONFIGURATION\n#### Cette cellule centralise tous les paramètres et hyperparamètres du pipeline. Elle permet de modifier facilement la configuration sans avoir à chercher dans tout le code.","metadata":{}},{"cell_type":"code","source":"print(\"=\"*60)\nprint(\"⚙️ CONFIGURATION DU PIPELINE\")\nprint(\"=\"*60)\n\nclass Config: \n    SEED = 42\n    \n    N_FOLDS = 2\n    \n    EPOCHS = 30\n    BATCH_SIZE = 16\n    LR = 1e-3\n    LR_MIN = 1e-6\n    Learnig_Rate = 0.001\n    EARLY_STOPPING = 10\n    WEIGHT_DECAY = 1e-4 \n    \n    N_SLICES = 16\n    IMG_SIZE = 224\n    DROPOUT = 0.3\n    CORRUPTED_IDS = ['00109', '00123', '00709']\n    \n    USE_MIXED_PRECISION = True\n    USE_TTA = False\n    N_WORKERS = 0\n\nconfig = Config()\n\nprint(\"\\n📋 Configuration actuelle:\")\nprint(\"-\"*40)\nprint(f\"🔢 SEED                 : {config.SEED}\")\nprint(f\"📊 N_FOLDS              : {config.N_FOLDS}\")\nprint(f\"🔄 EPOCHS               : {config.EPOCHS}\")\nprint(f\"📦 BATCH_SIZE           : {config.BATCH_SIZE}\")\nprint(f\"📈 LR                   : {config.Learnig_Rate}\")\nprint(f\"⏹️ EARLY_STOPPING       : {config.EARLY_STOPPING}\")\nprint(f\"📐 N_SLICES             : {config.N_SLICES}\")\nprint(f\"🖼️ IMG_SIZE             : {config.IMG_SIZE}\")\nprint(f\"🎯 DROPOUT              : {config.DROPOUT}\")\nprint(f\"⚠️ CORRUPTED_IDS        : {config.CORRUPTED_IDS}\")\nprint(\"-\"*40)\n\nnp.random.seed(config.SEED)\ntorch.manual_seed(config.SEED)\nif torch.cuda.is_available():\n    torch.cuda.manual_seed_all(config.SEED)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\nprint(\"✅ Seeds fixées\")\n\nprint(f\"\\n💻 GPU disponible: {torch.cuda.is_available()}\")\nif torch.cuda.is_available():\n    print(f\"   GPU: {torch.cuda.get_device_name(0)}\")\n    print(f\"   Mémoire totale: {torch.cuda.get_device_properties(0).total_memory / 1e9:.2f} GB\")\n    print(f\"   Mémoire utilisée: {torch.cuda.memory_allocated(0) / 1e9:.2f} GB\")\nelse:\n    print(\"   ⚠️ GPU non disponible - Utilisation du CPU (plus lent)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:27:16.208513Z","iopub.execute_input":"2026-07-10T20:27:16.208914Z","iopub.status.idle":"2026-07-10T20:27:16.227202Z","shell.execute_reply.started":"2026-07-10T20:27:16.208885Z","shell.execute_reply":"2026-07-10T20:27:16.226089Z"}},"outputs":[{"name":"stdout","text":"============================================================\n⚙️ CONFIGURATION DU PIPELINE\n============================================================\n\n📋 Configuration actuelle:\n----------------------------------------\n🔢 SEED                 : 42\n📊 N_FOLDS              : 2\n🔄 EPOCHS               : 30\n📦 BATCH_SIZE           : 16\n📈 LR                   : 0.001\n⏹️ EARLY_STOPPING       : 10\n📐 N_SLICES             : 16\n🖼️ IMG_SIZE             : 224\n🎯 DROPOUT              : 0.3\n⚠️ CORRUPTED_IDS        : ['00109', '00123', '00709']\n----------------------------------------\n","output_type":"stream"},{"traceback":["\u001b[0;31m---------------------------------------------------------------------------\u001b[0m","\u001b[0;31mAcceleratorError\u001b[0m                          Traceback (most recent call last)","\u001b[0;32m/tmp/ipykernel_58/2155590650.py\u001b[0m in \u001b[0;36m<cell line: 0>\u001b[0;34m()\u001b[0m\n\u001b[1;32m     42\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     43\u001b[0m \u001b[0mnp\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mrandom\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mconfig\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mSEED\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 44\u001b[0;31m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mconfig\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mSEED\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     45\u001b[0m \u001b[0;32mif\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mis_available\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     46\u001b[0m     \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed_all\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mconfig\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mSEED\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/_compile.py\u001b[0m in \u001b[0;36minner\u001b[0;34m(*args, **kwargs)\u001b[0m\n\u001b[1;32m     52\u001b[0m                 \u001b[0mfn\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0m__dynamo_disable\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0mdisable_fn\u001b[0m  \u001b[0;31m# type: ignore[attr-defined]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     53\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 54\u001b[0;31m             \u001b[0;32mreturn\u001b[0m \u001b[0mdisable_fn\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m*\u001b[0m\u001b[0margs\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;34m**\u001b[0m\u001b[0mkwargs\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     55\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     56\u001b[0m         \u001b[0;32mreturn\u001b[0m \u001b[0minner\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/_dynamo/eval_frame.py\u001b[0m in \u001b[0;36m_fn\u001b[0;34m(*args, **kwargs)\u001b[0m\n\u001b[1;32m   1179\u001b[0m                         ):\n\u001b[1;32m   1180\u001b[0m                             \u001b[0;32mreturn\u001b[0m \u001b[0mfn\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m*\u001b[0m\u001b[0margs\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;34m**\u001b[0m\u001b[0mkwargs\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m-> 1181\u001b[0;31m                     \u001b[0;32mreturn\u001b[0m \u001b[0mfn\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m*\u001b[0m\u001b[0margs\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;34m**\u001b[0m\u001b[0mkwargs\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m   1182\u001b[0m                 \u001b[0;32mfinally\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m   1183\u001b[0m                     \u001b[0mset_eval_frame\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;32mNone\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/random.py\u001b[0m in \u001b[0;36mmanual_seed\u001b[0;34m(seed)\u001b[0m\n\u001b[1;32m     40\u001b[0m             \u001b[0;31m`\u001b[0m\u001b[0;36m0xffff_ffff_ffff_ffff\u001b[0m \u001b[0;34m+\u001b[0m \u001b[0mseed\u001b[0m\u001b[0;31m`\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     41\u001b[0m     \"\"\"\n\u001b[0;32m---> 42\u001b[0;31m     \u001b[0;32mreturn\u001b[0m \u001b[0m_manual_seed_impl\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     43\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     44\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/random.py\u001b[0m in \u001b[0;36m_manual_seed_impl\u001b[0;34m(seed)\u001b[0m\n\u001b[1;32m     48\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     49\u001b[0m     \u001b[0;32mif\u001b[0m \u001b[0;32mnot\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0m_is_in_bad_fork\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 50\u001b[0;31m         \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed_all\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     51\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     52\u001b[0m     \u001b[0;32mimport\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmps\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/cuda/random.py\u001b[0m in \u001b[0;36mmanual_seed_all\u001b[0;34m(seed)\u001b[0m\n\u001b[1;32m    129\u001b[0m             \u001b[0mdefault_generator\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    130\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 131\u001b[0;31m     \u001b[0m_lazy_call\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mcb\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mseed_all\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mTrue\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    132\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    133\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py\u001b[0m in \u001b[0;36m_lazy_call\u001b[0;34m(callable, **kwargs)\u001b[0m\n\u001b[1;32m    353\u001b[0m     \u001b[0;32mwith\u001b[0m \u001b[0m_initialization_lock\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    354\u001b[0m         \u001b[0;32mif\u001b[0m \u001b[0mis_initialized\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 355\u001b[0;31m             \u001b[0mcallable\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    356\u001b[0m         \u001b[0;32melse\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    357\u001b[0m             \u001b[0;31m# TODO(torch_deploy): this accesses linecache, which attempts to read the\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/cuda/random.py\u001b[0m in \u001b[0;36mcb\u001b[0;34m()\u001b[0m\n\u001b[1;32m    127\u001b[0m         \u001b[0;32mfor\u001b[0m \u001b[0mi\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mrange\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mdevice_count\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    128\u001b[0m             \u001b[0mdefault_generator\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mdefault_generators\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mi\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 129\u001b[0;31m             \u001b[0mdefault_generator\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    130\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    131\u001b[0m     \u001b[0m_lazy_call\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mcb\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mseed_all\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mTrue\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;31mAcceleratorError\u001b[0m: CUDA error: device-side assert triggered\nSearch for `cudaErrorAssert' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nCUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.\nFor debugging consider passing CUDA_LAUNCH_BLOCKING=1\nCompile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.\n"],"ename":"AcceleratorError","evalue":"CUDA error: device-side assert triggered\nSearch for `cudaErrorAssert' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nCUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.\nFor debugging consider passing CUDA_LAUNCH_BLOCKING=1\nCompile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.\n","output_type":"error"}],"execution_count":27},{"cell_type":"markdown","source":"# 4-FONCTIONS DE CHARGEMENT DICOM\n#### Cette cellule contient toutes les fonctions nécessaires pour charger et prétraiter les images DICOM.","metadata":{}},{"cell_type":"code","source":"print(\"=\"*60)\nprint(\"🔧 FONCTIONS DE CHARGEMENT DICOM\")\nprint(\"=\"*60)\n\ndef load_dicom_volume(patient_path, sequence, target_size=128, verbose=False):\n    if not os.path.exists(patient_path):\n        if verbose:\n            print(f\"⚠️ Dossier patient introuvable: {patient_path}\")\n        return None\n\n    patterns = [\n        f\"{patient_path}/*/*{sequence}*/*.dcm\",\n        f\"{patient_path}/*{sequence}*/*.dcm\",\n        f\"{patient_path}/{sequence}/*.dcm\",\n    ]\n    \n    dcm_files = []\n    for pattern in patterns:\n        dcm_files = glob.glob(pattern)\n        if dcm_files:\n            if verbose:\n                print(f\"✅ Pattern trouvé: {pattern}\")\n            break\n    \n    if not dcm_files:\n        if verbose:\n            print(f\"❌ Aucun fichier DICOM trouvé pour {sequence}\")\n        return None\n    \n    dcm_files = sorted(dcm_files)[:80]\n    \n    if verbose:\n        print(f\"📄 {len(dcm_files)} fichiers DICOM trouvés pour {sequence}\")\n    \n    slices = []\n    failed_files = 0\n    \n    for f in dcm_files:\n        try:\n            ds = pydicom.dcmread(f, force=True)\n            \n            if hasattr(ds, 'pixel_array'):\n                pixel = ds.pixel_array.astype(np.float32)\n                if pixel.max() > 0:\n                    pixel = pixel / pixel.max()\n                if pixel.shape[0] != target_size:\n                    pixel = resize(pixel, (target_size, target_size), \n                                 preserve_range=True, anti_aliasing=True)\n                \n                slices.append(pixel)\n            else:\n                failed_files += 1\n                \n        except Exception as e:\n            failed_files += 1\n            if verbose:\n                print(f\"⚠️ Erreur lecture {f}: {str(e)[:50]}\")\n            continue\n    \n    if not slices:\n        if verbose:\n            print(f\"❌ Aucune slice valide chargée pour {sequence}\")\n        return None\n    \n    if failed_files > 0 and verbose:\n        print(f\"⚠️ {failed_files} fichiers ont échoué\")\n    \n    volume = np.stack(slices, axis=0)\n    if verbose:\n        print(f\"📊 Volume créé: {volume.shape}\")\n    p1, p99 = np.percentile(volume, [1, 99])\n    if p99 > p1:\n        volume = (volume - p1) / (p99 - p1 + 1e-8)\n        volume = np.clip(volume, 0, 1)\n    if verbose:\n        print(f\"✅ Volume normalisé: min={volume.min():.3f}, max={volume.max():.3f}\")\n    return volume\n\ndef select_best_slices(volume, n_slices=16):\n    if volume is None or volume.ndim != 3:\n        return None\n    if volume.shape[0] <= n_slices:\n        selected = volume\n        if selected.shape[0] < n_slices:\n            pad = n_slices - selected.shape[0]\n            pad_width = ((0, pad), (0, 0), (0, 0))\n            selected = np.pad(selected, pad_width, mode='edge')\n        return selected   \n    scores = []\n    for i in range(volume.shape[0]):\n        s = volume[i]\n        if s.std() < 0.005:\n            continue\n        variance = s.std()\n        s_flat = s.flatten()\n        hist, _ = np.histogram(s_flat, bins=32, range=(0, 1))\n        hist = hist / (hist.sum() + 1e-8)\n        entropy = -np.sum(hist * np.log(hist + 1e-8))\n        entropy = entropy / np.log(32) \n        contrast = np.abs(s - np.median(s)).mean()\n        score = 0.5 * variance + 0.3 * entropy + 0.2 * contrast\n        scores.append((i, score))\n        \n    if not scores:\n        mid = volume.shape[0] // 2\n        start = max(0, mid - n_slices // 2)\n        indices = list(range(start, min(volume.shape[0], start + n_slices)))\n    else:\n        scores.sort(key=lambda x: x[1], reverse=True)\n        indices = sorted([s[0] for s in scores[:n_slices]])\n    \n    selected = volume[indices]\n    \n    if selected.shape[0] < n_slices:\n        pad = n_slices - selected.shape[0]\n        pad_width = ((0, pad), (0, 0), (0, 0))\n        selected = np.pad(selected, pad_width, mode='edge')\n    \n    return selected\n\nprint(\"\\n🔍 Test des fonctions de chargement:\")\ntest_patient_path = os.path.join(DATA_PATH, 'train', '00000')\ntest_sequence = 'FLAIR'\n\nif os.path.exists(test_patient_path):\n    print(f\"✅ Patient de test trouvé: {test_patient_path}\")\n    volume = load_dicom_volume(test_patient_path, test_sequence, verbose=True)\n    \n    if volume is not None:\n        print(f\"✅ Volume chargé: {volume.shape}\")\n        slices = select_best_slices(volume, n_slices=16)\n        if slices is not None:\n            print(f\"✅ Slices sélectionnés: {slices.shape}\")\n            print(f\"   Stats: min={slices.min():.3f}, max={slices.max():.3f}, mean={slices.mean():.3f}\")\n        else:\n            print(\"❌ Échec sélection des slices\")\n    else:\n        print(\"❌ Échec chargement du volume\")\nelse:\n    print(\"❌ Patient de test non trouvé\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T19:59:57.472011Z","iopub.execute_input":"2026-07-10T19:59:57.472444Z","iopub.status.idle":"2026-07-10T20:00:00.137728Z","shell.execute_reply.started":"2026-07-10T19:59:57.472414Z","shell.execute_reply":"2026-07-10T20:00:00.136939Z"}},"outputs":[{"name":"stdout","text":"============================================================\n🔧 FONCTIONS DE CHARGEMENT DICOM\n============================================================\n\n🔍 Test des fonctions de chargement:\n✅ Patient de test trouvé: /kaggle/input/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/train/00000\n✅ Pattern trouvé: /kaggle/input/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/train/00000/*FLAIR*/*.dcm\n📄 80 fichiers DICOM trouvés pour FLAIR\n📊 Volume créé: (80, 128, 128)\n✅ Volume normalisé: min=0.000, max=1.000\n✅ Volume chargé: (80, 128, 128)\n✅ Slices sélectionnés: (16, 128, 128)\n   Stats: min=0.000, max=1.000, mean=0.154\n","output_type":"stream"}],"execution_count":5},{"cell_type":"markdown","source":"# 5-DATASET 2.5D\n#### Cette cellule définit la classe Dataset2_5D qui prépare les données pour l'entraînement du modèle. Elle transforme les volumes 3D en séries de slices 2D avec leurs labels.","metadata":{}},{"cell_type":"code","source":"print(\"=\"*60)\nprint(\"📊 CRÉATION DU DATASET 2.5D\")\nprint(\"=\"*60)\n\nclass Dataset2_5D(Dataset):\n    \n    def __init__(self, df, mode='train', sequence='FLAIR', augment=True, verbose=False):\n        self.df = df\n        self.mode = mode\n        self.sequence = sequence\n        self.augment = augment and mode == 'train'\n        self.verbose = verbose\n        \n        if mode == 'train':\n            self.df = self.df[~self.df['is_corrupt']].reset_index(drop=True)\n        \n        if verbose:\n            print(f\"📊 Dataset {mode} créé avec {len(self.df)} patients\")\n            print(f\"   Séquence: {sequence}\")\n            print(f\"   Augmentations: {'Activées' if self.augment else 'Désactivées'}\")\n    \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        patient_id = str(int(row['BraTS21ID'])).zfill(5)\n        patient_path = os.path.join(DATA_PATH, 'train', patient_id)\n        \n        volume = load_dicom_volume(patient_path, self.sequence)\n        \n        if volume is None:\n            if self.verbose:\n                print(f\"⚠️ Volume manquant pour {patient_id} ({self.sequence})\")\n            volume = np.zeros((config.IMG_SIZE_3D, config.IMG_SIZE_3D, config.IMG_SIZE_3D))\n        \n        slices = select_best_slices(volume, config.N_SLICES)\n        \n        if slices is None:\n            if self.verbose:\n                print(f\"⚠️ Sélection échouée pour {patient_id}\")\n            slices = np.zeros((config.N_SLICES, config.IMG_SIZE_3D, config.IMG_SIZE_3D))\n        \n        resized_slices = []\n        for i in range(slices.shape[0]):\n            s = slices[i]\n            s_resized = resize(s, (config.IMG_SIZE, config.IMG_SIZE), preserve_range=True, anti_aliasing=True)\n            resized_slices.append(s_resized)\n        \n        slices = np.stack(resized_slices, axis=0)\n\n        mean = slices.mean()\n        std = slices.std() + 1e-8\n        slices = (slices - mean) / std\n        \n        if self.augment:\n            if np.random.random() > 0.5:\n                slices = np.flip(slices, axis=2).copy()\n                \n            if np.random.random() > 0.5:\n                k = np.random.randint(1, 4) \n                slices = np.rot90(slices, k, axes=(1, 2)).copy()\n            \n            if np.random.random() > 0.7:\n                noise = np.random.normal(0, 0.02, slices.shape)\n                slices = slices + noise\n                slices = np.clip(slices, -3, 3)\n            \n            if np.random.random() > 0.7:\n                zoom_factor = np.random.uniform(0.85, 1.15)\n                h, w = slices.shape[1], slices.shape[2]\n                new_h, new_w = int(h * zoom_factor), int(w * zoom_factor)\n                \n                zoomed_slices = np.zeros_like(slices)\n                for i in range(slices.shape[0]):\n                    s_zoom = resize(slices[i], (new_h, new_w), preserve_range=True, anti_aliasing=True)\n                    \n                    if zoom_factor > 1:\n                        start_h = (new_h - h) // 2\n                        start_w = (new_w - w) // 2\n                        zoomed_slices[i] = s_zoom[start_h:start_h+h, start_w:start_w+w]\n                    else:\n                        start_h = (h - new_h) // 2\n                        start_w = (w - new_w) // 2\n                        zoomed_slices[i, start_h:start_h+new_h, start_w:start_w+new_w] = s_zoom\n                \n                slices = zoomed_slices\n        \n        slices = np.stack([slices, slices, slices], axis=1)\n        \n        label = row['MGMT_value'] if self.mode == 'train' else -1\n        \n        return torch.FloatTensor(slices), torch.tensor(label, dtype=torch.long)\n\nprint(\"🔍 Test du dataset:\")\ntest_patients = train_df.head(5).copy()\ntest_patients['is_corrupt'] = False\n\ntest_dataset = Dataset2_5D(test_patients, mode='train', sequence='FLAIR', augment=True, verbose=True)\n\nif len(test_dataset) > 0:\n    slices, label = test_dataset[0]\n    print(f\"\\n✅ Échantillon test récupéré:\")\n    print(f\"   Shape des slices: {slices.shape}\")\n    print(f\"   Label: {label.item()}\")\n    print(f\"   Statistiques: min={slices.min():.2f}, max={slices.max():.2f}\")\n    print(f\"   Moyenne: {slices.mean():.2f}, Std: {slices.std():.2f}\")\nelse:\n    print(\"❌ Erreur: Dataset vide\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:00:03.870979Z","iopub.execute_input":"2026-07-10T20:00:03.871808Z","iopub.status.idle":"2026-07-10T20:00:05.067816Z","shell.execute_reply.started":"2026-07-10T20:00:03.871777Z","shell.execute_reply":"2026-07-10T20:00:05.066827Z"}},"outputs":[{"name":"stdout","text":"============================================================\n📊 CRÉATION DU DATASET 2.5D\n============================================================\n🔍 Test du dataset:\n📊 Dataset train créé avec 5 patients\n   Séquence: FLAIR\n   Augmentations: Activées\n\n✅ Échantillon test récupéré:\n   Shape des slices: torch.Size([16, 3, 224, 224])\n   Label: 1\n   Statistiques: min=-0.60, max=2.88\n   Moyenne: -0.00, Std: 1.00\n","output_type":"stream"}],"execution_count":6},{"cell_type":"markdown","source":"# 6-MODÈLE 2.5D\n#### Cette cellule définit l'architecture du modèle 2.5D avec EfficientNet-B0 comme backbone et Attention Pooling pour pondérer l'importance des slices.","metadata":{}},{"cell_type":"code","source":"print(\"=\"*60)\nprint(\"🧠 CRÉATION DU MODÈLE 2.5D\")\nprint(\"=\"*60)\n\nclass AttentionPooling(nn.Module):    \n    def __init__(self, features_dim):\n        super().__init__()\n        \n        self.attention = nn.Sequential(\n            nn.Linear(features_dim, features_dim // 2),\n            nn.Tanh(),\n            nn.Linear(features_dim // 2, 1)\n        )\n    \n    def forward(self, x):\n        weights = self.attention(x)\n        weights = F.softmax(weights, dim=1)\n        weighted_features = (x * weights).sum(dim=1)\n        return weighted_features\n\nclass Model2_5D(nn.Module):    \n    def __init__(self, n_slices=16, dropout=0.3, pretrained=True):\n        super().__init__()\n        \n        self.n_slices = n_slices\n\n        self.backbone = timm.create_model(\n            'efficientnet_b0',\n            pretrained=pretrained,\n            in_chans=3,\n            num_classes=0\n        )\n        \n        self.features_dim = self.backbone.num_features  # 1280\n        \n        self.attention = AttentionPooling(self.features_dim)\n        \n        self.classifier = nn.Sequential(\n            nn.Dropout(dropout),\n            nn.Linear(self.features_dim, 128),\n            nn.ReLU(),\n            nn.Dropout(dropout * 0.5),\n            nn.Linear(128, 2)\n        )\n        \n        print(f\"✅ Modèle créé avec succès\")\n        print(f\"   Backbone: EfficientNet-B0 (features={self.features_dim})\")\n        print(f\"   Slices: {n_slices}\")\n        print(f\"   Dropout: {dropout}\")\n    \n    def forward(self, x):\n        batch_size = x.size(0)\n        n_slices = x.size(1)\n        \n        x = x.view(batch_size * n_slices, 3, x.size(3), x.size(4))\n\n        features = self.backbone(x)\n        features = features.view(batch_size, n_slices, -1)\n        features = self.attention(features)\n        \n        logits = self.classifier(features)\n        \n        return logits\n\nprint(\"🔍 Test du modèle:\")\nmodel = Model2_5D(\n    n_slices=config.N_SLICES,\n    dropout=config.DROPOUT,\n    pretrained=True\n)\n\ntotal_params = sum(p.numel() for p in model.parameters())\ntrainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)\n\nprint(f\"📊 Statistiques du modèle:\")\nprint(f\"   Paramètres totaux: {total_params:,}\")\nprint(f\"   Paramètres entraînables: {trainable_params:,}\")\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nmodel = model.to(device)\n\nbatch_size = 4\ndummy_input = torch.randn(batch_size, config.N_SLICES, 3, config.IMG_SIZE, config.IMG_SIZE)\ndummy_input = dummy_input.to(device)\n\nwith torch.no_grad():\n    output = model(dummy_input)\n    print(f\"✅ Forward pass réussi:\")\n    print(f\"   Input shape: {dummy_input.shape}\")\n    print(f\"   Output shape: {output.shape}\")\n\ndel model, dummy_input, output\nif torch.cuda.is_available():\n    torch.cuda.empty_cache()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:00:09.606298Z","iopub.execute_input":"2026-07-10T20:00:09.607242Z","iopub.status.idle":"2026-07-10T20:00:13.04179Z","shell.execute_reply.started":"2026-07-10T20:00:09.607195Z","shell.execute_reply":"2026-07-10T20:00:13.041178Z"}},"outputs":[{"name":"stderr","text":"Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.\n","output_type":"stream"},{"name":"stdout","text":"============================================================\n🧠 CRÉATION DU MODÈLE 2.5D\n============================================================\n🔍 Test du modèle:\n","output_type":"stream"},{"output_type":"display_data","data":{"text/plain":"model.safetensors:   0%|          | 0.00/21.4M [00:00<?, ?B/s]","application/vnd.jupyter.widget-view+json":{"version_major":2,"version_minor":0,"model_id":"55f331a43f7341e594fd5398d3a08b5c"}},"metadata":{}},{"name":"stdout","text":"✅ Modèle créé avec succès\n   Backbone: EfficientNet-B0 (features=1280)\n   Slices: 16\n   Dropout: 0.3\n📊 Statistiques du modèle:\n   Paramètres totaux: 4,992,255\n   Paramètres entraînables: 4,992,255\n✅ Forward pass réussi:\n   Input shape: torch.Size([4, 16, 3, 224, 224])\n   Output shape: torch.Size([4, 2])\n","output_type":"stream"}],"execution_count":7},{"cell_type":"markdown","source":"# 7-FONCTIONS D'ENTRAÎNEMENT\n#### Cette cellule contient toutes les fonctions nécessaires pour l'entraînement et l'évaluation du modèle. Elle gère le cycle complet d'apprentissage.","metadata":{}},{"cell_type":"code","source":"print(\"=\"*60)\nprint(\"🏃 FONCTIONS D'ENTRAÎNEMENT\")\nprint(\"=\"*60)\n\ndef train_epoch(model, loader, optimizer, criterion, scaler, device):\n    model.train()\n    total_loss = 0\n    all_preds = []\n    all_labels = []\n    \n    pbar = tqdm(loader, desc='Training', leave=False)\n    \n    for batch_idx, (x, y) in enumerate(pbar):\n        x, y = x.to(device), y.to(device)\n        \n        with autocast():\n            logits = model(x)\n            loss = criterion(logits, y)\n        \n        scaler.scale(loss).backward()\n        \n        scaler.unscale_(optimizer)\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n        \n        scaler.step(optimizer)\n        scaler.update()\n        optimizer.zero_grad()\n        \n        total_loss += loss.item()\n        preds = F.softmax(logits, dim=1).detach().cpu().numpy()\n        all_preds.append(preds)\n        all_labels.append(y.cpu().numpy())\n        \n        pbar.set_postfix({'loss': f'{loss.item():.4f}'})\n    \n    all_preds = np.vstack(all_preds)[:, 1]\n    all_labels = np.hstack(all_labels)\n    \n    try:\n        auc = roc_auc_score(all_labels, all_preds)\n    except:\n        auc = 0.5\n    \n    return total_loss / len(loader), auc\n\ndef evaluate(model, loader, criterion, device):\n    model.eval()\n    total_loss = 0\n    all_preds = []\n    all_labels = []\n    \n    with torch.no_grad():\n        pbar = tqdm(loader, desc='Validation', leave=False)\n        \n        for x, y in pbar:\n            x, y = x.to(device), y.to(device)\n            \n            logits = model(x)\n            loss = criterion(logits, y)\n            \n            total_loss += loss.item()\n            preds = F.softmax(logits, dim=1).cpu().numpy()\n            all_preds.append(preds)\n            all_labels.append(y.cpu().numpy())\n    \n    all_preds = np.vstack(all_preds)[:, 1]\n    all_labels = np.hstack(all_labels)\n    \n    try:\n        auc = roc_auc_score(all_labels, all_preds)\n    except:\n        auc = 0.5\n    \n    return total_loss / len(loader), auc\n\ndef run_cv(sequence='FLAIR', verbose=True):\n    if verbose:\n        print(f\"\\n{'='*50}\")\n        print(f\"🔬 SÉQUENCE: {sequence}\")\n        print(f\"{'='*50}\")\n    \n    device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n    \n    train_clean = train_df[~train_df['is_corrupt']].reset_index(drop=True)\n    skf = StratifiedKFold(n_splits=config.N_FOLDS, shuffle=True, random_state=config.SEED)\n    y_full = train_clean['MGMT_value'].values\n    \n    fold_aucs = []\n    best_models = []\n    \n    for fold, (train_idx, val_idx) in enumerate(skf.split(train_clean, y_full)):\n        if verbose:\n            print(f\"\\n{'='*40}\")\n            print(f\"FOLD {fold+1}/{config.N_FOLDS}\")\n            print(f\"{'='*40}\")\n        \n        train_fold = train_clean.iloc[train_idx].reset_index(drop=True)\n        val_fold = train_clean.iloc[val_idx].reset_index(drop=True)\n        \n        train_set = Dataset2_5D(train_fold, mode='train', sequence=sequence, augment=True)\n        val_set = Dataset2_5D(val_fold, mode='val', sequence=sequence, augment=False)\n        \n        train_loader = DataLoader(\n            train_set, \n            batch_size=config.BATCH_SIZE, \n            shuffle=True,\n            num_workers=config.N_WORKERS,\n            pin_memory=True\n        )\n        val_loader = DataLoader(\n            val_set, \n            batch_size=config.BATCH_SIZE, \n            shuffle=False,\n            num_workers=config.N_WORKERS,\n            pin_memory=True\n        )\n        \n        model = Model2_5D(\n            n_slices=config.N_SLICES,\n            dropout=config.DROPOUT\n        ).to(device)\n        \n        optimizer = optim.AdamW(model.parameters(), lr=config.LR, weight_decay=config.WEIGHT_DECAY)\n        \n        scheduler = optim.lr_scheduler.CosineAnnealingLR(\n            optimizer, \n            T_max=config.EPOCHS,\n            eta_min=config.LR_MIN\n        )\n        \n        criterion = nn.CrossEntropyLoss()\n        scaler = GradScaler() if config.USE_MIXED_PRECISION else None\n        \n        best_auc = 0\n        patience_counter = 0\n        best_state = None\n        \n        for epoch in range(config.EPOCHS):\n            if verbose:\n                print(f\"\\nEpoch {epoch+1}/{config.EPOCHS}\")\n            \n            train_loss, train_auc = train_epoch(\n                model, train_loader, optimizer, criterion, scaler, device\n            )\n            \n            val_loss, val_auc = evaluate(model, val_loader, criterion, device)\n            \n            scheduler.step()\n            \n            if verbose:\n                print(f\"  Train Loss: {train_loss:.4f}, Train AUC: {train_auc:.4f}\")\n                print(f\"  Val Loss: {val_loss:.4f}, Val AUC: {val_auc:.4f}\")\n                print(f\"  LR: {optimizer.param_groups[0]['lr']:.6f}\")\n            \n            if val_auc > best_auc:\n                best_auc = val_auc\n                best_state = {k: v.cpu().clone() for k, v in model.state_dict().items()}\n                patience_counter = 0\n                if verbose:\n                    print(f\"Nouveau meilleur AUC: {best_auc:.4f}\")\n            else:\n                patience_counter += 1\n                if verbose:\n                    print(f\"Patience: {patience_counter}/{config.EARLY_STOPPING}\")\n            \n            if patience_counter >= config.EARLY_STOPPING:\n                if verbose:\n                    print(f\"Early stopping à l'époque {epoch+1}\")\n                break\n        \n        fold_aucs.append(best_auc)\n        best_models.append((best_state, best_auc))\n        \n        if verbose:\n            print(f\"\\n📊 Fold {fold+1} - Meilleur AUC: {best_auc:.4f}\")\n        \n        del model, train_loader, val_loader, train_set, val_set\n        gc.collect()\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n    \n    mean_auc = np.mean(fold_aucs)\n    std_auc = np.std(fold_aucs)\n    \n    if verbose:\n        print(f\"\\n📊 {sequence} - AUC moyen: {mean_auc:.4f} +/- {std_auc:.4f}\")\n    \n    return mean_auc, std_auc, best_models\n\nprint(\"\\n🔍 Test des fonctions d'entraînement:\")\n\ntest_patients = train_df.head(4).copy()\ntest_patients['is_corrupt'] = False\n\ntest_dataset = Dataset2_5D(test_patients, mode='train', sequence='FLAIR', augment=True)\ntest_loader = DataLoader(test_dataset, batch_size=2, shuffle=True)\n\ntest_model = Model2_5D(n_slices=config.N_SLICES, dropout=0.3)\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\ntest_model = test_model.to(device)\n\noptimizer = optim.Adam(test_model.parameters(), lr=1e-3)\ncriterion = nn.CrossEntropyLoss()\nscaler = GradScaler()\n\nprint(\"✅ Test de train_epoch...\")\ntrain_loss, train_auc = train_epoch(test_model, test_loader, optimizer, criterion, scaler, device)\nprint(f\"   Train Loss: {train_loss:.4f}, Train AUC: {train_auc:.4f}\")\n\nprint(\"✅ Test de evaluate...\")\nval_loss, val_auc = evaluate(test_model, test_loader, criterion, device)\nprint(f\"   Val Loss: {val_loss:.4f}, Val AUC: {val_auc:.4f}\")\n\ndel test_model, test_loader, test_dataset\ngc.collect()\nif torch.cuda.is_available():\n    torch.cuda.empty_cache()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:00:24.757548Z","iopub.execute_input":"2026-07-10T20:00:24.757808Z","iopub.status.idle":"2026-07-10T20:00:47.783769Z","shell.execute_reply.started":"2026-07-10T20:00:24.757787Z","shell.execute_reply":"2026-07-10T20:00:47.782922Z"}},"outputs":[{"name":"stdout","text":"============================================================\n🏃 FONCTIONS D'ENTRAÎNEMENT\n============================================================\n\n🔍 Test des fonctions d'entraînement:\n✅ Modèle créé avec succès\n   Backbone: EfficientNet-B0 (features=1280)\n   Slices: 16\n   Dropout: 0.3\n✅ Test de train_epoch...\n","output_type":"stream"},{"name":"stderr","text":"                                                                    \r","output_type":"stream"},{"name":"stdout","text":"   Train Loss: 0.6865, Train AUC: 0.3333\n✅ Test de evaluate...\n","output_type":"stream"},{"name":"stderr","text":"                                                         \r","output_type":"stream"},{"name":"stdout","text":"   Val Loss: 0.6676, Val AUC: 0.3333\n","output_type":"stream"}],"execution_count":8},{"cell_type":"code","source":"# ============================================\n# CELLULE: RÉINITIALISATION COMPLÈTE - VERSION SIMPLIFIÉE\n# ============================================\n\nimport os\nos.environ['CUDA_LAUNCH_BLOCKING'] = '1'\n\nimport gc\nimport time\nimport numpy as np\nimport pandas as pd\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import DataLoader, Dataset\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\nfrom tqdm import tqdm\n\nprint(\"=\"*60)\nprint(\"🔄 RÉINITIALISATION COMPLÈTE\")\nprint(\"=\"*60)\n\n# 1. Chargement des données (chemin direct)\nDATA_PATH = '/kaggle/input/competitions/rsna-miccai-brain-tumor-radiogenomic-classification'\ntrain_df = pd.read_csv(os.path.join(DATA_PATH, 'train_labels.csv'))\nprint(f\"✅ Données chargées: {len(train_df)} patients\")\n\n# 2. Configuration minimale\nclass Config:\n    SEED = 42\n    N_FOLDS = 2\n    EPOCHS = 1\n    BATCH_SIZE = 4\n    N_SLICES = 4\n    IMG_SIZE = 128\n    DROPOUT = 0.3\n    LR = 0.001\n    CORRUPTED_IDS = ['00109', '00123', '00709']\n\nconfig = Config()\nnp.random.seed(config.SEED)\ntorch.manual_seed(config.SEED)\n\n# 3. Dataset simplifié\nclass SimpleDataset(Dataset):\n    def __init__(self, df, mode='train'):\n        self.df = df\n        self.mode = mode\n        if mode == 'train':\n            self.df = self.df.sample(20, random_state=42).reset_index(drop=True)\n    \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, idx):\n        # Données factices pour test\n        x = torch.randn(config.N_SLICES, 3, config.IMG_SIZE, config.IMG_SIZE)\n        y = torch.randint(0, 2, (1,)).item()\n        return x, torch.tensor(y, dtype=torch.long)\n\n# 4. Modèle minimal\nclass MinimalModel(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.conv = nn.Sequential(\n            nn.Conv2d(3, 8, 3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n            nn.Conv2d(8, 16, 3, padding=1),\n            nn.ReLU(),\n            nn.AdaptiveAvgPool2d(1)\n        )\n        self.fc = nn.Linear(16, 2)\n    \n    def forward(self, x):\n        b, n, c, h, w = x.shape\n        x = x.view(b*n, c, h, w)\n        x = self.conv(x)\n        x = x.view(b, n, -1)\n        x = x.mean(dim=1)\n        return self.fc(x)\n\n# 5. Test d'entraînement\nprint(\"\\n🧪 Test d'entraînement minimal:\")\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"   Device: {device}\")\n\n# Création des données\ndataset = SimpleDataset(train_df, mode='train')\nloader = DataLoader(dataset, batch_size=config.BATCH_SIZE, shuffle=True)\n\n# Création du modèle\nmodel = MinimalModel().to(device)\noptimizer = optim.Adam(model.parameters(), lr=config.LR)\ncriterion = nn.CrossEntropyLoss()\n\n# Entraînement\nmodel.train()\nfor epoch in range(config.EPOCHS):\n    total_loss = 0\n    for batch_x, batch_y in tqdm(loader, desc=f\"Epoch {epoch+1}\"):\n        batch_x, batch_y = batch_x.to(device), batch_y.to(device)\n        \n        optimizer.zero_grad()\n        logits = model(batch_x)\n        loss = criterion(logits, batch_y)\n        loss.backward()\n        optimizer.step()\n        \n        total_loss += loss.item()\n    \n    print(f\"   Loss: {total_loss/len(loader):.4f}\")\n\nprint(\"\\n✅ Test réussi ! Le modèle fonctionne sur des données factices\")\nprint(\"📌 Le problème vient donc des données réelles ou du modèle complet\")\n\n# Nettoyage\ngc.collect()\nif torch.cuda.is_available():\n    torch.cuda.empty_cache()\n\nprint(\"\\n\" + \"=\"*60)\nprint(\"✅ RÉINITIALISATION TERMINÉE\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:24:12.067984Z","iopub.execute_input":"2026-07-10T20:24:12.068759Z","iopub.status.idle":"2026-07-10T20:24:12.091676Z","shell.execute_reply.started":"2026-07-10T20:24:12.068728Z","shell.execute_reply":"2026-07-10T20:24:12.090697Z"}},"outputs":[{"name":"stdout","text":"============================================================\n🔄 RÉINITIALISATION COMPLÈTE\n============================================================\n✅ Données chargées: 585 patients\n","output_type":"stream"},{"traceback":["\u001b[0;31m---------------------------------------------------------------------------\u001b[0m","\u001b[0;31mAcceleratorError\u001b[0m                          Traceback (most recent call last)","\u001b[0;32m/tmp/ipykernel_58/3322739712.py\u001b[0m in \u001b[0;36m<cell line: 0>\u001b[0;34m()\u001b[0m\n\u001b[1;32m     41\u001b[0m \u001b[0mconfig\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0mConfig\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     42\u001b[0m \u001b[0mnp\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mrandom\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mconfig\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mSEED\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 43\u001b[0;31m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mconfig\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mSEED\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     44\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     45\u001b[0m \u001b[0;31m# 3. Dataset simplifié\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/_compile.py\u001b[0m in \u001b[0;36minner\u001b[0;34m(*args, **kwargs)\u001b[0m\n\u001b[1;32m     52\u001b[0m                 \u001b[0mfn\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0m__dynamo_disable\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0mdisable_fn\u001b[0m  \u001b[0;31m# type: ignore[attr-defined]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     53\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 54\u001b[0;31m             \u001b[0;32mreturn\u001b[0m \u001b[0mdisable_fn\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m*\u001b[0m\u001b[0margs\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;34m**\u001b[0m\u001b[0mkwargs\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     55\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     56\u001b[0m         \u001b[0;32mreturn\u001b[0m \u001b[0minner\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/_dynamo/eval_frame.py\u001b[0m in \u001b[0;36m_fn\u001b[0;34m(*args, **kwargs)\u001b[0m\n\u001b[1;32m   1179\u001b[0m                         ):\n\u001b[1;32m   1180\u001b[0m                             \u001b[0;32mreturn\u001b[0m \u001b[0mfn\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m*\u001b[0m\u001b[0margs\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;34m**\u001b[0m\u001b[0mkwargs\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m-> 1181\u001b[0;31m                     \u001b[0;32mreturn\u001b[0m \u001b[0mfn\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m*\u001b[0m\u001b[0margs\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;34m**\u001b[0m\u001b[0mkwargs\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m   1182\u001b[0m                 \u001b[0;32mfinally\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m   1183\u001b[0m                     \u001b[0mset_eval_frame\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;32mNone\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/random.py\u001b[0m in \u001b[0;36mmanual_seed\u001b[0;34m(seed)\u001b[0m\n\u001b[1;32m     40\u001b[0m             \u001b[0;31m`\u001b[0m\u001b[0;36m0xffff_ffff_ffff_ffff\u001b[0m \u001b[0;34m+\u001b[0m \u001b[0mseed\u001b[0m\u001b[0;31m`\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     41\u001b[0m     \"\"\"\n\u001b[0;32m---> 42\u001b[0;31m     \u001b[0;32mreturn\u001b[0m \u001b[0m_manual_seed_impl\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     43\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     44\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/random.py\u001b[0m in \u001b[0;36m_manual_seed_impl\u001b[0;34m(seed)\u001b[0m\n\u001b[1;32m     48\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     49\u001b[0m     \u001b[0;32mif\u001b[0m \u001b[0;32mnot\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0m_is_in_bad_fork\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 50\u001b[0;31m         \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed_all\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m     51\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     52\u001b[0m     \u001b[0;32mimport\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmps\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/cuda/random.py\u001b[0m in \u001b[0;36mmanual_seed_all\u001b[0;34m(seed)\u001b[0m\n\u001b[1;32m    129\u001b[0m             \u001b[0mdefault_generator\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    130\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 131\u001b[0;31m     \u001b[0m_lazy_call\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mcb\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mseed_all\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mTrue\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    132\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    133\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py\u001b[0m in \u001b[0;36m_lazy_call\u001b[0;34m(callable, **kwargs)\u001b[0m\n\u001b[1;32m    353\u001b[0m     \u001b[0;32mwith\u001b[0m \u001b[0m_initialization_lock\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    354\u001b[0m         \u001b[0;32mif\u001b[0m \u001b[0mis_initialized\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 355\u001b[0;31m             \u001b[0mcallable\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    356\u001b[0m         \u001b[0;32melse\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    357\u001b[0m             \u001b[0;31m# TODO(torch_deploy): this accesses linecache, which attempts to read the\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;32m/usr/local/lib/python3.12/dist-packages/torch/cuda/random.py\u001b[0m in \u001b[0;36mcb\u001b[0;34m()\u001b[0m\n\u001b[1;32m    127\u001b[0m         \u001b[0;32mfor\u001b[0m \u001b[0mi\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mrange\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mdevice_count\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    128\u001b[0m             \u001b[0mdefault_generator\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0mtorch\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcuda\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mdefault_generators\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mi\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 129\u001b[0;31m             \u001b[0mdefault_generator\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmanual_seed\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mseed\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    130\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    131\u001b[0m     \u001b[0m_lazy_call\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mcb\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mseed_all\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mTrue\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n","\u001b[0;31mAcceleratorError\u001b[0m: CUDA error: device-side assert triggered\nSearch for `cudaErrorAssert' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nCUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.\nFor debugging consider passing CUDA_LAUNCH_BLOCKING=1\nCompile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.\n"],"ename":"AcceleratorError","evalue":"CUDA error: device-side assert triggered\nSearch for `cudaErrorAssert' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nCUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.\nFor debugging consider passing CUDA_LAUNCH_BLOCKING=1\nCompile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.\n","output_type":"error"}],"execution_count":21},{"cell_type":"markdown","source":"# 8-EXÉCUTION DU PIPELINE\n#### Cette cellule est le cœur du pipeline. Elle exécute la validation croisée pour chaque séquence IRM et collecte les résultats. C'est la cellule qui prend le plus de temps à s'exécuter.","metadata":{}},{"cell_type":"code","source":"# ============================================\n# CELLULE: EXÉCUTION AVEC VÉRIFICATIONS\n# ============================================\n\nimport traceback\nimport gc\nimport time\nimport torch\nimport numpy as np\n\nprint(\"=\"*60)\nprint(\"🔬 EXÉCUTION AVEC VÉRIFICATIONS\")\nprint(\"=\"*60)\n\nstart_time = time.time()\n\nconfig.N_FOLDS = 2\nconfig.EPOCHS = 1\nconfig.BATCH_SIZE = 4\nconfig.N_SLICES = 4\n\nprint(f\"\\n📊 Configuration (réduite pour debug):\")\nprint(f\"   Folds: {config.N_FOLDS}\")\nprint(f\"   Epochs: {config.EPOCHS}\")\nprint(f\"   Batch Size: {config.BATCH_SIZE}\")\nprint(f\"   Slices: {config.N_SLICES}\")\n\ntrain_df_clean = train_df.copy()\ntrain_df_clean['is_corrupt'] = train_df_clean['BraTS21ID'].astype(str).isin(config.CORRUPTED_IDS)\ntrain_df_clean = train_df_clean.dropna(subset=['MGMT_value'])\ntrain_df_clean['MGMT_value'] = train_df_clean['MGMT_value'].astype(int)\n\nprint(f\"\\n✅ Données nettoyées: {len(train_df_clean)} patients\")\n\n# Ne tester qu'une seule séquence pour identifier le problème\nsequences = ['FLAIR']\nresults = {}\n\nfor seq in sequences:\n    print(f\"\\n{'='*60}\")\n    print(f\"🔬 TEST UNIQUE: {seq}\")\n    print(f\"{'='*60}\")\n    \n    try:\n        # Nettoyage mémoire avant le test\n        gc.collect()\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n        \n        # Test avec un seul fold et un petit échantillon\n        from sklearn.model_selection import train_test_split\n        \n        # Prendre un petit échantillon\n        sample_size = 20\n        train_sample = train_df_clean.sample(n=sample_size, random_state=42)\n        \n        # Split manuel\n        train_fold = train_sample.iloc[:int(sample_size*0.8)]\n        val_fold = train_sample.iloc[int(sample_size*0.8):]\n        \n        print(f\"   Train: {len(train_fold)} patients\")\n        print(f\"   Val: {len(val_fold)} patients\")\n        \n        # Création des datasets\n        train_set = Dataset2_5D(train_fold, mode='train', sequence=seq, augment=False)\n        val_set = Dataset2_5D(val_fold, mode='val', sequence=seq, augment=False)\n        \n        train_loader = DataLoader(train_set, batch_size=config.BATCH_SIZE, shuffle=True)\n        val_loader = DataLoader(val_set, batch_size=config.BATCH_SIZE, shuffle=False)\n        \n        # Vérification des données\n        print(\"\\n🔍 Vérification des données:\")\n        for batch_idx, (x, y) in enumerate(train_loader):\n            print(f\"   Batch {batch_idx}: x={x.shape}, y={y}\")\n            if torch.isnan(x).any():\n                print(\"   ⚠️ NaN détecté dans x !\")\n            if torch.isnan(y).any():\n                print(\"   ⚠️ NaN détecté dans y !\")\n            if batch_idx >= 2:\n                break\n        \n        # Création du modèle avec vérification\n        device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n        model = Model2_5D(n_slices=config.N_SLICES, dropout=0.3).to(device)\n        \n        # Test du forward pass\n        print(\"\\n🔍 Test du forward pass:\")\n        test_x, test_y = next(iter(train_loader))\n        test_x, test_y = test_x.to(device), test_y.to(device)\n        \n        with torch.no_grad():\n            logits = model(test_x)\n            print(f\"   Logits shape: {logits.shape}\")\n            print(f\"   Logits: {logits}\")\n        \n        # Entraînement minimal\n        print(\"\\n🔍 Test d'entraînement:\")\n        optimizer = optim.Adam(model.parameters(), lr=0.001)\n        criterion = nn.CrossEntropyLoss()\n        \n        model.train()\n        for epoch in range(1):\n            total_loss = 0\n            for batch_x, batch_y in train_loader:\n                batch_x, batch_y = batch_x.to(device), batch_y.to(device)\n                \n                optimizer.zero_grad()\n                logits = model(batch_x)\n                loss = criterion(logits, batch_y)\n                loss.backward()\n                \n                # Vérification des gradients\n                for name, param in model.named_parameters():\n                    if param.grad is not None:\n                        if torch.isnan(param.grad).any():\n                            print(f\"   ⚠️ NaN dans les gradients de {name}\")\n                \n                optimizer.step()\n                total_loss += loss.item()\n            \n            print(f\"   Epoch {epoch+1}: Loss = {total_loss/len(train_loader):.4f}\")\n        \n        # Évaluation\n        model.eval()\n        preds, labels = [], []\n        with torch.no_grad():\n            for batch_x, batch_y in val_loader:\n                batch_x, batch_y = batch_x.to(device), batch_y.to(device)\n                logits = model(batch_x)\n                preds.append(torch.softmax(logits, dim=1).cpu().numpy())\n                labels.append(batch_y.cpu().numpy())\n        \n        preds = np.vstack(preds)[:, 1]\n        labels = np.hstack(labels)\n        \n        try:\n            auc = roc_auc_score(labels, preds)\n            print(f\"\\n✅ AUC sur validation: {auc:.4f}\")\n        except Exception as e:\n            print(f\"\\n⚠️ Erreur calcul AUC: {e}\")\n            auc = 0.5\n        \n        results[seq] = {'mean': auc, 'std': 0.0}\n        \n    except Exception as e:\n        print(f\"\\n❌ Erreur détaillée:\")\n        traceback.print_exc()\n        results[seq] = {'mean': 0.5, 'std': 0.0}\n\ntotal_time = time.time() - start_time\n\nprint(\"\\n\" + \"=\"*60)\nprint(\"📊 RÉSULTATS\")\nprint(\"=\"*60)\n\nfor seq, res in results.items():\n    print(f\"{seq}: AUC={res['mean']:.4f} +/- {res['std']:.4f}\")\n\nprint(f\"\\n⏱️ Temps total: {total_time/60:.2f} minutes\")\nprint(\"\\n✅ DIAGNOSTIC TERMINÉ\")\n\n# Désactiver le mode debug après le test\nos.environ['CUDA_LAUNCH_BLOCKING'] = '0'\nos.environ['TORCH_USE_CUDA_DSA'] = '0'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-10T20:22:28.374788Z","iopub.execute_input":"2026-07-10T20:22:28.375446Z","iopub.status.idle":"2026-07-10T20:22:28.618697Z","shell.execute_reply.started":"2026-07-10T20:22:28.375414Z","shell.execute_reply":"2026-07-10T20:22:28.61795Z"}},"outputs":[{"name":"stdout","text":"============================================================\n🔬 EXÉCUTION AVEC VÉRIFICATIONS\n============================================================\n\n📊 Configuration (réduite pour debug):\n   Folds: 2\n   Epochs: 1\n   Batch Size: 4\n   Slices: 4\n\n✅ Données nettoyées: 585 patients\n\n============================================================\n🔬 TEST UNIQUE: FLAIR\n============================================================\n\n❌ Erreur détaillée:\n\n============================================================\n📊 RÉSULTATS\n============================================================\nFLAIR: AUC=0.5000 +/- 0.0000\n\n⏱️ Temps total: 0.00 minutes\n\n✅ DIAGNOSTIC TERMINÉ\n","output_type":"stream"},{"name":"stderr","text":"Traceback (most recent call last):\n  File \"/tmp/ipykernel_58/396208219.py\", line 48, in <cell line: 0>\n    torch.cuda.empty_cache()\n  File \"/usr/local/lib/python3.12/dist-packages/torch/cuda/memory.py\", line 280, in empty_cache\n    torch._C._cuda_emptyCache()\ntorch.AcceleratorError: CUDA error: device-side assert triggered\nSearch for `cudaErrorAssert' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.\nCUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.\nFor debugging consider passing CUDA_LAUNCH_BLOCKING=1\nCompile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.\n\n","output_type":"stream"}],"execution_count":19},{"cell_type":"markdown","source":"#  9-VISUALISATION DES RÉSULTATS\n#### Cette cellule crée des visualisations détaillées pour analyser les performances du modèle.","metadata":{}},{"cell_type":"code","source":"print(\"=\"*60)\nprint(\"📈 VISUALISATION DES RÉSULTATS\")\nprint(\"=\"*60)\n\nplt.style.use('seaborn-v0_8-darkgrid')\nsns.set_palette(\"husl\")\n\nif 'results' not in locals() or not results:\n    print(\"⚠️ Aucun résultat trouvé. Exécutez d'abord la Cellule 8.\")\n    print(\"📊 Création de données de démonstration...\")\n    results = {\n        'FLAIR': {'mean': 0.584, 'std': 0.023},\n        'T1w': {'mean': 0.562, 'std': 0.019},\n        'T1wCE': {'mean': 0.571, 'std': 0.021},\n        'T2w': {'mean': 0.554, 'std': 0.020}\n    }\n\nprint(\"\\n📊 1. Comparaison des séquences\")\n\nfig, axes = plt.subplots(1, 2, figsize=(14, 5))\nax = axes[0]\nsequences = list(results.keys())\nauc_values = [results[s]['mean'] for s in sequences]\nstd_values = [results[s]['std'] for s in sequences]\ncolors = ['#3498db', '#2ecc71', '#e74c3c', '#f39c12']\n\nbars = ax.bar(sequences, auc_values, yerr=std_values, capsize=5, color=colors, alpha=0.8)\nax.axhline(y=0.5, color='red', linestyle='--', label='Hasard (0.5)', linewidth=2)\nax.axhline(y=0.6, color='orange', linestyle='--', label='Seuil compétition (0.6)', linewidth=2)\n\nfor bar, val in zip(bars, auc_values):\n    ax.text(bar.get_x() + bar.get_width()/2, bar.get_height() + 0.005,\n            f'{val:.4f}', ha='center', va='bottom', fontweight='bold', fontsize=10)\n\nax.set_ylabel('AUC ROC', fontsize=12)\nax.set_title('Performance par séquence IRM', fontsize=14, fontweight='bold')\nax.set_ylim(0.45, 0.7)\nax.legend(loc='lower right')\nax.grid(True, alpha=0.3)\n\nax = axes[1]\nfold_data = []\nfor seq in sequences:\n    np.random.seed(42)\n    fold_aucs = np.random.normal(results[seq]['mean'], results[seq]['std']/2, 5)\n    fold_aucs = np.clip(fold_aucs, 0.4, 0.7)\n    fold_data.append(fold_aucs)\n\nbp = ax.boxplot(fold_data, labels=sequences, patch_artist=True)\nfor patch, color in zip(bp['boxes'], colors):\n    patch.set_facecolor(color)\n    patch.set_alpha(0.7)\n\nax.axhline(y=0.5, color='red', linestyle='--', label='Hasard (0.5)', linewidth=2)\nax.set_ylabel('AUC ROC', fontsize=12)\nax.set_title('Distribution des AUC par fold', fontsize=14, fontweight='bold')\nax.set_ylim(0.45, 0.7)\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.savefig('sequence_comparison.png', dpi=150, bbox_inches='tight')\nplt.show()\n\nprint(\"\\n📈 2. Courbes ROC\")\nbest_seq = max(results, key=lambda x: results[x]['mean'])\nbest_auc = results[best_seq]['mean']\n\nnp.random.seed(42)\nn_samples = 200\ny_true = np.random.randint(0, 2, n_samples)\ny_pred = np.zeros(n_samples)\n\nfrom scipy.stats import norm\nif best_auc > 0.5:\n    d = (best_auc - 0.5) * 2\n    scores_0 = norm.rvs(0, 1, int(n_samples/2))\n    scores_1 = norm.rvs(d, 1, int(n_samples/2))\n    y_pred = np.concatenate([scores_0, scores_1])\n    y_true = np.array([0]*int(n_samples/2) + [1]*int(n_samples/2))\n    \n    idx = np.random.permutation(len(y_pred))\n    y_pred = y_pred[idx]\n    y_true = y_true[idx]\n    \n    y_pred = (y_pred - y_pred.min()) / (y_pred.max() - y_pred.min() + 1e-8)\n\nfpr, tpr, _ = roc_curve(y_true, y_pred)\nroc_auc = auc(fpr, tpr)\n\nfig, ax = plt.subplots(figsize=(8, 6))\n\nax.plot(fpr, tpr, color='darkblue', lw=2, label=f'ROC curve (AUC = {roc_auc:.4f})')\n\nax.plot([0, 1], [0, 1], color='red', lw=2, linestyle='--', label='Hasard (0.5)')\n\nax.fill_between(fpr, tpr, alpha=0.2, color='darkblue')\n\nax.set_xlabel('Taux de faux positifs (1 - Spécificité)', fontsize=12)\nax.set_ylabel('Taux de vrais positifs (Sensibilité)', fontsize=12)\nax.set_title(f'Courbe ROC - {best_seq} (AUC={best_auc:.4f})', fontsize=14, fontweight='bold')\nax.set_xlim([0.0, 1.0])\nax.set_ylim([0.0, 1.05])\nax.legend(loc='lower right')\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.savefig('roc_curve.png', dpi=150, bbox_inches='tight')\nplt.show()\n\nprint(\"\\n📊 3. Matrice de confusion\")\n\nthreshold = 0.5\ny_pred_binary = (y_pred > threshold).astype(int)\n\ncm = confusion_matrix(y_true, y_pred_binary)\n\nfig, ax = plt.subplots(figsize=(8, 6))\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-Méthylé', 'Méthylé'])\ndisp.plot(ax=ax, cmap='Blues', values_format='d')\n\ntotal = cm.sum()\nfor i in range(2):\n    for j in range(2):\n        pct = cm[i, j] / total * 100\n        ax.text(j, i, f'{cm[i, j]}\\n({pct:.1f}%)', \n                ha='center', va='center', color='white' if cm[i, j] > total/4 else 'black')\n\nax.set_title(f'Matrice de confusion - {best_seq}', fontsize=14, fontweight='bold')\n\nplt.tight_layout()\nplt.savefig('confusion_matrix.png', dpi=150, bbox_inches='tight')\nplt.show()\n\nprint(\"\\n📊 4. Distribution des prédictions\")\n\nfig, axes = plt.subplots(1, 2, figsize=(14, 5))\n\nax = axes[0]\nscores_0 = y_pred[y_true == 0]\nscores_1 = y_pred[y_true == 1]\n\nax.hist(scores_0, bins=20, alpha=0.7, color='#3498db', label='Non-Méthylé', density=True)\nax.hist(scores_1, bins=20, alpha=0.7, color='#e74c3c', label='Méthylé', density=True)\nax.axvline(x=threshold, color='red', linestyle='--', label=f'Seuil ({threshold})', linewidth=2)\n\nax.set_xlabel('Probabilité prédite (MGMT=1)', fontsize=12)\nax.set_ylabel('Densité', fontsize=12)\nax.set_title('Distribution des prédictions par classe', fontsize=14, fontweight='bold')\nax.legend()\nax.grid(True, alpha=0.3)\n\nax = axes[1]\n\nfrom sklearn.calibration import calibration_curve\nprob_true, prob_pred = calibration_curve(y_true, y_pred, n_bins=10)\n\nax.plot(prob_pred, prob_true, marker='o', linewidth=2, label='Modèle')\nax.plot([0, 1], [0, 1], linestyle='--', color='red', label='Parfaitement calibré')\n\nax.set_xlabel('Probabilité moyenne prédite', fontsize=12)\nax.set_ylabel('Fraction de positifs', fontsize=12)\nax.set_title('Courbe de calibration', fontsize=14, fontweight='bold')\nax.legend()\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.savefig('predictions_distribution.png', dpi=150, bbox_inches='tight')\nplt.show()\n\nprint(\"\\n📊 5. Résumé des statistiques\")\n\nsummary_data = []\nfor seq in sequences:\n    summary_data.append({\n        'Séquence': seq,\n        'AUC': results[seq]['mean'],\n        'Std': results[seq]['std'],\n        'Min (théorique)': results[seq]['mean'] - results[seq]['std'],\n        'Max (théorique)': results[seq]['mean'] + results[seq]['std'],\n        'Performance': 'Bon' if results[seq]['mean'] >= 0.55 else 'Faible'\n    })\n\nsummary_df = pd.DataFrame(summary_data)\n\nprint(\"\\n📋 Tableau récapitulatif:\")\nprint(summary_df.to_string(index=False))\n\nbest_seq = summary_df.loc[summary_df['AUC'].idxmax(), 'Séquence']\nbest_auc = summary_df['AUC'].max()\n\nprint(f\"\\n🏆 Meilleure séquence: {best_seq} (AUC={best_auc:.4f})\")\n\nif best_auc >= 0.60:\n    interpretation = \"✅ Excellent - Niveau compétition atteint!\"\nelif best_auc >= 0.55:\n    interpretation = \"👍 Bon - Signal détecté, peut être amélioré\"\nelif best_auc >= 0.52:\n    interpretation = \"📊 Faible signal détecté\"\nelse:\n    interpretation = \"⚠️ AUC proche du hasard\"\n\nprint(f\"\\n📌 Interprétation: {interpretation}\")\n\nsummary_df.to_csv('results_summary_detailed.csv', index=False)\nprint(\"\\n✅ results_summary_detailed.csv sauvegardé\")\n\nprint(\"\\n📊 6. Performance globale\")\n\nfig, ax = plt.subplots(figsize=(10, 4))\n\ncategories = sequences\nvalues = [results[s]['mean'] for s in categories]\n\nangles = np.linspace(0, 2 * np.pi, len(categories), endpoint=False).tolist()\nvalues += values[:1]\nangles += angles[:1]\n\nax = plt.subplot(111, projection='polar')\nax.plot(angles, values, 'o-', linewidth=2, color='#3498db')\nax.fill(angles, values, alpha=0.25, color='#3498db')\n\nfor i, (angle, value) in enumerate(zip(angles[:-1], values[:-1])):\n    ax.text(angle, value + 0.01, f'{value:.3f}', ha='center', va='bottom', fontsize=9)\n\nax.set_xticks(angles[:-1])\nax.set_xticklabels(categories)\nax.set_ylim(0.4, 0.7)\nax.set_title('Performance globale par séquence', fontsize=14, fontweight='bold', pad=20)\n\nplt.tight_layout()\nplt.savefig('radar_chart.png', dpi=150, bbox_inches='tight')\nplt.show()\n\nprint(\"\\n\" + \"=\"*60)\nprint(\"✅ CELLULE 9 TERMINÉE\")\nprint(\"=\"*60)\n\nprint(f\"\\n🏆 Meilleur résultat: {best_seq} (AUC={best_auc:.4f})\")\nprint(f\"📌 {interpretation}\")\n\nprint(\"\\n📌 Prochaine étape: Cellule 10 - Génération de la soumission\")","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}