{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a href=\"https://www.kaggle.com/code/yannicksteph/rsna-miccai-brain-tumor-classification\" target=\"_blank\"><img align=\"left\" alt=\"Kaggle\" title=\"Open in Kaggle\" src=\"https://kaggle.com/static/images/open-in-kaggle.svg\"></a>","metadata":{"_uuid":"db314319-d6c3-4adb-8c2e-45f4523acceb","_cell_guid":"41cf6070-bb6e-4ef4-93d1-bb3ca5e070b0","_kg_hide-input":false,"_kg_hide-output":false,"trusted":true}},{"cell_type":"markdown","source":"**<center><font size=5>RSNA-MICCAI Brain Tumor Classification</font></center>**\n\n<center><img src=\"https://lingualab.ca/fr/project/language-recovery-psa/featured_hu67ab33455cf328a3b8dbb37d23762824_484672_720x0_resize_lanczos_2.png\" alt=\"equal-2495950-1920\" border=\"0\" width=\"700\"></center>\n\n----\n# Abstract\n\nGlioblastoma, the most common and aggressive form of brain cancer in adults, poses significant challenges in diagnosis and treatment. The presence of MGMT promoter methylation in the tumor has been identified as a crucial prognostic factor and an indicator of chemotherapy responsiveness. However, the current genetic analysis of brain cancer requires invasive procedures and time-consuming processes.\n\nThe objective of this project is to improve the diagnosis and treatment strategies for glioblastoma patients, minimizing the need for invasive procedures and streamlining the genetic analysis process. By leveraging radiogenomics, the aim is to develop a non-invasive method to predict the genetic profile of the tumor solely through imaging.\n\nTo address these issues, the Radiological Society of North America (RSNA) and the Medical Image Computing and Computer Assisted Intervention Society (MICCAI Society) have collaborated on a competition focusing on glioblastoma diagnosis and treatment planning. The competition involves using MRI scans to develop a model that can accurately predict the genetic subtype of glioblastoma by detecting the presence of MGMT promoter methylation.\n\nSuccessful outcomes from this competition will contribute to less invasive diagnostic procedures and more tailored treatment approaches for brain cancer patients. This abstract provides an overview of the project's objectives, the competition's context, and the potential impact on the management and survival rates of individuals affected by glioblastoma.\n<br><br>\n\n----\n\n**Table of Contents**\n- <a href='#1'>1. Project overview and objectives</a> \n  - <a href='#1_1'>1.1. Contributors</a>\n  - <a href='#1_2'>1.2. Data overview</a>\n  - <a href='#1_3'>1.3. Abbreviations list</a>\n  - <a href='#1_4'>1.4. Imports, Methods, Paths, Reading definitions</a>\n- <a href='#2'>2. Exploratory Data</a>\n  - <a href='#2_1'>2.1. MRI (FLAIR, T1w, T1wCE, T2w)</a>\n  - <a href='#2_2'>2.2. Highlight features or areas of interest MRI</a>\n  - <a href='#2_3'>2.3. Exploratory quality and interpretation of the MRI</a>  \n- <a href='#3'>3. Motivation and Choice of Segmentation</a>\n  - <a href='#3_1'>3.1. Motivation for Segmentation</a>\n  - <a href='#3_2'>3.2. Choice of 3D Segmentation</a>\n  - <a href='#3_3'>3.3. Complementing with U-Net Model and RadiomicsShape3D Class</a>  \n- <a href='#4'>4. Utilizing the 3D U-Net Model for Brain MRI Segmentation and Shape Analysis with RadiomicsShape3D</a>\n  - <a href='#4_1'>4.1. Architecture of the 3D U-Net Model for Brain MRI Segmentation</a>\n- <a href='#5'>5. Features Extracted from Brain MRI Images</a>\n- <a href='#6'>6. Data Exploration</a>\n- <a href='#7'>7. Analysis</a>\n  - <a href='#7_1'>7.1. Group Analysis</a> \n    - <a href='#7_1_1'>7.1.1. Hierarchical by Clustering Dendrogram</a>\n    - <a href='#7_1_2'>7.1.2. Average Patient MGMT Comparison</a>\n  - <a href='#7_2'>7.2. Univariate Analysis</a>\n    - <a href='#7_2_1'>7.2.1. Assessing Normality and Distributions</a>   \n    - <a href='#7_2_2'>7.2.2. Assessing Normality, Skewness, and Kurtosis</a> \n    - <a href='#7_2_3'>7.2.3. Comparison of Kurtosis Values for Different Features</a>  \n  - <a href='#7_3'>7.3. Outliers & Essential Variables</a>\n    - <a href='#7_3_1'>7.3.1. Outlier Detection</a>\n    - <a href='#7_3_2'>7.3.2. RFE Best features selection</a>\n  - <a href='#7_4'>7.4. Bivariate Analysis</a>    \n    - <a href='#7_4_1'>7.4.1. Correlation Matrix</a>\n    - <a href='#7_4_2'>7.4.2. Feature Selection and Redundancy Prevention for Avoiding Overfitting</a>\n    - <a href='#7_4_3'>7.4.3. Density Comparison</a>\n","metadata":{"_uuid":"2118c681-7e49-48f6-89a0-ccdb8a0312f5","_cell_guid":"a966507d-9558-42ea-813a-d84643cc1bee","trusted":true}},{"cell_type":"markdown","source":"----\n\n# <a id='1'>1. Project overview and objectives</a>\n\n### Overview:\n\nA malignant brain tumor is a life-threatening condition, specifically glioblastoma, which is the most common and has the poorest prognosis among adult brain cancers, with a median survival of less than a year. The presence of a specific genetic sequence called MGMT promoter methylation in the tumor has been identified as a favorable prognostic factor and a strong predictor of chemotherapy responsiveness.\n\nCurrently, the genetic analysis of cancer requires a surgical procedure to obtain a tissue sample, followed by a time-consuming process of determining the genetic characteristics of the tumor, which can take several weeks. Depending on the results and the chosen initial treatment, additional surgeries may be necessary. Developing an accurate method to predict the genetic profile of the cancer solely through imaging (known as radiogenomics) would potentially reduce the number of surgeries and allow for a more tailored therapy approach.\n\nThe Radiological Society of North America (RSNA) and the Medical Image Computing and Computer Assisted Intervention Society (MICCAI Society) have collaborated to enhance the diagnosis and treatment planning for glioblastoma patients.\n\n### Competition:\n\nIn this competition, participants are tasked with using MRI (magnetic resonance imaging) scans to train and test a model that can predict the genetic subtype of glioblastoma by detecting the presence of MGMT promoter methylation.\n\nSuccessful outcomes from this competition could significantly contribute to less invasive diagnoses and treatments for brain cancer patients. Introducing new and personalized treatment strategies before surgery holds the potential to improve the management, survival rates, and overall prospects of individuals affected by brain cancer.\n\n### Acknowledgments:\n\nThe Radiological Society of North America (RSNA®) is a non-profit organization representing 31 radiologic subspecialties from 145 countries worldwide. RSNA promotes excellence in patient care and healthcare delivery through education, research, and technological innovation.\n\nRSNA provides high-quality educational resources, publishes five top peer-reviewed journals, hosts the world's largest radiology conference, and is dedicated to shaping the future of the profession through the RSNA Research & Education (R&E) Foundation, which has funded $66 million in grants since its establishment. Additionally, RSNA actively supports and facilitates research in medical imaging artificial intelligence (AI) by sponsoring ongoing AI challenge competitions.\n\nThe Medical Image Computing and Computer Assisted Intervention Society (MICCAI Society) is committed to advancing research, education, and practice in the field of medical image computing, computer-assisted interventions, biomedical imaging, and medical robotics. The society achieves this objective by organizing high-quality international conferences, workshops, tutorials, and publications that promote the exchange and dissemination of advanced knowledge, expertise, and experiences produced by leading institutions, scientists, physicians, and educators worldwide.\n\nA complete list of acknowledgments can be found on this page.\n\n[RSNA-MICCAI Brain Tumor Radiogenomic Classification](https://www.kaggle.com/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/data?select=train_labels.csv)\n\n## <a id='1_1'>1.1. Contributors</a>\n\n- [David Goudard](https://www.kaggle.com/goudgoud)\n- [Louis-Marie Renaud](https://www.kaggle.com/louismarierenaud)\n- [Yannick Stephan](https://github.com/YanSteph)\n\n\n## <a id='1_2'>1.2. Data overview</a>\n\n<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n   <div class=\"card-body\">\n      <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n      <p style=\"color: white;\">\n         This \"Note\" block contains information and conclusions based on the analysis of previous results.<br>It provides important remarks and points to consider for proper interpretation of the data and conclusions drawn.\n      </p>\n   </div>\n</div>\n<br>\n<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n   <div class=\"card-body\">\n      <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n      <p style=\"color: white;\">\n         This \"Global\" block contains global information and settings relevant to the project.<br>It provides details on the configuration and global aspects to consider when analyzing data and developing models.\n      </p>\n   </div>\n</div>\n\n## <a id='1_3'>1.3. Abbreviations List</a>\n\n- **ADC** : Apparent Diffusion Coefficient\n- **ASL** : Arterial Spin Labelling\n- **AUC** : Area Under the Curve \n- **DSC** : Dynamic Susceptibility Contrast \n- **DTI** : Diffusion-Tensor Imaging \n- **DWI** : Diffusion-Weighted Imaging \n- **FLAIR** : Fluid-attenuated Inversion Recovery\n- **GLCM** : Gray Level Co-occurrence Matrix\n- **GLRLM** : Grey Level Run-Length Matrix \n- **MD** : Mean Diffusivity\n- **MNI** : Montreal Neurological Institute\n- **ROC** : Receiver Operating Characteristic\n- **ROI**: Region Of Interest","metadata":{"_uuid":"4339bd0b-a195-4387-8b52-b7a1e5e1e1e5","_cell_guid":"d486cb1b-8208-47e6-a774-527c33a1d615","execution":{"iopub.status.busy":"2023-05-14T19:19:16.412499Z","iopub.execute_input":"2023-05-14T19:19:16.41328Z","iopub.status.idle":"2023-05-14T19:19:17.264665Z","shell.execute_reply.started":"2023-05-14T19:19:16.413115Z","shell.execute_reply":"2023-05-14T19:19:17.263011Z"},"trusted":true}},{"cell_type":"markdown","source":"## <a id='1_4'>1.4. Imports, Methods, Paths, Reading definitions</a>","metadata":{"_uuid":"3482f72a-bcc3-4b15-b40a-0e0cdab1ad5a","_cell_guid":"5917792b-b306-4f91-b295-c0efa1c11a20","trusted":true}},{"cell_type":"markdown","source":"### Imports","metadata":{"_uuid":"5bd316fd-abe4-494d-be6d-879ae4fb4e0a","_cell_guid":"63f21d42-4549-41eb-a66a-cf76bdb7f60e","trusted":true}},{"cell_type":"code","source":"# Operating System and File System\nimport os\n\n# Scikit-learn\nfrom sklearn.feature_selection import RFECV  # Recursive Feature Elimination with Cross-Validation\nfrom sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor, AdaBoostRegressor\nfrom sklearn.linear_model import LinearRegression, Lasso, ElasticNet, Ridge\nfrom sklearn.svm import SVR\nfrom sklearn.datasets import make_regression  # Generate synthetic regression datasets\nfrom sklearn.decomposition import PCA  # Principal Component Analysis\nfrom sklearn.model_selection import KFold  # K-Fold cross-validation\n\n# XGBoost and LightGBM\nfrom xgboost import XGBRegressor\nfrom lightgbm import LGBMRegressor\n\n# Basic\nimport math  # Mathematical functions\nfrom enum import Enum  # Enumeration (enum) types\nfrom itertools import chain, permutations, combinations  # Iteration tools\n\n# Data Manipulation and Analysis\nimport numpy as np\nimport pandas as pd\n\n# Data Visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.animation as anim  # Animation in Matplotlib\nimport seaborn as sns  # Statistical data visualization\nfrom mpl_toolkits.mplot3d import Axes3D  # 3D plotting\n\n# Warnings\nimport warnings  # For suppressing warnings\n\n# JSON Handling\nimport json  # For working with JSON data\n\n# Encoding and Decoding Binary Data\nimport base64  # For encoding and decoding binary data\n\n# Interactive Widgets and Display\nimport ipywidgets as widgets  # For creating interactive widgets in Jupyter Notebook\nfrom IPython.display import HTML, display  # For displaying HTML content\nfrom IPython.display import Image as show_gif  # GIF display\n\n# Deep Learning Framework\nimport torch  # For working with PyTorch deep learning framework\n\n# DICOM File Handling\nimport pydicom  # For reading DICOM files\nfrom pydicom import dcmread  # For reading DICOM files\n\n# Image Processing and Filtering\nimport SimpleITK as sitk  # For image filtering\nfrom PIL import Image  # For image processing using the Python Imaging Library (PIL)\n\n# Machine Learning and Data Splitting\nfrom sklearn.model_selection import train_test_split  # For splitting data into training and testing sets\nfrom sklearn.preprocessing import RobustScaler, StandardScaler  # Data scaling for machine learning models\n\n# scipy\nimport scipy.cluster.hierarchy as sch\nfrom scipy.cluster.hierarchy import fcluster, linkage  # Hierarchical clustering\nfrom scipy import stats\n\n# statsmodels\nfrom statsmodels.stats.diagnostic import lilliefors\nimport statsmodels.api as sm\n\n# Additional Libraries\nimport glob  # File pattern matching\nimport random  # Random number generation\n\n!pip install pyradiomics > /dev/null  # Installing the pyradiomics library for radiomics feature extraction\nimport radiomics  # For extracting radiomics features from medical images","metadata":{"_uuid":"7ee05952-0cc5-464f-ad60-40778f96252e","_cell_guid":"6fe6a9d6-9a01-4516-bd63-de3387ea5362","_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:09:40.138716Z","iopub.execute_input":"2023-08-24T06:09:40.139573Z","iopub.status.idle":"2023-08-24T06:10:45.634783Z","shell.execute_reply.started":"2023-08-24T06:09:40.139533Z","shell.execute_reply":"2023-08-24T06:10:45.633108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Paths","metadata":{}},{"cell_type":"code","source":"# ------------------------------------#    \n# Definition Version\n# ------------------------------------#    \n\nclass DatasetVersion(Enum):\n    V_2D = '2D'\n    V_3D = '3D'\n    \n# ------------------------------------#  \n# Paths\n# ------------------------------------#  \n    \n# ------------#\n# Folders\n# ------------# \nbase_folder_path = \"../input/rsna-miccai-brain-tumor-radiogenomic-classification/\"\ntrain_folder_path = f\"{base_folder_path}train/\"\ndataset_folder_path = \"../input/githut-rsna-miccai-brain-tumor-classification-ai/dataset/\"\n\n# ------------#\n# Train / files\n# ------------# \ntrain_labels_path = f\"{base_folder_path}train_labels.csv\"\nsubmission_labels_path = f\"{base_folder_path}sample_submission.csv\"\n\ndataset_3d_path = f\"{dataset_folder_path}3d_rsna_miccai_brain_tumor_brain_segmentation_pytorch_unet.csv\"\ndataset_2d_path = f\"{dataset_folder_path}2d_rsna_miccai_brain_tumor_brain_segmentation_pytorch_unet.csv\"","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:45.637902Z","iopub.execute_input":"2023-08-24T06:10:45.639948Z","iopub.status.idle":"2023-08-24T06:10:45.654660Z","shell.execute_reply.started":"2023-08-24T06:10:45.639887Z","shell.execute_reply":"2023-08-24T06:10:45.651926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Train folder of RSNA: <b>`train_folder_path`</b></li>\n            <li>Generated Train labels Segmentation 3D CSV: <b>`train_labels_path`</b></li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Constants","metadata":{}},{"cell_type":"code","source":"# ------------------------------------#   \n# Configuration\n# ------------------------------------#    \n    \n# ℹ️ Flag to skip brain segmentation with PyTorch UNet\n# If set to True, we will import the dataset that has already been generated\nskip_brain_segmentation_pytorch_unet = True\n\n# ℹ️ Set the dataset version to use when <skip_brain_segmentation_pytorch_unet = True>\ndataset_version =  DatasetVersion.V_3D\n\n#Color\ndefault_palette_color = \"deep\"","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:45.657079Z","iopub.execute_input":"2023-08-24T06:10:45.657912Z","iopub.status.idle":"2023-08-24T06:10:45.679120Z","shell.execute_reply.started":"2023-08-24T06:10:45.657862Z","shell.execute_reply":"2023-08-24T06:10:45.677585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Flag to skip brain segmentation with PyTorch UNet: <b>`skip_brain_segmentation_pytorch_unet`</b></li>\n            <li>Flag to set the dataset version 2D or 3D: <b>`dataset_version`</b></li>\n            <li>👆🏻 Configuration of the project</li>\n        </ul>\n    </div>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"### Configure","metadata":{"_uuid":"228f3522-329f-4672-9102-108bfece2be9","_cell_guid":"2bc00e09-7471-4c5a-9608-0208530070ef","trusted":true}},{"cell_type":"code","source":"# ------------------------------------#\n# Configuration options\n# ------------------------------------#\n\n# Pandas configuration\npd.set_option('display.max_columns', None)\npd.set_option('display.expand_frame_repr', False)\npd.set_option('max_colwidth', -1)\npd.set_option('display.max_rows', None)\n\n# Suppressing Warnings\nwarnings.filterwarnings('ignore')\n\n# Define \"reader\"\n# Read serie of image files into a SimpleTK image\nsitk_reader = sitk.ImageSeriesReader()\nsitk_reader.LoadPrivateTagsOn()\n\n# Suppressing warnings SimpleITK\nsitk.ProcessObject.SetGlobalWarningDisplay(False)\n\n# Default color\nsns.set_theme(palette=default_palette_color)\n\n# ------------\n# Segmentation\n# ------------\n\n# Load mateuszbuda/brain-segmentation-pytorch, U-Net with batch normalization for biomedical image segmentation with pretrained weights for abnormality segmentation in brain MRI\nsegmentation_model = torch.hub.load('mateuszbuda/brain-segmentation-pytorch', \n                                    'unet', \n                                    in_channels=3, \n                                    out_channels=1, \n                                    init_features=32, \n                                    pretrained=True, \n                                    trust_repo=True)\n\n# ------------\n# Dataset\n# ------------\n# Dataset of the project, explanation in next section.    \npreview_dataset_df = pd.read_csv(train_labels_path)\nsamp_subm = pd.read_csv(submission_labels_path)","metadata":{"_uuid":"7caa653c-b826-4e99-852e-c22b28757274","_cell_guid":"ddf3ae13-9bc9-47e5-a6be-7c343074b818","_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:45.683168Z","iopub.execute_input":"2023-08-24T06:10:45.684650Z","iopub.status.idle":"2023-08-24T06:10:50.915579Z","shell.execute_reply.started":"2023-08-24T06:10:45.684590Z","shell.execute_reply":"2023-08-24T06:10:50.914565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Dataset of the competition: <b>`preview_dataset_df`</b></li>\n        </ul>\n    </div>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"### Methods","metadata":{"_uuid":"5f027225-6e9c-4b3a-9bdb-9b32287648ad","_cell_guid":"4b803e1c-e1f6-46f4-8d39-a2b4eb2900d7","trusted":true}},{"cell_type":"code","source":"# ===============================================\n# Base methods\n# ===============================================\n\ndef is_even(nombre):\n    return nombre % 2 == 0\n\ndef map_dataframe_to_tuples(df, columns):\n    mapped_tuples = []\n    for row in df.itertuples(index=False):\n        mapped_tuple = tuple(getattr(row, col) for col in columns)\n        mapped_tuples.append(mapped_tuple)\n    return mapped_tuples\n\ndef extract_number_from_filename(filename):\n    # Extrayez les chiffres à la fin du nom de fichier\n    number_str = ''.join(filter(str.isdigit, filename))\n    if number_str:\n        return int(number_str)\n    else:\n        return -1  # Si aucun chiffre n'est trouvé, renvoyez une valeur par défaut\n    import os\n    \ndef remove_values(values_to_remove, source_array):\n    \"\"\"\n    Removes elements from the source_array that are present in the values_to_remove.\n    \n    Args:\n        values_to_remove (list): The array containing the values to be removed.\n        source_array (list): The source array containing the values.\n\n    \n    Returns:\n        list: The source array with the matching values removed.\n    \"\"\"\n    for value in values_to_remove:\n        if value in source_array:\n            source_array.remove(value)\n    return source_array\n\n# ===============================================\n# Images methods\n# ===============================================\n\ndef get_processed_image(patient_id, dataset_version):\n    \"\"\"\n    Retrieves and processes the images for a given patient, grouping them for segmentation.\n\n    Args:\n        patient_id (str): The ID of the patient (BraTS21ID).\n\n    Returns:\n        numpy.ndarray: A processed image composed of the different images of the patient.\n    \"\"\"\n    # SEGMENTATION MODEL LIMITED TO 3 LAYERS\n    # T2W SKIPPED\n    patient_id = int(patient_id)\n\n    # Paths for image sequences\n    t1w_path = f'{train_folder_path}/{str(patient_id).zfill(5)}/T1w'\n    #t2w_path = f'{train_folder_path}/{str(patient_id).zfill(5)}/T2w'\n\n    flair_path = f'{train_folder_path}/{str(patient_id).zfill(5)}/FLAIR'\n    t1wce_path = f'{train_folder_path}/{str(patient_id).zfill(5)}/T1wCE'\n\n    # Retrieve image sequences\n    if dataset_version == DatasetVersion.V_2D:\n        t1w_image = sequence_filenames(t1w_path)\n        flair_image = sequence_filenames(flair_path)\n        t1wce_image = sequence_filenames(t1wce_path)\n        t2w_image = None\n    else:\n        t1w_image = sequence_filenames(t1w_path)\n        flair_image = sequence_filenames(flair_path)\n        t1wce_image = sequence_filenames(t1wce_path)\n        t2w_image = None\n        \n    # Resampling\n    if dataset_version == DatasetVersion.V_2D:\n        re_sampled_flair = re_sample_image(flair_image, t1w_image)\n        re_sampled_t1wce = re_sample_image(t1wce_image, t1w_image)\n        t2w_array = None\n    else:\n        re_sampled_flair = re_sample_image(flair_image, t1w_image)\n        re_sampled_t1wce = re_sample_image(t1wce_image, t1w_image)\n        t2w_array = None\n\n    \n    # Normalization\n    if dataset_version == DatasetVersion.V_2D:\n        t1w_array = normalize(sitk.GetArrayFromImage(t1w_image))\n        re_sampled_flair_array = normalize(sitk.GetArrayFromImage(re_sampled_flair))\n        re_sampled_t1wce_array = normalize(sitk.GetArrayFromImage(re_sampled_t1wce))\n        re_sampled_t2w_array = None\n    else:\n        t1w_array = normalize(sitk.GetArrayFromImage(t1w_image))\n        re_sampled_flair_array = normalize(sitk.GetArrayFromImage(re_sampled_flair))\n        re_sampled_t1wce_array = normalize(sitk.GetArrayFromImage(re_sampled_t1wce))\n        re_sampled_t2w_array = None\n\n    # Stack sequences\n    if dataset_version == DatasetVersion.V_2D:\n        sequence_stacked = np.stack([t1w_array, re_sampled_flair_array, re_sampled_t1wce_array])\n        central_slice = t1w_array.shape[0] // 2\n    else:\n        sequence_stacked = np.stack([t1w_array, re_sampled_flair_array, re_sampled_t1wce_array])\n        central_slice = t1w_array.shape[0] // 2\n\n    rvb = sequence_stacked[:, central_slice, :, :].transpose(1, 2, 0)\n    image = Image.fromarray((rvb * 255).astype(np.uint8))\n    return np.array([np.moveaxis(np.array(image.resize((256, 256))), -1, 0)])\n\n\n    \ndef sequence_filenames(path) :\n    \"\"\"\n    Retrieves a sequence of images for a given directory.\n\n    Args:\n        path (str): The path to the directory containing the DICOM data set.\n\n    Returns:\n        SimpleITK.Image: A sequence of images corresponding to the DICOM files in the directory.\n\n    Raises:\n        FileNotFoundError: If the specified path does not exist.\n    \"\"\"\n    filenames = sitk_reader.GetGDCMSeriesFileNames(path)\n    sitk_reader.SetFileNames(filenames)\n    image = sitk_reader.Execute()\n    \n    return image    \n\ndef normalize(dataset) :\n    \"\"\"\n    Normalizes the data obtained from the images.\n\n    Args:\n        dataset (numpy.ndarray): The dataset to be normalized.\n\n    Returns:\n        numpy.ndarray: The normalized dataset.\n    \"\"\"\n    return (dataset - np.min(dataset)) / (np.max(dataset) - np.min(dataset))\n\n\ndef re_sample_image(image, ref_img):\n    \"\"\"\n    Resamples the image to match the dimensions and properties of the reference image.\n\n    Args:\n        image (SimpleITK.Image): The image to be resampled.\n        ref_img (SimpleITK.Image): The reference image used for resampling.\n\n    Returns:\n        SimpleITK.Image: The resampled image.\n    \"\"\"\n    re_sampler = sitk.ResampleImageFilter()\n    re_sampler.SetReferenceImage(ref_img)\n    re_sampler.SetDefaultPixelValue(image.GetPixelIDValue())\n    re_sampler.SetInterpolator(sitk.sitkLinear)\n    re_sampler.SetOutputSpacing(ref_img.GetSpacing())\n    re_sampler.SetOutputDirection(ref_img.GetDirection())\n    re_sampler.SetOutputOrigin(ref_img.GetOrigin())\n    re_sampler.SetSize(ref_img.GetSize())\n    re_sampler.SetTransform(sitk.AffineTransform(image.GetDimension()))\n    re_sampled_image = re_sampler.Execute(image)\n    \n    return re_sampled_image\n\ndef segmentation_process(image_resized):\n    \"\"\"\n    Obtains the segmented image.\n\n    Args:\n        image_resized (numpy.ndarray): The resized image.\n\n    Returns:\n        numpy.ndarray: The segmented image.\n    \"\"\"\n    segmentation = segmentation_model(torch.Tensor(image_resized))\n    return segmentation\n    \n# ===============================================\n# Dataset creation methods\n# ===============================================\n\ndef init_dataset_radiomics() :\n    \"\"\"\n    Initializes the DataFrame structures for radiomics data acquisition.\n\n    Returns:\n        None\n    \"\"\"\n    global df_shapes\n    global df_textures\n    global df_first_orders_features\n    \n    if dataset_version == DatasetVersion.V_3D:\n        df_shapes_columns = ['ID','BraTS21ID','MeshVolume','VoxelVolume','SurfaceArea','SurfaceVolumeRatio','Sphericity',\n                             'Compactness1','Compactness2','SphericalDisproportion','Maximum3DDiameter','Maximum2DDiameterRow',\n                             'Maximum2DDiameterColumn','MajorAxisLength','MinorAxisLenth','LeastAxisLength','Elongation','Flatness']\n\n    else :\n        df_shapes_columns = ['ID','BraTS21ID','MeshSurface','PixelSurface','Perimeter','PerimeterSurfaceRatio','Sphericity',\n                             'SphericalDisproportion','MaximumDiameter','MajorAxisLength','MinorAxisLenth','Elongation']\n\n    \n    df_shapes = pd.DataFrame(columns=df_shapes_columns) \n\n    # GLCM features\n    df_textures_columns = ['ID','Autocorrelation', 'ClusterProminence', 'ClusterShade', 'ClusterTendency', 'Contrast', \n                           'Correlation', 'DifferenceAverage', 'DifferenceEntropy', 'DifferenceVariance', 'Id', 'Idm', \n                           'Idmn', 'Idn', 'Imc1', 'Imc2', 'InverseVariance', 'JointAverage', 'JointEnergy', 'JointEntropy', \n                           'MCC', 'MaximumProbability', 'SumAverage', 'SumEntropy', 'SumSquares']\n    df_textures = pd.DataFrame(columns=df_textures_columns) \n\n\n    # FIRST ORDERS features\n    df_first_orders_features_columns=['ID','10Percentile', '90Percentile', 'Energy', 'Entropy', 'InterquartileRange', 'Kurtosis', \n                                      'Maximum', 'MeanAbsoluteDeviation', 'Mean', 'Median', 'Minimum', 'Range', 'RobustMeanAbsoluteDeviation', \n                                      'RootMeanSquared', 'Skewness', 'TotalEnergy', 'Uniformity', 'Variance']\n    df_first_orders_features = pd.DataFrame(columns=df_first_orders_features_columns) \n\n    # gzlm features\n    df_gzlm_features_columns = ['ID','SmallAreaEmphasis','LargeAreaEmphasis','GrayLevelNonUniformity','GrayLevelNonUniformityNormalized',\n                                'SizeZoneNonUniformity','SizeZoneNonUniformityNormalized','ZonePercentage','GrayLevelVariance','ZoneVariance',\n                                'ZoneEntropy','LowGrayLevelZoneEmphasis','HighGrayLevelZoneEmphasis','SmallAreaLowGrayLevelEmphasis',\n                                'SmallAreaHighGrayLevelEmphasis','LargeAreaLowGrayLevelEmphasis','LargeAreaHighGrayLevelEmphasis']\n    df_gzlm_features = pd.DataFrame(columns=df_gzlm_features_columns)\n    \n    #glrlm_features\n    df_glrlm_features_columns = ['ID','ShortRunEmphasis','LongRunEmphasis','GrayLevelNonUniformity','RunLengthNonUniformity','RunLengthNonUniformityNormalized'\n                                 ,'RunPercentage','GrayLevelVariance','RunVariance','RunEntropy','LowGrayLevelRunEmphasis','HighGrayLevelRunEmphasis'\n                                 ,'ShortRunLowGrayLevelEmphasis','ShortRunHighGrayLevelEmphasis','LongRunLowGrayLevelEmphasis','LongRunHighGrayLevelEmphasis']\n    df_glrlm_features = pd.DataFrame(columns=df_glrlm_features_columns)\n    # ngtdm features\n    df_ngtdm_features_columns = ['ID','Coarseness','Contrast','Busyness','Complexity','Strength']\n    df_ngtdm_features = pd.DataFrame(columns=df_ngtdm_features_columns)\n    \n    # gldm features\n    df_gldm_features_columns = ['ID','SmallDependenceEmphasis','LargeDependenceEmphasis','GrayLevelNonUniformity','GrayLevelNonUniformityNormalized','DependenceNonUniformity'\n                               , 'DependenceNonUniformityNormalized', 'GrayLevelVariance', 'DependenceVariance', 'DependenceEntropy', 'DependencePercentage'\n                                , 'LowGrayLevelEmphasis', 'HighGrayLevelEmphasis', 'SmallDependenceLowGrayLevelEmphasis', 'SmallDependenceHighGrayLevelEmphasis'\n                               , 'LargeDependenceLowGrayLevelEmphasis', 'LargeDependenceHighGrayLevelEmphasis']\n    df_gldm_features = pd.DataFrame(columns=df_gldm_features_columns)\n    \n\ndef add_patient_data2D(ID,img_resized,segmentation) :\n    \"\"\"\n    Adds data from the specified patient's images to the analysis dataset.\n\n    Args:\n        ID (int): The ID of the patient.\n        img_resized (numpy.ndarray): Resized image of the patient.\n        segmentation (torch.Tensor): Segmentation of the patient's image.\n\n    Returns:\n        None\n    \"\"\"\n    global df_shapes\n    global df_textures\n    global df_first_orders_features\n    \n    # shape\n    results = radiomics.shape2D.RadiomicsShape2D(\n        sitk.GetImageFromArray(img_resized), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8)),\n        force2D=True\n    )\n    \n    shape2D = {}\n    shape2D['ID'] = int(ID)\n    shape2D['BraTS21ID'] = int(ID)\n    shape2D['MeshSurface'] = results.getMeshSurfaceFeatureValue()\n    shape2D['PixelSurface'] = results.getPixelSurfaceFeatureValue()\n    shape2D['Perimeter'] = results.getPerimeterFeatureValue()\n    shape2D['PerimeterSurfaceRatio'] = results.getPerimeterSurfaceRatioFeatureValue()\n    shape2D['Sphericity'] = results.getSphericityFeatureValue()\n    shape2D['SphericalDisproportion'] = results.getSphericalDisproportionFeatureValue()\n    shape2D['MaximumDiameter'] = results.getMaximumDiameterFeatureValue()\n    shape2D['MajorAxisLength'] = results.getMajorAxisLengthFeatureValue()\n    shape2D['MinorAxisLenth'] = results.getMinorAxisLengthFeatureValue()\n    shape2D['Elongation'] = results.getElongationFeatureValue()\n    \n    df_shapes=df_shapes.append(shape2D,ignore_index=True)\n    \n    # GLCM\n    results=radiomics.glcm.RadiomicsGLCM(\n        sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8)),\n        force2D=True\n    )\n\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n\n    df_textures=df_textures.append(res,ignore_index=True)\n    \n    # First-orders features\n    results =  radiomics.firstorder.RadiomicsFirstOrder(\n        sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8)),\n        force2D=True\n    )\n\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n\n    df_first_orders_features=df_first_orders_features.append(res,ignore_index=True)\n    \n\ndef add_patient_data3D(ID,img_resized,segmentation) :\n    \"\"\"\n    Adds data from the specified patient's images to the analysis dataset.\n\n    Args:\n        ID (int): The ID of the patient.\n        img_resized (numpy.ndarray): Resized image of the patient.\n        segmentation (torch.Tensor): Segmentation of the patient's image.\n\n    Returns:\n        None\n    \"\"\"\n    global df_shapes\n    global df_textures\n    global df_first_orders_features\n    global df_gzlm_features\n    global df_glrlm_features\n    global df_ngtdm_features\n    global df_gldm_features\n\n    # shape\n    results = radiomics.shape.RadiomicsShape(\n        sitk.GetImageFromArray(img_resized), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8))#,\n        #force2D=True\n    )\n    \n    shape3D = {}\n    shape3D['ID'] = int(ID)\n    shape3D['BraTS21ID'] = int(ID)\n    shape3D['MeshVolume'] = results.getMeshVolumeFeatureValue()\n    shape3D['VoxelVolume'] = results.getVoxelVolumeFeatureValue()\n    shape3D['SurfaceArea'] = results.getSurfaceAreaFeatureValue()\n    shape3D['SurfaceVolumeRatio']=results.getSurfaceVolumeRatioFeatureValue()\n    shape3D['Sphericity'] = results.getSphericityFeatureValue()\n    shape3D['Compactness1']=results.getCompactness1FeatureValue()\n    shape3D['Compactness2']=results.getCompactness2FeatureValue()\n    shape3D['SphericalDisproportion']=results.getSphericalDisproportionFeatureValue()\n    shape3D['Maximum3DDiameter']=results.getMaximum3DDiameterFeatureValue()\n    shape3D['Maximum2DDiameterRow']=results.getMaximum2DDiameterRowFeatureValue()\n    shape3D['Maximum2DDiameterColumn']=results.getMaximum2DDiameterColumnFeatureValue()\n    shape3D['MajorAxisLength'] = results.getMajorAxisLengthFeatureValue()\n    shape3D['MinorAxisLenth'] = results.getMinorAxisLengthFeatureValue()\n    shape3D['LeastAxisLength'] = results.getLeastAxisLengthFeatureValue()\n    shape3D['Elongation'] = results.getElongationFeatureValue()\n    shape3D['Flatness'] = results.getFlatnessFeatureValue()\n  \n    df_shapes=df_shapes.append(shape3D,ignore_index=True)\n    \n    # GLCM\n    results=radiomics.glcm.RadiomicsGLCM(\n        sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8))#,\n        #force2D=True\n    )\n\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n    \n    df_first_orders_features=df_first_orders_features.append(res,ignore_index=True)  \n    \n    # GLSZM features\n    results = radiomics.glszm.RadiomicsGLSZM(\n    \n        sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8))#,\n        #force2D=True\n    )\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n    df_gzlm_features = df_gzlm_features.append(res,ignore_index=True)\n    \n    # GLRLM features\n    results = radiomics.glrlm.RadiomicsGLRLM(\n         sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8))#,\n        #force2D=True\n    )\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n    df_glrlm_features = df_glrlm_features.append(res,ignore_index=True)\n    \n    # NGTDM features\n    results = radiomics.ngtdm.RadiomicsNGTDM(\n        sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8))#,\n        #force2D=True\n    )\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n    df_ngtdm_features = df_ngtdm_features.append(res,ignore_index=True)\n    \n    # GLDM features\n    results=radiomics.gldm.RadiomicsGLDM(\n        sitk.GetImageFromArray(img_resized[0,0,:,:].reshape(1, 256, 256)), \n        sitk.GetImageFromArray(np.array([\n            segmentation[0][0].detach().cpu().numpy() > 0.5\n        ]).astype(np.uint8))#,\n        #force2D=True\n    )\n    results.enableAllFeatures()\n    res = results.execute()\n    res['ID']=int(ID)\n    df_gldm_features = df_gldm_features.append(res,ignore_index=True) \n\n    \n# ===============================================\n# Show methods\n# ===============================================\n    \ndef show_brains(patient_id_mgmt_tuple_list, figsize, dataset_version, scans_cmap=['red', 'yellow'], show_segmentation=True, segmentation_cmap=\"Reds\"):\n    \"\"\"\n    Display a preview of patients from the given dataset.\n\n    Args:\n        patient_id_mgmt_tuples (list): List of tuples containing patient IDs and MGMT values. (Id, MGMT)\n        dataset_version (str): Version of the dataset (2D or 3D).\n    \"\"\"\n    # Définir le thème seaborn avec la palette personnalisée\n    custom_palette = sns.color_palette(scans_cmap)\n    sns.set_palette(custom_palette)\n    \n    # Loading\n    loading_max = len(patient_id_mgmt_tuple_list)\n    progress_bar = widgets.IntProgress(min=0, max=loading_max, description='Processing patients:', bar_style='info')\n    display(progress_bar)\n\n    # Process and display images for each patient\n    for i, patient_id_mgmt_value in enumerate(patient_id_mgmt_tuple_list, start=1):\n        # Process images for patient with MGMT value\n        patient_id = patient_id_mgmt_value[0]\n        patient_mgmt = patient_id_mgmt_value[1]\n        img_resized = get_processed_image(patient_id, dataset_version)\n        \n        if dataset_version == DatasetVersion.V_3D:\n            titles = ['T1w', 'FLAIR', 'T1wce', 'Segmentation'] # T2w\n        else:\n            titles = ['T1w', 'FLAIR', 'T1wce', 'Segmentation'] # T2w\n        \n        # Create the main figure\n        fig = plt.figure(figsize = figsize)\n\n        # Adjust top margin for the main figure\n        fig.subplots_adjust(top=0.85)\n\n        # Set the global title\n        fig.suptitle(f\"Patient {patient_id} and MGMT = {patient_mgmt}\", y=0.75)\n\n        x_size_subplot = 4 if show_segmentation else 3  \n        # Iterate over the source images\n        for i in range(3):\n            ax = fig.add_subplot(1, x_size_subplot, 1+i, )\n            \n            ax.imshow(img_resized[0, i])\n            ax.set_title(titles[i])\n            ax.set_xticks([])\n            ax.set_yticks([])\n            \n        if show_segmentation:\n            segmentation = segmentation_process(img_resized)\n            # Add the segmentation image\n            ax_segmentation = fig.add_subplot(1, 4, 4)\n            ax_segmentation.imshow(segmentation.detach().numpy()[0, 0], cmap=segmentation_cmap)\n            ax_segmentation.set_title(titles[3])\n            ax_segmentation.set_xticks([])\n            ax_segmentation.set_yticks([])\n   \n        plt.tight_layout()\n        plt.show()\n        \n        progress_bar.value = i + 1\n\n    sns.set_theme(palette=default_palette_color)\n    progress_bar.close()\n\n    \ndef show_segmentation(title, figsize, img_src, dataset_version, segmentation):\n    \"\"\"\n    Displays the resized source images and the segmentation image in a single line.\n\n    Args:\n        title (str): Global title for the plot.\n        img_src (numpy.ndarray): Resized source images.\n        segmentation (torch.Tensor): Segmentation image.\n\n    Returns:\n        None\n    \"\"\"\n    \ndef show_download_link(df, title = \"Download CSV file\", filename = \"data.csv\"):\n    \"\"\"\n    Displays a download link for a DataFrame as a CSV file.\n\n    Args:\n        df (pandas.DataFrame): The DataFrame to be downloaded.\n        title (str): The title of the download link (default: \"Download CSV file\").\n        filename (str): The name of the downloaded file (default: \"data.csv\").\n\n    Returns:\n        None\n    \"\"\"\n    csv = df.to_csv()\n    b64 = base64.b64encode(csv.encode())\n    payload = b64.decode()\n    html = '<a download=\"{filename}\" href=\"data:text/csv;base64,{payload}\" target=\"_blank\">{title}</a>'\n    html = html.format(payload=payload, title=title, filename=filename)\n    display(HTML(html))\n    \ndef fonc_test_normality(df, quantitative_var) : \n    fig = plt.figure(figsize=(12, 4))\n    plt.subplot(1,2,1)\n    sns.histplot(data=df[df['MGMT_value'] == 1][['MGMT_value', quantitative_var]], x=quantitative_var, kde=True, stat='density', color='orange', alpha=0.4)\n    sns.histplot(data=df[df['MGMT_value'] == 0][['MGMT_value', quantitative_var]], x=quantitative_var, kde=True, stat='density', color='blue', alpha=0.4)\n    plt.title('histogram of '+quantitative_var,fontsize=8)\n    plt.xlabel('Value',fontsize=7)\n    plt.ylabel('Frequence',fontsize=7)\n    plt.subplot(1,2,2)\n    sns.boxplot(data=df, y=quantitative_var, x='MGMT_value')\n    plt.title('Boxplot of '+quantitative_var,fontsize=8)\n    #plt.xlabel('Valeur')\n    plt.show();\n    \ndef create_kde_mgmt_pos_neg(df, x, y) : \n    \"\"\"\n    Displays 3 KDE plots that show the bivariate density using the columns x and y from a dataframe.\n\n    Args:\n        df: DataFrame, the dataframe whose 2 variables are going to be used for the KDE plot.\n        x, y: String, the two columns\n        \n    \"\"\"\n    fig = plt.figure(figsize=(15, 3))\n    fig.suptitle(\"Bivariate density of \"+y+\" considering \"+x, fontsize=10)\n    \n    plt.subplot(1,3,1)\n    sns.kdeplot(data=df, x=x, y=y, hue=\"MGMT_value\")\n\n    plt.subplot(1,3,2)\n    plt.title(\"Bivariate density for MGMT positive\", fontsize = 6)\n    sns.kdeplot(data=df[df[\"MGMT_value\"] == 1], x=x, y=y, color=\"orange\")\n\n    plt.subplot(1,3,3)\n    plt.title(\"Bivariate density for MGMT negative\", fontsize = 6)\n    sns.kdeplot(data=df[df[\"MGMT_value\"] == 0], x=x, y=y)\n    plt.show();\n\n\ndef show_explore_distribution_normality(df, quantitative_var, figsize=(20, 5)):\n    fig, axes = plt.subplots(nrows=1, ncols=6, figsize=figsize)\n\n    # Histogram with density estimation\n    histplot = sns.histplot(data=df, x=quantitative_var, kde=True, hue='MGMT_value', ax=axes[0])\n    axes[0].set_title('Distribution of Data')\n    axes[0].set_xlabel(quantitative_var)\n    axes[0].set_ylabel('Frequency')\n\n    legend = histplot.get_legend()\n    legend.set_title('MGMT_value')\n    legend.set_bbox_to_anchor((1, 1))\n\n    # Density plot with normal distribution\n    sns.kdeplot(data=df, x=quantitative_var, ax=axes[1], label='Density')\n    x = np.linspace(np.min(df[quantitative_var]), np.max(df[quantitative_var]), 100)\n    y = stats.norm.pdf(x, np.mean(df[quantitative_var]), np.std(df[quantitative_var]))\n    axes[1].plot(x, y, linewidth=2, label='Normal Distribution')\n    axes[1].set(title='Density Plot with Normal Distribution', xlabel=quantitative_var, ylabel='Density')\n    axes[1].legend(loc='upper right')\n\n    \n    # Boxplot\n    sns.boxplot(data=df, y=quantitative_var, x='MGMT_value', hue='MGMT_value', ax=axes[2])\n    axes[2].set_title('Comparison of Distributions')\n    axes[2].set_xlabel('MGMT_value')\n    axes[2].set_ylabel(quantitative_var)\n    axes[2].legend(title='MGMT_value', loc='upper center')\n\n    # Violin plot\n    sns.violinplot(data=df, y=quantitative_var, x='MGMT_value', hue='MGMT_value', ax=axes[3])\n    axes[3].set_title('Distribution and Density')\n    axes[3].set_xlabel('MGMT_value')\n    axes[3].set_ylabel(quantitative_var)\n    axes[3].legend(title='MGMT_value', loc='upper center')\n\n    # Q-Q plot\n    stats.probplot(df[quantitative_var], plot=axes[4])\n    axes[4].set_title('Quantile-Quantile (Q-Q) Comparison')\n    axes[4].set_xlabel('Theoretical Quantiles')\n    axes[4].set_ylabel('Ordered Values')\n    \n    # Scatter plot of values vs. mean\n    grouped = df.groupby('MGMT_value')\n    for group_name, group_data in grouped:\n        axes[5].scatter(range(len(group_data)), group_data[quantitative_var], label=group_name, alpha=0.5)\n    mean_value = np.mean(df[quantitative_var])\n    axes[5].axhline(y=mean_value, color='r', linestyle='--')\n    axes[5].set_title('Values vs. Mean')\n    axes[5].set_xlabel('Index')\n    axes[5].set_ylabel(quantitative_var)\n    axes[5].legend(title='MGMT_value', loc='upper right')\n\n    plt.suptitle(f'Exploratory Analysis of Data Distribution and Normality Assessment of {quantitative_var}', fontsize=14)\n    plt.tight_layout()\n    plt.show()\n\n\ndef show_test_normality(df):\n    new_df = pd.DataFrame(index=['skewness', 'kurtosis', 'excess_kurtosis', 'shapiro_test', 'normalite'])\n\n    for col in df.columns:\n        skewness = stats.skew(df[col])\n        kurtosis = stats.kurtosis(df[col], fisher=False)\n        excess_kurtosis = stats.kurtosis(df[col])\n        shapiro_test = stats.shapiro(df[col])[1]\n        normalite = 'Yes' if shapiro_test > 0.05 else 'No'\n\n\n        new_df[col] = [0.000001 if skewness == 0 else skewness,\n                       0.000001 if kurtosis == 0 else kurtosis,\n                       0.000001 if excess_kurtosis == 0 else excess_kurtosis,\n                       shapiro_test,\n                       normalite]\n\n    return new_df\n\n\n\n\ndef show_3D_scatter_plots(\n    combinations_party,\n    dataset, \n    title=None,\n    figsize=(14, 25),\n    elev_angle=30, \n    azimuth_angle=15, \n    show_legend=False,\n    legend_loc='lower center',\n    cmap='coolwarm',\n    alpha=1.0\n):\n    \"\"\"\n    Display 3D scatter plots for each combination of variables in combinations_party using the given dataset.\n\n    Args:\n        combinations_party (list): List of combinations of variables to plot (x, y, z).\n        dataset (DataFrame): The dataset containing the variables.\n        figsize (tuple, optional): Figure size in inches. Defaults to (14, 25).\n        elev_angle (float, optional): Elevation angle in degrees. Defaults to 30.\n        azimuth_angle (float, optional): Azimuth angle in degrees. Defaults to 15.\n        title (str, optional): Title for the plot. Defaults to None.\n        show_legend (bool, optional): Whether to show the legend. Defaults to True.\n        cmap (str, optional): Color map for the MGMT values. Defaults to 'coolwarm'.\n        alpha (float, optional): Transparency of the data points. Defaults to 1.0.\n    \"\"\"\n    num_combinations = len(combinations_party)\n    num_rows = int(math.ceil(num_combinations / 2))\n    num_cols = min(2, num_combinations)\n\n    fig, axes = plt.subplots(num_rows, num_cols, subplot_kw={'projection': '3d'}, figsize=figsize)\n\n    # Loop to create scatter plots for each combination\n    for i, combo in enumerate(combinations_party):\n        if num_combinations == 1:\n            ax = axes\n        elif num_rows == 1:\n            ax = axes[i]\n        else:\n            row = i // num_cols\n            col = i % num_cols\n            ax = axes[row, col]\n\n        value_x = combo[0]\n        value_y = combo[1]\n        value_z = combo[2]\n        value_c = value_z\n        \n        x = dataset[value_x]\n        y = dataset[value_y]\n        z = dataset[value_z]\n        c = dataset[value_c]\n            \n\n        # Display 3D scatter plot\n        scatter = ax.scatter(x, y, z, c=c, cmap=cmap, alpha=alpha)\n\n        # Add labels to axes\n        ax.set_xlabel(value_x)\n        ax.set_ylabel(value_y)\n        ax.set_zlabel(value_z)\n        ax.view_init(elev=elev_angle, azim=azimuth_angle)\n        ax.set_facecolor('white')\n        scatter.set_label(title)\n\n        if show_legend:\n            unique_values = np.unique(c)\n            colors = plt.get_cmap(cmap)(np.linspace(0, 1, len(unique_values)))\n            labels = [f\"{value_c} = {value}\" for value in unique_values]\n                    \n            # Ajouter une légende personnalisée\n            custom_legend = [plt.Line2D([0], [0], marker='o', color='w', markerfacecolor=color, markersize=10) for color in colors]\n            ax.legend(custom_legend, labels, loc=legend_loc, ncol=1)\n\n\n\n    # Hide the last subplot if the number of combinations is odd\n    if num_combinations % 2 == 1 and num_combinations > 1:\n        axes[num_rows-1, num_cols-1].axis('off')\n\n    fig.suptitle(title)  # Set the overall title for the plot\n\n    fig.tight_layout()\n    plt.show()\n    plt.close()\n\n\ndef show_triangle_correlation_matrix(data, title=\"Correlation Matrix\", filter_threshold=0, figsize=(50, 20)):\n    corr_df = data.corr()\n    if filter_threshold > 0:\n        corr_df = filter_correlation_matrix(corr_df, filter_threshold)\n\n    plt.figure(figsize=figsize)\n    \n    mask = np.triu(np.ones_like(corr_df))\n    ax = sns.heatmap(corr_df, annot=True, mask=mask, cmap=\"coolwarm\")  # Ajout de linewidths\n    ax.set_facecolor('white')\n    plt.title(title)\n    plt.setp(ax.get_xticklabels(), rotation=45, ha=\"right\", rotation_mode=\"anchor\")\n    \n    plt.show()\n    \ndef show_brain_slices(segmentation_path_folder, split_by = None, figsize=(50, 10)):\n    file_list = sorted(os.listdir(segmentation_path_folder), key=extract_number_from_filename)\n        \n    if split_by is not None:\n        file_list = file_list[::split_by]\n    \n    if len(file_list) == 0:\n        print(\"Aucune image à afficher.\")\n        return\n    \n    slice_images = []\n    \n    for file_name in file_list:\n        if file_name.endswith('.dcm'):\n            segmentation_path = os.path.join(segmentation_path_folder, file_name)\n            segmented_image = sitk.ReadImage(segmentation_path)\n            segmented_array = sitk.GetArrayFromImage(segmented_image)\n            slice_image = segmented_array[0]\n            slice_images.append(slice_image)\n    \n    \n    num_rows = len(slice_images) // 10 \n    num_columns =  len(slice_images) // num_rows  \n    fig, axes = plt.subplots(num_rows, num_columns, figsize=figsize)\n    \n    for ax, image in zip(axes.ravel(), slice_images):\n        ax.imshow(image, cmap='gray')\n        ax.axis('off')\n\n    plt.show()\n# ===============================================\n# Best Features\n# ===============================================\nfrom sklearn.feature_selection import RFECV\nfrom sklearn.feature_selection import RFECV\n\ndef select_best_features_RFE(df, target, models, features_scalers, reduction_step=0.001, cv=5, scoring='r2'):\n    \"\"\"\n    This function selects features using Recursive Feature Elimination (RFE) method.\n\n    Parameters:\n    df (DataFrame): The input dataframe.\n    features_and_scaler_tuple_list (list): List of tuples containing feature names and corresponding scaler.\n    target (str): The name of the target variable.\n    model (estimator): The estimator supports either fit_transform method or feature_importances_ attribute.\n    n_features_to_select (int, optional): The number of features to select. If None, half of the features are selected. Defaults to None.\n    reduction_step (float or int, optional): If greater than or equal to 1, then step corresponds to the (integer) number of features to remove at each iteration. If within (0.0, 1.0), then step corresponds to the percentage (rounded down) of features to remove at each iteration. Defaults to 0.001.\n\n    Returns:\n    list: The list of selected feature names ordered by importance.\n    \"\"\"\n    all_features = [feature for features, scaler in features_scalers for feature in features]\n\n    # Get the feature matrix and the target variable\n    X = df[all_features]\n    X_scaled = X.copy()\n    y = df[target]\n\n    # Iterate over the feature-scaler tuples\n    for features, scaler in features_scalers:\n        # Apply the scaler on the features\n        for feature in features:\n            X_scaled[feature] = scaler.fit_transform(X_scaled[feature].values.reshape(-1, 1))\n\n    results = []\n    for name, model, apply_scalers in models:\n        rfe = RFECV(estimator=model, step=reduction_step, cv=cv, scoring=scoring)\n        X_fit = X_scaled if apply_scalers else X\n        rfe.fit(X_fit, y)\n        selected_features = X_fit.iloc[:, rfe.support_]\n        score = rfe.score(X_fit, y)\n        results.append((name, score, selected_features))\n\n    return sorted(results, key=lambda x: x[1], reverse=True)\n\n\nfrom boruta import BorutaPy\nfrom sklearn.model_selection import cross_val_score\n\ndef select_best_features_Boruta(df, target, models, features_scalers, max_iter=100, n_estimators='auto', cv=5, scoring='r2', random_state=0):\n    \"\"\"\n    This function selects features using BorutaPy method.\n\n    Parameters:\n    df (DataFrame): The input dataframe.\n    target (str): The name of the target variable.\n    models (list): List of tuples containing model names and corresponding model instances.\n    features_scalers (list): List of tuples containing feature names and corresponding scaler instances.\n    max_iter (int, optional): Maximum number of iterations for Boruta. Defaults to 100.\n    cv (int or cross-validation generator, optional): Determines the cross-validation splitting strategy. Defaults to 5.\n    random_state (int, optional): Random seed for reproducibility. Defaults to 0.\n\n    Returns:\n    list: The list of selected feature names ordered by importance.\n    \"\"\"\n    all_features = [feature for features, scaler in features_scalers for feature in features]\n    \n    # Get the feature matrix and the target variable\n    X = df[all_features]\n    y = df[target]\n\n    # Iterate over the feature-scaler tuples\n    for features, scaler in features_scalers:\n        # Apply the scaler on the features\n        for feature in features:\n            X[feature] = scaler.fit_transform(X[feature].values.reshape(-1, 1))\n\n    results = []\n    for name, model in models:\n        boruta = BorutaPy(estimator=model, n_estimators=n_estimators, max_iter=max_iter, random_state=random_state)\n        # Perform cross-validation with BorutaPy\n        scores = cross_val_score(boruta, X.values, y.values, cv=cv, scoring=scoring)\n        boruta.fit(X.values, y.values)\n        selected_features = X.iloc[:, boruta.support_]\n        results.append((name, scores.mean(), selected_features.columns.tolist()))\n\n    return results\n\n# ===============================================\n# OutlierDetector\n# ===============================================\n\nclass OutlierDetector:\n    def __init__(self, threshold=1.5):\n        self.threshold = threshold\n\n    def detect_outliers(self, dataset_df, target_column, filter_column=None, filter_value=None):\n        # Calculate quartiles for each column\n        quartiles = dataset_df.quantile([0.25, 0.75])\n\n        # Calculate the interquartile range (IQR) for each column\n        iqr = quartiles.loc[0.75] - quartiles.loc[0.25]\n\n        # Filter the DataFrame based on the filter_column and filter_value if provided\n        if filter_column and filter_value:\n            dataset_filtered_df = dataset_df[dataset_df[filter_column] == filter_value]\n        else:\n            dataset_filtered_df = dataset_df\n\n        # Find columns with outliers\n        outlier_columns = dataset_filtered_df.drop(target_column, axis=1).apply(\n            lambda x: any((x > quartiles.loc[0.75, x.name] + self.threshold * iqr[x.name]) | \n                          (x < quartiles.loc[0.25, x.name] - self.threshold * iqr[x.name])),\n            axis=0\n        )\n\n        # Create a DataFrame with columns containing outliers\n        df_outliers = pd.DataFrame({\n            'Features': outlier_columns.index,\n            'Outliers present': outlier_columns.values\n        })\n\n        # Set the index of the DataFrame\n        df_outliers.set_index(\"Features\", inplace=True)\n\n        # Count the occurrences of outliers\n        count = df_outliers['Outliers present'].value_counts()\n\n        return df_outliers, count\n    \n# ===============================================\n# Correlation\n# ===============================================\n\ndef generate_correlated_features_table(dataset, correlation_thresholds):\n    \"\"\"\n    Generate a table of correlated features based on the given correlation thresholds.\n\n    Args:\n        dataset (pd.DataFrame): The dataset to analyze.\n        correlation_thresholds (list): List of correlation thresholds to consider.\n\n    Returns:\n        pd.DataFrame: The table of correlated features.\n    \"\"\"\n\n    groups_correlated = {}\n\n    for correlation_threshold in correlation_thresholds:\n        groups_correlated[correlation_threshold] = find_highly_correlated_groups(dataset.corr(), correlation_threshold)\n\n    # Préparer les données pour le tableau\n    correlated_features = {}\n\n    for correlation_threshold, groups_correlated_threshold in groups_correlated.items():\n        correlated_features[correlation_threshold] = [\", \".join(group_correlated) for group_correlated in groups_correlated_threshold if len(group_correlated) > 1]\n\n    # Vérifier et ajuster les longueurs des listes\n    max_length = max(len(features) for features in correlated_features.values())\n\n    for features in correlated_features.values():\n        features.extend([\"\"] * (max_length - len(features)))\n\n    # Créer un DataFrame avec les données\n    df_correlated = pd.DataFrame(correlated_features)\n    # Ajouter un titre au DataFrame\n    df_correlated.columns.name = \"Correlation Thresholds\"\n\n    return df_correlated\n\ndef extract_correlated_features(dataset, correlation_thresholds):\n    \"\"\"\n    Extracts correlated features from a dataset based on given correlation thresholds.\n    \n    Args:\n        dataset (pandas.DataFrame): The input dataset.\n        correlation_thresholds (list): A list of correlation thresholds.\n    \n    Returns:\n        numpy.ndarray: An array of correlated features.\n    \"\"\"\n    correlation_features = generate_correlated_features_table(dataset.copy(), [correlation_thresholds])\n    correlation_features = correlation_features.T.to_numpy()\n    correlation_features = np.concatenate(correlation_features).ravel()\n    correlation_features = [x for x in correlation_features if x != \"\"]\n    correlation_features = [string.split(\", \") for string in correlation_features]\n    correlation_features = np.concatenate(correlation_features).ravel()\n    \n    return correlation_features\n\n\ndef filter_correlation_matrix(correlation_matrix, correlation_threshold):\n    \"\"\"\n    Filters a correlation matrix by keeping only the absolute values greater than or equal to the correlation threshold.\n\n    Args:\n        correlation_matrix (pd.DataFrame): The correlation matrix.\n        correlation_threshold (float): The correlation threshold for filtering the matrix.\n\n    Returns:\n        pd.DataFrame: The filtered correlation matrix.\n\n    \"\"\"\n    filtered_correlation_matrix = correlation_matrix.copy()\n    filtered_correlation_matrix[abs(filtered_correlation_matrix) < correlation_threshold] = np.nan\n\n    return filtered_correlation_matrix\n\n\ndef find_highly_correlated_groups(correlation_matrix, correlation_threshold = 0.8, filter_duplicated_group = True, convert_indices_to_column_names = True):\n    \"\"\"\n    Finds groups of highly correlated variables from a correlation matrix.\n\n    Args:\n        correlation_matrix (pd.DataFrame): The correlation matrix.\n        correlation_threshold (float): The correlation threshold to consider as highly correlated.\n        filter_duplicated_group (bool): Indicates whether to filter out duplicated values in the correlated groups.\n        convert_indices_to_column_names (bool): Indicates whether to convert indices to column names.\n\n    Returns:\n        List[List[str]]: A list of groups, where each group contains the names of variables that are highly correlated.\n\n    \"\"\"\n    n = correlation_matrix.shape[0]  # Number of variables in the correlation matrix\n    groups_correlated = []  # List to store the correlated groups\n    \n    # Retrieve column names\n    column_names = correlation_matrix.columns\n    \n    # Traverse each variable\n    for i in range(n):\n        if column_names[i] not in [column_names[v] for g in groups_correlated for v in g]:  # Check if the variable has already been added to a group\n            group = [i]  # Create a new group containing the current variable (i)\n            for j in range(i+1, n):\n                if column_names[j] not in [column_names[v] for g in groups_correlated for v in g]:  # Check if the variable has already been added to a group\n                    correlation = correlation_matrix.iloc[i, j]  # Retrieve the correlation between variables i and j\n                    if abs(correlation) >= correlation_threshold:  # Strong correlation condition (adjust as needed)\n                        group.append(j)  # Add variable j to the group\n            \n            groups_correlated.append(group)  # Add the group to the list of correlated groups\n    \n    # Filter out duplicated values in the correlated groups\n    if filter_duplicated_group:\n        filtered_groups_correlated = []\n        for group in groups_correlated:\n            filtered_group = list(set(group))  # Convert to a set to eliminate duplicates, then convert back to a list\n            filtered_groups_correlated.append(filtered_group)\n        groups_correlated = filtered_groups_correlated # Reset by new one\n    \n    # Convert indices to column names\n    if convert_indices_to_column_names:\n        groups_correlated = [[column_names[i] for i in group] for group in groups_correlated]\n    \n    return groups_correlated\n\n","metadata":{"_uuid":"5a760e1d-74f2-4c96-8949-7e6daf4bcff8","_cell_guid":"afc69598-2d6f-42f4-904d-7e836dc83f2f","_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:50.919006Z","iopub.execute_input":"2023-08-24T06:10:50.919474Z","iopub.status.idle":"2023-08-24T06:10:51.102573Z","shell.execute_reply.started":"2023-08-24T06:10:50.919438Z","shell.execute_reply":"2023-08-24T06:10:51.100987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"----\n\n# <a id='2'>2. Exploratory Data</a>\n\nIn this section, we will conduct exploratory data analysis to gain initial insights into the brain MRI dataset.","metadata":{"_uuid":"2f518b02-410b-40f1-8287-8fdedd6945cd","_cell_guid":"7551c0dc-5900-4dd9-af62-b9489a79d556","trusted":true}},{"cell_type":"markdown","source":"The **\"train/\"** directory contains the training files for the competition. Each top-level directory represents a subject, and the **\"train_labels.csv\"** file contains the corresponding targets for each subject, indicating the presence of MGMT promoter methylation.","metadata":{}},{"cell_type":"code","source":"total = len(preview_dataset_df)  # Nombre total d'échantillons\ntitle = \"Total number of files\"\ndata = {title: [total]}\n\ndf = pd.DataFrame(data)\ndf.set_index(title, inplace=True)\ndf.head()","metadata":{"_uuid":"9e68198c-5bce-4c61-a33b-5ee3d6451821","_cell_guid":"b5bfcec0-ef1f-430c-829b-5524f221a045","_kg_hide-output":false,"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.104985Z","iopub.execute_input":"2023-08-24T06:10:51.105517Z","iopub.status.idle":"2023-08-24T06:10:51.142587Z","shell.execute_reply.started":"2023-08-24T06:10:51.105473Z","shell.execute_reply":"2023-08-24T06:10:51.141246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preview_dataset_df.head(10)","metadata":{"_uuid":"dfabdc6d-078c-4d65-a86c-3cd1546c0127","_cell_guid":"7788dfe2-782c-4786-99e4-2dbc5a2083a2","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.144546Z","iopub.execute_input":"2023-08-24T06:10:51.145153Z","iopub.status.idle":"2023-08-24T06:10:51.155591Z","shell.execute_reply.started":"2023-08-24T06:10:51.145120Z","shell.execute_reply":"2023-08-24T06:10:51.154303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(6, 6))\nsns.countplot(data=preview_dataset_df, x=\"MGMT_value\")\nplt.title(\"Distribution of MGMT values\")\nplt.xlabel(\"MGMT Value\")\nplt.xticks([0, 1], [\"Not Present\", \"Present\"])\nplt.show()","metadata":{"_uuid":"cf02cf8f-6ceb-4e45-a61e-6ed420d78166","_cell_guid":"f0249ce2-c619-4dde-a7fb-5439377e9ec1","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.158020Z","iopub.execute_input":"2023-08-24T06:10:51.158439Z","iopub.status.idle":"2023-08-24T06:10:51.456166Z","shell.execute_reply.started":"2023-08-24T06:10:51.158356Z","shell.execute_reply":"2023-08-24T06:10:51.455047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **\"test/\"** directory contains the test files. For each subject in the test data, there is no file containing the methylation targets, so these values must be predicted. The **\"sample_submission.csv\"** file is an example of a correctly formatted submission file, with MGMT values of **0.5** for each subject.\n\nOverall, the task of the competition is to predict the presence of MGMT (%) promoter methylation for each subject in the test data.","metadata":{"_uuid":"0c274f99-0b27-40e3-8383-70327d2d3398","_cell_guid":"df694fbf-9457-4a20-8dfd-43508944fe97","trusted":true}},{"cell_type":"code","source":"samp_subm.head(1)","metadata":{"_uuid":"e5d32f7a-462f-4b99-bee3-5997ca8fe212","_cell_guid":"9abe5753-0094-4f8a-b46b-c27c0fd39c15","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.457573Z","iopub.execute_input":"2023-08-24T06:10:51.457867Z","iopub.status.idle":"2023-08-24T06:10:51.468475Z","shell.execute_reply.started":"2023-08-24T06:10:51.457839Z","shell.execute_reply":"2023-08-24T06:10:51.467156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataset we will be working with consists of MRI data provided by the Radiological Society of North America (RSNA®) and the Medical Imaging Computation and Computer Assistance Society (MICCAI Society). \nThe images are provided in DICOM format and are accompanied by a CSV file containing radiomic features extracted from the images.","metadata":{}},{"cell_type":"code","source":"first_folder = str(preview_dataset_df.loc[0, 'BraTS21ID']).zfill(5) + \"/\"\ntitle = \"Folders content for all patients\"\n# Folders content\nfolder_path = train_folder_path + first_folder  # Replace train_folder_path with the actual path\nfolder_content = os.listdir(folder_path)\n\ndf = pd.DataFrame({\n        title: folder_content\n    })\ndf.set_index(title, inplace=True)\ndf.head()","metadata":{"_uuid":"16b496d4-fea4-4af0-84c5-5fd26747fbd1","_cell_guid":"25a41fda-4818-4491-b61e-b6841540cf70","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.475948Z","iopub.execute_input":"2023-08-24T06:10:51.476310Z","iopub.status.idle":"2023-08-24T06:10:51.497949Z","shell.execute_reply.started":"2023-08-24T06:10:51.476281Z","shell.execute_reply":"2023-08-24T06:10:51.496803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the first Dataset of the patient (ID: \"00000\") , we will explore the images contained in **['T2w', 'T1wCE', 'T1w', 'FLAIR']** of the first patient.","metadata":{}},{"cell_type":"code","source":"first_folder = str(preview_dataset_df.loc[0, 'BraTS21ID']).zfill(5) + \"/\"\nfolder_path = train_folder_path + first_folder  # Replace train_folder_path with the actual path\n\ntitle = 'Image Type'\nflair_count = len(os.listdir(folder_path + 'FLAIR'))\nt1w_count = len(os.listdir(folder_path + 'T1w'))\nt1wce_count = len(os.listdir(folder_path + 'T1wCE'))\nt2w_count = len(os.listdir(folder_path + 'T2w'))\n\ndata = {\n    title: ['FLAIR', 'T1w', 'T1wCE', 'T2w'],\n    'Count': [flair_count, t1w_count, t1wce_count, t2w_count]\n}\n\n\ndf = pd.DataFrame(data)\ndf.set_index(title, inplace=True)\n\n\ndf.head()","metadata":{"_uuid":"b4582e6f-229c-4e4b-8118-ea10eaef37db","_cell_guid":"76951d5c-e505-4e36-af00-f70a897a26fb","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.499490Z","iopub.execute_input":"2023-08-24T06:10:51.500304Z","iopub.status.idle":"2023-08-24T06:10:51.919111Z","shell.execute_reply.started":"2023-08-24T06:10:51.500262Z","shell.execute_reply":"2023-08-24T06:10:51.917918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='2_1'>2.1. MRI (FLAIR, T1w, T1wCE, T2w)</a>","metadata":{}},{"cell_type":"markdown","source":"\n\nThe files are mpMRI scans, this includes:\n\n- **Fluid Attenuated Inversion Recovery (FLAIR)**\n    * **What it is:** These are images that detect brain abnormalities, such as edema and inflammatory lesions. These images are sensitive to the detection of anomalies related to inflammatory and infectious diseases of the central nervous system.\n    * **What it highlights:** It helps to detect anomalies in the brain that might not be visible in other MRI sequences.\n    * These images allow for the detection of brain abnormalities related to inflammatory and infectious diseases of the central nervous system.","metadata":{}},{"cell_type":"code","source":"show_brain_slices(train_folder_path + \"00000/FLAIR\", split_by = 13)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:51.920658Z","iopub.execute_input":"2023-08-24T06:10:51.921053Z","iopub.status.idle":"2023-08-24T06:10:56.290964Z","shell.execute_reply.started":"2023-08-24T06:10:51.921020Z","shell.execute_reply":"2023-08-24T06:10:56.289217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **T1-weighted pre-contrast (T1w)**\n    * **What it is:** These are images that highlight soft tissues, such as muscles and nerves, and are useful for visualizing normal brain structures.\n    * **What it highlights:** It allows the visualization of the normal brain structures and also helps in the detection of tumors and lesions.\n    * These images allow for the detection of brain tumors and lesions.\n","metadata":{}},{"cell_type":"code","source":"show_brain_slices(train_folder_path + \"00000/T1w\")","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:10:56.292725Z","iopub.execute_input":"2023-08-24T06:10:56.293593Z","iopub.status.idle":"2023-08-24T06:11:00.942415Z","shell.execute_reply.started":"2023-08-24T06:10:56.293517Z","shell.execute_reply":"2023-08-24T06:11:00.940974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-  **T1-weighted post-contrast (T1wCE or T1Gd)**\n    * **What it is:** These are images that use a contrast agent to detect vascular anomalies, such as tumors and lesions, which are more visible after contrast agent administration.\n    * **What it highlights:** It enhances the visibility of vascular anomalies, such as tumors and lesions, making it easier to detect them.\n    * These images allow for the detection of vascular anomalies, such as tumors and lesions.\n","metadata":{}},{"cell_type":"code","source":"show_brain_slices(train_folder_path + \"00000/T1wCE\", split_by = 4)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:11:00.944886Z","iopub.execute_input":"2023-08-24T06:11:00.945643Z","iopub.status.idle":"2023-08-24T06:11:05.802077Z","shell.execute_reply.started":"2023-08-24T06:11:00.945557Z","shell.execute_reply":"2023-08-24T06:11:05.800702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **T2-weighted (T2w or T2)**\n    * **What it is:** These images detect abnormalities related to demyelination, such as multiple sclerosis, as well as brain tumors and lesions.\n    * **What it highlights:** It helps in the detection of anomalies related to cerebrospinal fluid, such as cysts and brain tumors.\n    * These images allow for the detection of anomalies related to demyelination, brain tumors, lesions, and cerebrospinal fluid.","metadata":{}},{"cell_type":"code","source":"show_brain_slices(train_folder_path + \"00000/T2w\", split_by = 13)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:11:05.803664Z","iopub.execute_input":"2023-08-24T06:11:05.804589Z","iopub.status.idle":"2023-08-24T06:11:09.948561Z","shell.execute_reply.started":"2023-08-24T06:11:05.804552Z","shell.execute_reply":"2023-08-24T06:11:09.946646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"MRI sequences: **FLAIR, T1wCE, T1w, T2w**","metadata":{}},{"cell_type":"code","source":"image_path = \"https://github.com/YanSteph/RSNA-MICCAI-Brain-Tumor-Classification-AI/blob/main/img/scan1.png?raw=true\"\nhtml_code = f'<img src=\"{image_path}\" style=\"width: 700px;\" />'\ndisplay(HTML(html_code))","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:11:09.950999Z","iopub.execute_input":"2023-08-24T06:11:09.951437Z","iopub.status.idle":"2023-08-24T06:11:09.963926Z","shell.execute_reply.started":"2023-08-24T06:11:09.951386Z","shell.execute_reply":"2023-08-24T06:11:09.962042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='2_2'>2.2. Highlight features or areas of interest MRI</a>\n\nThe \"Hot\" palette palette with Red and Yellow is commonly used in medicine to highlight areas of high activity or high intensity in brain images.\n\nThis palette uses warm colors, ranging from black to bright red, to visually represent regions of intense brain activity or high intensity values. These warm colors help highlight areas of interest and clearly distinguish them from areas of low activity or intensity.\n\nIn medicine, this can be particularly useful for visualizing aspects such as neuronal activity, areas of blood flow concentration, the intensity of certain characteristics or measurements, or any other relevant information related to activity or intensity in the brain.\n\nUsing the \"Hot\" palette makes it easier to quickly identify regions of high activity or intensity, which can be critical for healthcare professionals in interpreting brain images and detecting features or areas of interest.\n\n## <a id='2_3'>2.3. Exploratory quality and interpretation of the MRI</a>\n\n\nHere is the list of examples of arguments supporting that an MRI is well performed and that the scans are appropriate for the study:\n\n1. **Image resolution and quality:** A well-executed MRI will produce high-resolution images with clear image quality. This enables healthcare professionals to observe anatomical structures accurately and detect any potential abnormalities.\n\n2.  **Proper contrast:** The use of contrast agents during an MRI can enhance the visualization of certain structures or pathologies. When a scan is correctly performed, contrast will be applied appropriately, facilitating the identification of areas of interest.\n\n3.  **Positioning and immobilization:** High-quality MRI requires precise patient positioning and good immobilization to prevent unwanted movements during the procedure. When the patient is properly positioned and kept still, the obtained images will be more reliable and interpretable.\n\n4.  **Appropriate acquisition protocol:** Each type of MRI requires a specific acquisition protocol based on the body part being studied and the purpose of the examination. When the protocol is followed correctly, the images will provide the necessary information for diagnosis or study.\n\n5.  **Accurate interpretation:** Lastly, for an MRI to be well-executed and the scans ready for study, it is essential that interpretation is carried out by a qualified professional, such as a radiologist. Accurate interpretation of the images allows for the identification of potential abnormalities, proper diagnosis, and the recommendation of necessary treatments.\n\nIt is important to note that these arguments are general and that there are many more.","metadata":{"_uuid":"50bad42c-2162-40a7-a22a-c8c7740b3894","_cell_guid":"1fadb689-cb08-4fa2-afb3-0da0b576ade8","trusted":true}},{"cell_type":"code","source":"number_of_extraction = 3\n\n\n# Filtre df mght 1 and 0\npatient_without_mgmt = preview_dataset_df.loc[preview_dataset_df[\"MGMT_value\"] == 0].sort_values('BraTS21ID')\npatient_with_mgmt = preview_dataset_df.loc[preview_dataset_df[\"MGMT_value\"] == 1].sort_values('BraTS21ID')\n\n# Collonne selected\npatient_with_mgmt_folder_ids = patient_with_mgmt[[\"BraTS21ID\", \"MGMT_value\"]].iloc[:number_of_extraction]\npatient_without_mgmt_folder_ids = patient_without_mgmt[[\"BraTS21ID\", \"MGMT_value\"]].iloc[:number_of_extraction]\n\n# map df into tupple\npatient_with_mgmt_mapped_tuples = map_dataframe_to_tuples(patient_with_mgmt_folder_ids, [\"BraTS21ID\", \"MGMT_value\"])\npatient_without_mgmt_mapped_tuples = map_dataframe_to_tuples(patient_without_mgmt_folder_ids, [\"BraTS21ID\", \"MGMT_value\"])\n\n# Merge\npatient_combined_with_and_without_mgmt_ids = list(chain.from_iterable(zip(patient_with_mgmt_mapped_tuples, patient_without_mgmt_mapped_tuples)))\n\n# Show\nshow_brains(patient_combined_with_and_without_mgmt_ids, figsize = (8, 5), dataset_version = dataset_version)","metadata":{"_uuid":"6fcd178a-3d53-43b7-bac0-03d2c70a3730","_cell_guid":"f90ac8c9-9fc4-4b35-bc2b-20e67bed7ef5","_kg_hide-input":true,"_kg_hide-output":false,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:11:09.965881Z","iopub.execute_input":"2023-08-24T06:11:09.966249Z","iopub.status.idle":"2023-08-24T06:12:08.926537Z","shell.execute_reply.started":"2023-08-24T06:11:09.966211Z","shell.execute_reply":"2023-08-24T06:12:08.924962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When an MRI is performed improperly or the scan is of poor quality, several issues can arise, potentially leading to adverse consequences for the patient. \n\nHere potential problems:\n\n1.  **Image artifacts:** A poorly performed MRI can result in artifacts, which are distortions or errors in the image. This can make interpretation difficult or even impossible, leading to incorrect or incomplete diagnoses.\n\n2.  **Insufficient image quality:** Improper execution of the MRI can lead to insufficient image quality. Inadequate resolution or excessive blurring can make it difficult to accurately identify anatomical structures or lesions, compromising the reliability of the diagnosis.\n\n3.  **Positioning errors:** Proper patient positioning is crucial during an MRI to obtain precise and consistent images. Incorrect positioning, such as poor immobilization or improper device calibration, can cause deformations and significant loss of information.\n\n4.  **Motion artifacts:** MRI often requires the patient to remain still during image acquisition. Any involuntary movement, such as tremors or respiratory motion, can introduce motion artifacts into the image. These artifacts can mask abnormalities or lead to incorrect interpretation of the results.\n\n5.  **Parameter setting errors:** The parameters used in an MRI, such as the acquisition sequence, echo time, or repetition time, need to be properly defined based on the clinical objective. Parameter setting errors can result in poor visualization of certain structures or inappropriate sensitivity to detect certain pathologies.\n\n6.  **Region of interest not covered:** When an MRI is poorly positioned or planning errors occur, it is possible that the region of interest may not be fully covered by the image. This can lead to a loss of essential information and incomplete evaluation of the area under examination, thereby compromising the accuracy of the diagnosis.\n\n\nThese are just an example. It is important to note that these problems are not exclusive to a single MRI scan. They can occur in exceptional circumstances or due to human error.\n\nHere a preview of wrong IRM:","metadata":{}},{"cell_type":"code","source":"excluded_patient_ids = [109, 123, 709] \n\npatient_with_wrong_mri_ids = preview_dataset_df.loc[(preview_dataset_df[\"BraTS21ID\"].isin(excluded_patient_ids))].sort_values('BraTS21ID')\npatient_with_wrong_mri_ids_tuples = map_dataframe_to_tuples(patient_with_wrong_mri_ids, [\"BraTS21ID\", \"MGMT_value\"])\n\nshow_brains(patient_with_wrong_mri_ids_tuples, figsize = (8, 5), dataset_version = dataset_version)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:08.928730Z","iopub.execute_input":"2023-08-24T06:12:08.929748Z","iopub.status.idle":"2023-08-24T06:12:37.062141Z","shell.execute_reply.started":"2023-08-24T06:12:08.929706Z","shell.execute_reply":"2023-08-24T06:12:37.060732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>Our training dataset consists of <b>585 patients</b>, which is relatively limited for machine learning. A larger dataset would provide more robust results.</li>\n            <li>Submissions are evaluated on the area under the <b>ROC curve</b> between the predicted probability and the observed target. <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\" style=\"color: white;\">🔗 Link: ROC (Receiver operating characteristic)</a>.</li>\n            <li>For the submission file each BraTS21ID in the test set, we must predict a <b>probability (%)</b> for the <b>target MGMT_value.</b></li>\n            <li>We need to split the given <b>train sets into two parts: one for train and the other for test.</b></li>\n            <li>Report on the competition page, there are unexpected issues with the following cases in the training dataset, <b>we will exclude these patient IDs:</b>\n                <ul>\n                    <li><b>[109, 123, 709]</b></li>\n                </ul>\n            </li>\n            <li>To ensure accurate prediction results, <b>we exclude the \"/test\" folder</b>, it only contains thresholded values of <b>0.5</b> and this is the part of submission to the competition.</li>\n        </ul>\n    </div>\n</div>\n</br>\n<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Excluded patient IDs: <b>`excluded_patient_ids`</b></li>\n            <li>We excluded this patient Ids for the rest of the project.</li>\n                <ul>\n                    <li><b>[109, 123, 709]</b></li>\n                </ul>\n        </ul>\n    </div>\n</div>","metadata":{"_uuid":"53d42ad4-da06-4b10-a7ef-0eb25337ebd5","_cell_guid":"821b1528-ae86-4ee8-bb98-f81bf996f3a3","trusted":true}},{"cell_type":"markdown","source":"----\n\n# <a id='3'>3. Motivation and Choice of Segmentation</a>\n\n## <a id='3_1'>3.1. Motivation for Segmentation</a>\n\nBrain segmentation is of paramount importance in the field of medical imaging as it enables the extraction of precise information about different brain regions and structures. This accurate segmentation is essential for various medical applications, such as tumor detection, anatomical analysis, and treatment planning. It also facilitates advanced medical research and clinical applications.\n\n## <a id='3_2'>3.2. Choice of 3D Segmentation</a>\n\nInitially, our project was based on 2D data, where we only considered the central slice of the images. However, given the size of our dataset and the limitations of 2D visualization in fully understanding the brain structure, it seemed appropriate to transition to a 3D approach. By working with 3D image volumes, we were able to increase the richness of our data and obtain a more comprehensive representation of the brain's features. This transition to 3D segmentation offers new perspectives for more detailed analysis and more accurate results.\n\n## <a id='3_3'>3.3. Complementing with U-Net Model and RadiomicsShape3D Class</a>\n\nTo extend the capabilities of our project to volumetric brain image analysis, we have incorporated the 3D U-Net model for brain MRI and the RadiomicsShape3D class. By combining these components, we can achieve more comprehensive segmentation and perform detailed shape analysis. The 3D U-Net model is designed to process volumetric brain images, capturing details at different scales, while the RadiomicsShape3D class provides functionalities for analyzing the 3D shape characteristics of segmented regions. This combination allows us to fully harness the benefits offered by volumetric image analysis, opening up new perspectives for advanced analysis in medical research and clinical applications.\n\n\n","metadata":{"_uuid":"2ecfbb0e-3a96-4742-9b28-1c526f0bb86b","_cell_guid":"f3106509-4ab2-4874-968f-4b40af192238","trusted":true}},{"cell_type":"markdown","source":"----\n\n# <a id='4'>4. Utilizing the 3D U-Net Model for Brain MRI Segmentation and Shape Analysis with RadiomicsShape3D</a>\n\nTo fully leverage the 3D capabilities of brain segmentation, we are utilizing the 3D U-Net model for brain MRI segmentation and the RadiomicsShape3D class for shape analysis. The 3D U-Net model extends the architecture from 2D to process volumetric brain images, allowing for more comprehensive segmentation. The RadiomicsShape3D class provides functionalities to analyze the 3D shape characteristics of segmented regions in medical images. By combining these components, we can achieve more accurate brain segmentation and perform detailed shape analysis, enhancing our ability to analyze volumetric images and facilitating advancements in medical research and clinical applications.\n\n## <a id='4_1'>4.1. Architecture of the 3D U-Net Model for Brain MRI Segmentation</a>\n\nThe 3D U-Net model for brain MRI segmentation extends the 2D architecture to process volumetric brain images. It follows a similar U-shaped architecture with skip connections, but the convolutional layers operate in a 3D manner. The model consists of encoder and decoder paths, with each path containing multiple levels of blocks. The number of filters in the convolutional layers varies across the levels, allowing the model to capture different levels of details in the volumetric data.","metadata":{}},{"cell_type":"markdown","source":"----\n\n# <a id='5'>5. Features Extracted from Brain MRI Images</a>\n\nSegmentation provides us with features that provide information about various properties of brain MRI images, such as shape, texture, and grayscale statistics. These features are commonly used for the analysis and classification of medical images to aid in the detection and characterization of brain pathologies.\n\nHere is a list of features extracted from brain MRI images:\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ General Features:</summary>\n  <ul>\n    <li><strong>ID:</strong> Sample identifier (Index).</li>\n    <li><strong>MGMT_value:</strong> Presence of MGMT.</li>\n  </ul>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\"><strong>▶️ First Order Statistics Features:</strong></summary>\n  <ul>\n    <li><strong>10Percentile:</strong> The 10th percentile of the image array.</li>\n    <li><strong>90Percentile:</strong> The 90th percentile of the image array.</li>\n    <li><strong>Energy:</strong> A measure of the magnitude of voxel values in an image. A larger value implies a greater sum of the squares of these values.</li>\n    <li><strong>Entropy:</strong> Specifies the uncertainty/randomness in the image values. It measures the average amount of information required to encode the image values.</li>\n    <li><strong>InterquartileRange:</strong> The difference between the 75th and 25th percentile of the image array.</li>\n    <li><strong>Kurtosis:</strong> A measure of the ‘peakedness’ of the distribution of values in the image ROI. A higher kurtosis implies that the mass of the distribution is concentrated towards the tail(s) rather than towards the mean. A lower kurtosis implies the reverse: that the mass of the distribution is concentrated towards a spike near the Mean value.</li>\n    <li><strong>Maximum:</strong> The maximum gray level intensity within the ROI.</li>\n    <li><strong>MeanAbsoluteDeviation:</strong> The mean distance of all intensity values from the Mean Value of the image array.</li>\n    <li><strong>Mean:</strong> The average gray level intensity within the ROI.</li>\n    <li><strong>Median:</strong> The median gray level intensity within the ROI.</li>\n    <li><strong>Minimum:</strong> The minimum gray level intensity within the ROI.</li>\n    <li><strong>Range:</strong> The range of gray values in the ROI.</li>\n    <li><strong>RobustMeanAbsoluteDeviation:</strong> The mean distance of all intensity values from the Mean Value calculated on the subset of image array with gray levels in between, or equal to the 10th and 90th percentile.</li>\n    <li><strong>RootMeanSquared:</strong> The square-root of the mean of all the squared intensity values. It is another measure of the magnitude of the image values.</li>\n    <li><strong>Skewness:</strong> Measures the asymmetry of the distribution of values about the Mean value. Depending on where the tail is elongated and the mass of the distribution is concentrated, this value can be positive or negative.</li>\n    <li><strong>TotalEnergy:</strong> The value of Energy feature scaled by the volume of the voxel in cubic mm.</li>\n    <li><strong>Uniformity:</strong> A measure of the sum of the squares of each intensity value. This is a measure of the homogeneity of the image array, where a greater uniformity implies a greater homogeneity or a smaller range of discrete intensity values.</li>\n    <li><strong>Variance:</strong> The mean of the squared distances of each intensity value from the Mean value. This is a measure of the spread of the distribution about the mean.</li>\n  </ul>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\"><strong>▶️ Shape (3D) Features:</strong></summary>\n  <ul>\n    <li><strong>MeshVolume:</strong> The volume of the ROI is calculated from the triangle mesh of the ROI. For each face in the mesh, the (signed) volume of the tetrahedron defined by that face and the origin of the image is calculated. The total volume of the ROI is obtained by taking the sum of all these volumes.</li>\n    <li><strong>VoxelVolume:</strong> The volume of the ROI is approximated by multiplying the number of voxels in the ROI by the volume of a single voxel. This is a less precise approximation of the volume and is not used in subsequent features.</li>\n    <li><strong>SurfaceArea:</strong> To calculate the surface area, first the surface area of each triangle in the mesh is calculated. The total surface area is then obtained by taking the sum of all calculated sub-areas.</li>\n    <li><strong>SurfaceVolumeRatio:</strong> This is the ratio of the surface area of the tumor region to the volume of the tumor region. A lower value indicates a more compact (sphere-like) shape.</li>\n    <li><strong>Sphericity:</strong> Sphericity is a measure of the roundness of the shape of the tumor region relative to a sphere. It is a dimensionless measure, independent of scale and orientation.</li>\n    <li><strong>Compactness1:</strong> Similar to Sphericity, Compactness 1 is a measure of how compact the shape of the tumor is relative to a sphere (most compact).</li>\n    <li><strong>Compactness2:</strong> Similar to Sphericity and Compactness 1, Compactness 2 is a measure of how compact the shape of the tumor is relative to a sphere (most compact).</li>\n    <li><strong>SphericalDisproportion:</strong> Spherical Disproportion is the ratio of the surface area of the tumor region to the surface area of a sphere with the same volume as the tumor region, and by definition, the inverse of Sphericity.</li>\n    <li><strong>Maximum3DDiameter:</strong> Maximum 3D diameter is defined as the largest pairwise Euclidean distance between tumor surface mesh vertices.</li>\n    <li><strong>Maximum2DDiameterSlice:</strong> Maximum 2D diameter (Slice) is defined as the largest pairwise Euclidean distance between tumor surface mesh vertices in the row-column (generally the axial) plane.</li>\n    <li><strong>Maximum2DDiameterColumn:</strong> Maximum 2D diameter (Column) is defined as the largest pairwise Euclidean distance between tumor surface mesh vertices in the row-slice (usually the coronal) plane.</li>\n    <li><strong>Maximum2DDiameterRow:</strong> Maximum 2D diameter (Row) is defined as the largest pairwise Euclidean distance between tumor surface mesh vertices in the column-slice (usually the sagittal) plane.</li>\n    <li><strong>MajorAxisLength:</strong> This feature yields the largest axis length of the ROI-enclosing ellipsoid and is calculated using the largest principal component.</li>\n    <li><strong>MinorAxisLength:</strong> This feature yields the second-largest axis length of the ROI-enclosing ellipsoid and is calculated using the second largest principal component.</li>\n    <li><strong>LeastAxisLength:</strong> This feature yields the smallest axis length of the ROI-enclosing ellipsoid and is calculated using the smallest principal component.</li>\n    <li><strong>Elongation:</strong> Elongation shows the relationship between the two largest principal components in the ROI shape.</li>\n    <li><strong>Flatness:</strong> Flatness shows the relationship between the largest and smallest principal components in the ROI shape.</li>\n    <li><strong>MinorAxisLength:</strong> Not found.</li>\n  </ul>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\"><strong>▶️ Gray Level Co-occurrence Matrix (GLCM) Features:</strong></summary>\n  <ul>\n    <li><strong>Autocorrelation</strong>: This is a measure of the magnitude of the fineness and coarseness of texture.</li>\n    <li><strong>ClusterProminence</strong>: A measure of the asymmetry and complexity of the texture.</li>\n    <li><strong>ClusterShade</strong>: A measure of skewness and uniformity of the texture.</li>\n    <li><strong>ClusterTendency</strong>: A measure of groupings of voxels with similar gray-level values.</li>\n    <li><strong>Contrast</strong>: A measure of the local variations in the gray-level co-occurrence matrix.</li>\n    <li><strong>Correlation</strong>: A measure of the joint probability occurrence of the specified pixel pairs.</li>\n    <li><strong>DifferenceAverage</strong>: Provides a measure of contrast in an image.</li>\n    <li><strong>DifferenceEntropy</strong>: A measure of randomness or variation in an image.</li>\n    <li><strong>DifferenceVariance</strong>: A measure of heterogeneity that places higher weights on differing intensity level pairs.</li>\n    <li><strong>Id</strong>: Inverse Difference (ID) is a measure of local homogeneity.</li>\n    <li><strong>Idm</strong>: Inverse Difference Moment (IDM) is another measure of local homogeneity.</li>\n    <li><strong>Idmn</strong>: Inverse Difference Moment Normalized (IDMN) is a measure of local homogeneity with a higher emphasis on smaller gray-level differences.</li>\n    <li><strong>Idn</strong>: Inverse Difference Normalized (IDN) is a measure of local homogeneity with a higher emphasis on larger gray-level differences.</li>\n    <li><strong>Imc1</strong>: Informational Measure of Correlation 1 (IMC1) is a measure of complexity.</li>\n    <li><strong>Imc2</strong>: Informational Measure of Correlation 2 (IMC2) is another measure of complexity.</li>\n    <li><strong>InverseVariance</strong>: A measure of the sum of squares of each intensity value.</li>\n    <li><strong>JointAverage</strong>: The average value of the joint histogram of the image.</li>\n    <li><strong>JointEnergy</strong>: A measure of similarity of the distribution of the joint histogram to that of the image.</li>\n    <li><strong>JointEntropy</strong>: A measure of uncertainty or randomness.</li>\n    <li><strong>MCC</strong>: Maximal Correlation Coefficient (MCC) is a measure of correlation between elements in the GLCM.</li>\n    <li><strong>MaximumProbability</strong>: The maximum probability value in the GLCM.</li>\n    <li><strong>SumAverage</strong>: The average sum of gray-level pairs in the GLCM.</li>\n    <li><strong>SumEntropy</strong>: A measure of randomness or complexity in the GLCM.</li>\n    <li><strong>SumSquares</strong>: A measure of the variance in the GLCM.</li>\n  </ul>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\"><strong>▶️ Gray Level Size Zone Matrix (GLSZM) Features:</strong></summary>\n  <ul>\n    <li><strong>SmallAreaEmphasis:</strong> Measures the proportion of small size zones, with a higher value indicating a greater proportion of small size zones.</li>\n    <li><strong>LargeAreaEmphasis:</strong> Measures the proportion of large size zones, with a higher value indicating a greater proportion of large size zones.</li>\n    <li><strong>GrayLevelNonUniformity:</strong> Measures the similarity of gray-level intensity values in the image, with a lower value indicating a greater similarity in intensity values.</li>\n    <li><strong>GrayLevelNonUniformityNormalized:</strong> Measures the similarity of gray-level intensity values in the image, with a lower value indicating a greater similarity in intensity values. This is normalized to account for differences in the size of the region of interest.</li>\n    <li><strong>SizeZoneNonUniformity:</strong> Measures the variability of size-zone lengths throughout the image, with a lower value indicating more homogeneity in size-zone lengths.</li>\n    <li><strong>SizeZoneNonUniformityNormalized:</strong> Measures the variability of size-zone lengths throughout the image, with a lower value indicating more homogeneity in size-zone lengths. This is normalized to account for differences in the size of the region of interest.</li>\n    <li><strong>ZonePercentage:</strong> Measures the number of size-zones per number of voxels in the region of interest.</li>\n    <li><strong>GrayLevelVariance:</strong> Measures the variance in gray-level intensity values.</li>\n    <li><strong>ZoneVariance:</strong> Measures the variance of size-zone lengths.</li>\n    <li><strong>ZoneEntropy:</strong> Measures the uncertainty/randomness of the size-zone length distribution.</li>\n    <li><strong>LowGrayLevelZoneEmphasis:</strong> Measures the proportion of low gray-level size zones.</li>\n    <li><strong>HighGrayLevelZoneEmphasis:</strong> Measures the proportion of high gray-level size zones.</li>\n    <li><strong>SmallAreaLowGrayLevelEmphasis:</strong> Measures the proportion of small size zones with lower gray-level values.</li>\n    <li><strong>SmallAreaHighGrayLevelEmphasis:</strong> Measures the proportion of small size zones with higher gray-level values.</li>\n    <li><strong>LargeAreaLowGrayLevelEmphasis:</strong> Measures the proportion of larger size zones with lower gray-level values.</li>\n    <li><strong>LargeAreaHighGrayLevelEmphasis:</strong> Measures the proportion of larger size zones with higher gray-level values.</li>\n  </ul>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\"><strong>▶️ Gray Level Run Length Matrix (GLRLM) Features:</strong></summary>\n  <ul>\n    <li><strong>ShortRunEmphasis:</strong> Measures the distribution of short runs, with a lower value indicating a greater prevalence of longer runs.</li>\n    <li><strong>LongRunEmphasis:</strong> Measures the distribution of long runs, with a higher value indicating a greater prevalence of longer runs.</li>\n    <li><strong>GrayLevelNonUniformity:</strong> Measures the similarity of run lengths throughout the image, with a lower value indicating a greater similarity in run lengths.</li>\n    <li><strong>RunLengthNonUniformity:</strong> Measures the similarity of gray-level intensity values in the image, with a lower value indicating more homogeneity.</li>\n    <li><strong>RunLengthNonUniformityNormalized:</strong> This is the normalized version of the Run Length Non-Uniformity feature.</li>\n    <li><strong>RunPercentage:</strong> This is a measure of the coarseness of the texture, with a higher value indicating a coarser texture.</li>\n    <li><strong>GrayLevelVariance:</strong> Measures the variance in gray-level intensity values in the image.</li>\n    <li><strong>RunVariance:</strong> Measures the variance of runs lengths in the image.</li>\n    <li><strong>RunEntropy:</strong> A measure of complexity or information that can be extracted from the GLRLM matrix.</li>\n    <li><strong>LowGrayLevelRunEmphasis:</strong> Measures the proportion of runs with lower gray-level values.</li>\n    <li><strong>HighGrayLevelRunEmphasis:</strong> Measures the proportion of runs with higher gray-level values.</li>\n    <li><strong>ShortRunLowGrayLevelEmphasis:</strong> Measures the joint distribution of shorter run lengths and lower gray-level values.</li>\n    <li><strong>ShortRunHighGrayLevelEmphasis:</strong> Measures the joint distribution of shorter run lengths and higher gray-level values.</li>\n    <li><strong>LongRunLowGrayLevelEmphasis:</strong> Measures the joint distribution of longer run lengths and lower gray-level values.</li>\n    <li><strong>LongRunHighGrayLevelEmphasis:</strong> Measures the joint distribution of longer run lengths and higher gray-level values.</li>\n    <li><strong>RunPercentage:</strong> The Run Percentage measures the proportion of runs in the image that have a length longer than the specified run length. It quantifies the coarseness or texture of the image, with a higher value indicating a coarser texture.</li>\n  </ul>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\"><strong>▶️ Gray Level Dependence Matrix (GLDM) Features:</strong></summary>\n  <ul>\n    <li><strong>SmallDependenceEmphasis (SDE):</strong> Measures the distribution of small dependencies with lower values indicating a greater proportion of small dependencies in the image. Small dependencies are considered those that are shorter than the specified dependence length.</li>\n    <li><strong>LargeDependenceEmphasis (LDE):</strong> Measures the distribution of large dependencies with higher values indicating a greater proportion of large dependencies in the image. Large dependencies are considered those that are longer than the specified dependence length.</li>\n    <li><strong>GrayLevelNonUniformity (GLN):</strong> Measures the similarity of gray-level intensity values in the image, with a lower value indicating a greater similarity in intensity values.</li>\n    <li><strong>GrayLevelNonUniformityNormalized (GLNN):</strong> Measures the similarity of gray-level intensity values in the image, with a lower value indicating a greater similarity in intensity values. This is a normalized version of the Gray Level Non-Uniformity feature.</li>\n    <li><strong>DependenceNonUniformity (DN):</strong> Measures the similarity of dependencies in the image, with a lower value indicating a greater similarity in dependencies.</li>\n    <li><strong>DependenceNonUniformityNormalized (DNN):</strong> Measures the similarity of dependencies in the image, with a lower value indicating a greater similarity in dependencies. This is a normalized version of the Dependence Non-Uniformity feature.</li>\n    <li><strong>GrayLevelVariance (GLV):</strong> Measures the variance in gray-level intensity values.</li>\n    <li><strong>DependenceVariance (DV):</strong> Measures the variance of dependencies in the image.</li>\n    <li><strong>DependenceEntropy (DE):</strong> Measures the uncertainty/randomness of the distribution of dependencies in the image. It measures the average amount of information required to encode the dependencies.</li>\n    <li><strong>DependencePercentage (DP):</strong> Measures the percentage of dependencies in the image that are longer than the specified dependence length.</li>\n    <li><strong>LowGrayLevelEmphasis (LGLE):</strong> Measures the proportion of low gray-level intensity values.</li>\n    <li><strong>HighGrayLevelEmphasis (HGLE):</strong> Measures the proportion of high gray-level intensity values.</li>\n    <li><strong>SmallDependenceLowGrayLevelEmphasis (SDLGLE):</strong> Measures the joint distribution of small dependencies and low gray-level values.</li>\n    <li><strong>SmallDependenceHighGrayLevelEmphasis (SDHGLE):</strong> Measures the joint distribution of small dependencies and high gray-level values.</li>\n    <li><strong>LargeDependenceLowGrayLevelEmphasis (LDLGLE):</strong> Measures the joint distribution of large dependencies and low gray-level values.</li>\n    <li><strong>LargeDependenceHighGrayLevelEmphasis (LDHGLE):</strong> Measures the joint distribution of large dependencies and high gray-level values.</li>\n    <li><strong>DependenceNonUniformity (DN):</strong> Measures the similarity of dependencies in the image, with a lower value indicating a greater similarity in dependencies.</li>\n    <li><strong>DependenceNonUniformityNormalized (DNN):</strong> Measures the similarity of dependencies in the image, with a lower value indicating a greater similarity in dependencies. This is a normalized version of the Dependence Non-Uniformity feature.</li>\n    <li><strong>DependenceVariance (DV):</strong> Measures the variance of dependencies in the image.</li>\n    <li><strong>DependenceEntropy (DE):</strong> Measures the uncertainty/randomness of the distribution of dependencies in the image. It measures the average amount of information required to encode the dependencies.</li>\n    <li><strong>DependencePercentage (DP):</strong> Measures the percentage of dependencies in the image that are longer than the specified dependence length.</li>\n    <li><strong>SmallDependenceLowGrayLevelEmphasis (SDLGLE):</strong> Measures the joint distribution of small dependencies and low gray-level values.</li>\n    <li><strong>SmallDependenceHighGrayLevelEmphasis (SDHGLE):</strong> Measures the joint distribution of small dependencies and high gray-level values.</li>\n    <li><strong>LargeDependenceLowGrayLevelEmphasis (LDLGLE):</strong> Measures the joint distribution of large dependencies and low gray-level values.</li>\n    <li><strong>LargeDependenceHighGrayLevelEmphasis (LDHGLE):</strong> Measures the joint distribution of large dependencies and high gray-level values.</li>\n    <li><strong>DependenceNonUniformity (DN):</strong> Measures the similarity of dependencies in the image, with a lower value indicating a greater similarity in dependencies. It quantifies the variation in the distribution of dependencies.</li>\n    <li><strong>DependenceNonUniformityNormalized (DNN):</strong> Measures the similarity of dependencies in the image, with a lower value indicating a greater similarity in dependencies. This is a normalized version of the Dependence Non-Uniformity feature that accounts for differences in the size of the region of interest.</li>\n    <li><strong>DependencePercentage (DP):</strong> Measures the percentage of dependencies in the image that are longer than the specified dependence length. It quantifies the proportion of dependencies in the image.</li>\n  </ul>\n</details>","metadata":{}},{"cell_type":"code","source":"if skip_brain_segmentation_pytorch_unet:\n    dataset_df = pd.read_csv(dataset_3d_path if dataset_version == DatasetVersion.V_3D else dataset_2d_path)\n    \nelse:\n    loader = widgets.IntProgress(min=0, max=len(preview_dataset_df), description='Loading:')\n    display(loader)\n    \n    # Empty creation of datasets \n    init_dataset_radiomics()\n\n    for i in preview_dataset_df.BraTS21ID :\n        loader.value += 1\n        img_resized = get_processed_image(i, dataset_version)\n        segmentation = segmentation_process(img_resized)\n\n        if dataset_version == DatasetVersion.V_3D:\n            add_patient_data3D(i, img_resized, segmentation)\n        else : \n            add_patient_data2D(i, img_resized, segmentation)\n\n    # Join the 7 datasets\n    df_shapes = df_shapes.set_index('ID')\n    df_textures = df_textures.set_index('ID')\n    df_first_orders_features = df_first_orders_features.set_index('ID')\n\n    df_gzlm_features = df_gzlm_features.set_index('ID').add_prefix('gzlm_')\n    df_glrlm_features = df_glrlm_features.set_index('ID').add_prefix('glrlm_')\n    df = df_shapes.join(df_textures).join(df_first_orders_features)\n    df_ngtdm_features = df_ngtdm_features.set_index('ID').add_prefix('ngtdm_')\n    df_gldm_features = df_gldm_features.set_index('ID').add_prefix('gldm_')\n    \n    df = df_shapes.join(df_textures).join(df_first_orders_features).join(df_gzlm_features).join(df_glrlm_features).join(df_ngtdm_features).join(df_gldm_features)\n\n    # Define 'BraTS21ID' column as integer IDs\n    df['BraTS21ID'] = df['BraTS21ID'].astype(int)\n    \n    # Merge the old dataset with the new one\n    dataset_df = pd.merge(preview_dataset_df, df, left_on='BraTS21ID', right_on='BraTS21ID')\n    dataset_df.rename(columns={'BraTS21ID': 'ID'}, inplace=True)\n\n# Patient BraTS21ID now is ID, and ID of Dataset\ndataset_df = dataset_df.set_index('ID')\n# Drop patient\ndataset_df = dataset_df.drop(index=excluded_patient_ids)\n\nshow_download_link(dataset_df)","metadata":{"_uuid":"3c7d4169-1b44-4576-bdbd-3482098642c6","_cell_guid":"98df2c38-fc59-4868-9f32-dda82317efca","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:37.064335Z","iopub.execute_input":"2023-08-24T06:12:37.064983Z","iopub.status.idle":"2023-08-24T06:12:37.313617Z","shell.execute_reply.started":"2023-08-24T06:12:37.064931Z","shell.execute_reply":"2023-08-24T06:12:37.312456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Segmentation Dataset 2D or 3D: <b>`dataset_df`</b></li> \n            <li>We excluded patient Ids with wrong data for the rest of the project.</li>\n        </ul>\n    </div>\n</div>","metadata":{"_uuid":"0de17d8c-17d0-4577-a91d-5bec55cbb907","_cell_guid":"bfb66090-3f88-47db-ae66-67a103dae18c","trusted":true}},{"cell_type":"markdown","source":"----\n\n# <a id='6'>6. Data Exploration</a>\n\nIn this section, we will perform an exploratory data analysis to gain an initial overview of the new brain MRI dataset generated by the segmentation.","metadata":{"_uuid":"146734cf-c101-4223-82ab-cbe6b5bcd05d","_cell_guid":"288af7d4-080b-43eb-9be2-4a979a3ecd07","trusted":true}},{"cell_type":"code","source":"dataset_df.head()","metadata":{"_uuid":"5ccde63a-9288-44ea-995f-3d04a83c3fea","_cell_guid":"b32ba107-fb49-45a4-a3bc-6dc4c7ea990d","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:37.314939Z","iopub.execute_input":"2023-08-24T06:12:37.315261Z","iopub.status.idle":"2023-08-24T06:12:37.430014Z","shell.execute_reply.started":"2023-08-24T06:12:37.315222Z","shell.execute_reply":"2023-08-24T06:12:37.428449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_df.info()","metadata":{"_uuid":"0ad26915-cc1f-4626-a0e7-65d4adac4646","_cell_guid":"7b82d05f-9b77-46f1-bb7c-1318037cf75b","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:37.431687Z","iopub.execute_input":"2023-08-24T06:12:37.432086Z","iopub.status.idle":"2023-08-24T06:12:37.454763Z","shell.execute_reply.started":"2023-08-24T06:12:37.432054Z","shell.execute_reply":"2023-08-24T06:12:37.453474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_df.describe()","metadata":{"_uuid":"c2528cda-f551-4897-8130-28f51a447494","_cell_guid":"644a88a8-43b1-4faf-88b9-5180fd04fd4d","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:37.456163Z","iopub.execute_input":"2023-08-24T06:12:37.456537Z","iopub.status.idle":"2023-08-24T06:12:37.871614Z","shell.execute_reply.started":"2023-08-24T06:12:37.456506Z","shell.execute_reply":"2023-08-24T06:12:37.869733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n   <div class=\"card-body\">\n      <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n      <p style=\"color: white;\">Based on the provided analysis, we can draw the following conclusions:</p>\n      <ul>\n         <li>\n            <strong>Distribution & Outliers:</strong>\n            <ul>\n               <li>Most have means that are significantly higher than the median, suggesting positive skew. Variables like <strong>'LeastAxisLength', 'Flatness', 'TotalEnergy', 'Maximum3DDiameter', 'Energy', 'Variance'</strong>, etc.</li>\n               <li>The variables have positive skewness with large values. This can complicate data analysis.</li>\n               <li>See also in the previous summary, have some wrong data to remove.</li>\n               <ul>\n                   <li>Apply feature scaling techniques to address the wide range of values and improve machine learning algorithm performance.</li>\n               </ul>\n            </ul>\n         </li>\n         <li>\n            <strong>Zero Values:</strong>\n            <ul>\n               <li>Variables such as <strong>\"LeastAxisLength\", \"Flatness\", \"gldm_GrayLevelNonUniformityNormalized\", \"gldm_DependencePercentage\"</strong> show a minimum value of 0 or NaN, which might indicate missing or incorrect data.</li>\n               <ul>\n                   <li>We now remove these variables for the rest of the project.</li>\n               </ul>\n            </ul>\n         </li>\n         <li>\n            <strong>Absence of Negative Values:</strong>\n            <ul>\n               <li>Most variables do not contain negative values, which makes sense considering they seem to represent physical measurements. (Brain MRI).</li>\n            </ul>\n         </li>\n         <li>\n            <strong>Potential Redundancy:</strong>\n            <ul>\n               <li>The <strong>'Id'</strong>, <strong>'Idm'</strong>, <strong>'Idmn'</strong>, and <strong>'Idn'</strong>, etc., variables appear to be very similar in terms of their statistical properties, which might indicate redundancy.</li>\n               <ul>\n                   <li>Potential redundancy, so high correlation. Potential variables to remove.</li>\n               </ul>\n            </ul>\n         </li>\n         <li>\n            <strong>Data Correlations:</strong>\n            <ul>\n               <li>This basic statistical summary doesn't provide any information about possible correlations between <strong>'MeshVolume'</strong>, <strong>'VoxelVolume'</strong>, and <strong>'MGMT_value'</strong>.</li>\n               <ul>\n                   <li>The brain size does not affect MGMT values.</li>\n               </ul>\n            </ul>\n         </li>\n         <li>\n            <strong>Statistical Variables:</strong>\n            <ul>\n               <li>The statistical variables included in the table, such as <strong>'Mean'</strong>, <strong>'Median'</strong>, <strong>'Maximum'</strong>, etc., are derived summaries or calculations based on other variables.</li>\n               <ul>\n                   <li>Potential for analysis and perhaps potential exclusion.</li>\n               </ul>\n            </ul>\n         </li>\n      </ul>\n      <p style=\"color: white;\">👉🏿 It's important to keep in mind that potential errors, biases, or other anomalies in how the data was gathered could influence these statistical properties.</p>\n   </div>\n</div>","metadata":{}},{"cell_type":"code","source":"# First Orders features\nfirst_orders_features_columns_names =['10Percentile', '90Percentile', 'Energy', 'Entropy', 'InterquartileRange', 'Kurtosis', \n                                      'Maximum', 'MeanAbsoluteDeviation', 'Mean', 'Median', 'Minimum', 'Range', 'RobustMeanAbsoluteDeviation', \n                                      'RootMeanSquared', 'Skewness', 'TotalEnergy', 'Uniformity', 'Variance']\n\n# Shape 3D  \nshapes_features_columns_names = ['MeshVolume','VoxelVolume','SurfaceArea','SurfaceVolumeRatio','Sphericity',\n                             'Compactness1','Compactness2','SphericalDisproportion','Maximum3DDiameter','Maximum2DDiameterRow',\n                             'Maximum2DDiameterColumn','MajorAxisLength','MinorAxisLenth','Elongation']\n\n# GLCM features\ntextures_features_columns_names = ['Autocorrelation', 'ClusterProminence', 'ClusterShade', 'ClusterTendency', 'Contrast', \n                           'Correlation', 'DifferenceAverage', 'DifferenceEntropy', 'DifferenceVariance', 'Id', 'Idm', \n                           'Idmn', 'Idn', 'Imc1', 'Imc2', 'InverseVariance', 'JointAverage', 'JointEnergy', 'JointEntropy', \n                           'MCC', 'MaximumProbability', 'SumAverage', 'SumEntropy', 'SumSquares']\n \n# GZLM features\ngzlm_features_columns_names = ['gzlm_SmallAreaEmphasis','gzlm_LargeAreaEmphasis','gzlm_GrayLevelNonUniformity','gzlm_GrayLevelNonUniformityNormalized',\n                                'gzlm_SizeZoneNonUniformity','gzlm_SizeZoneNonUniformityNormalized','gzlm_ZonePercentage','gzlm_GrayLevelVariance','gzlm_ZoneVariance',\n                                'gzlm_ZoneEntropy','gzlm_LowGrayLevelZoneEmphasis','gzlm_HighGrayLevelZoneEmphasis','gzlm_SmallAreaLowGrayLevelEmphasis',\n                                'gzlm_SmallAreaHighGrayLevelEmphasis','gzlm_LargeAreaLowGrayLevelEmphasis','gzlm_LargeAreaHighGrayLevelEmphasis']\n\n# GLRLM\nglrlm_features_columns_names = ['glrlm_ShortRunEmphasis','glrlm_LongRunEmphasis','glrlm_GrayLevelNonUniformity','glrlm_RunLengthNonUniformity','glrlm_RunLengthNonUniformityNormalized'\n                                 ,'glrlm_RunPercentage','glrlm_GrayLevelVariance','glrlm_RunVariance','glrlm_RunEntropy','glrlm_LowGrayLevelRunEmphasis','glrlm_HighGrayLevelRunEmphasis'\n                                 ,'glrlm_ShortRunLowGrayLevelEmphasis','glrlm_ShortRunHighGrayLevelEmphasis','glrlm_LongRunLowGrayLevelEmphasis','glrlm_LongRunHighGrayLevelEmphasis']\n\n# NGTM features\nngtdm_features_columns_names = ['ngtdm_Coarseness','ngtdm_Contrast','ngtdm_Busyness','ngtdm_Complexity','ngtdm_Strength']\n\n    \n# GLDM features\ngldm_features_columns_names = ['gldm_SmallDependenceEmphasis','gldm_LargeDependenceEmphasis','gldm_GrayLevelNonUniformity','gldm_DependenceNonUniformity'\n                               , 'gldm_DependenceNonUniformityNormalized', 'gldm_GrayLevelVariance', 'gldm_DependenceVariance', 'gldm_DependenceEntropy'\n                                , 'gldm_LowGrayLevelEmphasis', 'gldm_HighGrayLevelEmphasis', 'gldm_SmallDependenceLowGrayLevelEmphasis', 'gldm_SmallDependenceHighGrayLevelEmphasis'\n                               , 'gldm_LargeDependenceLowGrayLevelEmphasis', 'gldm_LargeDependenceHighGrayLevelEmphasis']\n\nall_features_column_names = {\n    'first': first_orders_features_columns_names,\n    'shapes': shapes_features_columns_names,\n    'textures': textures_features_columns_names,\n    'gzlm': gzlm_features_columns_names,\n    'glrlm': glrlm_features_columns_names,\n    'ngtdm': ngtdm_features_columns_names,\n    'gldm': gldm_features_columns_names,\n}\n\n# Drop\nexcluded_features_column_names = [\"LeastAxisLength\", \"Flatness\", \"gldm_GrayLevelNonUniformityNormalized\", \"gldm_DependencePercentage\"]\ndataset_df.drop(excluded_features_column_names, axis=1, inplace=True)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:37.874602Z","iopub.execute_input":"2023-08-24T06:12:37.875500Z","iopub.status.idle":"2023-08-24T06:12:37.892700Z","shell.execute_reply.started":"2023-08-24T06:12:37.875427Z","shell.execute_reply":"2023-08-24T06:12:37.891170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>First orders Features: <b>`first_orders_features_columns_names`</b></li>\n            <li>Shape 3D Features: <b>`shapes_features_columns_names`</b></li>\n            <li>textures Features: <b>`textures_features_columns_names`</b></li>      \n            <li>gzlm Features: <b>`gzlm_features_columns_names`</b></li>      \n            <li>glrlm Features: <b>`glrlm_features_columns_names`</b></li>    \n            <li>ngtdm Features: <b>`ngtdm_features_columns_names`</b></li>\n            <li>gldm Features: <b>`ngtdm_features_columns_names`</b></li>\n            <li>All Features: <b>`all_features_column_names`</b></li>\n            <li>We exluded features with <b>zero or NaN:</b></li>\n            <ul>\n                <li><b>[\"LeastAxisLength\", \"Flatness\", \"gldm_GrayLevelNonUniformityNormalized\", \"gldm_DependencePercentage]</b></li>\n            </ul>\n            </li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"----\n\n# <a id='7'>7. Analysis</a>\n\nIn this section, we present a detailed analysis of the dataset, aiming to gain deeper insights into its underlying patterns and characteristics. Through various analytical techniques and visualizations, we explore the data to uncover meaningful information and identify key factors that may influence the target variable. The analysis encompasses both univariate and multivariate approaches, allowing us to examine variables individually as well as their relationships and interactions.\n\nBy conducting a thorough examination of the dataset, we strive to uncover hidden patterns, trends, and potential outliers. This analysis serves as a crucial step in understanding the data and informing subsequent decision-making processes. Through a combination of statistical measures, data visualization techniques, and exploratory data analysis, we aim to provide a comprehensive understanding of the dataset's structure and potential insights it may offer.\n\nThe results and findings obtained from this analysis will serve as a solid foundation for further investigations, model development, and decision-making in subsequent stages of the project.\n\n## <a id='7_1'>7.1. Group Analysis</a>\n\nIn this section, we delve into a group analysis. We aim to uncover meaningful patterns and similarities among the data points within the dataset.\n\nThrough this group analysis, we aim to understand the underlying structure and organization of the data, identifying subgroups and assessing the similarities between different segments. This exploration helps in uncovering hidden patterns and relationships, contributing to a deeper understanding of the dataset.\n\n### <a id='7_1_1'>7.1.1. Hierarchical by Clustering Dendrogram</a>\n\nThe Hierarchical Clustering Dendrogram is a powerful method used to group brain images based on their structural and functional similarities. By applying this technique, we aim to uncover underlying patterns and relationships within the dataset.\n\nThe resulting dendrogram provides a visual representation of the hierarchical structure of the grouped brain images. It enables us to identify distinct subgroups and assess the similarities between different brain segmentations. This analysis helps in understanding the complex patterns and relationships that exist within the dataset, shedding light on the underlying structure of the brain images.\n\nBy leveraging the Hierarchical Clustering Dendrogram, we gain valuable insights into the clustering behavior and organization of the brain images. This approach serves as a robust tool for analyzing and comprehending the intricate patterns and connections that exist within the brain dataset, facilitating a deeper understanding of the data and its potential implications.","metadata":{"_uuid":"9a500165-8624-4966-9706-0d8993b923d8","_cell_guid":"cc2fbfa2-c63b-44e6-af99-8928aa29c48e","trusted":true}},{"cell_type":"code","source":"# Extract features from the DataFrame\nfeatures = dataset_df.drop(\"MGMT_value\", axis=1)\n\n# Calculate the distances between features\ndistances = sch.distance.pdist(features)\n\n# Perform hierarchical clustering and obtain the linkage matrix\nlinkage_matrix = sch.linkage(distances, method='ward')\n\n# Plot the dendrogram\nplt.figure(figsize=(30, 8))\ndendrogram = sch.dendrogram(linkage_matrix, no_labels=True)\n# Adjust the spacing of y-axis tick labels\n\nplt.xlabel('Scans ID')\nplt.ylabel('Distance')\n\nplt.title('Hierarchical Clustering Dendrogram of Scans ID')\n\nplt.tight_layout()\nplt.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:37.894558Z","iopub.execute_input":"2023-08-24T06:12:37.895748Z","iopub.status.idle":"2023-08-24T06:12:38.438659Z","shell.execute_reply.started":"2023-08-24T06:12:37.895705Z","shell.execute_reply":"2023-08-24T06:12:38.437329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_clusters = 3\nlabels = sch.fcluster(linkage_matrix, num_clusters, criterion='maxclust')\n\n# Create a dictionary to group features by label\ncluster_three_groups = {}\nfor i, label in enumerate(labels):\n    if label not in cluster_three_groups:\n        cluster_three_groups[label] = []\n    cluster_three_groups[label].append(features.index[i])","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:38.440126Z","iopub.execute_input":"2023-08-24T06:12:38.440482Z","iopub.status.idle":"2023-08-24T06:12:38.450923Z","shell.execute_reply.started":"2023-08-24T06:12:38.440453Z","shell.execute_reply":"2023-08-24T06:12:38.450019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With the dentogram we can clearly emerge 3 distinct groups. (Distance 3)","metadata":{}},{"cell_type":"code","source":"def check_and_display_group(key, patients, anomaly_type):\n    if np.isin(patients, cluster_three_groups[key]).any():\n        display(HTML(f\"<h3>Group {key} ({anomaly_type})</h3>\"))\n        patient_with_wrong_mri_ids = preview_dataset_df.loc[preview_dataset_df[\"BraTS21ID\"].isin(patients)]\n        patient_with_wrong_mri_ids_tuples = map_dataframe_to_tuples(patient_with_wrong_mri_ids, [\"BraTS21ID\", \"MGMT_value\"])\n        show_brains(patient_with_wrong_mri_ids_tuples, figsize=(8, 5), dataset_version=dataset_version)\n\nkey = 1\npatients = [111, 456]\ncheck_and_display_group(key, patients, \"ROI with compact shape\")\n\nkey = 2\npatients = [35, 380, 305]\ncheck_and_display_group(key, patients, \"ROI with compact and dispersed shape with noisy\")\n\npatients = [351, 442, 589]\ncheck_and_display_group(key, patients, \"Imperfect, anomaly\")\n\nkey = 3\npatients = [810, 483]\ncheck_and_display_group(key, patients, \"ROI with dispersed shape\")\n","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:12:38.452589Z","iopub.execute_input":"2023-08-24T06:12:38.453123Z","iopub.status.idle":"2023-08-24T06:13:51.506907Z","shell.execute_reply.started":"2023-08-24T06:12:38.453089Z","shell.execute_reply":"2023-08-24T06:13:51.505356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>        \n        <p style=\"color: white;\">From analyzing the results of the dendrogram, here are some conclusions that can be drawn:</p>\n        <ul>\n            <li><b>Group 1:</b> Shows a group with potential scans with ROI with compact shape.</li>\n            <li><b>Group 2:</b> Shows a group with two potential types.</li>\n            <ul>\n                <li>A group where the segmentation is more dispersed or compact with artifacts, the scans are potentially less perfect, but still useful.</li>\n                <li>In this group with can see some anomaly scans that can be removed (Does not respect the quality advice seen upstream, Positioning errors, Parameter setting errors, Region of interest not covered, Insufficient image quality).</li>\n            </ul>\n            <li><b>Group 3:</b> Shows a group with potential scans with ROI with dispersed shape.</li>            \n        </ul>\n        <p>We do not see any significant values or correlations with group and MGMT.</p>\n    </div>\n</div>\n","metadata":{}},{"cell_type":"code","source":"patients_with_wrong_scans_ids = [11, 351, 353, 442, 445, 554, 556, 568, 570, 572, 578, 581, 584, 587, 589, 593, 594, 596, 610]\ndataset_df.drop(index=patients_with_wrong_scans_ids, inplace=True)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:13:51.509170Z","iopub.execute_input":"2023-08-24T06:13:51.510111Z","iopub.status.idle":"2023-08-24T06:13:51.519377Z","shell.execute_reply.started":"2023-08-24T06:13:51.510052Z","shell.execute_reply":"2023-08-24T06:13:51.518479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Cluster group level 3: <b>`cluster_three_groups`</b></li>\n            <li>We drop anomaly scans where does not fit with the quality guidance: <b>`[11, 351, 353, 442, 445, 554, 556, 568, 570, 572, 578, 581, 584, 587, 589, 593, 594, 596, 610]`</b></li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### <a id='7_1_2'>7.1.2. Average Patient MGMT Comparison</a>\n\nIn this section, we present a comparative analysis between the average patient with MGMT 1 (Mean) and the average patient with MGMT 0 (Mean). By examining the differences in values between these two groups, we aim to identify any notable variations and understand their potential impact.\n\nThrough a detailed analysis, we calculate and summarize the differences in values between the average MGMT 1 patient and the average MGMT 0 patient. For each variable, we indicate whether there is an increase or decrease in the values and provide the corresponding percentage change. This analysis allows us to quantify and assess the magnitude of differences between the two groups.\n\nBy highlighting these differences and their potential impact in terms of percentage change, we gain insights into the distinct characteristics and variations associated with different MGMT values. This knowledge can be valuable in understanding the factors that contribute to MGMT-related outcomes and inform decision-making processes in the context of the dataset.","metadata":{}},{"cell_type":"code","source":"def fun_significant(x):\n    if float(x[:-1]) >= 1:\n        return True\n    elif float(x[:-1]) <= 1:\n        return True\n    else:\n        return False\n    \ndef fun_significant_symbol(x):\n    value = float(x[:-1])\n    if value > 1:\n        return \"⬆️\"\n    elif value < -1:\n        return \"⬇️\"\n    else:\n        return \"-\"\n\n# Select the first row for patients with MGMT_value = 1\npatient_with_mgmt = dataset_df.loc[dataset_df[\"MGMT_value\"] == 1].mean()\n\n# Select the first row for patients with MGMT_value = 0\npatient_without_mgmt = dataset_df.loc[dataset_df[\"MGMT_value\"] == 0].mean()\n\n# Transpose the dataframes\npatient_with_mgmt = patient_with_mgmt.T\npatient_without_mgmt = patient_without_mgmt.T\n\n# Concatenate the transposed dataframes horizontally\nsignificant_values_df = pd.concat([patient_without_mgmt, patient_with_mgmt, patient_without_mgmt - patient_with_mgmt], axis=1)\n\n# Rename the columns\nsignificant_values_df.columns = ['Patient MGMT 0 (Mean)', 'Patient MGMT 1 (Mean)', 'Difference']\n\n# Calculate the percentage difference\nsignificant_values_df['Difference (%)'] = (significant_values_df['Patient MGMT 1 (Mean)'] - significant_values_df['Patient MGMT 0 (Mean)']) / significant_values_df['Patient MGMT 0 (Mean)'] * 100\n\nsignificant_values_df = significant_values_df.sort_values(by=['Difference (%)'], ascending=False)\n\n# Set the format of the 'Difference (%)' column\nsignificant_values_df['Difference (%)'] = np.where((np.isinf(significant_values_df['Difference (%)'])) | (significant_values_df['Difference (%)'].isna()), '0%', significant_values_df['Difference (%)'].apply(lambda x: f\"{x:.2f}%\"))\n\n#Drop MGMT\nsignificant_values_df = significant_values_df.drop(significant_values_df.index[0])\n\n# Add the 'Significant' and 'No Significant' columns\nsignificant_values_df['Significant'] = significant_values_df['Difference (%)'].apply(fun_significant_symbol)\n\n\nsignificant_values_df","metadata":{"_uuid":"c94296ca-9434-4f8a-9db9-c5807654a5fa","_cell_guid":"0d6d20ca-3061-436a-8b21-195682b1e9f4","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:13:51.527766Z","iopub.execute_input":"2023-08-24T06:13:51.528222Z","iopub.status.idle":"2023-08-24T06:13:51.596112Z","shell.execute_reply.started":"2023-08-24T06:13:51.528185Z","shell.execute_reply":"2023-08-24T06:13:51.594832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>        \n        <p style=\"color: white;\">From analyzing the results of the table, here are some conclusions that can be drawn:</p>\n        <ul>\n        <li>Among other variables, some also show significant differences between patients with and without the MGMT value:</li>\n            <ul>\n                <li>Such as <b>\n                                'ClusterShade', 'Kurtosis', 'ngtdm_Strength', 'gzlm_LargeAreaLowGrayLevelEmphasis',\n                                'Skewness' etc..\n                             </b>.\n                    </li> \n                <li>These differences suggest that these variables are potentially important in distinguishing patients with and without the MGMT value.</li>\n            </ul>\n        <li>Some variables do not show significant differences:\n            <ul>\n                <li>Such as \n                    <b>'Maximum', 'ngtdm_Contrast', 'Idmn' etc..</b>.\n                </li> \n                <li>This suggests that these variables have no significant impact on the presence or absence of the MGMT value.</li>\n            </ul>\n        <li>It is important to note that the percentage differences can vary significantly across variables. Some variables show relatively small differences\n            <ul>\n                <li>Such as <b>\"SurfaceVolumeRatio\"</b> with 0.01%, while others exhibit more substantial differences.</li> \n                <li>Such as <b>\"ClusterShade\"</b> with 65.77% or <b>\"Kurtosis\"</b> with 38.62%.</li>\n                <li>These percentage variations indicate the magnitude of the impact of each variable on the presence or absence of the MGMT value.</li>\n            </ul>\n        </ul>        \n        <p style=\"color: white;\">👉🏿 Overall, these results suggest that some variables have a significant influence on the MGMT value, while others do not have a substantial impact. This information can be valuable in understanding the factors associated with the MGMT value and in developing predictive models or appropriate treatment approaches.</p>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Significant values: <b>`significant_values_df`</b></li>            \n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## <a id='7_2'>7.2. Univariate Analysis</a>\n\nIn this section, we undertake a thorough univariate analysis to explore the characteristics of individual variables within the dataset. The primary objective of this analysis is to identify potentially significant variables that exhibit distinct characteristics or trends.\n\nEach variable is examined individually, allowing us to gain insights into its distribution, descriptive statistics, and patterns. By scrutinizing the variables in isolation, we aim to uncover any noteworthy characteristics that could indicate their importance within the dataset. This meticulous analysis provides a foundation for further investigations and supports the selection of variables for subsequent modeling or analysis phases.\n\nDuring the univariate analysis, we consider various factors such as the shape of the distribution, measures of central tendency, dispersion, and any potential outliers or peculiar patterns. These insights enable us to understand the unique properties and relevance of each variable, informing subsequent data-driven decisions.\n\nThe findings from this comprehensive univariate analysis serve as a critical step in understanding the dataset's composition and identifying variables that may play a significant role in further analyses or modeling endeavors.\n\n### <a id='7_2_1'>7.2.1 Assessing Normality and Distributions</a>\n\nIn this section, we perform a thorough assessment of the normality of a quantitative variable, compare its distribution, and evaluate any deviations from normality. These visualizations play a vital role in statistical analysis, exploratory data analysis, and modeling, providing valuable insights into data distributions, identifying outliers, and assessing goodness-of-fit.\n\nThrough comprehensive visual analysis, we aim to gain a deeper understanding of the underlying distribution of the variable of interest. This includes examining various graphical representations such as histograms, kernel density plots, and quantile-quantile (Q-Q) plots. These techniques allow us to visually assess the shape, skewness, and presence of outliers in the distribution.\n\nAssessing the normality and distribution of the variable is crucial for ensuring appropriate modeling and analysis techniques. By identifying departures from normality, we can make informed decisions on data transformations, robust statistical methods, or the need for further investigation into potential underlying factors contributing to non-normality.\n\nThese visualizations serve as valuable tools for gaining insights into the data's distributional characteristics and informing subsequent analyses. They aid in building accurate statistical models and making reliable inferences from the data.\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ Histogram with Density Estimation:</summary>\n  <p>\n    A histogram is a graphical representation of the distribution of data. It is commonly used to visualize the frequency or count of observations falling into different intervals or bins. In addition to the histogram bars, a probability density curve can be overlaid on the histogram. This curve represents the theoretical distribution that the data may follow. It provides insights into the shape and characteristics of the data distribution.\n  </p>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ Density Plot with Normal Distribution:</summary>\n  <p>\n    The density plot with a normal distribution overlays a kernel density estimation of the data along with a curve representing the normal distribution. It allows for visual comparison between the data distribution and the theoretical normal distribution. Deviations between the two curves can indicate departures from normality in the data.\n  </p>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ Boxplot:</summary>\n  <p>\n    A boxplot, also known as a box-and-whisker plot, provides a visual summary of the distribution of the data through quartiles. The box represents the interquartile range (IQR) and the line inside the box represents the median. The whiskers extend to the minimum and maximum values, excluding outliers, which are often plotted as individual points. It is useful for comparing multiple datasets or analyzing the spread and skewness of a single dataset.\n  </p>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ Violin Plot:</summary>\n  <p>\n    A violin plot combines aspects of a boxplot and a kernel density plot. It displays the distribution of data by featuring a rotated kernel density plot on each side of a central axis. The width of the violin at a specific point represents the density or concentration of data at that value. The violin plot provides insights into both the summary statistics (like a boxplot) and the underlying distribution of the data.\n  </p>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ Q-Q Plot (Quantile-Quantile Plot):</summary>\n  <p>\n    A Q-Q plot is a graphical tool used to assess how well a given sample of data matches a theoretical distribution. It compares the quantiles of the data to the quantiles of a theoretical distribution. If the points on the plot roughly follow a straight line, it indicates a good fit between the data and the theoretical distribution. Deviations from the straight line suggest departures from the assumed distribution. Q-Q plots are especially useful for assessing normality assumptions.\n  </p>\n</details>\n\n<details>\n  <summary style=\"font-weight: bold; font-size: 1.2em;\">▶️ Scatter Plot of Values vs. Mean:</summary>\n  <p>\n    The scatter plot shows the relationship between individual values of the quantitative variable and the mean value. Each group defined by the \"MGMT_value\" category is represented by different colors or markers. The horizontal line represents the mean value of the variable. This plot provides insights into how the individual values are distributed around the mean.\n  </p>\n</details>\n<br>\n","metadata":{"_uuid":"9b9a51f2-5847-4cde-afe0-417787afd0fa","_cell_guid":"e47622c2-7b4b-498a-bcc0-b66eae7641fb","trusted":true}},{"cell_type":"code","source":"display(HTML(\"<h3>3 first relevant ⬆️</h3>\"))\nfirst_5 = significant_values_df.head(3)\nfor column_name in first_5.index:\n    show_explore_distribution_normality(dataset_df, column_name)    \n\n\ndisplay(HTML(\"<h3>3 last relevant ⬇️</h3>\"))\nlast_5 = significant_values_df.tail(3)\nfor column_name in last_5.index:\n    show_explore_distribution_normality(dataset_df, column_name)  \n    \ndisplay(HTML(\"<h3>Sample Follow normaly law (Correct skewness) ⏺</h3>\"))    \nshow_explore_distribution_normality(dataset_df, \"MeshVolume\")\nshow_explore_distribution_normality(dataset_df, \"SurfaceArea\")\n\ndisplay(HTML(\"<h3>Sample Follow normaly law (Good kurtosis) ⏺</h3>\"))\nshow_explore_distribution_normality(dataset_df, \"Sphericity\")\nshow_explore_distribution_normality(dataset_df, \"Compactness1\")\n\n\ndisplay(HTML(\"<h3>Sample Follow normaly law (Good Q-Q) ⏺</h3>\"))\nshow_explore_distribution_normality(dataset_df, \"Elongation\")\nshow_explore_distribution_normality(dataset_df, \"DifferenceAverage\")\n\n","metadata":{"_uuid":"8355d298-9bce-470c-9e09-6267a5bad61a","_cell_guid":"94d3fc65-ecb3-423a-932b-9a630d304bda","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:13:51.597758Z","iopub.execute_input":"2023-08-24T06:13:51.598196Z","iopub.status.idle":"2023-08-24T06:14:27.452837Z","shell.execute_reply.started":"2023-08-24T06:13:51.598155Z","shell.execute_reply":"2023-08-24T06:14:27.451202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>The dataset exhibits homogeneity in the target values.</li>\n            <li>When conducting the Shapiro-Wilk test, many explanatory variables do not follow a normal distribution. However, some variables show close approximation to a normal distribution.</li>\n            <li>Variables with correct skewness: <b>MeshVolume, SurfaceArea, MajorAxisLength, Contrast, etc.</b></li>\n            <li>Variables with good kurtosis: <b>Sphericity, Compactness1, Elongation, DifferenceAverage, DifferenceEntropy, Idn, JointAverage, SumAverage, Mean, etc.</b></li>\n            <li>Upon analyzing the graphs, some variables exhibit a good fit on the Q-Q plot without significant extreme values: <b>Elongation, DifferenceAverage, DifferenceEntropy, JointEntropy, SumEntropy, Entropy, Maximum, Range, etc.</b></li>  \n            <li>The distribution of values generally appears similar, but noticeable differences can be observed between patients with different MGMT values.</li>\n        </ul>\n        <p style=\"color: white;\">It is important to note that these observations do not necessarily imply that the dataset is flawed. Instead, they highlight the presence of variables with extreme values.</p>\n    </div>\n</div>","metadata":{"_uuid":"694e0daa-0f72-4c41-b32a-fb872b492b89","_cell_guid":"a6816578-a68f-4031-93c4-5849e45d1745","trusted":true}},{"cell_type":"markdown","source":"### <a id='7_2.2'>7.2.2. Assessing Normality, Skewness, and Kurtosis</a>\n\nIn this section, we delve into the analysis of normality, skewness, and kurtosis to gain a comprehensive understanding of the distributional characteristics of the variables. These tests provide valuable insights into the shape, symmetry, and tail behavior of the data.\n\n#### Normality Test: Shapiro-Wilk\nThe normality test employed in this analysis is the Shapiro-Wilk test. This test is well-suited for small sample sizes and provides precise results. It evaluates whether the dataset follows a normal distribution. The Shapiro-Wilk test assesses the hypothesis that the data is normally distributed based on the sample data, yielding a p-value that indicates the level of deviation from normality.\n\n#### Skewness Test\nThe skewness test measures the asymmetry of a series. A skewness value of 0 indicates a perfectly symmetric distribution following the normal law. However, non-zero skewness values indicate the presence and direction of asymmetry. Positive skewness suggests a longer right tail, while negative skewness suggests a longer left tail.\n\n#### Kurtosis Test\nThe kurtosis test measures the shape and flatness of the distribution. A kurtosis value of 3 corresponds to Laplace's normal law, indicating a distribution with the same tail thickness as a standard normal distribution. However, we also consider the excess kurtosis.\n\n- If the kurtosis is greater than 3, the dataset is leptokurtic, indicating thicker tails than a normal distribution. This suggests a grouping of outliers or extreme values.\n- If the kurtosis is less than 3, the dataset is platykurtic, indicating thinner tails than a normal distribution. This implies a negative excess of outliers, with most of the data clustering around the mean.\n- When the kurtosis is equal to 3, the dataset is mesokurtic, signifying that the tails of the distribution have the same thickness as a normal distribution.\n\nThe excess kurtosis is calculated by subtracting 3 from the kurtosis value, providing further insights into the deviation from a normal distribution.\n\nBy conducting these tests, we gain a deeper understanding of the normality, skewness, and kurtosis properties of the variables. These insights inform subsequent analyses, decision-making processes, and modeling approaches, ensuring the appropriate handling of the data based on its distributional characteristics.\n\n","metadata":{}},{"cell_type":"code","source":"describe = show_test_normality(dataset_df.drop('MGMT_value',axis=1))\ndisplay(describe)\n\ndescribe_transposed = describe.loc[['skewness','kurtosis','excess_kurtosis']].transpose()\ndescribe_transposed['absolute_value'] = np.abs(describe_transposed['excess_kurtosis']).astype('float')\ntop_50_closest = describe_transposed.nsmallest(50, 'absolute_value').sort_values('absolute_value')\n\ndescribe_transposed = describe.loc[['skewness', 'kurtosis', 'excess_kurtosis']].transpose()\ndescribe_transposed['absolute_value'] = np.abs(describe_transposed['excess_kurtosis']).astype(float)\n\n# Calculer la différence entre la colonne 'skewness' et la valeur idéale de 1\ndescribe_transposed['skewness_diff'] = np.abs(describe_transposed['skewness'] - 1).astype(float)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:14:27.454212Z","iopub.execute_input":"2023-08-24T06:14:27.454598Z","iopub.status.idle":"2023-08-24T06:14:27.713951Z","shell.execute_reply.started":"2023-08-24T06:14:27.454566Z","shell.execute_reply":"2023-08-24T06:14:27.712864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\nSome variables have a better skewness with an MGMT of 1 than compared to the global dataset and vice versa, these same variables are therefore sensitive to markers. Similarly Kurtosis is sensitive too. Depending on the variables, the correlation is positive or negative between skewness and kurtosis.\n        <ul>\n            <li>None of the variables pass the Shapiro-Wilk test, indicating that the explanatory variables do not follow a normal distribution.</li>\n            <li>Some variables exhibit different skewness and kurtosis values between MGMT 1 and the global dataset, indicating sensitivity to markers.</li>\n            <li>The correlation between skewness and kurtosis varies depending on the variables.</li>\n            <li>The estimated density function tends to have a higher maximum value for patients with MGMT value 1, except for the kurtosis variable.</li>\n            <li>The estimated density function for patients with MGMT value 1 tends to overlap with the one for patients with MGMT value 0.</li>\n            <li>The most normal variables are <b>`MajorAxisLength`,`glrlm_RunLengthNonUniformity`, `Elongation`</b></li>\n        </ul>\n    </div>\n</div>","metadata":{"_uuid":"17fe8764-e4e4-4619-b907-45fa39b3984a","_cell_guid":"8b102231-f537-4e52-a935-aef3f8b36f27","trusted":true}},{"cell_type":"markdown","source":"### <a id='7_2_3'>7.2.3. Comparison of Kurtosis Values for Different Features</a>\n\nIn this section, we perform a comparison of kurtosis values for different features within the dataset. Kurtosis is a statistical measure that quantifies the shape and flatness of a distribution. By analyzing the kurtosis values, we gain insights into the varying degrees of tail behavior exhibited by different features.\n\nThe comparison allows us to identify features that deviate significantly from a normal distribution. We consider the excess kurtosis, which measures the deviation from a normal distribution by subtracting 3 from the kurtosis value. This provides a clearer understanding of the nature of the tails in each feature's distribution.\n\nBy assessing and comparing the kurtosis values, we can identify features with leptokurtic distributions, characterized by thicker tails and potentially higher frequencies of extreme values. Conversely, features with platykurtic distributions exhibit thinner tails and a lower frequency of outliers. Understanding the kurtosis values aids in identifying features with distinct distributional properties and tail behavior.\n\nThis analysis provides valuable insights into the distributional characteristics of different features, enabling us to make informed decisions in subsequent modeling or analysis phases. By considering the kurtosis values, we can tailor our data handling and modeling approaches to appropriately accommodate the distributional properties of each feature.","metadata":{}},{"cell_type":"code","source":"describe_MGMT_1=show_test_normality(dataset_df[dataset_df.MGMT_value == 1].drop('MGMT_value',axis=1))\ndescribe_MGMT_0=show_test_normality(dataset_df[dataset_df.MGMT_value == 0].drop('MGMT_value',axis=1))\n\nT_describe = describe.drop(['Kurtosis'], axis=1).T\nT_MGMT_1 = describe_MGMT_1.drop(['Kurtosis', 'Maximum'], axis=1).T\nT_MGMT_0 = describe_MGMT_0.drop(['Kurtosis', 'Maximum'], axis=1).T\n\nplt.figure(figsize=(20, 5))\nT_describe[T_describe['kurtosis'] < 10]['kurtosis'].plot(label='Complet', color='blue')\nT_MGMT_1[T_MGMT_1['kurtosis'] < 10]['kurtosis'].plot(label='MGMT_1', color='red', linestyle='-.')\nT_MGMT_0[T_MGMT_0['kurtosis'] < 10]['kurtosis'].plot(label='MGMT_0', color='tab:orange', linestyle='--')\n\nplt.title('Comparison of Kurtosis Values for Different Datasets')\nplt.ylabel('Kurtosis')\nplt.xticks(range(len(T_describe.index)), T_describe.index, rotation=90)  # Afficher toutes les lignes sur l'axe X avec une rotation de 90 degrés\nplt.legend()\nplt.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:14:27.715378Z","iopub.execute_input":"2023-08-24T06:14:27.716420Z","iopub.status.idle":"2023-08-24T06:14:29.933298Z","shell.execute_reply.started":"2023-08-24T06:14:27.716368Z","shell.execute_reply":"2023-08-24T06:14:29.931929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>A first conclusion could be that a glioma with MGMT value = 0 would have variables containing more outliers.</li>\n            <li>There are exceptions with seemingly absurd variations, such as the excess of kurtosis for the MajorAxisLength variable which increases by almost 1232%.</li>\n            <li>These variations could indicate that certain variables have a greater impact compared to the target value. For example, kurtosis shows fewer outliers with MGMT value of 1, indicating better normality of Idn.</li>\n        </ul>\n    </div>\n</div>\n","metadata":{"_uuid":"96c47c41-d28e-44e0-8e42-f376b80a028b","_cell_guid":"a9e13f73-2cd5-47cc-b7f1-e1a7f20a7500","trusted":true}},{"cell_type":"markdown","source":"## <a id='7_3'>7.3. Outliers & Essential Variables</a> \n\nIn this section, we embark on a comprehensive exploration of outliers and the selection of essential variables within our dataset. By meticulously analyzing outliers and carefully selecting the most impactful variables, we aim to enhance the quality, reliability, and effectiveness of our statistical analyses and predictive models. These crucial steps enable us to uncover valuable insights, improve model performance, and make informed decisions based on robust data analysis.\n\n### <a id='7_3_1'>7.3.1. Outlier Detection</a>\n\nIn this section, we focus on detecting outliers within the dataset. Outliers are observations that deviate significantly from the majority of the data points and can have a substantial impact on statistical analyses and predictive models.\n\nTo identify outliers, we employ an approach based on the interquartile range (IQR). The IQR method allows us to measure the central variability of the data by calculating the quartiles. By defining a threshold, typically 1.5 times the IQR, we identify values that fall outside this range and classify them as outliers.\n\nAdditionally, we provide the option to filter the data based on the MGMT column. This allows for detecting outliers within specific subsets of the data, providing a more targeted analysis based on the MGMT values.\n\nBy conducting outlier detection, we can pinpoint potential data points that exhibit unusual behavior and may require further investigation. Detecting and handling outliers appropriately is crucial for maintaining data integrity and ensuring the validity of subsequent analyses and modeling efforts.","metadata":{"_uuid":"6af89121-b7cc-4add-be33-4d193ef03e55","_cell_guid":"d96a66c3-9af1-4598-9cd0-4f5fdf15953d","trusted":true}},{"cell_type":"code","source":"# Instantiate the OutlierDetector\ndetector = OutlierDetector(threshold=2)\n# Call the detect_outliers method with an additional filter\noutliers_mgmt_zeros_df, count = detector.detect_outliers(dataset_df, \"MGMT_value\", filter_column=\"MGMT_value\", filter_value=1)\n\n\nwith_outliers_desc = \"With outliers:\"\nwithout_outliers_desc = \"Without outliers:\"\n\n# Print the count of outliers present\ndisplay(HTML(\"<h3>Outliers with MGMT 1</h3>\"))\n# Print the count of outliers present with descriptions\ncount_str = count.to_string()\noutput_str = f\"{with_outliers_desc} {count[True]}\\n{without_outliers_desc} {count[False]}\"\nprint(output_str)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:14:29.935626Z","iopub.execute_input":"2023-08-24T06:14:29.936103Z","iopub.status.idle":"2023-08-24T06:14:30.022209Z","shell.execute_reply.started":"2023-08-24T06:14:29.936071Z","shell.execute_reply":"2023-08-24T06:14:30.020935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Call the detect_outliers method with an additional filter\noutliers_mgmt_one_df, count = detector.detect_outliers(dataset_df, \"MGMT_value\", filter_column=\"MGMT_value\", filter_value=0)\n\n\ndisplay(HTML(\"<h3>Outliers with MGMT 0</h3>\"))\n# Print the count of outliers present with descriptions\ncount_str = count.to_string()\noutput_str = f\"{with_outliers_desc} {count[True]}\\n{without_outliers_desc} {count[False]}\"\nprint(output_str)","metadata":{"_uuid":"8a44ad53-ee36-4edf-947f-05ce076d7a16","_cell_guid":"7a047191-7bf8-4cf2-893a-694f98bb57e7","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:14:30.023331Z","iopub.execute_input":"2023-08-24T06:14:30.023666Z","iopub.status.idle":"2023-08-24T06:14:30.101853Z","shell.execute_reply.started":"2023-08-24T06:14:30.023636Z","shell.execute_reply":"2023-08-24T06:14:30.100418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Call the detect_outliers method with an additional filter\noutliers_df, count = detector.detect_outliers(dataset_df, \"MGMT_value\")\n\n\ndisplay(HTML(\"<h3>Outliers with MGMT 0 and 1</h3>\"))\n# Print the count of outliers present with descriptions\ncount_str = count.to_string()\noutput_str = f\"{with_outliers_desc} {count[True]}\\n{without_outliers_desc} {count[False]}\"\nprint(output_str)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:14:30.103685Z","iopub.execute_input":"2023-08-24T06:14:30.104036Z","iopub.status.idle":"2023-08-24T06:14:30.179288Z","shell.execute_reply.started":"2023-08-24T06:14:30.104006Z","shell.execute_reply":"2023-08-24T06:14:30.178053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>More than half of the population contains outliers.</li>\n            <li>A possible conclusion is with MGMT value = 0 may have variables with more outliers.</li>\n    </ul>\n    </div>\n</div>\n","metadata":{"_uuid":"92fe41c8-b200-4c41-b994-efdde6b86265","_cell_guid":"73ba0475-b1a9-4b23-9489-605caae0aefe","trusted":true}},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Dataframe represents a collection of column names with a boolean for outliers detected in a dataset: <b>`outliers_df`</b></li>      \n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### <a id='7_3_2'>7.3.2. RFE Best features selection</a>    \n\nWe performed feature selection using the Recursive Feature Elimination (RFE) method to select the most important features in our medical dataset based on brain scans for predicting survival.\n\nWe explored several regression models for our analysis:\n\n- The **Random Forest Regressor** model was employed, known for its robustness and ability to handle complex data. The random forest is an ensemble model that combines multiple decision trees to make predictions. It is less sensitive to outliers and can effectively handle them. Additionally, the random forest's ensemble nature reduces its sensitivity to feature correlations. Therefore, we did not apply any data scaling or specific preprocessing techniques for handling feature correlations in this model.\n\n- The **Gradient Boosting Regressor** model was also used, leveraging the boosting algorithm to sequentially train trees and correct errors. The gradient boosting algorithm builds trees in a sequential manner, with each subsequent tree learning from the mistakes of the previous ones. While the Gradient Boosting Regressor can be influenced by feature correlations, it generally handles them reasonably well. However, to further improve the model's performance, we applied data scaling using either the StandardScaler or RobustScaler. This scaling helps accelerate the convergence of the boosting algorithm and addresses potential issues related to feature correlation.\n\n- The **AdaBoost Regressor** model is another boosting algorithm that focuses on previously mispredicted samples during the training process. Similar to the Gradient Boosting Regressor, the AdaBoost Regressor can be affected by feature correlations. To improve the model's convergence and performance, we applied data scaling using either the StandardScaler or RobustScaler.\n\n- Additionally, we utilized the **XGBoost Regressor** model, known for its efficiency and regularization capabilities. XGBoost is an optimized implementation of the gradient boosting algorithm. Like the Gradient Boosting Regressor, the XGBoost Regressor can be sensitive to feature correlations. To improve the model's convergence and performance, we applied data scaling using either the StandardScaler or RobustScaler, considering potential correlations among the features.\n\n- Lastly, we employed the **LightGBM Regressor** model, optimized for speed and efficiency. LightGBM is a gradient boosting framework that excels in large-scale and high-dimensional problems. Similar to the other tree-based models, the LightGBM Regressor can be influenced by feature correlations. To improve the model's convergence and performance, we applied data scaling using either the StandardScaler or RobustScaler.\n\nBy considering multiple tree-based regression models and applying data scaling when necessary, we can leverage the strengths of these algorithms while addressing potential challenges associated with feature correlations. This approach allows us to make accurate predictions of survival based on our brain scan dataset.","metadata":{}},{"cell_type":"code","source":"random_state=42\nfeature_importance_with_outliers = {}\n\noutlier_feature_names = outliers_df.loc[outliers_df[\"Outliers present\"] == True].index.values\nnon_outlier_feature_names = outliers_df.loc[outliers_df[\"Outliers present\"] == False].index.values\n\nfeatures_and_scaler_tuple_list = [\n    (outlier_feature_names, RobustScaler()),\n    (non_outlier_feature_names, StandardScaler()),\n]\n\n\nmodels = [\n    (\"Random Forest Regressor\", RandomForestRegressor(random_state=random_state), False),\n    (\"Gradient Boosting Regressor\", GradientBoostingRegressor(random_state=random_state), True),\n    (\"AdaBoost Regressor\", AdaBoostRegressor(random_state=random_state), True),\n    (\"XGBoost Regressor\", XGBRegressor(random_state=random_state), True),\n    (\"LightGBM Regressor\", LGBMRegressor(random_state=random_state), True),\n]\n\ncounter_df = pd.DataFrame(outliers_df.index.values, columns=[\"features\"])\n\ncounter_df[\"count\"] = 0\n\ndef check_feature_importance(feature):\n     return feature in selected_features\n\ncv = KFold(n_splits=5)    \n\nresults = select_best_features_RFE(\n        dataset_df, \n        \"MGMT_value\",\n        models, \n        features_and_scaler_tuple_list,\n        reduction_step=0.01,\n        cv=cv\n    )\n    \nfor name, score, selected_features in results:\n    model_name = f\"{name} - {score}\" \n    counter_df[model_name] = counter_df[\"features\"].apply(check_feature_importance)\n    counter_df[\"count\"] += counter_df[model_name].astype(int) \n    \nbest_selected_features_df = counter_df.sort_values(by=\"count\", ascending=False)\ndisplay(HTML(\"<h3>Best features:</h3>\"))\ndisplay(best_selected_features_df)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:14:30.181232Z","iopub.execute_input":"2023-08-24T06:14:30.181627Z","iopub.status.idle":"2023-08-24T06:45:55.103080Z","shell.execute_reply.started":"2023-08-24T06:14:30.181586Z","shell.execute_reply":"2023-08-24T06:45:55.101724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>We perform feature selection using RFECV (Recursive Feature Elimination with Cross-Validated Selection): RFECV is a method that combines recursive feature selection with cross-validation. It assesses the importance of features by gradually eliminating them and using cross-validation to estimate model performance. RFECV is a robust approach for feature selection, as it takes correlations into account and is able to handle outliers. It also helps to find the optimal number of features to select using cross-validation.</li>\n            <li>This analysis helps us identify the most relevant features to predict the target variable, taking into account the presence or absence of outliers.</li> \n            <li>By exploring different models, scales, and reduction steps, we can better understand the impact of outliers on feature importance and optimize feature selection for our predictive models.</li> \n            <li>We can also validate that we find certain features to study before.</li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Array of important Selected Features:<b>`best_selected_features_df`</b></li>           \n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## <a id='7_4'>7.4. Bivariate Analysis</a>\n\nIn this section, we delve into the fascinating world of bivariate analysis, where we explore the intricate relationships and interactions between two variables. By examining the interplay between these variables, we gain a deeper understanding of how they influence each other and uncover valuable insights into their joint behavior. Through various analytical techniques and visualizations, we aim to unravel the complex connections and uncover meaningful patterns that exist between these variables. This in-depth exploration allows us to uncover valuable insights and make informed decisions based on a comprehensive understanding of their interrelationships.\n\n### <a id='7_4_1'>7.4.1. Correlation Matrix</a>\n\nIn this section, we delve into the analysis of the correlation matrix, which allows us to examine the relationships between variables within our dataset. The correlation coefficient measures the strength and direction of the linear relationship between two variables.\n\nTo interpret the correlation coefficients, we can use the following rule of thumb:\n* 0.00-0.19: Very weak correlation.\n* 0.20-0.39: Weak correlation.\n* 0.40-0.59: Moderate correlation.\n* 0.60-0.79: Strong correlation.\n* 0.80-1.00: Very strong correlation.\n\nBy analyzing the correlation matrix, we gain insights into the strength and nature of the relationships between different variables. This knowledge helps us identify potential patterns, dependencies, and associations that can guide further analysis and decision-making processes.","metadata":{"_uuid":"e92549fe-6384-46de-bba5-0b3023871410","_cell_guid":"b1e57578-e79f-4df0-be9d-98971dbd65e1","trusted":true}},{"cell_type":"code","source":"show_triangle_correlation_matrix(dataset_df, figsize=(100, 50))","metadata":{"_uuid":"69352ff6-51b4-44a2-81ff-fe3cf02ff276","_cell_guid":"1204c4cd-dcc5-4c94-8574-a3818ac72c38","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:45:55.104548Z","iopub.execute_input":"2023-08-24T06:45:55.104889Z","iopub.status.idle":"2023-08-24T06:46:14.760520Z","shell.execute_reply.started":"2023-08-24T06:45:55.104861Z","shell.execute_reply":"2023-08-24T06:46:14.757594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The target variable exhibits no significant correlation with the explanatory variables, indicating the absence of an evident linear relationship between them. This suggests that the traits or features being analyzed may not directly influence or predict the target variable in a linear manner.","metadata":{"execution":{"iopub.status.busy":"2023-06-12T13:22:37.668834Z","iopub.execute_input":"2023-06-12T13:22:37.669264Z","iopub.status.idle":"2023-06-12T13:22:37.705268Z","shell.execute_reply.started":"2023-06-12T13:22:37.669229Z","shell.execute_reply":"2023-06-12T13:22:37.702949Z"}}},{"cell_type":"markdown","source":"### Correlation group with threshold strength\n\nFor correlation threshold equal to **0.7, strong.**","metadata":{"_uuid":"60639114-c5b0-49dd-8607-00c011e62c62","_cell_guid":"d1d98454-739c-411c-958f-168af7652473","execution":{"iopub.status.busy":"2023-05-29T09:22:54.2922Z","iopub.execute_input":"2023-05-29T09:22:54.292577Z","iopub.status.idle":"2023-05-29T09:22:54.299079Z","shell.execute_reply.started":"2023-05-29T09:22:54.29255Z","shell.execute_reply":"2023-05-29T09:22:54.29756Z"},"trusted":true}},{"cell_type":"code","source":"# Filtrer by correlation threshold\nshow_triangle_correlation_matrix(dataset_df,filter_threshold = 0.7, figsize=(100, 50))","metadata":{"_uuid":"b22be0c4-edcf-46d5-abd8-503c6fa95531","_cell_guid":"63d56950-1566-4c45-99cb-0fdaae2b4982","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:14.762702Z","iopub.execute_input":"2023-08-24T06:46:14.763318Z","iopub.status.idle":"2023-08-24T06:46:22.412161Z","shell.execute_reply.started":"2023-08-24T06:46:14.763261Z","shell.execute_reply":"2023-08-24T06:46:22.410825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlation group with threshold strength\n\nFor correlation threshold equal to **0.9, very strong.**","metadata":{"_uuid":"d73597f6-88e2-4e57-addc-7db9766c65e1","_cell_guid":"bd92a22b-976c-4a58-9bf0-0363f64a655b","trusted":true}},{"cell_type":"code","source":"show_triangle_correlation_matrix(dataset_df,filter_threshold = 0.9, figsize=(100, 50))","metadata":{"_uuid":"4cccb58e-3d39-4662-8ca4-a444931be6f3","_cell_guid":"d55e4a9d-6e91-444c-be13-80d879c324b3","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:22.413533Z","iopub.execute_input":"2023-08-24T06:46:22.413897Z","iopub.status.idle":"2023-08-24T06:46:28.711592Z","shell.execute_reply.started":"2023-08-24T06:46:22.413867Z","shell.execute_reply":"2023-08-24T06:46:28.710455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the provided analysis with 3D, we can valided the conclusions:","metadata":{}},{"cell_type":"code","source":"hight_correlated_features_tupples = [\n    ('JointAverage', 'SumAverage', 'MGMT_value'),\n    ('TotalEnergy', 'Energy', 'MGMT_value'),\n    ('MeshVolume', 'VoxelVolume', 'MGMT_value'),\n    ('MajorAxisLength', 'MinorAxisLenth', 'MGMT_value')\n]\n\n\nshow_3D_scatter_plots(\n    hight_correlated_features_tupples, \n    dataset_df, \n    title=\"Samples Correlation Thresholds >0.99 without angles\",\n    figsize=(10, 10), \n    elev_angle=90,\n    azimuth_angle=90, \n    show_legend=True,\n    alpha=0.2\n)\n\nshow_3D_scatter_plots(\n    hight_correlated_features_tupples, \n    dataset_df, \n    figsize=(10, 10), \n    elev_angle=45, \n    azimuth_angle=90, \n    title=\"Samples Correlation Thresholds >0.99\", \n    show_legend=True,\n    legend_loc=\"upper right\",\n    alpha=0.2\n)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:28.713108Z","iopub.execute_input":"2023-08-24T06:46:28.713452Z","iopub.status.idle":"2023-08-24T06:46:31.399281Z","shell.execute_reply.started":"2023-08-24T06:46:28.713423Z","shell.execute_reply":"2023-08-24T06:46:31.398152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>        \n        <p style=\"color: white;\">Based on the provided analysis, we can draw the following correlation conclusions:</p>  \n        <h2 class=\"card-title\" style=\"color: white;\"><strong>Note START 👇🏻</strong></h2> \n    </div>\n</div>\n","metadata":{"_uuid":"2341fefa-7b12-4c1e-af09-eca26c60f0a1","_cell_guid":"7e32f71b-cd6b-42c3-aaf9-812ce203e819","trusted":true}},{"cell_type":"code","source":"correlation_thresholds = [0.5, 0.6, 0.7, 0.8, 0.9, 0.99, 1]\ncorrelated_features_table = generate_correlated_features_table(dataset_df, correlation_thresholds)\n\ncorrelated_features_table","metadata":{"_uuid":"5d3fa7f9-e02e-468e-ad16-17359cc55d3e","_cell_guid":"5e774ae9-f7ce-4790-891f-f87907f2552d","_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:31.400951Z","iopub.execute_input":"2023-08-24T06:46:31.401658Z","iopub.status.idle":"2023-08-24T06:46:33.444290Z","shell.execute_reply.started":"2023-08-24T06:46:31.401614Z","shell.execute_reply":"2023-08-24T06:46:33.443214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Table listing the features correlated by thresholds: <b>`correlated_features_table`</b></li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### <a id='7_4_2'>7.4.2. Feature Selection and Redundancy Prevention for Avoiding Overfitting</a>\n\nDuring our analysis, we observed that not all features in our dataset are equally relevant. To prevent overfitting and improve the performance of our model, we aim to select the most significant variables while avoiding redundancy.\n\nIn this section, we employ the Recursive Feature Elimination (RFE) method to identify the best features. We consider a variable significant if it is deemed relevant by four or more algorithms. By leveraging the results obtained through RFE, we can pinpoint the most influential variables for our analysis.\n\nFurthermore, we address the issue of redundant features. Some variables may exhibit high correlation with other variables in the dataset. To avoid redundancy, we group correlated variables and retain only one variable from each group if their correlation score exceeds 0.90. This approach ensures that we capture the essential information while eliminating redundant features.\n\nFinally, we prioritize selecting features that are not among the first-order features, particularly if there are other types of variables within the same correlation group. This strategy allows us to include a diverse set of features that collectively contribute to a comprehensive and robust analysis.\n\nBy implementing these steps of feature selection and redundancy prevention, we enhance the quality and generalization capability of our model, reducing the risk of overfitting and enabling more accurate predictions.","metadata":{}},{"cell_type":"code","source":"feats_list = best_selected_features_df.loc[best_selected_features_df[\"count\"] >= 4][\"features\"].tolist()\nfeats_list = list(dict.fromkeys(feats_list))\n\ndef create_significant_vars_list(df_correlated, significant_values, threshold):\n    df_first_orders_features_columns=['ID','10Percentile', '90Percentile', 'Energy', 'Entropy', 'InterquartileRange', 'Kurtosis',\n                                      'Maximum', 'MeanAbsoluteDeviation', 'Mean', 'Median', 'Minimum', 'Range', 'RobustMeanAbsoluteDeviation',\n                                      'RootMeanSquared', 'Skewness', 'TotalEnergy', 'Uniformity', 'Variance']\n    significant_vars_to_plot = []\n    vars_to_exclude = []\n    groups_correlated_threshold = df_correlated[threshold].tolist()\n    groups_correlated_threshold = filter(lambda x: x!= \"\", groups_correlated_threshold)\n    groups_correlated_threshold = list(groups_correlated_threshold)\n    groups_correlated_threshold = list(map((lambda x : x.split(', ')), groups_correlated_threshold))\n            \n    for var in significant_values :\n        if var not in vars_to_exclude :\n            \n            if var in df_first_orders_features_columns :\n                non_fst_order_found = False\n                \n                for corr_vars in groups_correlated_threshold :\n                    if var in corr_vars :\n                        for i in corr_vars :\n                            if i not in df_first_orders_features_columns and not non_fst_order_found :\n                                significant_vars_to_plot.append(i)\n                                non_fst_order_found = True\n                            else :\n                                if i != var :\n                                    vars_to_exclude.append(i)\n                                    \n                        if non_fst_order_found :\n                            vars_to_exclude.append(var)\n                        else :\n                            significant_vars_to_plot.append(var)\n                \n            else :\n                significant_vars_to_plot.append(var)\n                \n                for corr_vars in groups_correlated_threshold :\n                    if var in corr_vars :\n                        for i in corr_vars :\n                            if i in significant_values :\n                                vars_to_exclude.append(i)\n    \n    return significant_vars_to_plot\n\nsignificant_values = significant_values_df[significant_values_df['Significant'] == \"⬆️\"].index.tolist()\nsignificant_values = list(set(significant_values) & set(feats_list))\n\nsignificant_features_avoid_overfitting_df = create_significant_vars_list(correlated_features_table, significant_values, 0.90)\n\nsignificant_features = pd.DataFrame({\"feature\":significant_features_avoid_overfitting_df})\nsignificant_features","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:33.446245Z","iopub.execute_input":"2023-08-24T06:46:33.446680Z","iopub.status.idle":"2023-08-24T06:46:33.473500Z","shell.execute_reply.started":"2023-08-24T06:46:33.446642Z","shell.execute_reply":"2023-08-24T06:46:33.472380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Amongst the significant variables that we determined, we chose to display the univariate analysis of the ones that seem the most relevant. In order to compare the patients with a different MGMT value, regarding those features, we chose to a normalized histogram with a density function, and a boxplot per variable.","metadata":{}},{"cell_type":"code","source":"significant_vars_to_plot_copy = (var for var in significant_features_avoid_overfitting_df if var not in ['ngtdm_Coarseness','glrlm_GrayLevelNonUniformity','Maximum2DDiameterRow','Maximum2DDiameterColumn','gzlm_LargeAreaLowGrayLevelEmphasis','ngtdm_Strength','glrlm_GrayLevelNonUniformity','gzlm_GrayLevelNonUniformity','glrlm_GrayLevelNonUniformityNormalized', 'glrlm_LongRunLowGrayLevelEmphasis'])\nsignificant_vars_to_plot_copy = list(significant_vars_to_plot_copy)\n\ndisplay(HTML(\"<h3>Exploring of Feature Selection and Redundancy Prevention for Avoiding Overfitting:</h3>\"))\n\nfor var in significant_vars_to_plot_copy :\n    fonc_test_normality(dataset_df, var)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:33.474932Z","iopub.execute_input":"2023-08-24T06:46:33.475334Z","iopub.status.idle":"2023-08-24T06:46:37.696959Z","shell.execute_reply.started":"2023-08-24T06:46:33.475303Z","shell.execute_reply":"2023-08-24T06:46:37.695785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>Our univariate analyses reveal no significant differences between our two groups of patients.</li>\n            <li>We observe intriguing patterns. For example, in the density functions of <b>Maximum3DDiameter and MajorAxisLength</b>, there appear to be local maxima alongside the global maximum for the group of patients with a positive MGMT value.</li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Table listing the most significant features : <b>`significant_features_avoid_overfitting_df`</b></li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### <a id='7_4_3'>7.4.3. Density Comparison</a>\n\nIn this section, we delve into the bivariate analysis by utilizing KDE (Kernel Density Estimation) plots to compare the density of variables.\n\nThe primary objective of this analysis is to identify any potential relationships or patterns between the variables for our two groups of patients. By visualizing the density distributions using KDE plots, we can gain insights into the overlap, differences, or similarities in the distributions of the variables across the patient groups.\n\nThis approach allows us to assess the density of variables and explore any potential correlations, clusters, or divergences that may exist between the groups. By examining the density comparisons, we can uncover valuable information about the underlying distributions and potentially identify significant variables that differentiate the patient groups.","metadata":{}},{"cell_type":"code","source":"display(HTML(\"<h3>Exploring of Density Comparison:</h3>\"))\n\nsignificant_values_combinations = []\n\nfor combination in combinations(significant_vars_to_plot_copy, 2):\n    significant_values_combinations.append(combination)\n    \n# Select the first 5 combinations\nselected_combinations = significant_values_combinations[:5]\n\nfor combination in selected_combinations :\n    create_kde_mgmt_pos_neg(dataset_df, combination[0], combination[1])","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:37.698687Z","iopub.execute_input":"2023-08-24T06:46:37.699137Z","iopub.status.idle":"2023-08-24T06:46:46.005692Z","shell.execute_reply.started":"2023-08-24T06:46:37.699095Z","shell.execute_reply":"2023-08-24T06:46:46.004490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>📝 Note:</strong></h2>\n        <ul>\n            <li>Our density analyses do not reveal significant differences between our two groups of patients.</li>          \n            <li>In density analyses, we often observe a common main zone for both groups of patients, along with at least one additional zone that differs between them. An example of this is seen in the bivariate analysis of <b>gzlm_GrayLevelVariance and Maximum3DDiameter.</b></li>\n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----\n----","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nfrom sklearn.feature_selection import SelectKBest, f_regression\nfrom sklearn.preprocessing import StandardScaler\nimport numpy as np\n\ntarget = dataset_df[\"MGMT_value\"]\nfeatures = dataset_df.drop(\"MGMT_value\", axis=1)\n\n# Supposons que vous ayez votre jeu de données dans une variable X et les étiquettes correspondantes dans une variable y\nX, X_test, y, y_test = train_test_split(features, target, test_size=0.2, random_state=123)\n# Réduire la dimension en utilisant PCA\nscaler = StandardScaler()\nX_scaled = scaler.fit_transform(X)  # Mise à l'échelle des caractéristiques\n\npca = PCA(n_components=0.9)  # Sélectionner le nombre de composantes qui capturent 95% de la variance\nX_reduced = pca.fit_transform(X_scaled)\n\n# Sélectionner les caractéristiques les plus importantes en utilisant la régression F\nselector = SelectKBest(score_func=f_regression, k=2)  # Sélectionner les 10 meilleures caractéristiques\nX_selected = selector.fit_transform(X_reduced, y)\n\n# Obtenir les indices des caractéristiques sélectionnées\nselected_feature_indices = selector.get_support(indices=True)\n\n# Obtenir les noms des caractéristiques sélectionnées\nselected_feature_names = np.array(list(X.columns))[selected_feature_indices]\n\n# Supprimer les caractéristiques hautement corrélées\ncorrelation_matrix = np.corrcoef(X_selected, rowvar=False)\ncorrelation_mask = np.abs(correlation_matrix) < 0.9  # Définir le seuil de corrélation ici (0.9 dans cet exemple)\ncorrelation_mask = np.all(correlation_mask, axis=0)  # Sélectionner les caractéristiques qui satisfont le seuil de corrélation\n\nX_final = X_selected[:, correlation_mask]\n\n# Utiliser X_final pour entraîner votre modèle\n\nselected_feature_names","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.007229Z","iopub.execute_input":"2023-08-24T06:46:46.007693Z","iopub.status.idle":"2023-08-24T06:46:46.111862Z","shell.execute_reply.started":"2023-08-24T06:46:46.007651Z","shell.execute_reply":"2023-08-24T06:46:46.109641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <a id='analysis_7_4'>7.4.2 Univariate Analysis: Top 50 Best Normal Distribution features selection</a>\n","metadata":{}},{"cell_type":"code","source":"# Sélectionner les variables avec les valeurs les plus proches de 1 pour 'skewness' et les plus faibles 'absolute_value'\nunivariate_analysis_normality_skewness_kurtosis_tests_top_50_df = describe_transposed.nsmallest(50, ['absolute_value', 'skewness_diff'])\nunivariate_analysis_normality_skewness_kurtosis_tests_top_50_df = univariate_analysis_normality_skewness_kurtosis_tests_top_50_df.drop(['skewness_diff', 'absolute_value'], axis=1)\n\ndisplay(HTML(\"<h3>Top 50 normal distribution:</h3>\"))\ndisplay(univariate_analysis_normality_skewness_kurtosis_tests_top_50_df)","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.114146Z","iopub.execute_input":"2023-08-24T06:46:46.119120Z","iopub.status.idle":"2023-08-24T06:46:46.208192Z","shell.execute_reply.started":"2023-08-24T06:46:46.119062Z","shell.execute_reply":"2023-08-24T06:46:46.206964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"card\" style=\"background-color: #007bff; border-radius: 8px; padding: 16px; color: white;\">\n    <div class=\"card-body\">\n        <h2 class=\"card-title\" style=\"color: white;\"><strong>🛠 Global:</strong></h2>\n        <ul>\n            <li>Top 50 of Univariate Analysis: Understanding Normality, Skewness, and Kurtosis Tests <b>`univariate_analysis_normality_skewness_kurtosis_tests_top_50_df`</b></li>            \n        </ul>\n    </div>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"markdown","source":"### Copy","metadata":{}},{"cell_type":"code","source":"dataset_copy = dataset_df.copy() ","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.209509Z","iopub.execute_input":"2023-08-24T06:46:46.209879Z","iopub.status.idle":"2023-08-24T06:46:46.216471Z","shell.execute_reply.started":"2023-08-24T06:46:46.209846Z","shell.execute_reply":"2023-08-24T06:46:46.215372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_copy.head()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.217858Z","iopub.execute_input":"2023-08-24T06:46:46.218256Z","iopub.status.idle":"2023-08-24T06:46:46.332593Z","shell.execute_reply.started":"2023-08-24T06:46:46.218222Z","shell.execute_reply":"2023-08-24T06:46:46.331284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import necessary libraries\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, cross_val_score, GridSearchCV\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.svm import OneClassSVM\nfrom sklearn.metrics import classification_report\nimport pandas as pd\nfrom sklearn.preprocessing import FunctionTransformer\nfrom sklearn_pandas import gen_features, DataFrameMapper\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.metrics import accuracy_score\n\nfrom sklearn.preprocessing import RobustScaler\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.336489Z","iopub.execute_input":"2023-08-24T06:46:46.336936Z","iopub.status.idle":"2023-08-24T06:46:46.375768Z","shell.execute_reply.started":"2023-08-24T06:46:46.336903Z","shell.execute_reply":"2023-08-24T06:46:46.374728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cleaning","metadata":{}},{"cell_type":"code","source":"#Contrast\n#DifferenceVariance \n#DifferenceEntropy\n#DifferenceVariance \n#DifferenceEntropy\n#DifferenceAverage\n\n#LargeAreaHighGrayLevelEmphasis\n#LargeDependenceHighGrayLevelEmphasis\n#LargeAreaHighGrayLevelEmphasis\n#Kurtosis \n#Idmn\n\n#Correlation\n#LargeDependenceHighGrayLevelEmphasis\n#SumEntropy\n#HighGrayLevelEmphasis\n#LargeAreaHighGrayLevelEmphasis\n#JointAverage\n\nfeature_outliers = [\n    'Kurtosis',\n    'glrlm_LongRunHighGrayLevelEmphasis',\n    'gzlm_SizeZoneNonUniformityNormalized',\n    'MajorAxisLength',\n    'Imc2',\n    'TotalEnergy',\n    'gzlm_LargeAreaLowGrayLevelEmphasis',\n    'Idn',\n    'gzlm_SmallAreaLowGrayLevelEmphasis',\n    'gzlm_GrayLevelVariance'\n]\n\nfeature_non_outliers = ['InverseVariance',\n    'Range',\n    'MaximumProbability',\n    'Maximum3DDiameter',\n    'Maximum2DDiameterColumn',\n    'Entropy',\n    'Maximum2DDiameterRow']\n\nfeature_to_keep = [\"MGMT_value\"]\n\ndataset_copy = dataset_copy[feature_outliers + feature_non_outliers + feature_to_keep]","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.377375Z","iopub.execute_input":"2023-08-24T06:46:46.377756Z","iopub.status.idle":"2023-08-24T06:46:46.386355Z","shell.execute_reply.started":"2023-08-24T06:46:46.377724Z","shell.execute_reply":"2023-08-24T06:46:46.385338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Drop features","metadata":{}},{"cell_type":"code","source":"dataset_copy.head()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.388038Z","iopub.execute_input":"2023-08-24T06:46:46.388416Z","iopub.status.idle":"2023-08-24T06:46:46.427698Z","shell.execute_reply.started":"2023-08-24T06:46:46.388365Z","shell.execute_reply":"2023-08-24T06:46:46.425913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Slipt train test (Features & Targets)","metadata":{}},{"cell_type":"code","source":"target = dataset_copy[\"MGMT_value\"]\nfeatures = dataset_copy.drop(\"MGMT_value\", axis=1)\n\nX_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.2, random_state=123)","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.429421Z","iopub.execute_input":"2023-08-24T06:46:46.429788Z","iopub.status.idle":"2023-08-24T06:46:46.441840Z","shell.execute_reply.started":"2023-08-24T06:46:46.429757Z","shell.execute_reply":"2023-08-24T06:46:46.440488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Pipeline","metadata":{}},{"cell_type":"code","source":"pipeline_steps = []","metadata":{"execution":{"iopub.status.busy":"2023-08-24T06:46:46.446466Z","iopub.execute_input":"2023-08-24T06:46:46.446875Z","iopub.status.idle":"2023-08-24T06:46:46.455053Z","shell.execute_reply.started":"2023-08-24T06:46:46.446828Z","shell.execute_reply":"2023-08-24T06:46:46.454193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocessing","metadata":{}},{"cell_type":"code","source":"#-------------------------------#\n# Name\n#-------------------------------#\npreprocessing_name = \"preprocessing\"\n\n#-------------------------------#\n# Map\n#-------------------------------#\n\ntransformers = []\n#-------------------------------#\n# Maps all features into transformers based on their categories\n#-------------------------------#\n#transformers = map_all_features_into_transformers(all_features)\ntransformers.append((f\"ze_a_transformer\", RobustScaler(), feature_outliers))\ntransformers.append((f\"ze_b_transformer\", StandardScaler(), feature_non_outliers))\n\n#-------------------------------#\n# Create ColumnTransformer\n#-------------------------------#\npreprocessor = ColumnTransformer(transformers)\n\n#-------------------------------#\n# Pipeline first step\n#-------------------------------#\npipeline_steps.append((preprocessing_name, preprocessor))\n","metadata":{"execution":{"iopub.status.busy":"2023-08-24T06:46:46.456467Z","iopub.execute_input":"2023-08-24T06:46:46.456787Z","iopub.status.idle":"2023-08-24T06:46:46.471519Z","shell.execute_reply.started":"2023-08-24T06:46:46.456751Z","shell.execute_reply":"2023-08-24T06:46:46.470210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model","metadata":{}},{"cell_type":"code","source":"#-------------------------------#\n# Name\n#-------------------------------#\nmodel_name = \"model\"\n\n#-------------------------------#\n# Model\n#-------------------------------#\n# KNeighborsClassifier\n#from sklearn.neighbors import KNeighborsClassifier\n#model = KNeighborsClassifier()\n\n# DecisionTreeClassifier\n#from sklearn.tree import DecisionTreeClassifier\n#model = DecisionTreeClassifier()\n\n# SVC\n#from sklearn.svm import SVC\n#model = SVC()\n\n# DBSCAN\n#from sklearn.cluster import DBSCAN\n#model = DBSCAN()\n\n# RandomForestClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n#model = RandomForestClassifier()\n\n# XGBoost\n#from xgboost import XGBClassifier\n#model = XGBClassifier()\n\n# LabelSpreading\n#from sklearn.semi_supervised import LabelSpreading\n#model = LabelSpreading()\n\n# LabelSpreading\n#from sklearn.cluster import KMeans\n#model = KMeans()\n\n# AgglomerativeClustering\n#from sklearn.cluster import AgglomerativeClustering\n#model = AgglomerativeClustering()\n\n# BaggingClassifier\nfrom sklearn.ensemble import BaggingClassifier\nmodel = BaggingClassifier(base_estimator=RandomForestClassifier())\n\n#-------------------------------#\n# Pipeline second step\n#-------------------------------#\npipeline_steps.append((model_name, model))","metadata":{"execution":{"iopub.status.busy":"2023-08-24T06:46:46.476038Z","iopub.execute_input":"2023-08-24T06:46:46.476497Z","iopub.status.idle":"2023-08-24T06:46:46.491426Z","shell.execute_reply.started":"2023-08-24T06:46:46.476462Z","shell.execute_reply":"2023-08-24T06:46:46.490144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pipeline","metadata":{}},{"cell_type":"code","source":"pipeline = Pipeline(steps=pipeline_steps)\npipeline","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.492953Z","iopub.execute_input":"2023-08-24T06:46:46.493320Z","iopub.status.idle":"2023-08-24T06:46:46.549006Z","shell.execute_reply.started":"2023-08-24T06:46:46.493287Z","shell.execute_reply":"2023-08-24T06:46:46.547821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Hyperparameter tuning, cross-validation, optimization","metadata":{}},{"cell_type":"code","source":"# Define the parameter grid for GridSearchCV\n#param_grid = {\n#    'model__nu': [0.1, 0.2, 0.3],  # Values to try for the 'nu' parameter\n#    'model__gamma': [0.1, 0.2, 0.3]  # Values to try for the 'gamma' parameter, if applicable\n#}\n\n# KNeighborsClassifier\n#param_grid = {\n#    f'{model_name}__n_neighbors': list(range(1, 200)),  # Valeurs possibles pour le nombre de voisins\n#    f'{model_name}__weights': ['uniform', 'distance']  # Valeurs possibles pour les poids\n#}\n\n# DecisionTreeClassifier\n#param_grid = {\n#    f'{model_name}__criterion': ['gini', 'entropy'],  # Critère pour mesurer la qualité des splits\n#    f'{model_name}__max_depth': [None, 5, 10],  # Profondeur maximale de l'arbre\n#    f'{model_name}__min_samples_split': [2, 5, 10],  # Nombre minimal d'échantillons requis pour effectuer un split\n#    f'{model_name}__min_samples_leaf': [1, 2, 3],  # Nombre minimal d'échantillons requis dans une feuille\n#    f'{model_name}__max_features': ['auto', 'sqrt', 'log2'],  # Nombre maximal de caractéristiques à considérer pour chaque split\n#    f'{model_name}__min_impurity_decrease': [0.0, 0.1, 0.2],  # Seuil minimal pour effectuer un split basé sur l'impureté\n#}\n\n# SVC\n#param_grid = {\n#    f'{model_name}__C': [0.1, 1.0, 10.0],  # Paramètre de régularisation C\n#    f'{model_name}__kernel': ['linear', 'rbf'],  # Noyau du modèle SVC\n#    f'{model_name}__gamma': ['scale', 'auto']  # Paramètre gamma pour les noyaux rbf et poly\n#}\n\n\n#RandomForestClassifier\n#param_grid = {\n#    f\"{model_name}__criterion\": [\"gini\", \"entropy\"],\n#    f\"{model_name}__n_estimators\":  [5, 100, 200, 300],  # Nombre d'estimateurs dans le RandomForestClassifier\n#    f\"{model_name}__max_depth\":  [1, 5, 15, 50],  # Profondeur maximale des arbres dans le RandomForestClassifier\n#}\n\n# XGBClassifier\n#param_grid = {\n#    f\"{model_name}__n_estimators\": [100, 200, 300, 400, 500, 600, 700],  # Nombre d'estimateurs\n#    f\"{model_name}__max_depth\": [3, 4, 5, 6],  # Profondeur maximale des arbres\n#    f\"{model_name}__learning_rate\": [0.3, 0.2, 0.1, 0.01, 0.001]  # Taux d'apprentissage\n#}\n#param_grid = {\n    # Nombre d'estimateurs (nombre d'arbres dans le modèle)\n   # f\"{model_name}__n_estimators\": [50, 100, 200, 300, 350],\n\n    # Profondeur maximale des arbres\n   # f\"{model_name}__max_depth\": [3, 4, 5, 6],\n\n    # Taux d'apprentissage (contrôle la contribution de chaque arbre dans le modèle)\n    #f\"{model_name}__learning_rate\": [0.3, 0.2, 0.1, 0.01, 0.001],\n\n    # Sous-échantillonnage des exemples (contrôle la proportion d'échantillons utilisée pour la construction de chaque arbre)\n   # f\"{model_name}__subsample\": [0.6, 0.8, 1.0],\n\n    # Sous-échantillonnage des colonnes (features) (contrôle la proportion de features utilisée pour la construction de chaque arbre)\n    #f\"{model_name}__colsample_bytree\": [0.6, 0.8, 1.0],\n\n    # Valeur de pénalité pour les coupes dans l'arbre (contrôle la complexité de l'arbre)\n    #f\"{model_name}__gamma\": [0, 0.1, 0.2],\n\n    # Valeur de pénalité L1 pour les poids de l'arbre (contrôle la régularisation L1)\n   # f\"{model_name}__reg_alpha\": [0, 0.1, 0.5],\n\n    # Valeur de pénalité L2 pour les poids de l'arbre (contrôle la régularisation L2)\n    #f\"{model_name}__reg_lambda\": [0, 0.1, 0.5]\n#}\n\n\n# LabelSpreading\n#param_grid = {\n#    f\"{model_name}__kernel\": ['knn', 'rbf'],\n#    f\"{model_name}__gamma\": [0.1, 1.0, 2, 3, 4, 5, 10, 15, 20 , 25, 30],\n#    f\"{model_name}__n_neighbors\": [3, 5, 7, 5, 10, 15, 30, 50, 60, 100],\n#    f\"{model_name}__max_iter\": [40, 50, 100, 200, 300],\n#}\n\n\n# kmeans\n#param_grid = {\n#    f\"{model_name}__n_clusters\": [2],  # Nombre de clusters à essayer\n#    f\"{model_name}__init\": ['k-means++', 'random'],  # Méthode d'initialisation des centroides\n#    f\"{model_name}__max_iter\": [100, 200, 300, 400, 500 , 600]  # Nombre maximum d'itérations\n#}\n\n# agglomerative\n#param_grid = {\n#    f\"{model_name}__n_clusters\": [2],  # Nombre de clusters à essayer\n   # f\"{model_name}__affinity\": ['euclidean', 'l1', 'l2', 'manhattan', 'cosine'],  # Métrique de distance utilisée\n   # f\"{model_name}__linkage\": ['ward', 'complete', 'average', 'single'],  # Méthode de liaison pour former les clusters\n   # 'agglomerative__connectivity': [None, nearest_neighbors_graph],  # Matrice de connectivité pour contraindre les regroupements\n   # 'agglomerative__distance_threshold': [None, 0.5, 1.0, 1.5]  # Seuil de distance pour former les clusters\n#}\n\n\n# f\"{model_name}__n_estimators\": [50, 100, 200],\n# Nombre d'estimateurs dans le BaggingClassifier.\n# Plus le nombre d'estimateurs est élevé, plus le modèle est robuste, mais cela augmente le temps d'entraînement.\n\n# f\"{model_name}__max_samples\": [0.5, 0.8, 1.0],\n# La proportion des échantillons d'entraînement à utiliser pour chaque estimateur.\n# Une valeur inférieure à 1.0 crée des échantillons d'entraînement bootstrap.\n\n# f\"{model_name}__max_features\": [0.5, 0.8, 1.0],\n# La proportion des caractéristiques à utiliser pour chaque estimateur.\n# Une valeur inférieure à 1.0 effectue une sélection aléatoire des caractéristiques.\n\n# f\"{model_name}__base_estimator__max_depth\": [None, 5, 10],\n# La profondeur maximale de chaque arbre de décision de RandomForestClassifier.\n# Une valeur None signifie que les arbres sont développés jusqu'à ce que toutes les feuilles soient pures ou jusqu'à ce que toutes les feuilles contiennent un nombre minimum d'échantillons.\n\n# f\"{model_name}__base_estimator__min_samples_split\": [2, 5, 10],\n# Le nombre minimum d'échantillons requis pour scinder un nœud interne dans chaque arbre de décision.\n# Une valeur plus élevée peut empêcher le surapprentissage en obligeant les arbres à avoir un nombre minimum d'échantillons dans chaque nœud.\n\n# f\"{model_name}__base_estimator__min_samples_leaf\": [1, 2, 4]\n# Le nombre minimum d'échantillons requis dans une feuille d'arbre.\n# Une valeur plus élevée peut également aider à prévenir le surapprentissage en limitant la croissance de l'arbre.\n\n# Define the parameter grid for GridSearchCV\nparam_grid = {\n    f\"{model_name}__n_estimators\": [50, 100, 200],\n    f\"{model_name}__max_samples\": [0.5, 0.8, 1.0],\n    f\"{model_name}__max_features\": [0.5, 0.8, 1.0],\n    f\"{model_name}__base_estimator__max_depth\": [None, 5, 10],\n    f\"{model_name}__base_estimator__min_samples_split\": [2, 5, 10],\n    f\"{model_name}__base_estimator__min_samples_leaf\": [1, 2, 4]\n}","metadata":{"execution":{"iopub.status.busy":"2023-08-24T06:46:46.550896Z","iopub.execute_input":"2023-08-24T06:46:46.551478Z","iopub.status.idle":"2023-08-24T06:46:46.569032Z","shell.execute_reply.started":"2023-08-24T06:46:46.551443Z","shell.execute_reply":"2023-08-24T06:46:46.567534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Search Hyperparame","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import KFold\n\ncross_validator = KFold(n_splits=10)\n\ngrid_search = GridSearchCV(pipeline, param_grid, cv=cross_validator, error_score='raise')","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.570959Z","iopub.execute_input":"2023-08-24T06:46:46.571685Z","iopub.status.idle":"2023-08-24T06:46:46.579340Z","shell.execute_reply.started":"2023-08-24T06:46:46.571649Z","shell.execute_reply":"2023-08-24T06:46:46.578163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.fit(X_train , y_train)","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-08-24T06:46:46.580954Z","iopub.execute_input":"2023-08-24T06:46:46.581329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Score","metadata":{}},{"cell_type":"code","source":"# Get the best parameters and best score from GridSearchCV\nbest_params = grid_search.best_params_\nbest_score = grid_search.best_score_\nprint(best_params, \"\\n\")\nprint(best_score, \"\\n\")\n \n# Obtenez les prédictions du meilleur modèle sur X_train\ny_train_pred = grid_search.best_estimator_.predict(X_train)\n# Générez le rapport de classification pour les prédictions sur X_train\nclassification_rep = classification_report(y_train, y_train_pred)\nprint(classification_rep)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Predict\n","metadata":{}},{"cell_type":"code","source":"# Evaluate the predictions\naccuracy = grid_search.score(X_test, y_test)\n# Affichez le score d'exactitude\nprint(\"Accuracy Score on Test Set:\", accuracy, \"\\n\")\n\n# Use the best model for predictions\nbest_model = grid_search.best_estimator_\npredictions = best_model.predict(X_test)\n\n# Evaluate the predictions\nprint(classification_report(y_test, predictions))\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"----","metadata":{}},{"cell_type":"code","source":"import itertools\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Dropout\nfrom tensorflow.keras.optimizers import Adam\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Dropout, BatchNormalization\nfrom tensorflow.keras.optimizers import Adamax\nfrom sklearn.metrics import accuracy_score, confusion_matrix\nfrom sklearn.preprocessing import OneHotEncoder\nfrom imblearn.over_sampling import RandomOverSampler\nfrom sklearn.model_selection import KFold\nfrom keras.utils import to_categorical\n\n'''\ndef build_model(layer_size, activation, dropout):\n    model = Sequential()\n    model.add(Dense(layer_size, activation=activation, input_shape=(input_dim,)))\n    model.add(Dropout(dropout))\n    model.add(Dense(num_classes, activation='softmax'))\n    return model\n'''\n\ndef build_model(layer_size, activation, dropout):\n    model = Sequential()\n    model.add(Dense(layer_size, activation=activation, input_shape=(input_dim,)))\n    model.add(Dropout(dropout))\n    model.add(BatchNormalization())\n    model.add(Dense(layer_size, activation=activation))\n    model.add(Dropout(dropout))\n    model.add(BatchNormalization())\n    #model.add(Dense(num_classes, activation='softmax'))\n    model.add(Dense(num_classes, activation='sigmoid'))\n    return model\n\nnum_classes = 2\n\ndataset_copy = dataset_df.copy() \n\nfeature_to_keep = [\"MGMT_value\"]\n\nfeature_outliers = ['glrlm_LongRunHighGrayLevelEmphasis','MajorAxisLength','gzlm_SizeZoneNonUniformityNormalized','Kurtosis','Idn',\n'Idmn','gldm_LargeDependenceHighGrayLevelEmphasis','Mean','MinorAxisLenth','glrlm_RunPercentage']\nfeature_non_outliers = ['InverseVariance','Maximum3DDiameter']\n\ndataset_copy = dataset_copy[feature_outliers + feature_non_outliers + feature_to_keep]\n\ntarget = dataset_copy[\"MGMT_value\"]\nfeatures = dataset_copy.drop(\"MGMT_value\", axis=1)\n\n# Encodage One-Hot des étiquettes\nonehot_encoder = OneHotEncoder(sparse=False)\ntarget_encoded = onehot_encoder.fit_transform(target.values.reshape(-1, 1))\n\n# Convertir le codage one-hot en classe binaire\ny_binary = np.argmax(target_encoded, axis=1)\n\n# Compter les occurrences de chaque classe dans y_binary\nclass_counts = np.bincount(y_binary)\n\n# Afficher les répartitions des classes\nfor cls, count in enumerate(class_counts):\n    print(f\"Classe {cls}: {count} échantillons\")\n\n# Diviser les données en ensembles d'entraînement et de test\nX_train, X_test, y_train, y_test = train_test_split(\n    features, y_binary, test_size=0.2, random_state=42\n)\n\n\n# Création d'une instance du sur-échantillonneur RandomOverSampler\noversampler = RandomOverSampler(random_state=42)\n\n# Appliquer le sur-échantillonnage sur vos données\nX_train, y_train = oversampler.fit_resample(X_train, y_train)\n\n# Normaliser les données d'entraînement et de test\nscaler = StandardScaler()\nX_train = scaler.fit_transform(X_train)\nX_test = scaler.transform(X_test)\n\n# Dimension d'entrée du modèle\ninput_dim = X_train.shape[1]\n\n# Liste des paramètres à tester\n# Meilleurs paramètres : (512, 'relu', 0.2, 256, 0.01)\n# Meilleure précision : 0.598313745856285\n\n# AUTRE\n#Accuracy: 0.5555555555555556 | Parameters: 512, relu, 1e-05, 512, 0.001\n#Best Parameters: {'layer_size': 512, 'activation': 'relu', 'dropout': 1e-05, 'batch_size': 512, 'learning_rate': 0.001}\n#Confusion Matrix:\n#array([[ 3, 49],\n#       [ 3, 62]])\n\n#Meilleurs paramètres : (512, 'relu', 1e-05, 256, 0.001)\n#Meilleure précision : 0.5630980551242828\n#Confusion Matrix:\n#array([[122, 132],\n#       [138, 116]])\n'''\nlayer_sizes = [512]\nactivations = ['relu']\ndropouts = [0.2]\nbatch_sizes = [256]\nlearning_rates = [0.01]\n'''\n'''\nlayer_sizes = [512]\nactivations = ['relu']\ndropouts = [1e-05]\nbatch_sizes = [512]\nlearning_rates = [0.001]\n'''\n\nlayer_sizes = [64]#, 128, 256, 512]\nactivations = ['relu','sigmoid']#,'leakyrelu']\ndropouts = [1e-05, 0.0001,0.001]#, 0.01, 0.1, 0.2,0.5,0.8]\nbatch_sizes = [16,32,64]#,128,256,512]\nlearning_rates = [0.0001,0.001,0.01]#, 0.1,0.2]\n\nepochs = 10\n\n# Effectuer une recherche par grille\nbest_accuracy = 0.0\nbest_params = {}\nconfusion_matrix_best = None\n\nkf = KFold(n_splits=10, shuffle=True)  # Utilisation de la validation croisée avec 10 plis\n\nfor params in itertools.product(layer_sizes, activations, dropouts, batch_sizes, learning_rates):\n    layer_size, activation, dropout, batch_size, learning_rate = params\n\n    accuracies = []  # Stocker les précisions pour chaque pli\n    y_pred_fold = []  # Stocker les prédictions pour chaque pli\n\n\n    for train_index, val_index in kf.split(X_train):\n        X_train_fold, X_val_fold = X_train[train_index], X_train[val_index]\n        y_train_fold, y_val_fold = y_train[train_index], y_train[val_index]\n\n        # Convertir les étiquettes en encodage one-hot\n        y_train_fold = to_categorical(y_train_fold)\n        y_val_fold = to_categorical(y_val_fold)\n        \n        # Construire et entraîner le modèle avec les paramètres actuels\n        model = build_model(layer_size, activation, dropout)\n        model.compile(optimizer=Adam(learning_rate=learning_rate), loss='categorical_crossentropy', metrics=['accuracy'])\n        #model.compile(optimizer=Adamax(learning_rate=learning_rate), loss='categorical_crossentropy', metrics=['accuracy'])\n        \n        history = model.fit(X_train_fold, y_train_fold, batch_size=batch_size, epochs=epochs, validation_data=(X_val_fold, y_val_fold), verbose=0)\n\n        # Évaluer les performances du modèle sur l'ensemble de validation actuel\n        accuracy = model.evaluate(X_val_fold, y_val_fold)[1]\n        accuracies.append(accuracy)\n        \n        # Obtenir les prédictions sur l'ensemble de validation actuel\n        y_pred = model.predict(X_val_fold)\n        y_pred = np.argmax(y_pred, axis=1)\n        y_pred_fold.extend(y_pred)\n\n    # Calculer la moyenne des précisions sur tous les plis\n    mean_accuracy = sum(accuracies) / len(accuracies)\n\n    # Vérifier si les performances actuelles sont les meilleures\n    if mean_accuracy > best_accuracy:\n        best_accuracy = mean_accuracy\n        best_params = params\n        y_pred_best = np.array(y_pred_fold)\n        confusion_matrix_best = confusion_matrix(y_train, y_pred_best)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Afficher les meilleurs paramètres et la meilleure précision\nprint(\"Meilleurs paramètres :\", best_params)\nprint(\"Meilleure précision :\", best_accuracy)\n\nlayer_size, activation, dropout, batch_size, learning_rate = best_params\n# Construire et entraîner le modèle avec les paramètres actuels\nmodel = build_model(layer_size, activation, dropout)\nmodel.compile(optimizer=Adam(learning_rate=learning_rate), loss='categorical_crossentropy', metrics=['accuracy'])\nmodel.fit(X_train, to_categorical(y_train), batch_size=batch_size, epochs=100)\n\ny_pred = np.argmax(model.predict(X_test), axis=1)\nprint(\"Confusion Matrix:\")\ndisplay(confusion_matrix(y_test, y_pred))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Juillet 2023 DAVID : Autre tentatives d'approches sur modèles ML\n----","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport itertools\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom matplotlib import offsetbox\nfrom mpl_toolkits import mplot3d\n\nfrom sklearn.model_selection import train_test_split, KFold, cross_val_score, GridSearchCV, StratifiedKFold, RandomizedSearchCV\nfrom sklearn.preprocessing import StandardScaler, RobustScaler, MinMaxScaler\nfrom sklearn.decomposition import PCA\nfrom sklearn.manifold import LocallyLinearEmbedding, MDS, Isomap, TSNE\nfrom sklearn.decomposition import PCA\nfrom sklearn.metrics import confusion_matrix, classification_report, make_scorer, accuracy_score, f1_score\nfrom sklearn.metrics import roc_curve, roc_auc_score\nfrom sklearn.metrics import make_scorer, precision_score, recall_score\n\nfrom sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier\nfrom sklearn.feature_selection import RFECV, SelectKBest, f_classif\nfrom imblearn.over_sampling import SMOTE\nfrom sklearn.calibration import CalibratedClassifierCV\nfrom sklearn.svm import SVC\nfrom xgboost.core import DMatrix\nimport xgboost as xgb\n\ndataset_copy = dataset_df.copy() \nexcluded_features_column_names = [\"LeastAxisLength\", \"Flatness\", \"gldm_GrayLevelNonUniformityNormalized\", \"gldm_DependencePercentage\"]\ndataset_df.drop(excluded_features_column_names, axis=1, inplace=True)\n\n# Patient BraTS21ID now is ID, and ID of Dataset\ndataset_copy = dataset_copy.set_index('ID')\n# Drop patient\nexcluded_patient_ids = [109, 123, 709] #, 11, 351, 353, 442, 445, 554, 556, 568, 570, 572, 578, 581, 584, 587, 589, 593, 594, 596, 610]\ndataset_copy = dataset_copy.drop(index=excluded_patient_ids)\n\nX = dataset_copy.drop(\"MGMT_value\",axis=1)\ny = dataset_copy['MGMT_value']\n\n\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"MODELE 1","metadata":{}},{"cell_type":"code","source":"pca = PCA()\npca.fit(X)\nplt.bar(range(len(pca.explained_variance_ratio_)), pca.explained_variance_ratio_)\nplt.xlabel('Composante principale')\nplt.ylabel('Ratio de variance expliquée')\nplt.show()\n\n#display(pca.n_components_)\ndisplay(pca.score(X))\npca = PCA(n_components=10)\n\nX = pca.fit_transform(X)\n\n# Division en ensembles de train et de test\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Utilisation de StratifiedKFold pour préserver les proportions des classes lors de la validation croisée\ncv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\n\n'''Utilisation de la classe StratifiedKFold pour effectuer une validation croisée plus fine \ntout en préservant les proportions des classes lors de la division des données. \nUtilisation de GridSearchCV avec une grille de paramètres plus spécifique pour trouver les \nmeilleurs hyperparamètres du modèle RandomForest. Ajout d'une étape de calibrage du modèle \navec CalibratedClassifierCV pour ajuster les seuils de décision et améliorer la performance \nde classification. Ce code ajoute la sélection des caractéristiques en utilisant SelectKBest \navec le test F pour choisir les meilleures caractéristiques. Les caractéristiques fortement \ncorrélées sont identifiées et exclues avant la sélection des meilleures caractéristiques. \nSMOTE pour augmenter artificiellement le jeu de données SelectKBest pour sélection des \nvariables \"pertinentes\"'''\n\n'''\n# Augmentation artificielle des données avec SMOTE\nsmote = SMOTE(random_state=42)\nX_train, y_train = smote.fit_resample(X_train, y_train)\n\n\n# Prétraitement des données : Standardisation des caractéristiques\n\nscaler = StandardScaler()\nX_train = scaler.fit_transform(X_train)\nX_test = scaler.transform(X_test)\n\n# Sélection des meilleures caractéristiques avec SelectKBest\nk = get_k(selected_features)\nselector = SelectKBest(score_func=f_classif, k=k)  # Définissez le nombre de caractéristiques à sélectionner\n\nX_train = selector.fit_transform(X_train, y_train)\nX_test = selector.transform(X_test)\n\n# Affichage des caractéristiques sélectionnées\nprint(\"Features sélectionnées :\")\nselected_features = [feature for feature, selected in zip(X.columns, selector.get_support()) if selected]\nprint(selected_features)\n'''\n\n# Définition des hyperparamètres à optimiser avec une validation croisée plus fine\n'''\nparam_grid = {\n  'kernel':['rbf'],#['linear','poly','rbf','sigmoid'],\n  'gamma': [0, 0.1, 0.2],\n}\n'''\n'''\n# RandomForestClassifier\nparam_grid = {\n    'n_estimators': [100], #[100,200,300,400,500],\n    'max_depth': [10], #,50,100,200],\n}\n'''\n# the best\nparam_grid = {\n 'max_depth': [20], #[10,20,30,50],\n 'max_features': ['sqrt'],\n 'min_samples_leaf': [1],#[1,2,5,10],\n 'min_samples_split': [5],#[1,2,5,10],\n 'n_estimators': [100]#[100,200,300]\n}\n\n'''\nparam_grid = {\n  'n_estimators': [100],\n  'learning_rate' : [0.1],\n  'max_depth':[2]\n}\n#'''\n'''\n# XGBoost\nparam_grid = {\n    'n_estimators': [100], #, 200, 300],\n    'learning_rate': [0.1, 0.3, 1.0],#0.01, 0.001],\n    'max_depth': [3, 5, 6, 7],\n    'min_child_weight': [1, 3, 5],\n    'subsample': [0.5, 0.8, 0.9, 1.0],\n    'colsample_bytree': [0.5, 0.8, 0.9, 1.0],\n    'gamma': [0, 0.1, 0.2],\n    'reg_alpha': [0, 0.1, 0.2],\n    'reg_lambda': [0, 0.1, 0.2],\n    #'objective': ['binary_logistic'],\n    'criterion': ['friedman_mse', 'squared_error'],\n    'min_weight_fraction_leaf': [0, 0.5],\n    'random_state' : [42]\n}\n#'''\n'''\nparam_grid = {\n 'colsample_bytree': [0.8],\n 'gamma': [0],\n 'learning_rate': [0.1],\n 'max_depth': [5],\n 'min_child_weight': [1],\n 'n_estimators': [100],\n 'reg_alpha': [0.2],\n 'reg_lambda': [0],\n 'subsample': [0.8]}\n#'''\n\n# Sélection récursive de variables avec validation croisée\nrf = RandomForestClassifier(random_state=42)\n#rf = xgb.XGBClassifier(random_state=42)\n#rf = GradientBoostingClassifier()\n#rf = SVC()\n# Optimisation des hyperparamètres avec validation croisée plus fine\ngrid_search = GridSearchCV(rf, param_grid, cv=cv, scoring='accuracy', error_score='raise')\n#grid_search = RandomizedSearchCV(rf, param_distributions=param_grid, n_iter=10, cv=cv)\n\ngrid_search.fit(X_train, y_train)\n\n# Meilleurs paramètres trouvés\nbest_params = grid_search.best_params_\ndisplay(\"Best params : \", best_params)\nbest_rf= grid_search.best_estimator_\n# Entraînement du modèle avec les meilleurs paramètres\n#best_rf = RandomForestClassifier(random_state=42, **best_params)\n#best_rf = xgb.XGBClassifier(random_state=42, **best_params)\nbest_rf.fit(X_train, y_train)\ny_pred = best_rf.predict(X_test)\n'''\n# Conversion des données d'entraînement en objet DMatrix\ndtrain = xgb.DMatrix(X_train, label=y_train)\n# Conversion des données de test en objet DMatrix\ndtest = xgb.DMatrix(X_test)\nmodel = xgb.train(best_params, dtrain)\ny_pred = model.predict(dtest)\n\n# Convertir les prédictions en classes binaires\ny_pred = (y_pred >= 0.5).astype(int)\n'''\n\n# ===========================\n# Resultats\n# ===========================\ndisplay('Accuracy train :', accuracy_score(y_train, best_rf.predict(X_train)))\n\n# Calcul de l'accuracy\naccuracy = accuracy_score(y_test, y_pred)\ndisplay('Accuracy : ', accuracy)\n\n# Calcul du F1-score\nf1 = f1_score(y_test, y_pred)\n\n# Calcul de la matrice de confusion\nconfusion = confusion_matrix(y_test, y_pred)\n\nreport = classification_report(y_test, y_pred)\nprint(report)\n\n\n# Création de la figure\nplt.figure(figsize=(8, 6))\n\n# Affichage de la matrice de confusion avec heatmap\nsns.heatmap(confusion, annot=True, fmt='d', cmap='Blues')\n\n# Ajout des labels des axes\nplt.xlabel('Prédictions')\nplt.ylabel('Vraies étiquettes')\nplt.title('Matrice de confusion')\n\n# Affichage de la figure\nplt.show()\n\n# Calcul de la courbe ROC\nfpr, tpr, thresholds = roc_curve(y_test, y_pred)\n\n# Calcul de l'aire sous la courbe ROC\nauc = roc_auc_score(y_test, y_pred)\n\n# Affichage de la courbe ROC\nplt.plot(fpr, tpr, label='Courbe ROC (AUC = {:.2f})'.format(auc))\nplt.plot([0, 1], [0, 1], 'k--')  # Diagonale aléatoire\nplt.xlabel('Taux de faux positifs')\nplt.ylabel('Taux de vrais positifs')\nplt.title('Courbe ROC')\nplt.legend(loc='lower right')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"MODELE 2\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import GridSearchCV, train_test_split\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import RobustScaler\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score\n\n# Créer un pipeline de prétraitement\npreprocessor = Pipeline([\n    ('scaler', RobustScaler())\n])\n\n# Définir les métriques d'évaluation pour optimiser le réglage du seuil\nscoring = {\n    'Precision': make_scorer(precision_score),\n    'Recall': make_scorer(recall_score),\n    'F1': make_scorer(f1_score)\n}\n\n\n# Créer un dictionnaire avec les modèles à tester et leurs grilles de paramètres\nmodels = {\n    'Logistic Regression': {\n        'model': LogisticRegression(),\n        'params': {\n            'model__C': [0.001,0.01,0.1, 1.0, 10.0],\n            'model__penalty': [None, 'l2'],\n            ##'random_state' : [42],\n            ##'model__threshold': [0.3, 0.4, 0.5, 0.6, 0.7]\n        }\n    },\n'Support Vector Machine': {\n    'model': SVC(probability=True),\n    'params': {\n        'model__C': [0.0001, 0.001, 0.01, 0.1, 1.0, 10.0, 20.0, 50.0],\n        'model__kernel': ['linear', 'rbf','poly','sigmoid'],\n        'model__gamma': ['scale', 'auto'],\n        'model__degree': [2, 3, 4],\n        'model__shrinking': [True, False],\n        #'random_state' : [42],\n        #'model__threshold': [0.3, 0.4, 0.5, 0.6, 0.7]\n    }\n},\n    'Random Forest': {\n        'model': RandomForestClassifier(),\n        'params': {\n            'model__n_estimators': [100, 200],\n            'model__max_depth': [3, 5, 7,20],\n            'model__min_samples_split': [2, 5, 10],\n            #'random_state' : [42],\n            #'model__threshold': [0.3, 0.4, 0.5, 0.6, 0.7]\n        }\n    },\n\n  # XGBoost\n   'XGBoost' : {\n      'model' : xgb.XGBClassifier(),\n      'params' : {\n        'n_estimators': [100, 200, 300],\n        #'learning_rate': ['memory', 'steps', 'verbose'], #[0.1, 0.3, 1.0, 0.01, 0.001],\n        'max_depth': [3, 5, 6, 7],\n\n      }\n  }\n}\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Effectuer la recherche des meilleurs hyperparamètres pour chaque modèle\nfor model_name, model_info in models.items():\n    print(\"Entraînement du modèle:\", model_name)\n\n    # Définir le modèle du pipeline\n    model_info['model'] = Pipeline([\n        ('preprocessor', preprocessor),\n        ('model', model_info['model'])\n    ])\n\n    # Effectuer la recherche des hyperparamètres optimaux\n    grid_search = GridSearchCV(model_info['model'], model_info['params'], cv=cv, error_score='raise', scoring=scoring, refit='F1')\n    grid_search.fit(X_train, y_train)\n\n    # Évaluer le modèle sur l'ensemble de test\n    y_pred = grid_search.predict(X_test)\n\n\n    # Afficher les meilleurs paramètres et scores\n    print(\"Meilleurs paramètres : \", grid_search.best_params_)\n    print(\"Meilleur score F1 : \", grid_search.best_score_)\n\n    # Utiliser le modèle avec les meilleurs paramètres sur les données de test\n    best_model = grid_search.best_estimator_\n    y_pred_proba = best_model.predict_proba(X_test)[:, 1]\n\n    # Appliquer un seuil manuellement pour la prédiction\n    threshold = 0.5  # Seuil par défaut\n    y_pred = (y_pred_proba >= threshold).astype(int)\n\n    accuracy = accuracy_score(y_test, y_pred)\n    precision = precision_score(y_test, y_pred)\n    recall = recall_score(y_test, y_pred)\n    f1 = f1_score(y_test, y_pred)\n    print(\"Accuracy train : \", accuracy_score(y_train, grid_search.predict(X_train)))\n    # Afficher les résultats\n    print(\"Meilleurs hyperparamètres:\", grid_search.best_params_)\n    print(\"Accuracy test:\", accuracy)\n    print(\"Précision:\", precision)\n    print(\"Rappel:\", recall)\n    print(\"Score F1:\", f1)\n    print(\"-------------------------------------\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"MODELE 3 : DL","metadata":{}},{"cell_type":"code","source":"def build_model_conv(layer_size, activation, dropout, input_shape):\n    model = Sequential()\n    model.add(Conv1D(layer_size, kernel_size=3, activation=activation, input_shape=input_shape))\n    model.add(MaxPooling1D(pool_size=2))\n    model.add(Conv1D(layer_size * 2, kernel_size=3, activation='relu'))\n    model.add(MaxPooling1D(pool_size=2))\n    model.add(Flatten())\n    model.add(Dense(layer_size, activation=activation))\n    model.add(Dropout(dropout))\n    model.add(Dense(1, activation='sigmoid'))\n    return model\n\ndef build_model(layer_size, activation, dropout, input_shape) :\n    # Définition de l'entrée du modèle\n    inputs = Input(shape=input_shape)\n\n    # Couches du modèle\n    x = Dense(layer_size)(inputs)\n    x = BatchNormalization()(x)\n    x = Activation(activation)(x)\n\n    x = Dense(layer_size*2)(x)\n    x = BatchNormalization()(x)\n    x = Activation(activation)(x)\n\n    x = Dense(layer_size*2)(x)\n    x = BatchNormalization()(x)\n    x = Activation(activation)(x)\n\n    # Connexion résiduelle\n    residual = Dense(layer_size*2)(inputs)\n    residual = BatchNormalization()(residual)\n\n    # Ajout de la connexion résiduelle\n    x = Add()([x, residual])\n    x = Activation('relu')(x)\n\n    # Couche de sortie\n    outputs = Dense(1, activation='sigmoid')(x)\n\n    # Création du modèle\n    model = Model(inputs=inputs, outputs=outputs)\n\n    return model\n\n# Paramètres du modèle\nepoch = 150\n\nlayer_sizes = [128]\nactivations = ['relu']#,'leakyrelu']\ndropouts = [0.1]#, 0.01, 0.1, 0.2,0.5,0.8]\nbatch_sizes = [128]#,128,256,512]\nlearning_rates = [0.1]#, 0.1,0.2]\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\nscaler = RobustScaler()\nX_train = scaler.fit_transform(X_train)\nX_test = scaler.transform(X_test)\n\ny_train = y_train.reset_index(drop=True)\n\nbest_accuracy = 0.0\nbest_params = {}\n\n\n\nkf = KFold(n_splits=10, shuffle=True)\n\nfor params in itertools.product(layer_sizes, activations, dropouts, batch_sizes, learning_rates):\n    layer_size, activation, dropout, batch_size, learning_rate = params\n    accuracies = []\n\n    for train_index, val_index in kf.split(X_train):\n        X_train_fold, X_val_fold = X_train[train_index], X_train[val_index]\n        y_train_fold, y_val_fold = y_train[train_index], y_train[val_index]\n\n        input_shape = (X_train_fold.shape[1], 1)\n        #model = build_model_conv(layer_size, activation, dropout, input_shape)\n        model = build_model(layer_size, activation, dropout, input_shape)\n        model.compile(optimizer=Adam(learning_rate=learning_rate), loss='binary_crossentropy', metrics=['accuracy'])\n        model.fit(X_train_fold, y_train_fold, batch_size=batch_size, epochs=epoch, verbose=0)\n        accuracy = model.evaluate(X_val_fold, y_val_fold)[1]\n        accuracies.append(accuracy)\n\n    mean_accuracy = sum(accuracies) / len(accuracies)\n\n    if mean_accuracy > best_accuracy:\n        best_accuracy = mean_accuracy\n        best_params = params\n        display(best_accuracy, best_params)\n\n\n\nlayer_size, activation, dropout, batch_size, learning_rate = best_params\n\ninput_shape = (X_train.shape[1], 1)\nbest_model = build_model_conv(layer_size, activation, dropout, input_shape)\nbest_model.compile(optimizer=Adam(learning_rate=learning_rate), loss='binary_crossentropy', metrics=['accuracy'])\n\n# Créer une instance du callback History\nhistory = History()\n\nbest_model.fit(X_train, y_train, batch_size=batch_size, epochs=epoch, verbose=1, callbacks=[history])\ntest_loss, test_accuracy = best_model.evaluate(X_test, y_test)\n\n# Afficher les résultats des epochs\nloss = history.history['loss']\naccuracy = history.history['accuracy']\n\n# Afficher les résultats\nepochs = range(1, len(loss) + 1)\nplt.plot(epochs, loss, 'b', label='Training loss')\nplt.plot(epochs, accuracy, 'r', label='Training accuracy')\nplt.title('Training Loss and Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Loss / Accuracy')\nplt.legend()\nplt.show()\n\n\n\ny_pred_prob = best_model.predict(X_test)\n# Restreindre les valeurs entre 0 et 1\ny_pred_prob = np.clip(y_pred_prob, 0, 1)\n\n# Convertir les valeurs en entiers (0 ou 1)\ny_pred = (y_pred_prob > 0.5).astype(int)\n\n# Vérifier les valeurs uniques dans y_test et y_pred\nunique_values_test = np.unique(y_test)\nunique_values_pred = np.unique(y_pred)\n\nprint(\"Unique values in y_test:\", unique_values_test)\nprint(\"Unique values in y_pred:\", unique_values_pred)\n\ncm = confusion_matrix(y_test, y_pred)\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues')\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.show()\n\ndisplay(best_accuracy, best_params)\n\n\nreport = classification_report(y_test, y_pred)\nprint(report)\n\n\n# Création de la figure\nplt.figure(figsize=(8, 6))\n\n# Affichage de la matrice de confusion avec heatmap\nsns.heatmap(confusion, annot=True, fmt='d', cmap='Blues')\n\n# Ajout des labels des axes\nplt.xlabel('Prédictions')\nplt.ylabel('Vraies étiquettes')\nplt.title('Matrice de confusion')\n\n# Affichage de la figure\nplt.show()\n\n\n\n# Calcul de la courbe ROC\nfpr, tpr, thresholds = roc_curve(y_test, y_pred)\n\n# Calcul de l'aire sous la courbe ROC\nauc = roc_auc_score(y_test, y_pred)\n\n# Affichage de la courbe ROC\nplt.plot(fpr, tpr, label='Courbe ROC (AUC = {:.2f})'.format(auc))\nplt.plot([0, 1], [0, 1], 'k--')  # Diagonale aléatoire\nplt.xlabel('Taux de faux positifs')\nplt.ylabel('Taux de vrais positifs')\nplt.title('Courbe ROC')\nplt.legend(loc='lower right')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}