{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA-MICCAI Brain Tumor Radiogenomic Classification\nTask: Detect the presence of MGMT promoter methylation from the MRI images\n\n### Domain Knowledge:\n* **Glioblastoma** is the most common form of brain cancer and considered the deadliest human cancer.\n* **MGMT promoter methylation** is the key mechanism of MGMT gene silencing and predicts a favorable outcome in patients with glioblastoma who are exposed to alkylating agent chemotherapy.\n* **MGMT gene** - ($ O_6$-methylguanine-DNA methyltransferase) is a DNA repair enzyme that protects normal and glioma cells from alkylating chemotherapeutic agents.\n* **Methylation** - presence of it means that the gene is repressed (not expressed)\n\nThus, the goal of this competition is not on detecting tumors (since all the images in train and test sets have tumors/glioblastomas) but on predicting whether or not chemotherapy will be an effective treatment by predicting if the brain is methylated (MGMT_value=0) or not (MGMT_value=1). If successful, doctors will have better clinical decisions on the type of treatments to be given to their cancer patients with glioblastoma. \n\n### Published Papers:\n* ([levner2009](https://pubmed.ncbi.nlm.nih.gov/20426152/)) - MGMT methylation of glioblastomas can be predicted from the MRI. Diffused tumor (texture) could be attributed to MGMT methylation. \n   * Thus, we can try to first localize the tumor in the MRI then extract its texture features then classify if it is diffused or not\n* ([crisi2020](https://pubmed.ncbi.nlm.nih.gov/32374045/)) - MGMT methylation of glioblastomas can be predicted from perfusion dynamic susceptibility contrast magnetic resonance imaging (DSC-MRI). 92 quantitative image features were obtained from relative cerebral blood volume and relative cerebral blood flow maps\n    * Can we extract these features from a mere MRI?","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2\nimport pydicom\nfrom tqdm import tqdm\nfrom multiprocessing import Pool\nfrom sklearn.preprocessing import OrdinalEncoder","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:54:32.280378Z","iopub.execute_input":"2021-08-31T11:54:32.280784Z","iopub.status.idle":"2021-08-31T11:54:33.747387Z","shell.execute_reply.started":"2021-08-31T11:54:32.280703Z","shell.execute_reply":"2021-08-31T11:54:33.746462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pd.read_csv('../input/rsna-miccai-brain-tumor-radiogenomic-classification/sample_submission.csv')\ntrain_labels = pd.read_csv('../input/rsna-miccai-brain-tumor-radiogenomic-classification/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:54:33.748955Z","iopub.execute_input":"2021-08-31T11:54:33.749361Z","iopub.status.idle":"2021-08-31T11:54:33.767294Z","shell.execute_reply.started":"2021-08-31T11:54:33.749299Z","shell.execute_reply":"2021-08-31T11:54:33.766207Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:54:33.769015Z","iopub.execute_input":"2021-08-31T11:54:33.769306Z","iopub.status.idle":"2021-08-31T11:54:33.796072Z","shell.execute_reply.started":"2021-08-31T11:54:33.769276Z","shell.execute_reply":"2021-08-31T11:54:33.795213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The train set seems balanced in terms of the target labels (MGMT_value).","metadata":{}},{"cell_type":"code","source":"with plt.xkcd():\n    sns.countplot(x=train_labels.MGMT_value)\n    \ntrain_labels.MGMT_value.value_counts()    ","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:54:33.799327Z","iopub.execute_input":"2021-08-31T11:54:33.799654Z","iopub.status.idle":"2021-08-31T11:54:33.991483Z","shell.execute_reply.started":"2021-08-31T11:54:33.799622Z","shell.execute_reply":"2021-08-31T11:54:33.990414Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The sample_submission is only composed of 87 data points which is 22% of the total test set. This means that the total test set is only composed of ~395 data points. As explained by some of the competitors in this [discussion thread](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/268066), one reason why this competition is not crowded is because the test set is very small such that the final standings could be a lottery.","metadata":{}},{"cell_type":"code","source":"sample_submission","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:54:33.993767Z","iopub.execute_input":"2021-08-31T11:54:33.99407Z","iopub.status.idle":"2021-08-31T11:54:34.008371Z","shell.execute_reply.started":"2021-08-31T11:54:33.99404Z","shell.execute_reply":"2021-08-31T11:54:34.007196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's explore the folder structure of the new dataset with .png files converted from the original .dcm files. The rsna-miccai-png folder contains 2 main subfolders: `test` and `train`. The test folder contains 87 main subfolders which correspond to the BraTS21ID or simply the case ID written as a 5-digit number (e.g. `00114`). Each case ID  folder contains 4 subfolders: `T2w`, `T1wCE`, `T1w`, and `FLAIR`, which corresponds to the image contrast. The tricky part is that the 4 subfolders contain varying number of images (roughly around 19 to 300 per folder).\n\n### Folder: rsna-miccai-png\n\n    |--test\n    |   |\n    |   |--00114\n    |   |   |\n    |   |   |--T2w\n    |   |   |   |\n    |   |   |   |--Image-4.png\n    |   |   |   |--Image-19.png\n    |   |   |   |...\n    |   |   |\n    |   |   |--T1wCE\n    |   |   |   |\n    |   |   |   |--Image-79.png\n    |   |   |   |--Image-64.png\n    |   |   |   |...\n    |   |   |\n    |   |   |--T1w\n    |   |   |   |\n    |   |   |   |--Image-4.png\n    |   |   |   |--Image-5.png\n    |   |   |   |...\n    |   |   |\n    |   |   |--FLAIR\n    |   |       |\n    |   |       |--Image-4.png\n    |   |       |--Image-19.png\n    |   |       |...\n    |   |       \n    |   |--00013\n    |   | ...\n    |   | ...\n    |   | ...\n    |    \n    |--train\n        |\n        |...\n        |...\n    ","metadata":{}},{"cell_type":"markdown","source":"### Count of images per subfolder of the test set\n\nLet's create a new dataframe `test_df` that has the count of images per subfolder.","metadata":{}},{"cell_type":"code","source":"TEST_DIR = '../input/rsna-miccai-png/test/'\n\ncaseID = []\nT2w = []\nT1wCE = []\nT1w = []\nFLAIR = []\nfor dirname, _, filenames in os.walk(f'{TEST_DIR}'):\n    for filename in filenames:\n        case_id = dirname.split('/')[-2]\n        folder = dirname.split('/')[-1]\n        count = len(os.listdir(dirname))\n\n        if case_id not in caseID:\n            caseID.append(case_id)\n        \n        if folder == 'T2w':\n            T2w.append(count)\n        elif folder == 'T1wCE':\n            T1wCE.append(count)\n        elif folder == 'T1w':\n            T1w.append(count)\n        elif folder == 'FLAIR':\n            FLAIR.append(count)\n        break\n\nsample_submission['caseID'] = sample_submission['BraTS21ID'].astype(str).str.zfill(5)      \n        \ntest_df = pd.DataFrame({\n                        'caseID':caseID,\n                        'T2w_count':T2w,\n                        'T1wCE_count':T1wCE,\n                        'T1w_count':T1w,\n                        'FLAIR_count':FLAIR,\n                       })        \n\ntest_df = test_df.merge(sample_submission[['caseID','BraTS21ID']], on='caseID',how='left')\ntest_df","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:54:34.010127Z","iopub.execute_input":"2021-08-31T11:54:34.010605Z","iopub.status.idle":"2021-08-31T11:54:43.625439Z","shell.execute_reply.started":"2021-08-31T11:54:34.01056Z","shell.execute_reply":"2021-08-31T11:54:43.62451Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Count of images per subfolder of the train set\n\nIn the same way, let's create a new dataframe `train_df` that has the count of images per subfolder. Note that according to the competition host ([discussion thread](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/262046)), there are three case ids (`00109`, `00123`, `00709`) in the train set that should be excluded because they contain unexpected errors (e.g. missing images). ","metadata":{}},{"cell_type":"code","source":"TRAIN_DIR = '../input/rsna-miccai-png/train/'\n\ncaseID = []\nT2w = []\nT1wCE = []\nT1w = []\nFLAIR = []\nfor dirname, _, filenames in os.walk(f'{TRAIN_DIR}'):\n    for filename in filenames:\n        case_id = dirname.split('/')[-2]\n        folder = dirname.split('/')[-1]\n\n        if case_id not in ('00109', '00123', '00709'):\n            count = len(os.listdir(dirname))\n\n            if case_id not in caseID:\n                caseID.append(case_id)\n\n            if folder == 'T2w':\n                T2w.append(count)\n            elif folder == 'T1wCE':\n                T1wCE.append(count)\n            elif folder == 'T1w':\n                T1w.append(count)\n            elif folder == 'FLAIR':\n                FLAIR.append(count)\n            break\n\ntrain_labels['caseID'] = train_labels['BraTS21ID'].astype(str).str.zfill(5)      \n        \ntrain_df = pd.DataFrame({\n                        'caseID':caseID,\n                        'T2w_count':T2w,\n                        'T1wCE_count':T1wCE,\n                        'T1w_count':T1w,\n                        'FLAIR_count':FLAIR,\n                       })        \n\ntrain_df = train_df.merge(train_labels, on='caseID',how='left')\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:54:43.626558Z","iopub.execute_input":"2021-08-31T11:54:43.626804Z","iopub.status.idle":"2021-08-31T11:55:35.239649Z","shell.execute_reply.started":"2021-08-31T11:54:43.62678Z","shell.execute_reply":"2021-08-31T11:55:35.238913Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlation between the number of images vs. the MGMT value\n\nThere seems to be no correlation between the number of images per MRI view per case id vs. the MGMT value.","metadata":{}},{"cell_type":"code","source":"corr = train_df.corr()\nsns.heatmap(corr, \n            xticklabels=corr.columns.values,\n            yticklabels=corr.columns.values)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:35.241487Z","iopub.execute_input":"2021-08-31T11:55:35.241927Z","iopub.status.idle":"2021-08-31T11:55:35.485698Z","shell.execute_reply.started":"2021-08-31T11:55:35.241883Z","shell.execute_reply":"2021-08-31T11:55:35.484961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Get the best image per MRI view per case_id\n\nThere is a lot of MRI images per folder but not all of them are useful. In fact, majority of the scans are mostly black. Let's try to choose only the image with the most number of nonzero (nonblack) pixels per folder per case id. Currently, I do not know how to make use of the other images so I will only use these chosen images for now. ","metadata":{}},{"cell_type":"code","source":"test_best_images = pd.read_csv(f'../input/images-with-the-most-nonzero-pixels/test_df.csv')\ntest_df = test_df.merge(test_best_images[['BraTS21ID','T1w','T1wCE','T2w','FLAIR']], on='BraTS21ID', how='left')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:35.486992Z","iopub.execute_input":"2021-08-31T11:55:35.48744Z","iopub.status.idle":"2021-08-31T11:55:35.511591Z","shell.execute_reply.started":"2021-08-31T11:55:35.487389Z","shell.execute_reply":"2021-08-31T11:55:35.510483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:55:35.512948Z","iopub.execute_input":"2021-08-31T11:55:35.513265Z","iopub.status.idle":"2021-08-31T11:55:35.534584Z","shell.execute_reply.started":"2021-08-31T11:55:35.513234Z","shell.execute_reply":"2021-08-31T11:55:35.533576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_best_images = pd.read_csv(f'../input/images-with-the-most-nonzero-pixels/train_df.csv')\ntrain_df = train_df.merge(train_best_images[['BraTS21ID','T1w','T1wCE','T2w','FLAIR']], on='BraTS21ID', how='left')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:35.536032Z","iopub.execute_input":"2021-08-31T11:55:35.536329Z","iopub.status.idle":"2021-08-31T11:55:35.56317Z","shell.execute_reply.started":"2021-08-31T11:55:35.536302Z","shell.execute_reply":"2021-08-31T11:55:35.562052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2021-08-31T11:55:35.564604Z","iopub.execute_input":"2021-08-31T11:55:35.565036Z","iopub.status.idle":"2021-08-31T11:55:35.58942Z","shell.execute_reply.started":"2021-08-31T11:55:35.56499Z","shell.execute_reply":"2021-08-31T11:55:35.588448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's look at some images with Methylation (MGMT_value=0)","metadata":{}},{"cell_type":"code","source":"case_id = '00688'\nfilename = train_df[train_df.caseID==case_id].T1w.iloc[0]\nt1w = plt.imread(f'{TRAIN_DIR}{case_id}/T1w/{filename}')\nfilename = train_df[train_df.caseID==case_id].T1wCE.iloc[0]\nt1wce = plt.imread(f'{TRAIN_DIR}{case_id}/T1wCE/{filename}')\nfilename = train_df[train_df.caseID==case_id].T2w.iloc[0]\nt2w = plt.imread(f'{TRAIN_DIR}{case_id}/T2w/{filename}')\nfilename = train_df[train_df.caseID==case_id].FLAIR.iloc[0]\nflair = plt.imread(f'{TRAIN_DIR}{case_id}/FLAIR/{filename}')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:35.590686Z","iopub.execute_input":"2021-08-31T11:55:35.590977Z","iopub.status.idle":"2021-08-31T11:55:35.635703Z","shell.execute_reply.started":"2021-08-31T11:55:35.590946Z","shell.execute_reply":"2021-08-31T11:55:35.63457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(26,6))\nplt.gray()\nax1 = fig.add_subplot(141)\nplt.imshow(t1w, aspect='auto')\nax2 = fig.add_subplot(142)\nplt.imshow(t1wce, aspect='auto')\nax3 = fig.add_subplot(143)\nplt.imshow(t2w, aspect='auto')\nax4 = fig.add_subplot(144)\nplt.imshow(flair, aspect='auto')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:35.637099Z","iopub.execute_input":"2021-08-31T11:55:35.637442Z","iopub.status.idle":"2021-08-31T11:55:36.494847Z","shell.execute_reply.started":"2021-08-31T11:55:35.637409Z","shell.execute_reply":"2021-08-31T11:55:36.493754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's look at some images without Methylation (MGMT_value=1)","metadata":{}},{"cell_type":"code","source":"case_id = '00058'\nfilename = train_df[train_df.caseID==case_id].T1w.iloc[0]\nt1w = plt.imread(f'{TRAIN_DIR}{case_id}/T1w/{filename}')\nfilename = train_df[train_df.caseID==case_id].T1wCE.iloc[0]\nt1wce = plt.imread(f'{TRAIN_DIR}{case_id}/T1wCE/{filename}')\nfilename = train_df[train_df.caseID==case_id].T2w.iloc[0]\nt2w = plt.imread(f'{TRAIN_DIR}{case_id}/T2w/{filename}')\nfilename = train_df[train_df.caseID==case_id].FLAIR.iloc[0]\nflair = plt.imread(f'{TRAIN_DIR}{case_id}/FLAIR/{filename}')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:36.49616Z","iopub.execute_input":"2021-08-31T11:55:36.496568Z","iopub.status.idle":"2021-08-31T11:55:36.540743Z","shell.execute_reply.started":"2021-08-31T11:55:36.496531Z","shell.execute_reply":"2021-08-31T11:55:36.539685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(26,6))\nplt.gray()\nax1 = fig.add_subplot(141)\nplt.imshow(t1w, aspect='auto')\nax2 = fig.add_subplot(142)\nplt.imshow(t1wce, aspect='auto')\nax3 = fig.add_subplot(143)\nplt.imshow(t2w, aspect='auto')\nax4 = fig.add_subplot(144)\nplt.imshow(flair, aspect='auto')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:36.543396Z","iopub.execute_input":"2021-08-31T11:55:36.543845Z","iopub.status.idle":"2021-08-31T11:55:37.25516Z","shell.execute_reply.started":"2021-08-31T11:55:36.543795Z","shell.execute_reply":"2021-08-31T11:55:37.254175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Extract Metadata from DICOM\n\nThanks to ([this code](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/252942)) by @pestipeti\n\nThere are two features (`WindowWidth` and `WindowCenter`) in the metadata that are difficult to include in the correlation matrix due to different data types: (string, float, int) and different values (1, 1.0, '1.0', '[1.0, 1.0]').","metadata":{}},{"cell_type":"code","source":"metadata = pd.read_csv('../input/extract-metadata-from-dicom/dicom_meta_train.csv')\n\nimage = []\ndicom_src = metadata['dicom_src'].tolist()\nfor src in dicom_src:\n    image.append(src.split('/')[-1].replace('dcm','png'))\n\nmetadata['image'] = image\n\n# Check if there's a data leak in timestamp\nmetadata['new_timestamp'] = metadata['timestamp'] - np.mean(metadata['timestamp'])\nmetadata = metadata.merge(train_df[['BraTS21ID','MGMT_value']], on='BraTS21ID', how='left')\nmetadata","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:37.256599Z","iopub.execute_input":"2021-08-31T11:55:37.257113Z","iopub.status.idle":"2021-08-31T11:55:45.878136Z","shell.execute_reply.started":"2021-08-31T11:55:37.257072Z","shell.execute_reply":"2021-08-31T11:55:45.877196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert object data into numerical\n\nobject_cols = metadata.select_dtypes('object').columns\nobject_cols = [col for col in object_cols if col not in ('WindowWidth', 'WindowCenter')]\n\nordinal_encoder = OrdinalEncoder()\nmetadata[object_cols] = metadata[object_cols].fillna('NA')\nmetadata[object_cols] = ordinal_encoder.fit_transform(metadata[object_cols])\nmetadata.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:45.87954Z","iopub.execute_input":"2021-08-31T11:55:45.879925Z","iopub.status.idle":"2021-08-31T11:55:56.270998Z","shell.execute_reply.started":"2021-08-31T11:55:45.879885Z","shell.execute_reply":"2021-08-31T11:55:56.269998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlation between MGMT value and Metadata features\n\nThere seems to be **no correlation** between MGMT value and any other metadata features. Note: MGMT_value is at the last row/column and the feature with the best correlation with it is `SpatialResolution` which mainly contains *NaN* values.","metadata":{}},{"cell_type":"code","source":"# Exclude metadata features that are empty and irrelevant\nnotnull_cols = [col for col in metadata.columns if col not in ('dataset',\n                                                               'ImageDimensions', \n                                                               'ImageLocation',\n                                                               'SOPClassUID',\n                                                               'SpacingBetweenSlices',\n                                                               'type',\n                                                               'dicom_src',\n                                                               'timestamp',\n                                                               )]\n\ncorr = metadata[notnull_cols].corr()\nsns.set(rc={'figure.figsize':(30,26)})\nsns.heatmap(corr, \n            xticklabels=corr.columns.values,\n            yticklabels=corr.columns.values,\n            )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-31T11:55:56.272068Z","iopub.execute_input":"2021-08-31T11:55:56.272317Z","iopub.status.idle":"2021-08-31T11:56:00.169204Z","shell.execute_reply.started":"2021-08-31T11:55:56.272292Z","shell.execute_reply":"2021-08-31T11:56:00.16826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}